Skip to main content

πŸ“ Convolutional Neural Network (CNN)

Description​

< What is it? >​

A convolutional neural network (CNN) is a neural network designed for grid-like data, especially images. Instead of connecting every pixel to every neuron, it applies small learnable filters across local regions of the input.

Those filters can learn patterns such as edges, textures, and shapes. Stacking convolutional layers lets a CNN combine simple local patterns into richer visual features.

image β†’ convolution + activation β†’ feature maps β†’ downsampling β†’ task head

The same filter is reused at every location. A CNN can therefore learn a visual pattern once and recognize it when it appears elsewhere in the image.

  • cnn
  • feature extraction
  • pooling layers

Key points​

< Local receptive fields and feature maps >​

Each output value is computed from a small input window, called its receptive field. A filter slides across the input and produces one feature map, which shows where that pattern was detected.

Early layers often detect edges or colors; later layers combine those signals into textures, parts, and higher-level objects.

< Local connectivity >​

Local connectivity means that one convolution output value connects only to a nearby input patch, not to every pixel in the image. For example, one value from a 3 Γ— 3 convolution on an RGB image uses 3 Γ— 3 Γ— 3 = 27 input values.

This matches image structure: nearby pixels often form related patterns such as edges, shapes or textures. It also reduces the number of connections and parameters. By stacking convolutional layers, later values can still receive information from a much larger effective region of the original image.

< What is a feature map? >​

A feature map is the 2D grid produced when one filter scans one input image. Each value says how strongly that filter’s learned pattern appears at one spatial location.

For output shape (N, C_out, H_out, W_out), each of the C_out channels is one feature map for each image in the batch. For example, a filter that responds to vertical edges produces large values where vertical edges appear. After an activation function, a feature map is also often called an activation map.

< Kernel (filter) >​

A kernel, also called a filter, is a small grid of learnable weights that scans across the input. For example, a 3 Γ— 3 kernel reads one 3 Γ— 3 input patch at a time, multiplies corresponding values, and sums them to produce one value in a feature map.

As the kernel slides, the same weights look for the same pattern at each location. Different filters can learn to respond to different patterns, such as vertical edges, corners, or textures.

In casual use, kernel and filter mean the same thing. More precisely, for multi-channel input, one filter contains one 2D kernel per input channel; their results are summed to produce one output channel.

< Weight sharing >​

A 3 Γ— 3 filter has nine weights per input channel and reuses them at every image location. This requires far fewer parameters than a fully connected layer and gives CNNs translation equivariance: shifting an input pattern shifts its feature response by the same amount.

Downsampling and a task head can make the final prediction less sensitive to small shifts, but a CNN is not automatically perfectly translation-invariant.

< Channels and filters >​

An RGB image has three input channels. One convolutional filter contains a small kernel for every input channel and produces one output feature map. Using many filters produces many output channels.

  • Input: (N, C_in, H, W)
  • Filter weights: (C_out, C_in, K_h, K_w)
  • Output: (N, C_out, H_out, W_out)

Here, N is the batch size, C is the number of channels, and H and W are height and width.

C_in is the number of channels entering the layer. For a normal RGB image, C_in = 3 because it has red, green, and blue channels.

C_out is the number of filters the layer learns, and therefore the number of feature maps it produces. For example, 64 filters applied to an RGB image use weights of shape (64, 3, 3, 3) and produce 64 output channels. In a later convolutional layer, C_in usually equals the previous layer’s C_out.

< Feature extraction, not necessarily dimensionality reduction >​

The main objective of a convolutional layer is feature extraction: it transforms input channels into feature maps that respond to useful patterns.

This does not necessarily reduce the size of the representation. A layer can reduce spatial dimensions H and W with stride or pooling while increasing C_out. For example, (N, 3, 224, 224) can become (N, 64, 112, 112): each feature map is smaller, but there are many more feature maps. Whether the total number of activations shrinks depends on the architecture.

< Padding, stride, and downsampling >​

  • Padding adds values, usually zeros, around the border so filters can use edge pixels and optionally preserve spatial size.
  • Stride is the number of pixels a filter moves at each step. A larger stride reduces the output height and width.
  • Pooling downsamples a feature map with a fixed operation such as max or average. Modern CNNs often use strided convolutions instead.

< Learning the filters >​

The filter weights and biases are learned from data using backpropagation and gradient descent. Convolutions are normally followed by a non-linear activation function, such as ReLU, so stacked layers can model complex patterns.

< What is convolution? >​

Convolution is a mathematical operation that combines two functions or vectors to produce a third. In the context of machine learning and image processing, it's used to extract features from input data (like images) by sliding a small filter (also called a kernel) over the input.

< A common CNN pattern >​

Image classifiers often repeat convolution β†’ normalization β†’ activation β†’ downsampling blocks, then use a classification head. CNNs are also widely used as feature extractors in detection, segmentation, medical imaging, and audio models using spectrograms.

< Why use CNNs for images >​

Q: Why use CNNs for images instead of fully connected networks?
A: CNNs are more efficient and effective for images because they exploit the spatial structure of the data. They use local receptive fields, weight sharing, and pooling to reduce the number of parameters and capture hierarchical features, making them better suited for image recognition tasks.

Comparison​

< What is a Vision Transformer? >​

A Vision Transformer (ViT) divides an image into fixed-size patches, turns those patches into tokens, and uses self-attention to let image regions exchange information. A CNN begins with local filters and builds broader context by stacking layers.

< CNN patch vs ViT patch >​

In a CNN, patch usually informally means the local receptive field: the small input window currently under a kernel. A 3 Γ— 3 kernel visits many such windows, often with overlap, and uses each one to calculate one feature-map value. The window is not normally kept as a token.

In a standard ViT, a patch is a fixed image block, often 16 Γ— 16, that is usually non-overlapping. It is flattened and projected into one token, which remains in the token sequence processed by self-attention.

  • CNN patch: a small sliding local window used by a learned kernel; stacked convolutions build larger receptive fields
  • ViT patch: an image block converted into a token; position embeddings and attention relate it to other patches

< Is a ViT better than a CNN? >​

No architecture is universally better. ViTs can be very strong when large-scale pretraining, transfer learning, or global relationships between image regions matter. CNNs have useful built-in locality and translation biases, so they are often strong choices with limited data or when efficient local processing matters.

  • CNN: local filters and useful locality/translation biases
  • ViT: patch tokens and attention, which can connect distant image regions directly
  • In practice: compare well-trained, pretrained models on the target data, accuracy, latency, memory, and deployment hardware rather than choosing by architecture name alone

Modern CNNs such as ConvNeXt remain competitive, and hybrid models also combine convolution with attention.

< Should I still learn CNNs? >​

Yes. CNNs remain widely used, and their core ideasβ€”kernels, feature maps, padding, stride, downsampling, and receptive fieldsβ€”are foundational for computer vision. Understanding them also makes it easier to understand ViT patch embeddings, hierarchical vision models, and convolutional components used alongside transformers.

Learn CNNs first for the visual-processing fundamentals, then learn ViTs to understand attention-based vision models. Neither has made the other obsolete.

🏊

Video Tutorial​

Reference​