π 2D Convolution (Conv2D)
Descriptionβ
< What is it? >β
A 2D convolution layer slides each output filter over a padded input, multiplies the filter and local input window element by element, sums across all input channels, and optionally adds one bias per output channel.
For multi-channel input, one output filter has shape (C_in, K_h, K_w) and contains one 2D kernel slice per input channel. The slices are applied to their corresponding channels and summed to produce one output channel. In deep-learning usage, filter and kernel are often used interchangeably.
For an NCHW input of shape (N, C_in, H, W) and filters of shape (C_out, C_in, K_h, K_w), the output shape is (N, C_out, H_out, W_out).
< What is NCHW? >β
NCHW is a tensor layout used in CNNs and deep-learning frameworks. It gives the dimension order for a batch of images.
| Letter | Meaning | Typical size |
|---|---|---|
N | Number of images in the batch | e.g., 32 |
C | Channels: C_in for the layer input and C_out for its output | e.g., 3 β 64 |
H | Height of each image or feature map | e.g., 224 |
W | Width of each image or feature map | e.g., 224 |
For example, an NCHW tensor with shape (32, 3, 224, 224) means:
32 images Γ 3 channels Γ 224 height Γ 224 width
C_in is the number of channels entering the layer, such as 3 for an RGB image. C_out is the number of learned filters and therefore the number of output feature maps. For example, a layer with C_in = 3 and C_out = 64 maps RGB input to 64 output channels; later layers usually use the previous layerβs C_out as their C_in.
Key pointsβ
< Tensor shapes >β
x: input tensor(N, C_in, H, W)w: filter tensor(C_out, C_in, K_h, K_w)K_h,K_w: kernel height and width;K_h = 3,K_w = 3means a3 Γ 3kernel, whileK_h = 3,K_w = 5means a non-square3 Γ 5kernelb: optional bias of lengthC_outs_h,s_w: height and width stridesp_h,p_w: height and width zero-padding