π 1D Convolution (Conv1D)
Descriptionβ
< What is it? >β
A 1D convolution slides learnable filters along one axis of a sequence, such as time, audio samples, sensor readings, or token positions. Each filter looks at a small local window and produces one value in an output feature sequence.
This page uses the common NCL layout: (N, C, L), where N is batch size, C is channels, and L is sequence length.
Key pointsβ
< Tensor shapes >β
x: input sequence(N, C_in, L_in)w: filter weights(C_out, C_in, K)b: optional bias of lengthC_outs: stride along the sequencep: zero-padding at each end of the sequence
One output filter has shape (C_in, K). It contains one length-K kernel slice per input channel; their results are summed to produce one output channel.
< Output length >β
< Local sequence patterns >β
With kernel size K = 3, each output value summarizes three nearby sequence positions. Different filters can learn patterns such as a short audio event, a local sensor trend, or a nearby combination of embedding features.
Stacking Conv1D layers expands the effective receptive field, allowing later layers to combine information from longer portions of the sequence.
< Conv1D and Conv2D >β
Conv1D slides across one sequence axis. Conv2D slides across height and width, so it is commonly used for images. Both use local connectivity, shared weights, padding, stride, and learnable filters.
< Valid output >β
In valid mode, the kernel is evaluated only where it fully overlaps the input. For an input of length n and a kernel of length k, stride 1 produces n - k + 1 outputs. An empty kernel or a kernel longer than the input has no valid placement, so this implementation returns an empty list.
< Cross-correlation >β
Deep-learning layers usually use cross-correlation: multiply the kernel in its given order. Mathematical convolution reverses the kernel first, but this function does not reverse it.
< Scalar bias >β
A scalar bias is added once to every output value after the windowβkernel products are summed.
Favoritesβ
βΈοΈ
Related ideasβ
- CNN introduces kernels, feature maps, and local connectivity.
- 2D Convolution (Conv2D) extends the operation to two spatial axes.
- Recurrent Neural Network (RNN) is another family of models for sequential data.
- LSTM adds gated memory to an RNN.