Skip to main content

πŸ“ 1D Convolution (Conv1D)

Description​

< What is it? >​

A 1D convolution slides learnable filters along one axis of a sequence, such as time, audio samples, sensor readings, or token positions. Each filter looks at a small local window and produces one value in an output feature sequence.

This page uses the common NCL layout: (N, C, L), where N is batch size, C is channels, and L is sequence length.

Key points​

< Tensor shapes >​

  • x: input sequence (N, C_in, L_in)
  • w: filter weights (C_out, C_in, K)
  • b: optional bias of length C_out
  • s: stride along the sequence
  • p: zero-padding at each end of the sequence

One output filter has shape (C_in, K). It contains one length-K kernel slice per input channel; their results are summed to produce one output channel.

< Output length >​

Lout=⌊Lin+2pβˆ’KsβŒ‹+1L_{\text{out}} = \left\lfloor \frac{L_{\text{in}} + 2p - K}{s} \right\rfloor + 1

< Local sequence patterns >​

With kernel size K = 3, each output value summarizes three nearby sequence positions. Different filters can learn patterns such as a short audio event, a local sensor trend, or a nearby combination of embedding features.

Stacking Conv1D layers expands the effective receptive field, allowing later layers to combine information from longer portions of the sequence.

< Conv1D and Conv2D >​

Conv1D slides across one sequence axis. Conv2D slides across height and width, so it is commonly used for images. Both use local connectivity, shared weights, padding, stride, and learnable filters.

< Valid output >​

In valid mode, the kernel is evaluated only where it fully overlaps the input. For an input of length n and a kernel of length k, stride 1 produces n - k + 1 outputs. An empty kernel or a kernel longer than the input has no valid placement, so this implementation returns an empty list.

< Cross-correlation >​

Deep-learning layers usually use cross-correlation: multiply the kernel in its given order. Mathematical convolution reverses the kernel first, but this function does not reverse it.

< Scalar bias >​

A scalar bias is added once to every output value after the window–kernel products are summed.

Favorites​

⛸️