Vision as a hierarchy of features

The Convolution & Cross-Correlation Operation
The key operation performed in CNN layers is that of 2D convolution. In fact in practice they are 4D convolutions as we try to learn many filters and we also consider many input images (mini-batch) in the iteration of our SGD optimizer. Convolution itself, in one and two dimensions, with its properties and boundary handling, is covered in Convolution and Linear Filters. Here you need only the 2D case and how a CNN layer uses it.2D Convolution
The 2D convolution of an input with a kernel (filter) , usually much smaller than the input, is Convolution is commutative, so you can flip either the input or the kernel; frameworks differ only in which form they write. Many ML frameworks don’t even implement convolution: they compute the very similar cross-correlation and call it convolution, which adds to the confusion. Both PyTorch’sConv2d and TensorFlow compute cross-correlation under the hood:

From fixed filters to learned filters
Classical image processing applies filters whose weights you choose by hand: a box or Gaussian kernel to blur an image (Smoothing an Image with Blur Filters), or a derivative kernel to find edges (Measuring Change with Image Derivatives). A CNN uses the same convolution, but it learns the filter weights from data. To keep the notation aligned with dense neural network layers, the filter is written : the weights that training adjusts.PyTorch reference
Key references: (Dumoulin & Visin, 2016; Mao et al., 2016; Kulkarni et al., 2015; Zeiler & Fergus, 2013; Simonyan & Zisserman, 2014)
References
- Dumoulin, V., Visin, F. (2016). A guide to convolution arithmetic for deep learning. arXiv [stat.ML].
- Kulkarni, T., Whitney, W., Kohli, P., Tenenbaum, J. (2015). Deep Convolutional Inverse Graphics Network. arXiv [cs.CV].
- Mao, X., Shen, C., Yang, Y. (2016). Image restoration using convolutional auto-encoders with symmetric skip connections. arXiv [cs.CV].
- Simonyan, K., Zisserman, A. (2014). Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv [cs.CV].
- Zeiler, M., Fergus, R. (2013). Visualizing and Understanding Convolutional Networks. arXiv [cs.CV].

