Grayscale images

Color images
A color image assigns a vector to each pixel instead of a single value: The third dimension holds the image channels. When the channels are the red, green, and blue primaries, the image is in the RGB color space, and combining the three channels gives the color of each pixel. In code, a color image is therefore a three-dimensional array, and the order of its dimensions is a convention, not something the array enforces. OpenCV, and NumPy code built around it, uses height by width by channels (HWC) and orders the channels blue, green, red. Most PyTorch vision APIs, including the torchvision models, expect channels by height by width (CHW) in red, green, blue order. Converting between the two is up to you:torch.from_numpy keeps the HWC order of an OpenCV image, so you still have to permute the dimensions, for example with permute(2, 0, 1), and swap blue and red.
Normalization
Before processing an image, you convert its pixel values to floating-point numbers in [0, 1] by dividing each value by the largest value its channel can hold. This is normalization. For an 8-bit RGB image:

