Backpropagation through a single neuron
Consider a sigmoid neuron with input vector , weight vector , and bias : For a scalar loss , the chain rule gives Define the local error signal The parameter gradients are therefore The weight gradient is proportional to the input vector. Its magnitude depends on both the input scale and the local error signal. For a sigmoid, approaches zero as grows, so saturation can suppress the gradient even when the input is large. The bias gradient contains no direct factor of .
Backpropagation through a dense layer
A dense layer followed by a ReLU computes Let be the upstream gradient and let be the ReLU gate. The error signal at the pre-activation is Applying the chain rule to each element gives and the complete parameter gradients are The weight gradient is an outer product. Every row is the layer input scaled by one component of the error signal. For a mini-batch, the update is the sum or mean of these outer products over the examples. Consequently, changes in the distribution and scale of the layer inputs change the distribution and scale of its weight gradients. The bias gradient again has no direct multiplicative dependence on the input.
The internal covariate shift hypothesis
For layer , An update to an earlier layer changes . The next optimization step therefore presents layer with a different input distribution and a different distribution of weight gradients. Ioffe and Szegedy called this moving target internal covariate shift. This term is specific to hidden activations during training and should not be confused with covariate shift between training and test data. The original argument proceeds as follows:- Earlier layers continually change the coordinates received by later layers.
- Later layers must adapt to those changing coordinates.
- Standardizing intermediate features should make their scale more stable, improve gradient flow, and allow larger learning rates.
Linear -> BatchNorm -> ReLU arrangement, batch normalization standardizes the pre-activation before the nonlinearity. This directly controls the coordinates delivered to the ReLU and indirectly controls the input delivered to the next layer. It does not directly normalize in the gradient formula for the current weight matrix, an important limitation of the simple argument.
If a wide layer has approximately Gaussian pre-activations, the learned affine transformation gives
After a ReLU, the probability of an output being clipped to zero is approximately
This calculation is useful, but it is not a definition of batch normalization. The layer fixes only the first two moments, not the shape of the distribution, and and are learned.

Evidence against the internal covariate shift explanation
The optimization benefit of batch normalization is well established. The claim that the benefit is caused by reduced internal covariate shift is not. Santurkar et al. (2018) tested the causal claim directly. Their experiments found that the distributional stability of layer inputs had little connection to successful training. They instead observed that batch normalization made the loss and gradients vary more smoothly with the parameters. A smoother objective makes a gradient step more predictive and supports larger learning rates. This result rejects a strong version of the original story. A moving hidden distribution may occur during training, but reducing that movement is neither a sufficient explanation nor the measured source of batch normalization’s benefit. Santurkar et al. instead support an optimization account based on a smoother loss landscape and more stable, predictive gradients.References
- Ioffe, S., and Szegedy, C. (2015). Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. ICML.
- Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A. (2018). How Does Batch Normalization Help Optimization?. NeurIPS.
- LeCun, Y., Bottou, L., Orr, G. B., and Muller, K.-R. (1998). Efficient BackProp.
- Data preprocessing and whitening.

