
- and are learnable parameters (of shape ),
- is a small constant for numerical stability,
- denotes element-wise multiplication.
PyTorch reference
Key references: (Bjorck et al., 2018; Santurkar et al., 2018; Zhang & Sennrich, 2019; Ioffe & Szegedy, 2015; Keskar et al., 2016)
References
- Bjorck, J., Gomes, C., Selman, B., Weinberger, K. (2018). Understanding batch normalization. arXiv [cs.LG].
- Ioffe, S., Szegedy, C. (2015). Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. arXiv [cs.LG].
- Keskar, N., Mudigere, D., Nocedal, J., Smelyanskiy, M., Tang, P. (2016). On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. arXiv [cs.LG].
- Santurkar, S., Tsipras, D., Ilyas, A., Madry, A. (2018). How does batch Normalization help optimization?. arXiv [stat.ML].
- Zhang, B., Sennrich, R. (2019). Root mean square layer normalization. arXiv [cs.LG].

