Skip to main content
Open In Colab

ROI Head: ROI Align + Classification and Box Regression

Section 4 of 6 in the Faster RCNN from-scratch series Given proposals from the RPN, we extract fixed-size features via ROI Align, then classify each proposal and refine its bounding box. Mask RCNN extension point: this section also demonstrates the 14×14 ROI Align variant used by the mask head (section 07).
Output from cell 5
Key references: (Redmon et al., 2015; -Scratch-Vision-Trans, n.d.; Zagoruyko & Komodakis, 2016; Wightman et al., 2021; Ren et al., 2015)

References

  • Redmon, J., Divvala, S., Girshick, R., Farhadi, A. (2015). You only look once: Unified, real-time object detection.
  • Ren, S., He, K., Girshick, R., Sun, J. (2015). Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.
  • (n.d.). PyTorch-Scratch-Vision-Transformer-ViT: Simple and easy to understand PyTorch implementation of Vision Transformer (ViT) from scratch, with detailed steps. Tested on common datasets like MNIST, CIFAR10, and more.
  • Wightman, R., Touvron, H., Jégou, H. (2021). ResNet strikes back: An improved training procedure in timm.
  • Zagoruyko, S., Komodakis, N. (2016). Wide Residual Networks.