Skip to main content
Visual explanation of YOLOv1 operation YOLOv1 divides the input image into an S×SS\times S grid. If the center of an object falls in cell ii, that cell is responsible for predicting the object. Each grid cell outputs BB bounding-box hypotheses and one set of probabilities over CC classes. The resulting tensor has dimensions S×S×(B⋅5+C)S\times S\times (B\cdot 5 + C). For VOC, S=7S=7, B=2B=2, C=20⇒7×7×30C=20 \Rightarrow 7\times7\times 30. Each bounding-box prediction contains five numbers: (x,y,w,h,confidence),(x, y, w, h, \text{confidence}), Here, (x,y)(x,y) are the box-center offsets relative to the responsible cell. The values (w,h)(w,h) are normalized by the image width and height. The confidence target equals the IoU between the predicted box and the closest ground-truth box. It is 00 when no object is present.

Class-specific confidence at test time

During inference, we combine the conditional class probabilities for each cell with the confidence for each box. This gives a score for every box and class: Pr⁡(Classi∣Object)⋅Pr⁡(Object)⋅IoUpredtruth  =  Pr⁡(Classi)⋅IoUpredtruth.(1)\Pr(\text{Class}_i \mid \text{Object}) \cdot \Pr(\text{Object})\cdot \text{IoU}_{\text{pred}}^{\text{truth}} \;=\; \Pr(\text{Class}_i)\cdot \text{IoU}_{\text{pred}}^{\text{truth}}. \tag{1} These class-specific confidence scores are used before NMS.

Network architecture and activations

YOLOv1 Architecture The detector is a single CNN with 24 convolutional layers and 2 fully connected layers. The early layers extract features. The fully connected layers map those features to the S×S×(B⋅5+C)S\times S\times (B\cdot5+C) output. A fast variant uses fewer convolutional layers. The final layer uses a linear activation. All other layers use leaky ReLU: ϕ(x)={x,x>0,0.1x,otherwise.(2)\phi(x)= \begin{cases} x, & x>0,\\ 0.1x, & \text{otherwise.} \end{cases} \tag{2} Coordinates are normalized as described above.

Training targets and responsibility

Each cell predicts BB boxes. For a given object, YOLO assigns responsibility to the predictor whose current box has the highest IoU with the ground-truth box. This specialization improves recall. The assignment determines the training targets:
  • Only the responsible predictor receives coordinate and objectness regression targets for the object.
  • The other predictors in the cell are trained toward “no object” confidence. This reduces spurious positives.
This assignment is made at each iteration using the model’s current boxes.

The multi-part loss

YOLOv1 uses a sum-squared error for location, size, objectness (confidence), and classification. The coefficients λcoord\lambda_{\text{coord}} and λnoobj\lambda_{\text{noobj}} balance these terms. To reduce sensitivity to scale, the loss regresses w,h\sqrt{w},\sqrt{h} instead of w,hw,h. L=  λcoord∑i=1S2∑j=1B1ijobj[(xi−x^i)2+(yi−y^i)2]+λcoord∑i=1S2∑j=1B1ijobj[(wi−w^i)2+(hi−h^i)2]+∑i=1S2∑j=1B1ijobj(Ci−C^i)2+λnoobj∑i=1S2∑j=1B1ijnoobj(Ci−C^i)2+∑i=1S21iobj∑c∈C(pi(c)−p^i(c))2.\begin{aligned} \mathcal{L} = \;& \lambda_{\text{coord}} \sum_{i=1}^{S^2} \sum_{j=1}^{B} \mathbf{1}^{\text{obj}}_{ij} \Big[(x_i-\hat{x}_i)^2 + (y_i-\hat{y}_i)^2\Big] \\ &+ \lambda_{\text{coord}} \sum_{i=1}^{S^2} \sum_{j=1}^{B} \mathbf{1}^{\text{obj}}_{ij} \Big[\big(\sqrt{w_i}-\sqrt{\hat{w}_i}\big)^2 + \big(\sqrt{h_i}-\sqrt{\hat{h}_i}\big)^2\Big] \\ &+ \sum_{i=1}^{S^2}\sum_{j=1}^{B} \mathbf{1}^{\text{obj}}_{ij}\big(C_i-\hat{C}_i\big)^2 + \lambda_{\text{noobj}} \sum_{i=1}^{S^2}\sum_{j=1}^{B} \mathbf{1}^{\text{noobj}}_{ij}\big(C_i-\hat{C}_i\big)^2 \\ &+ \sum_{i=1}^{S^2}\mathbf{1}^{\text{obj}}_{i} \sum_{c\in\mathcal{C}}\big(p_i(c)-\hat{p}_i(c)\big)^2. \end{aligned} Here 1ijobj=1\mathbf{1}^{\text{obj}}_{ij}=1 if and only if predictor jj in cell ii is responsible for an object. The value 1ijnoobj=1\mathbf{1}^{\text{noobj}}_{ij}=1 for “no object” cases. The coefficients are λcoord=5\lambda_{\text{coord}}=5 and λnoobj=0.5\lambda_{\text{noobj}}=0.5. The classification loss applies only when a cell contains an object.

Optimization details

A typical VOC training schedule uses about 135 epochs, a batch size of 64, momentum of 0.9, and weight decay of 5 ⁣× ⁣10−45\!\times\!10^{-4}. The learning rate warms up from 10−310^{-3} to 10−210^{-2}. It then remains at 10−210^{-2} for 75 epochs, 10−310^{-3} for 30 epochs, and 10−410^{-4} for 30 epochs. Regularization includes dropout at a rate of 0.5 after the first fully connected layer. Data augmentation includes random scaling and translation of up to 20%, along with exposure and saturation changes in HSV of up to 1.5×.

End-to-end inference

  1. Preprocess the image. Resize it, for example to 448×448448\times 448, and pass it through the CNN once.
  2. Decode the raw outputs.
    • For each cell ii and predictor jj, convert the normalized (x,y,w,h)(x,y,w,h) values to image coordinates. Use the predicted confidence CijC_{ij}.
    • Combine the confidence with the class probabilities pi(c)p_i(c) using Eq. (1). This gives the class-specific scores sijc=pi(c)⋅Cijs_{ijc} = p_i(c)\cdot C_{ij}.
  3. Filter and suppress the predictions.
    • Discard low-score boxes.
    • Perform non-max suppression for each class. NMS is less important here than in proposal-based pipelines. It removes duplicate predictions from neighboring cells and adds about 2 to 3 mAP points.

Strengths and limitations

  • YOLO processes the image in one pass and uses information from the entire image. This makes inference very fast.
  • Compared with the R-CNN family, YOLO produces fewer background false positives but more localization errors.
  • The fixed grid limits performance on crowded scenes with small objects. Downsampling produces coarse features, and small-box localization is sensitive to errors.
  • A grid cell is responsible for an object if the object’s center falls inside the cell.
  • One predictor learns the geometry of each object assigned to a cell. Responsibility is based on IoU.
  • Confidence == objectness ×\times IoU. Class probabilities apply to the cell. Eq. (1) combines them into a score for each class.
  • The loss balances localization, objectness, and classification using λcoord,λnoobj\lambda_{\text{coord}},\lambda_{\text{noobj}}. It uses the square roots of box dimensions to reduce scale sensitivity. Eq. (3).

PyTorch sections

The following sections build a complete YOLOv11 anchor-free detector from scratch in PyTorch.

Data Pipeline

Load COCO data and apply letterbox resizing, mosaic augmentation, and multi-scale target encoding for anchor-free detection.

Backbone

Build Conv-BN-SiLU blocks, Bottleneck, C3k2 (CSP), SPPF, and the backbone that produces P3/P4/P5 features.

Neck and Head

Implement FPN top-down and PAN bottom-up feature aggregation, C2PSA attention, and a decoupled anchor-free head with DFL.

Loss and Training

Implement IoU variants (GIoU/DIoU/CIoU), Task-Aligned Learning, the BCE + CIoU + DFL loss, and the training loop.

Inference and Evaluation

Decode predictions, implement NMS from scratch, evaluate COCO mAP, and produce Grad-CAM visualizations.

PyTorch reference

Key references: (Redmon et al., 2015; Redmon & Farhadi, 2016; Liu et al., 2015; Canziani et al., 2016; Godard et al., 2016)

References

  • Canziani, A., Paszke, A., Culurciello, E. (2016). An Analysis of Deep Neural Network Models for Practical Applications. arXiv [cs.CV].
  • Godard, C., Aodha, O., Brostow, G. (2016). Unsupervised Monocular Depth Estimation with Left-Right Consistency. arXiv [cs.CV].
  • Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., et al. (2015). SSD: Single Shot MultiBox Detector. arXiv [cs.CV].
  • Redmon, J., Divvala, S., Girshick, R., Farhadi, A. (2015). You only look once: Unified, real-time object detection. arXiv [cs.CV].
  • Redmon, J., Farhadi, A. (2016). YOLO9000: Better, Faster, Stronger. arXiv [cs.CV].