
Class-specific confidence at test time
During inference, we combine the conditional class probabilities for each cell with the confidence for each box. This gives a score for every box and class: These class-specific confidence scores are used before NMS.Network architecture and activations

Training targets and responsibility
Each cell predicts boxes. For a given object, YOLO assigns responsibility to the predictor whose current box has the highest IoU with the ground-truth box. This specialization improves recall. The assignment determines the training targets:- Only the responsible predictor receives coordinate and objectness regression targets for the object.
- The other predictors in the cell are trained toward “no object” confidence. This reduces spurious positives.
The multi-part loss
YOLOv1 uses a sum-squared error for location, size, objectness (confidence), and classification. The coefficients and balance these terms. To reduce sensitivity to scale, the loss regresses instead of . Here if and only if predictor in cell is responsible for an object. The value for “no object” cases. The coefficients are and . The classification loss applies only when a cell contains an object.Optimization details
A typical VOC training schedule uses about 135 epochs, a batch size of 64, momentum of 0.9, and weight decay of . The learning rate warms up from to . It then remains at for 75 epochs, for 30 epochs, and for 30 epochs. Regularization includes dropout at a rate of 0.5 after the first fully connected layer. Data augmentation includes random scaling and translation of up to 20%, along with exposure and saturation changes in HSV of up to 1.5×.End-to-end inference
- Preprocess the image. Resize it, for example to , and pass it through the CNN once.
-
Decode the raw outputs.
- For each cell and predictor , convert the normalized values to image coordinates. Use the predicted confidence .
- Combine the confidence with the class probabilities using Eq. (1). This gives the class-specific scores .
-
Filter and suppress the predictions.
- Discard low-score boxes.
- Perform non-max suppression for each class. NMS is less important here than in proposal-based pipelines. It removes duplicate predictions from neighboring cells and adds about 2 to 3 mAP points.
Strengths and limitations
- YOLO processes the image in one pass and uses information from the entire image. This makes inference very fast.
- Compared with the R-CNN family, YOLO produces fewer background false positives but more localization errors.
- The fixed grid limits performance on crowded scenes with small objects. Downsampling produces coarse features, and small-box localization is sensitive to errors.
- A grid cell is responsible for an object if the object’s center falls inside the cell.
- One predictor learns the geometry of each object assigned to a cell. Responsibility is based on IoU.
- Confidence objectness IoU. Class probabilities apply to the cell. Eq. (1) combines them into a score for each class.
- The loss balances localization, objectness, and classification using . It uses the square roots of box dimensions to reduce scale sensitivity. Eq. (3).
PyTorch sections
The following sections build a complete YOLOv11 anchor-free detector from scratch in PyTorch.Data Pipeline
Load COCO data and apply letterbox resizing, mosaic augmentation, and multi-scale target encoding for anchor-free detection.
Backbone
Build Conv-BN-SiLU blocks, Bottleneck, C3k2 (CSP), SPPF, and the backbone that produces P3/P4/P5 features.
Neck and Head
Implement FPN top-down and PAN bottom-up feature aggregation, C2PSA attention, and a decoupled anchor-free head with DFL.
Loss and Training
Implement IoU variants (GIoU/DIoU/CIoU), Task-Aligned Learning, the BCE + CIoU + DFL loss, and the training loop.
Inference and Evaluation
Decode predictions, implement NMS from scratch, evaluate COCO mAP, and produce Grad-CAM visualizations.
PyTorch reference
Key references: (Redmon et al., 2015; Redmon & Farhadi, 2016; Liu et al., 2015; Canziani et al., 2016; Godard et al., 2016)
References
- Canziani, A., Paszke, A., Culurciello, E. (2016). An Analysis of Deep Neural Network Models for Practical Applications. arXiv [cs.CV].
- Godard, C., Aodha, O., Brostow, G. (2016). Unsupervised Monocular Depth Estimation with Left-Right Consistency. arXiv [cs.CV].
- Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., et al. (2015). SSD: Single Shot MultiBox Detector. arXiv [cs.CV].
- Redmon, J., Divvala, S., Girshick, R., Farhadi, A. (2015). You only look once: Unified, real-time object detection. arXiv [cs.CV].
- Redmon, J., Farhadi, A. (2016). YOLO9000: Better, Faster, Stronger. arXiv [cs.CV].

