YOLO: You Only Look Once: Unified, Real-Time Object Detection
A single network pass turns an image into all of its object boxes and class labels.
YOLO divides the image into a 7×7 grid. Each cell predicts two boxes, a confidence for each box, and one set of 20 class probabilities, so the whole detection result is a single 7×7×30 tensor that a convolutional network computes in about 22 milliseconds.
Explaining the paperYou Only Look Once: Unified, Real-Time Object DetectionA 2015 paper from the University of Washington, the Allen Institute for AI and Facebook AI Research that made object detection fast enough for live video, and measured exactly how much localization accuracy that cost.
Joseph Redmon, Santosh Divvala, Ross Girshick and Ali Farhadi posted YOLO to arXiv in June 2015, and the fifth version appeared at CVPR 2016. The base model detects the 20 Pascal VOC object classes at 45 frames per second on a Titan X GPU with 63.4 mAP; a smaller Fast YOLO runs at 155 frames per second with 52.7 mAP. The most accurate detector of the time, Faster R-CNN with a VGG-16 backbone, scored 73.2 mAP at 7 frames per second.
The page covers, in order:
- what a detector must output, and how the pipelines before YOLO produced it;
- the 7×7 grid, what each cell predicts, and how a labeled box becomes a training target;
- the confidence score, intersection over union (IoU), and the class-specific score of equation (1);
- the network and the five-part squared-error loss of equation (3), with the weights and the square root it adds to plain squared error;
- how 98 raw boxes become a detection list, and how mAP is computed from that list;
- the results, the paper's error analysis, and the limitations, including one the code makes stricter than the text.
What an object detector has to output
A classifier maps an image to one label. A detector maps an image to a list of variable length: for every object, a bounding box (an axis-aligned rectangle, here written as a center , a width and a height ), a class, and a score. On Pascal VOC there are 20 classes (person, car, dog, bottle, chair and so on) and an image holds anywhere from one object to a dozen or more.
Before YOLO, the standard approach reused a classifier. Deformable parts models (DPM) slide a classifier over evenly spaced positions and scales and keep the windows that score high. R-CNN first runs Selective Search, a hand-designed grouping algorithm that proposes about 2000 candidate boxes per image, then warps each box's pixels to a fixed size, runs a convolutional network on it, scores the features with per-class SVMs, refines the box with a linear regressor, and finally removes duplicates with non-maximum suppression. The paper quotes R-CNN at more than 40 seconds per image. Fast R-CNN shares the convolutional work across proposals and speeds up the classification stage, but Selective Search still takes about 2 seconds per image, so the full system runs at 0.5 frames per second.
Each stage in those pipelines is trained separately against its own objective. In R-CNN the proposal method is not trained at all, the network is fine-tuned to classify warped regions, the SVMs are then fit to its features, and the box regressor is fit after that, so no stage is trained against the final detection score. YOLO replaces all of them with one network whose output is the detection list, trained with one loss.
How YOLO works: the grid and the output tensor
The input image is resized to 448×448 pixels and conceptually divided into an grid, with for Pascal VOC, so each cell covers 64×64 pixels. The rule that turns detection into regression is this: if the center of an object falls into a grid cell, that cell is responsible for detecting the object.
Each cell predicts:
- boxes, each as five numbers: (the box center as an offset inside the cell, each in ), (width and height as fractions of the whole image, so also in ), and a confidence;
- class probabilities, conditioned on the cell containing an object. There is one set per cell, shared by both boxes.
That is numbers per cell and numbers per image. The network's last layer outputs exactly 1470 values and the code reshapes them into the grid; nothing else is computed to produce the detections, apart from a threshold and non-maximum suppression at the end.
Take the dog in Figure 1. Its labeled box has center in normalized image coordinates (134, 296 in pixels), width 0.26 and height 0.42 of the image. Multiplying the center by 7 gives , so the center lies in column 2, row 4 (counting from 0), and the in-cell offsets are , . The target for that cell is those two offsets, and (the network regresses square roots of the size, for reasons covered two sections down), a confidence target, and a one-hot class vector with a 1 in the "dog" slot. The other 46 cells in this image have no target object, and the loss pushes their confidences toward 0.
In Figure 1 the shaded cells are the responsible ones. Click any cell, or move between cells with the arrow keys, to read its target; drag a center dot to move an object across cell boundaries, for instance the dog's center into the bicycle's cell.
When two centers share a cell, the paper's text does not say what happens. The authors' Darknet code does: it shuffles the image's boxes, writes the first one into the cell's target, and skips any later box whose cell is already filled. The skipped object contributes nothing to the loss for that image. Each cell therefore trains on at most one object, and an image can supervise at most 49. The second box predictor in a cell is not a slot for a second object; the section on the responsible predictor shows what it is for.
Confidence, IoU and the class-specific score
Intersection over union measures how well two boxes agree: the area they share divided by the area they cover together. Identical boxes score 1, disjoint boxes 0. A predicted box that is the right size but shifted by half its width scores 1/3 (the overlap is half of each box, and the union is one and a half boxes).
YOLO defines each box's confidence as . If no object is centered in the cell the target is 0; otherwise the target is the IoU between the predicted box and the ground truth. A single number therefore encodes both whether an object is there and how well the box fits it. The code computes that IoU from the network's current box at every training step and uses it as a fixed target, so the confidence learns to predict the network's own localization quality.
Figure 2 shows the IoU of two boxes. The bands under the field are the paper's error categories from Section 4.2, used later on this page: for a box with the right class, IoU above 0.5 is correct, between 0.1 and 0.5 a localization error, and below 0.1 against every object a background error.
IoU depends on errors relative to the box size. Shifting a box sideways by 10 pixels leaves a 300-pixel-wide box at IoU 0.94 and drops a 20-pixel-wide box to 0.33. The loss, two sections down, has to account for this.
At test time YOLO multiplies each box's confidence by its cell's conditional class probabilities:
The right side follows from the product rule, under the assumption that a class can be present only if an object is:
This gives one score per box per class: 98 boxes times 20 classes, 1960 scores. If the dog's responsible box has confidence 0.40 and the cell says "dog" with probability 0.6, its dog score is 0.24 and its cat score is 0.40 times whatever the cell gives "cat". Equation (1) defines what the outputs should mean; nothing forces the network to satisfy it, since both factors are unconstrained regression outputs.
The network
The detector is a convolutional network with 24 convolutional layers and 2 fully connected layers, modeled loosely on GoogLeNet but with 1×1 "reduction" convolutions (which cut the channel count cheaply) followed by 3×3 convolutions instead of inception modules. The layer list, read from the paper's Figure 3:
layer group output (H x W x channels)
input image 448 x 448 x 3
conv 7x7x64 stride 2, maxpool 2x2 stride 2 112 x 112 x 64
conv 3x3x192, maxpool 56 x 56 x 192
conv 1x1x128, 3x3x256, 1x1x256, 3x3x512, maxpool 28 x 28 x 512
(conv 1x1x256, 3x3x512) x4, 1x1x512, 3x3x1024, pool 14 x 14 x 1024
(conv 1x1x512, 3x3x1024) x2 14 x 14 x 1024
--- the first 20 conv layers above are pretrained on ImageNet at 224 x 224 ---
conv 3x3x1024, 3x3x1024 stride 2 7 x 7 x 1024
conv 3x3x1024, 3x3x1024 7 x 7 x 1024
fully connected (dropout 0.5 after it) 4096
fully connected, linear, reshaped 7 x 7 x 30Every layer except the last uses a leaky rectified linear activation; the last is linear, because the outputs are regression targets that a ReLU would clip:
The spatial size falls from 448 to 7 through one stride-2 convolution, four 2×2 max-pools and one more stride-2 convolution, a factor of 64 in each direction. A 7×7 feature map lines up with the 7×7 grid of cells, but the detection outputs do not come from a 1×1 convolution on that map. They come from two fully connected layers, and the first of them connects every one of the features to each of its 4096 units. Every one of the 1470 outputs therefore depends on every pixel of the image. The paper describes this as reasoning "globally about the full image", and it accounts for YOLO's low background error rate in the results below.
The first fully connected layer also holds most of the weights. Counting from Figure 3, the 24 convolutional layers hold 60.2 million weights and biases, the first fully connected layer 205.5 million (50,176 × 4,096), and the second 6.0 million, for 271.7 million in total, so 75.6% of the parameters sit in that one matrix. YOLOv2 removed both fully connected layers.
Fast YOLO uses 9 convolutional layers instead of 24, with fewer filters, and is otherwise trained and tested identically. The paper also trains a YOLO on a VGG-16 backbone, for comparison with detectors built on it.
The YOLO loss function, term by term
YOLO is trained with sum-squared error on its 1470 outputs. The paper chooses it because it is easy to optimize and states that it does not match the goal of maximizing average precision. Unweighted squared error over all outputs fails in two ways that the paper names. First, most cells contain no object: in an image with three objects, 46 of 49 cells have only "no object" targets, and their confidence errors outweigh the few cells that matter. Second, it weights an error of 0.05 in width the same for a box covering half the image and for a box covering a twentieth of it. The loss YOLO actually uses adds two weights, two indicator functions and a square root:
The symbols: runs over the 49 cells and over the 2 predictors in a cell (the paper drops the subscript on , but each predictor has its own). One of each pair of symbols is the target and the other the prediction; the paper does not say which carries the hat, and squared error is symmetric, so it does not matter. The indicators decide which terms exist:
- is 1 if an object's center is in cell ;
- is 1 if predictor of that cell is responsible for the object, meaning its current box has the higher IoU with the truth of the two;
- is 1 for every other predictor, including the non-responsible one in an object's cell (the paper does not define it; this is what the code does).
Read top to bottom, the five lines are: the center error of responsible predictors, weighted by ; their size error on square roots, also weighted by 5; their confidence error against the IoU target, weight 1; the confidence error of all other predictors against 0, weighted by ; and the class error, only in cells that contain an object. A cell without an object contributes nothing to the class or coordinate terms, which is why the class outputs are conditional probabilities: they are only ever trained on cells where an object is present.
A worked example: the dog's cell
Return to the dog: target , , , . Suppose predictor 0 outputs with confidence 0.40, and predictor 1 outputs with confidence 0.20. Decoded into image coordinates, predictor 0 is a 0.20 by 0.49 box with IoU 0.69 against the dog; predictor 1 is a 0.09 by 0.09 square with IoU 0.07. Predictor 0 is responsible. The class outputs are 0.6 for dog, 0.3 for cat and 0 elsewhere.
The five terms (the size term uses the unrounded square roots):
This cell contributes 0.460. The class term is the largest, and the target of the confidence term is not fixed: if predictor 0's box improves to IoU 0.85, its confidence target rises to 0.85 with it. The pseudocode below computes the same thing for a whole image, in the order Darknet does it.
# YOLO v1 loss for one image, S = 7, B = 2, C = 20, as Darknet computes it.
# pred[i]: 2 boxes (x, y, sqrt_w, sqrt_h, conf) + 20 class scores for cell i
# target[i]: None, or one object (x, y in the cell; w, h of the image; class)
loss = 0
for i in range(49):
for b in range(2): # every predictor starts as "no object"
loss += 0.5 * pred[i].conf[b] ** 2 # lambda_noobj = 0.5, target 0
t = target[i]
if t is None:
continue
loss += sum((onehot(t.cls) - pred[i].cls) ** 2) # 20 class terms
ious = [iou(decode(i, pred[i].box[b]), t.box) for b in range(2)]
r = argmax(ious) # the responsible predictor
p = pred[i].box[r]
loss -= 0.5 * pred[i].conf[r] ** 2 # undo its "no object" term
loss += (ious[r] - pred[i].conf[r]) ** 2 # IoU is a constant target here
loss += 5 * ((t.x - p.x) ** 2 + (t.y - p.y) ** 2)
loss += 5 * ((sqrt(t.w) - p.sqrt_w) ** 2 + (sqrt(t.h) - p.sqrt_h) ** 2)
# Gradients reach every weight through the 1470 outputs; nothing flows through
# ious[r] or argmax. If no predictor overlaps the truth, Darknet picks the one
# with the smallest coordinate RMSE instead of the highest IoU.Why the empty cells get a weight of 0.5
At the start of training, suppose every confidence output is near 0.5 and the image has three objects. Of the 98 predictors, 95 have target 0. The derivative of with respect to is , so with each of them pushes its confidence down with gradient 1.0, a total of 95. The three responsible predictors, with IoU targets around 0.3, each push with gradient , a total of 1.2. All 98 outputs come from the same shared weights, so early updates mostly push every confidence down. The paper says this imbalance "can lead to model instability, causing training to diverge early on."
Halving the weight of empty predictors () and multiplying the coordinate terms by 5 shifts the balance toward the cells with objects. The imbalance also shrinks on its own: the gradient fades as the empty-cell confidences approach 0, so once the network has learned that most cells are empty, those terms stop dominating. The learning-rate warm-up in the training recipe is the paper's other response to divergence early in training.
Why YOLO predicts square roots of width and height
Detection quality is measured by IoU, and IoU depends on errors relative to the box size. Take a square box of side and a prediction that is too wide by , with the height and center right. Then
For a fixed , the IoU penalty falls roughly like : a 0.03 error costs a box of width 0.1 an IoU of 0.77, and a box of width 0.8 an IoU of 0.96. Squared error on itself gives both the same . Squared error on the square roots behaves differently because the square root is steeper near 0. To first order,
which also falls like . So for a fixed absolute error, the square-root loss charges small boxes more in about the same proportion that IoU does. The paper calls this a partial fix, and Figure 3 shows why. If the error is a fixed fraction of the box (say , so the IoU is the same 0.87 at every size), the square-root loss still grows in proportion to , while plain squared error grows like . Large boxes are still penalized more than IoU would penalize them, just much less than before.
With the absolute-error toggle, the teal square-root curve runs nearly parallel to the white IoU curve while the amber plain-error curve stays flat. With the relative-error toggle, IoU is flat and both squared errors slope upward, the square-root one at half the slope, which is why the paper still lists the loss among its limitations.
One responsible predictor per object
Each cell has two box predictors but at most one object. Why train only one of them on it, instead of both? If both predictors were pulled toward every object they saw, they would receive identical gradients on average and end up predicting the same box, and the second predictor would be wasted. YOLO assigns the object to whichever predictor currently has the higher IoU with it, trains that predictor's coordinates and confidence toward the object, and trains the other one's confidence toward 0.
A small initial difference then grows. A predictor that happens to output slightly wider boxes wins the wide objects, gets trained on wide objects, and gets wider; the other wins the tall and small ones. The paper reports that the predictors specialize in "certain sizes, aspect ratios, or classes of object, improving overall recall."
Figure 4 reproduces that dynamic in a toy where each predictor outputs a single fixed shape. In YOLO the predicted shape depends on the image, so the specialization shows up as a tendency of the learned function. Press Play with the IoU rule selected, then switch to "every object trains both" and play again.
Under the IoU rule the predictors separate within about fifteen steps, one settling near 0.59 by 0.33 (the cars) and the other near 0.16 by 0.36 (the people and the small objects). Training both on everything leaves them at the same average shape, 0.28 by 0.35, which fits none of the three groups well. YOLOv2 later used the same assignment idea, run as k-means over the training boxes with distance , to pick its anchor box shapes before training.
Training recipe
Training has two stages. First, the first 20 convolutional layers, followed by an average-pooling layer and a fully connected layer, are trained as an ImageNet classifier on 224×224 inputs for about a week, reaching 88% single-crop top-5 accuracy on the ImageNet 2012 validation set. Then four convolutional layers and the two fully connected layers are added with random weights and the input resolution is doubled to 448×448, because "detection often requires fine-grained visual information." All of it is implemented in Darknet, Redmon's own C and CUDA framework.
The detector is trained for about 135 epochs on the VOC 2007 and 2012 train and validation sets (plus VOC 2007 test when evaluating on 2012), with batch size 64, momentum 0.9 and weight decay 0.0005. The learning rate rises from to over the first epochs, because starting high made the model diverge, then stays at for 75 epochs, for 30 and for 30. Dropout with rate 0.5 follows the first fully connected layer. Data augmentation scales and translates the image by up to 20% of its size and scales exposure and saturation in HSV space by up to a factor of 1.5.
The bounding box targets move with the augmentation. A translated image shifts every object center, which can move an object into a different cell, so the same object is trained on several different cells and offsets over the course of training.
From 98 boxes to detections: thresholds and non-maximum suppression
At test time the network runs once and produces 98 boxes, each with 20 class scores from equation (1). Decoding a predictor back into pixels inverts the target encoding:
# Decode one predictor (cell row, col; outputs x, y, sqrt_w, sqrt_h, conf)
# into a box in pixels, and score it for every class (equation 1).
cx = (col + x) / 7 * image_width
cy = (row + y) / 7 * image_height
w = sqrt_w ** 2 * image_width
h = sqrt_h ** 2 * image_height
scores = conf * class_scores[row, col] # 20 numbers, one per class
keep = scores > threshold # Darknet: 0.2 demo, 0.001 for mAP
# then per class: sort by score, delete any box with IoU > nms with a kept boxTwo steps turn those 1960 scores into a short list. A score threshold drops boxes that are unlikely to contain an object of that class. Non-maximum suppression (NMS) removes duplicates: within each class, sort the surviving boxes by score, keep the top one, delete every remaining box whose IoU with a kept box exceeds a set limit, and repeat down the list.
The grid already gives most objects a single confident box, since only one cell is responsible for each. Duplicates come from large objects and from objects near cell borders, where neighboring cells have also learned to predict a box for the object. In Figure 5, each person, the dog and the car has one tight box from its responsible cell and several looser ones from the cells it covers.
At the default settings, 15 boxes pass the 0.2 threshold and NMS at 0.5 keeps 4, one per object. The two people's best boxes overlap with an IoU of about 0.2. Lower the NMS limit below 0.2 and NMS deletes the lower-scoring person as well. A strict limit removes more duplicates and also more real neighbors of the same class.
Duplicates matter because of how Pascal VOC scores a detector. For one class, all detections across the test set are sorted by score. Walking down the list, a detection is a true positive if it overlaps a not-yet-matched ground-truth object of that class with IoU above 0.5, and a false positive otherwise, which includes a second detection of an object already matched. Precision (the fraction of detections so far that are true positives) and recall (the fraction of all objects found so far) trace a curve. VOC 2007 average precision (AP) is the mean, over the recall levels 0, 0.1, …, 1, of the highest precision reached at that recall or beyond; mAP is the mean of AP over the 20 classes. A high-scoring duplicate is a false positive near the top of the list, where it lowers precision at every recall level after it. For YOLO, the paper reports that NMS adds 2 to 3 mAP, and describes it as "not critical to performance as it is for R-CNN or DPM", whose overlapping proposals and sliding windows produce many more duplicates.
One network evaluation per image makes YOLO's inference cost constant: 98 boxes, whether the image holds one object or ten. R-CNN classifies about 2000 proposals per image.
Speed and accuracy
Table 1 of the paper compares detectors on VOC 2007 test, timed on a Titan X with no batching. Only Sadeghi et al.'s GPU implementation of DPM had reached real time (30 frames per second or more) before, at 26.1 mAP when run at 30 Hz and 16.0 mAP at 100 Hz. Fast YOLO at 155 frames per second scores 52.7 mAP, twice the 30 Hz DPM, and YOLO at 45 frames per second scores 63.4. At 45 frames per second a frame takes 22 milliseconds, which is the "less than 25 milliseconds of latency" the paper quotes for streaming video.
Below real time, Faster R-CNN with VGG-16 is 9.8 mAP more accurate than YOLO and 6.4 times slower (7 against 45 frames per second). Faster R-CNN with the smaller Zeiler-Fergus network runs at 18 frames per second, 2.5 times slower than YOLO, with 1.3 mAP less. YOLO trained on VGG-16 reaches 66.4 mAP at 21 frames per second, which the paper mentions for comparison and then sets aside because it is not real time.
On the VOC 2012 test set YOLO scores 57.9 mAP, near the original R-CNN with VGG-16 (59.2 without box regression, 62.4 with it) and below the state of the art of November 2015 (73.9). Its per-class numbers show where it loses: 22.7 AP on bottles, 28.9 on potted plants and 52.2 on sheep, classes that tend to be small or appear in groups, against 81.4 on cats and 73.9 on trains.
What kind of mistakes YOLO makes
Section 4.2 breaks down errors with the method of Hoiem et al. For each class, take the top detections, where is the number of objects of that class in the test set, and sort each into the categories of Figure 2 (plus "similar class" and "other class" for a box on an object of the wrong class). Figure 7 shows the shares averaged over the 20 classes for YOLO and Fast R-CNN.
Localization errors are 19.0% of YOLO's top detections, more than its similar-class, other-class and background errors combined (15.5%), against 8.6% for Fast R-CNN. Background errors, boxes on no object, are 13.6% for Fast R-CNN and 4.75% for YOLO, a factor of 2.9. The paper attributes the first gap to the coarse grid and the direct regression of coordinates, and the second to each output depending on the whole image.
Section 4.3 exploits the difference. For every box Fast R-CNN predicts, check whether YOLO predicts a similar box, and if so raise the Fast R-CNN score by an amount based on YOLO's probability and the overlap of the two boxes (the paper gives no formula). The best Fast R-CNN model goes from 71.8 to 75.0 mAP on VOC 2007. Combining it with three other Fast R-CNN variants instead gains 0.3 to 0.6 mAP, and the paper attributes YOLO's larger gain to it making different kinds of mistakes. On VOC 2012 the combination scores 70.7, 2.3 above Fast R-CNN alone, and was fourth on the public leaderboard. It runs both models, so it is no faster than Fast R-CNN.
Section 4.5 tests generalization on paintings. Trained on natural photos, YOLO detects people in the Picasso dataset with 53.3 AP, against 59.2 on VOC 2007 people, a drop of 5.9. R-CNN falls from 54.2 to 10.4 and DPM from 43.2 to 37.8. On the People-Art dataset YOLO scores 45, DPM 32 and R-CNN 26. The paper's explanation is that R-CNN's Selective Search proposals are tuned to natural images and its classifier sees only small regions, while YOLO, like DPM, models the size, shape and layout of whole objects, which carry over from photographs to paintings better than pixel statistics do.
Limitations, and what YOLOv2 changed
Section 2.4 of the paper lists these limitations:
- Nearby small objects. Each cell predicts two boxes and has one set of class probabilities, and the training code gives each cell at most one object. Two birds whose centers fall in the same 64×64-pixel cell cannot both be learned. The paper names flocks of birds as the failure case.
- Unusual shapes. Box shapes are learned from data, so objects in aspect ratios or configurations rare in training are boxed poorly.
- Coarse features. The boxes are regressed from a 7×7 map after six downsampling steps, so the network has little fine spatial detail to place box edges with.
- The loss. Even with square roots, squared error penalizes a given relative error in a large box more than in a small one, while IoU treats them alike. The paper names incorrect localization as its main source of error.
YOLOv2, published by Redmon and Farhadi a year and a half later as part of the YOLO9000 paper, addressed several of these directly. It removed the fully connected layers and predicted boxes as offsets from anchor boxes, a fixed set of prior shapes per cell, chosen by k-means over the training boxes with distance . Each anchor has its own class prediction, so a cell is no longer limited to one class, and the model outputs more than a thousand boxes per image instead of 98. Adding batch normalization to every convolutional layer gave more than 2 mAP and let it drop dropout. In that paper's ablation, anchors lowered mAP slightly (69.5 to 69.2) and raised recall from 81% to 88%.
Implementation notes from Darknet
The paper's code is the Darknet repository, and several details there matter for anyone reimplementing YOLO v1 or reading its loss curves.
- The 1470 outputs are not stored cell by cell. The detection layer reads them as 980 class scores (49 cells × 20), then 98 confidences, then 392 coordinates (98 boxes × 4). The 7×7×30 tensor in the paper is a logical view.
- The class outputs pass through no softmax in the released configuration (
softmax=0), and the last layer is linear, so they are trained by squared error toward a one-hot vector and are not constrained to be non-negative or to sum to 1. - The confidence target for the responsible predictor is the IoU of its current box (
rescore=1), computed each step and treated as a constant. With rescoring off it would be 1. - When none of a cell's predictors overlaps the truth, which is common early in training, the responsible predictor is the one with the smallest root-mean-square error between its box and the truth.
- For mAP evaluation the score threshold is 0.001, so nearly every box enters the ranked list, and NMS runs per class at IoU 0.5. The webcam demo uses a threshold of 0.2 and NMS at 0.4.
- The released
yolov1.cfgis a later variant of the paper's model: three boxes per cell (1715 outputs) and a locally connected layer before a single connected layer. The paper's results are for and the 4096-unit layer described above.
The intersection-over-union quantity used throughout, and a model trained to predict its own IoU, also appear in the Segment Anything explainer, where the mask decoder scores each mask by an estimated IoU.
Questions you might still have
Why is it called "you only look once"?
Earlier detectors evaluated a classifier many times per image: DPM slides one over a grid of positions and scales, and R-CNN runs a convolutional network on each of about 2000 proposed regions. YOLO runs one network once on the full 448 by 448 image and reads every box and class score off its 7 by 7 by 30 output.
How many objects can YOLO v1 detect in one image?
At most 49, one per grid cell, in practice fewer. The training code assigns each cell at most one ground-truth object and skips any other object whose center falls in the same 64 by 64 pixel cell. Each cell also has a single set of class probabilities, so its two boxes share one class. Crowds of small objects, such as a flock of birds, lose most of their members.
Are the class outputs real probabilities?
Not strictly. The final layer is linear and the released configuration has softmax switched off, so the 20 class outputs are trained by squared error toward a one-hot vector and are not forced to be non-negative or to sum to 1. They behave like probabilities only to the extent training makes them. The confidence is likewise a regression toward the IoU, not a calibrated probability.
Why does YOLO make fewer background errors than Fast R-CNN?
Each YOLO output depends on the whole image, because the first fully connected layer connects all 50,176 features of the last conv layer to every unit. Fast R-CNN classifies each proposed region from features pooled inside that region, so a patch of texture that looks like an object in isolation can get a high score. On VOC 2007, 13.6% of Fast R-CNN's top detections were background, against 4.75% for YOLO.
Why does YOLO localize worse?
The boxes come from a 7 by 7 grid of cells, each 64 pixels wide at the 448 input, after six stride-2 downsampling steps, and a fully connected layer regresses the coordinates directly with no proposal to refine. The squared-error loss also treats a given width error in a large and a small box more alike than IoU does, even after the square-root change. Localization errors were 19.0% of YOLO's top detections, more than all its other error types combined.
Is YOLO v1 still used?
Rarely. YOLOv2 (December 2016) removed the fully connected layers and added anchor boxes and batch normalization, and the later detectors that carry the YOLO name build on that design rather than on the 7 by 7 by 30 output of v1. The v1 paper is still the reference for the idea of reading every box and class off one forward pass.
Does the 45 frames per second include the whole pipeline?
The paper times the base network at 45 frames per second on a Titan X with no batch processing. Section 5 connects YOLO to a webcam and reports that it stays real time including the time to fetch frames from the camera and draw the detections, without giving a separate number.
What happens to images that are not square?
They are resized to 448 by 448, which stretches them. Because x and w are fractions of the image width and y and h fractions of its height, the decoded boxes are mapped back by multiplying with the original width and height, so the stretch is undone in the output coordinates.
What does the NMS step do, and why does YOLO need it less than R-CNN?
Non-maximum suppression sorts the boxes of a class by score and deletes any box that overlaps a higher-scoring one by more than a set IoU. YOLO's grid already assigns each object to one cell, so most objects get one confident box, but large objects and objects near cell borders get duplicates from neighboring cells. The paper reports that NMS adds 2 to 3 mAP for YOLO and calls it critical for R-CNN and DPM, whose overlapping proposals and windows produce many more duplicates.
Footnotes & further reading
- The paper: Redmon, Divvala, Girshick & Farhadi, You Only Look Once: Unified, Real-Time Object Detection (CVPR 2016; arXiv v5, May 2016). Project page and code: pjreddie.com and the Darknet repository, files
src/detection_layer.c,src/data.c,examples/yolo.candcfg/yolov1.cfg. - The pipelines it replaced: Felzenszwalb et al., Object Detection with Discriminatively Trained Part Based Models (PAMI 2010); Girshick et al., Rich feature hierarchies for accurate object detection and semantic segmentation (R-CNN, CVPR 2014); Girshick, Fast R-CNN (ICCV 2015); Ren et al., Faster R-CNN (NeurIPS 2015).
- The evaluation protocol: Everingham et al., The Pascal Visual Object Classes (VOC) Challenge (IJCV 2010), Section 4.2 for the 11-point AP. The error categories: Hoiem, Chodpathumwan & Dai, Diagnosing Error in Object Detectors (ECCV 2012).
- The successor: Redmon & Farhadi, YOLO9000: Better, Faster, Stronger (CVPR 2017), Section 2 for anchor boxes, dimension clusters and batch normalization.
- The artwork datasets: Ginosar et al., Detecting People in Cubist Art (ECCV Workshops 2014), and Cai et al., The cross-depiction problem (2015). Dropout, used after the first fully connected layer: Hinton et al., Improving neural networks by preventing co-adaptation of feature detectors (2012).
How could this explainer be improved? Found an error, or something unclear? I read every message.