DETR: End-to-End Object Detection with Transformers
A transformer predicts a fixed set of 100 boxes, trained by one-to-one matching to the real objects, so detection needs no anchors and no duplicate removal.
DETR runs a ResNet and a standard transformer encoder-decoder over an image and outputs 100 (class, box) pairs in one pass. During training the Hungarian algorithm pairs each labeled object with exactly one of those outputs, and every unpaired output is trained to predict "no object".
Explaining the paperEnd-to-End Object Detection with TransformersA 2020 paper from Facebook AI Research that removed the hand-designed parts of an object detector and still matched a tuned Faster R-CNN on COCO, at 42.0 AP with 41 million parameters.
Nicolas Carion and Francisco Massa (equal contribution), with Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov and Sergey Zagoruyko, posted DETR (DEtection TRansformer) to arXiv in May 2020 and presented it at ECCV 2020. On COCO it scores 42.0 AP, the same as a Faster R-CNN with a feature pyramid network trained with the same augmentations and a longer schedule, at half the GFLOPS: 86 against 180. It is 7.7 points better on large objects and 6.1 points worse on small ones, and it needs a 500-epoch training schedule to get there.
The page covers, in order (equation numbers follow the paper):
- what earlier detectors do to produce a list of boxes, and why that list contains duplicates;
- detection as set prediction: a fixed number of output slots and a "no object" class;
- the bipartite matching of equation (1), its cost, and the Hungarian algorithm that solves it;
- the loss of equation (2) and the box loss of equations (9) and (10), L1 plus generalized IoU;
- the network: backbone, encoder, decoder, object queries, positional encodings, auxiliary losses;
- the training recipe, the results, the ablations, panoptic segmentation, and the limitations;
- the official code for the matcher, the loss and the post-processing.
Anchors, proposals and non-maximum suppression
An object detector maps an image to a list of variable length. Each entry is a class (one of COCO's 80 object categories, such as person, car, giraffe), an axis-aligned bounding box, and a confidence score. A COCO training image holds 7 labeled objects on average and at most 63.
Neural networks output tensors of fixed shape, so detectors before DETR turned the list into a dense grid of guesses. Faster R-CNN, the baseline in this paper, places a set of anchor boxes (several sizes and aspect ratios) at every position of a convolutional feature map. A first stage scores each anchor for "is this an object" and regresses an offset from the anchor to a tighter proposal; a second stage classifies the best proposals and refines their boxes again. One-stage detectors such as YOLO skip the proposal stage and predict directly from grid cells or anchors.
Training needs a rule that says which guess should predict which labeled object. The anchor-based rule is many-to-one: every anchor that overlaps a labeled box by more than a threshold, measured in intersection over union (IoU, the shared area divided by the combined area), is trained to predict that box. A person who overlaps twelve anchors gets twelve positive training examples, which helps learning but means the trained network also outputs about twelve boxes around that person at test time. A post-processing step, non-maximum suppression (NMS), sorts the boxes by score and deletes every box that overlaps a higher-scoring box of the same class by more than a fixed IoU, such as 0.5.
Both pieces are hand-tuned. The anchor sizes and ratios, the IoU thresholds for positives and negatives, and the NMS threshold all have to be set per dataset, and the final accuracy depends on them. NMS is not part of the training loss, so the network is never trained against the output it is evaluated on, and when two real objects overlap heavily (two people hugging) NMS can delete one of them. DETR replaces both with a loss that assigns each labeled object to exactly one output.
Detection as set prediction
DETR outputs a fixed number of predictions per image, always, in one pass. Each prediction is a probability distribution over the classes plus one extra class written ("no object"), and a box given as center , center , width and height, each as a fraction of the image size. is chosen to be much larger than the number of objects in a typical image; on an image with 7 objects, 93 of the 100 outputs should say .
The ground truth is a set: the order of the 7 labeled objects in the annotation file means nothing. The paper pads it to size with entries so both sides have 100 elements. Each ground-truth element is a pair , a class (possibly ) and a box in the same normalized format.
To compute a loss you have to decide which output is compared with which object. Fixing the order does not work. If slot 1 always had to predict the leftmost object, two people standing side by side would swap slots whenever one moved a few pixels past the other, so the target of each slot would jump discontinuously with the image. A loss that is fair to a set has to be permutation-invariant: reordering the 100 predictions must not change it. DETR gets that by choosing, for each image during training, the pairing of predictions with objects that the current network already fits best, and computing the loss under that pairing.
How the Hungarian matching works
A pairing is a permutation of the indices: ground-truth element is compared with prediction . Equation (1) picks the permutation with the lowest total cost:
where is the set of all permutations. The pairwise cost scores how well prediction fits object :
The first term rewards a prediction that gives a high probability to the object's class; the second, defined in equation (9) below, penalizes a box that is far from the object's box. Both terms carry the indicator , which is 1 for a real object and 0 for padding. A padding element therefore costs 0 against every prediction, so it does not matter which leftover predictions the padding takes. The problem reduces to assigning the real objects to distinct predictions out of 100, and the official code solves exactly that: a 100 by cost matrix per image, with the padding never built.
A worked example with objects and, to keep it small, 4 slots instead of 100. The image holds a dog with box and a frisbee with box . Slot 1 predicts a box 0.08 away from the dog in summed L1 distance, with GIoU 0.821 (generalized IoU, defined below; 1 means identical boxes) and . With the paper's weights, 5 on L1 and 2 on , its cost against the dog is −0.70 + 5(0.080) + 2(0.179) = 0.057. Slot 2 is a looser second box on the dog with , cost 0.631. Slot 3 sits on the frisbee with , cost 0.759. Slot 4 predicts a box covering most of the image, and costs 7.7 and 12.0 against the two objects. Of the 12 possible pairings, the cheapest is dog to slot 1 and frisbee to slot 3, total 0.817; the next best, dog to slot 2, totals 1.390. Slots 2 and 4 are assigned .
Trying every pairing is out of the question at full size: with 7 objects and 100 slots there are 100 × 99 × ... × 94, about 8 × 10¹³, of them. This is the assignment problem, and the Hungarian method (Kuhn, 1955) solves it in polynomial time; standard implementations take steps for an matrix. It keeps a price for every row and column, adjusted so that the cheapest edges become "tight", and grows the matching one object at a time along augmenting paths, reassigning earlier objects when a later one needs their slot. DETR calls SciPy's linear_sum_assignment, a modified Jonker-Volgenant variant of the same idea that accepts the rectangular 100 by matrix directly, on the CPU.
Picture a taxi dispatcher: each passenger needs one taxi, each taxi carries one passenger, and the dispatcher minimizes total driving distance. Letting the first passenger take the nearest taxi, then the next passenger, and so on, can leave a later passenger with a taxi across town. The analogy stops at what happens to idle taxis: in DETR the unmatched predictions are not ignored, they are trained to output .
Figure 1 shows the greedy failure on three objects and five slots. In the default layout the dog grabs slot 1, its cheapest option at 2.79, and the cat, whose only good box is also slot 1, is left with slot 3 at 6.34. Matching the dog to slot 2 instead costs 0.09 more and saves the cat 3.94: the Hungarian total is 6.96, the greedy one 10.80.
The matching cost uses the probability where the loss below uses its logarithm. The paper chose this so the class term has the same scale as the box terms, which range over several units, and reports better results with it. A log-probability would let a class the network currently rates at 0.001 add 6.9 to a cost and override the box geometry. The official code writes the GIoU term as , dropping the constant 2 per object from ; that adds the same amount to every complete assignment and does not change which one is cheapest.
The Hungarian loss
With the assignment fixed, the loss of equation (2) is computed over all pairs:
The class term covers every slot. A slot matched to a real object is trained toward that object's class with the usual negative log-likelihood; a slot matched to padding is trained toward . The box term applies only to the matched real objects, because a slot that should predict nothing has no box to be compared with. With 100 slots and about 7 objects, the targets outnumber the real ones by more than 10 to 1, so the paper multiplies the log-probability term by 0.1 when , the role that sampling a fixed ratio of positive and negative proposals plays in Faster R-CNN.
Back to the worked example. Slot 1 contributes −log 0.70 = 0.357 for the class and 5(0.080) + 2(0.179) = 0.757 for the box. Slot 3 contributes −log 0.40 = 0.916 and 5(0.060) + 2(0.430) = 1.159. Slot 2, the duplicate on the dog, puts probability 0.30 on and contributes 0.1 × (−log 0.30) = 0.120, and its gradient raises that 0.30. Slot 2 is penalized even though its box is good: the dog already has a slot, so a second box on it counts as a false positive.
The matching itself receives no gradient. It returns integer indices, computed under torch.no_grad; given those indices the loss is an ordinary differentiable function of the 100 outputs, and backpropagation runs through it into the prediction heads, the transformer and the backbone. The matching is recomputed from scratch at every training step, for every image, so a slot can win an object early in training and lose it later to a slot that has become a better fit.
Figure 2 trains eight free boxes on one toy image under two rules: DETR's one-to-one matching, and the many-to-one rule of anchor-based detectors, here "every box that overlaps an object with IoU at least 0.1 is a positive for the object it overlaps most". The boxes are free parameters rather than network outputs, and the class is reduced to one probability of "object"; a box counts as confident when .
At step 200 the one-to-one run has exactly 3 confident boxes, one per object, with the other five at . The many-to-one run has 7: three on the person, three on the dog, one on the kite, all at . A real anchor-based detector ends the same way and relies on NMS to delete the four extras. In DETR the slots are outputs of one network, so suppressing duplicates is something the network has to compute: each slot needs to know what the other slots are predicting, which the decoder's self-attention (two sections down) provides.
Box loss: L1 plus generalized IoU
DETR predicts boxes directly, as absolute normalized coordinates, rather than as offsets from an anchor. The obvious loss for that is L1, the summed absolute difference of the four numbers, and L1 grows with box size: a box shifted by 10% of its own size has an L1 error eight times larger if the box is eight times larger, so large objects dominate the gradient while small boxes with the same relative error contribute little. The paper adds the generalized IoU loss of Rezatofighi et al. (2019), which depends only on relative error:
with and , and the generalized IoU loss
Here is area and is the smallest axis-aligned box containing both. The first fraction is plain IoU. The second is the share of the enclosing box that neither box covers. When the boxes overlap well the enclosing box is barely larger than their union and the correction is small. When they do not overlap at all, IoU is 0 whatever the distance, so a pure IoU loss is flat at 1 and gives no gradient, while the second fraction keeps growing as the boxes move apart. GIoU therefore ranges from −1 to 1 and the loss from 0 to 2. Every quantity is a min or max of linear functions of the box coordinates, so the loss is differentiable almost everywhere.
Figure 3 shifts a prediction diagonally by box sizes ( moves it 30% of its width right and 30% of its height down) for a small box and a large one, and plots each loss term against . Watch the two amber L1 curves separate by a factor of 8 while the teal GIoU curve serves both boxes, and watch the grey IoU curve go flat at .
At the weighted L1 term is 0.165 for the small box and 1.32 for the large one, and 2(1 − GIoU) = 1.56 for both. The ablation in Table 4 of the paper measures what each term contributes. Removing GIoU (class plus L1 only) drops COCO AP from 40.6 to 35.8, and AP on small objects from 19.9 to 13.7. Removing L1 (class plus GIoU) costs only 0.7 AP. Both box losses are summed over the matched pairs and divided by the number of objects in the batch, counted across all GPUs, so they are averages per object.
How DETR works: backbone, encoder, decoder
The network has three parts: a convolutional backbone that turns the image into a grid of feature vectors, a transformer encoder-decoder that turns 100 learned query vectors into 100 object descriptions, and a small head that reads a class and a box off each description. The paper's inference listing in PyTorch is under 50 lines.
Backbone: A ResNet-50 pretrained on ImageNet, with its classification layer removed and its batch-normalization statistics frozen. It downsamples by 32. During training images are resized so the shorter side is between 480 and 800 pixels (at most 1333 on the longer side); a typical 640 by 480 COCO photo evaluated at shorter side 800 becomes 800 by 1066, and the backbone outputs a feature map: 25 rows, 34 columns, 2048 channels.
Encoder: A 1×1 convolution reduces the 2048 channels to , and the grid is flattened into a sequence of 850 tokens of width 256. Six standard transformer encoder layers (multi-head self-attention with 8 heads, then a feed-forward network of width 2048, each with a residual connection and layer normalization, dropout 0.1) process the sequence. Self-attention lets each of the 850 positions attend to all the others, so after the encoder each feature vector depends on the whole image, not only on its receptive field. The cost is quadratic in the number of tokens: the appendix gives per layer, with .
Decoder: Six decoder layers take vectors of width 256. Each layer runs self-attention among the 100 vectors, then cross-attention from the 100 vectors to the 850 encoder outputs, then a feed-forward network. The original Transformer decoder generates its output one token at a time, each conditioned on the previous ones. DETR's decoder has no causal mask and produces all 100 outputs in one pass, since a set has no order to generate in.
Prediction heads: Each of the 100 output vectors goes through a linear layer that gives class logits (softmax over the classes plus ) and a 3-layer perceptron with ReLU and hidden width 256 that gives 4 numbers, passed through a sigmoid to land in . The same heads, with a shared layer norm in front, are applied to the output of every decoder layer during training.
# DETR (ResNet-50) on one 800 x 1066 image, shapes as in the official code
f = resnet50_conv(image) # 2048 x 25 x 34 (3 x 800 x 1066 in)
z0 = conv1x1(f) # 256 x 25 x 34
src = z0.flatten(2).T # 850 x 256 one token per cell
pos = sine_2d(25, 34) # 850 x 256 fixed, not learned
memory = encoder(src, pos) # 850 x 256 6 layers, 8 heads
q_pos = query_embed.weight # 100 x 256 learned object queries
tgt = zeros(100, 256) # decoder input starts at zero
hs = decoder(tgt, memory, pos, q_pos) # 6 x 100 x 256, every layer kept
logits = class_embed(hs) # 6 x 100 x 92 91 class ids + no-object
boxes = bbox_mlp(hs).sigmoid() # 6 x 100 x 4 (cx, cy, w, h) in [0, 1]With 6 encoder and 6 decoder layers of width 256 the model has 41.3 million parameters: 23.5M in the ResNet-50 and 17.8M in the transformer, about the size of Faster R-CNN with a feature pyramid network (42M).
A stride of 32 is coarse. The feature map has one vector per 32 by 32 pixel cell, and a small object gets about one cell. The DC5 variant ("dilated C5") removes the stride from the backbone's last stage and dilates its convolutions instead, which doubles the resolution: 50 by 67 cells, 3,350 tokens. The paper reports 16 times the encoder self-attention cost and twice the total computation (187 GFLOPS against 86). Figure 4 shows both grids over the 800 by 1066 image.
At the default 40 pixels (24 in the original photo, a COCO "small" object) the object spans 1.25 cells at stride 32 and 2.5 at stride 16. In Table 1, DC5 raises DETR's AP on small objects from 20.5 to 22.5, so resolution accounts for part of the small-object gap in the results below. Faster R-CNN with a feature pyramid network reads small objects off a map with stride 4, where quadratic self-attention over every position would be far too expensive.
Object queries and positional encodings
Attention by itself does not see order or position: permuting its inputs permutes its outputs and changes nothing else. In the decoder, if the 100 input vectors were identical, all 100 outputs would be identical too. DETR gives each slot a learned 256-dimensional vector, the object query, stored as a embedding table trained along with everything else. In the official code the decoder's input is a tensor of zeros, and the object queries are added to the queries and keys of every attention layer of every decoder layer, never to the values. The paper calls them output positional encodings: they say which slot a vector belongs to, while the content of the slot is built up layer by layer from the image.
In the encoder, the 850 tokens have lost their place in the image when the grid was flattened. DETR adds a fixed 2D sine encoding: 128 channels encode the row with sines and cosines at geometrically spaced frequencies, as in the Transformer, and 128 encode the column. Like the object queries, it is added to queries and keys at every attention layer, following equation (7) of the appendix:
where and are the query and key-value sequences, their positional encodings and the projection matrices of one head. Adding positions inside every layer rather than once at the input is worth 1.4 AP: Table 3 reports 39.2 AP for sine encodings added once at the input against 40.6 for the default. Removing the spatial encodings entirely costs 7.8 AP (32.8), and removing them from the encoder only costs 1.3 (39.3). Learned spatial encodings in place of the sines give 39.6.
After training, the slots specialize by location and size. The paper plots the boxes that 20 of the 100 slots predict over the 5,000 validation images. Each slot has a few preferred regions and box sizes, and every slot also predicts image-wide boxes, which the authors relate to the distribution of objects in COCO. The paper finds no strong specialization by class. COCO's training set has no image with more than 13 giraffes, and DETR finds all 24 on a synthetic image of 24 giraffes.
Parallel decoding, auxiliary losses, and why NMS is unnecessary
Under the set loss only one slot is matched to each object, so a second slot that predicts the same object is trained toward , and avoiding that requires each slot's output to depend on what the other slots output. Decoder self-attention provides that dependence: in every decoder layer each of the 100 vectors attends to the other 99, so a slot whose vector already resembles a confident description of the dog can push the others toward . A single decoder layer cannot do this, because its self-attention runs before any slot's vector has passed through cross-attention to the image. The first layer's self-attention is applied to the all-zero input plus object queries, which carry no image content.
The paper measures this. Because a prediction head is trained on every decoder layer's output (below), each layer's output can be evaluated as a detector. AP rises after every layer, by 8.2 AP (9.5 AP50) from the first to the sixth. Running standard NMS on these outputs raises AP for the first layer, whose outputs contain duplicates; the gain shrinks with each layer, and on the last layers NMS lowers AP slightly, because the boxes it deletes are correct detections of separate objects.
For the auxiliary losses, after each of the 6 decoder layers the shared prediction heads produce 100 (class, box) pairs, the Hungarian matching is run on them separately, and the loss of equation (2) is added. The total loss is the sum over the 6 layers. The paper found this helps the model output the correct number of objects of each class. At test time only the last layer's predictions are used.
The decoder's cross-attention maps, visualized in the paper, concentrate on the extremities of each object (heads, feet, the edges of a car), while the encoder's self-attention maps already separate individual instances. The authors' hypothesis is that the encoder does the separation of objects and the decoder only needs the boundary regions to place the box and decide the class.
Training recipe
- Optimizer: AdamW with weight decay , learning rate for the transformer and for the backbone, gradient norm clipped at 0.1. The paper reports that a backbone learning rate about ten times smaller than the rest is important for stability in the first epochs.
- Initialization: Xavier for the transformer; ImageNet-pretrained ResNet from torchvision with batch normalization frozen (its statistics and affine weights are not updated).
- Schedule: 300 epochs with the learning rate divided by 10 after 200, for the ablations; 500 epochs with the drop at 400 for the main comparison, which adds 1.5 AP. On 16 V100 GPUs with 4 images each (batch 64), 300 epochs take 3 days.
- Augmentation: random resizing (shorter side 480 to 800), and with probability 0.5 a random rectangular crop that is resized again. The crop adds about 1 AP; the paper's explanation is that it helps the encoder learn global relationships.
- Losses: weights 1 (class), 5 (L1), 2 (GIoU); no-object weight 0.1; dropout 0.1 in the transformer.
At inference every slot whose most likely class is is reported anyway, with its highest-scoring real class and that class's probability. COCO AP rewards ranking extra low-confidence detections below the confident ones rather than dropping them, and this gains 2 AP over discarding the empty slots. Every image therefore yields exactly 100 scored detections.
# Inference (models/detr.py, PostProcess): no NMS, no threshold for COCO AP
prob = logits[-1].softmax(-1) # last decoder layer, 100 x 92
scores, labels = prob[:, :-1].max(-1) # best REAL class, even if the
# slot's top class is no-object
xyxy = cxcywh_to_xyxy(boxes[-1]) * [W, H, W, H]
# 100 scored detections per image go to the COCO evaluator. For a picture,
# keep the slots whose score clears a threshold.Results against Faster R-CNN
COCO reports average precision (AP), the area under the precision-recall curve averaged over the 80 classes and over ten IoU thresholds from 0.50 to 0.95. AP50 and AP75 use a single threshold each; AP_S, AP_M and AP_L restrict the evaluation to small, medium and large objects.
The baselines are Faster R-CNN models from the Detectron2 library. The authors made them stronger in the same ways DETR is trained (generalized IoU in the box loss, the same random crops, and a 9× schedule of about 109 epochs), which adds 1 to 2 AP; these are marked with "+". Figure 5 compares each DETR model with the enhanced Faster R-CNN on the same backbone, from Table 1.
With ResNet-50 the two models tie at 42.0 AP. DETR is 7.7 points higher on large objects (61.1 against 53.4) and 6.1 points lower on small ones (20.5 against 26.6). The paper attributes the large-object gain to the encoder's global attention; the small-object deficit matches the stride-32 grid of Figure 4. DETR uses 86 GFLOPS against 180 and runs at 28 frames per second against 26. With ResNet-101, DETR-R101 reaches 43.5 AP against 44.0, and the DC5 models trade speed for accuracy: DETR-DC5-R101 is the best DETR at 44.9 AP but runs at 10 frames per second.
What the ablations show
The ablations use the ResNet-50 model on the 300-epoch schedule (40.6 AP), reporting the median of the last 10 epochs.
| Change | Params | AP | AP_S | AP_L |
|---|---|---|---|---|
| Baseline (6 encoder, 6 decoder layers) | 41.3M | 40.6 | 19.9 | 60.2 |
| 0 encoder layers | 33.4M | 36.7 | 16.8 | 54.2 |
| 3 encoder layers | 37.4M | 40.1 | 18.5 | 58.6 |
| 12 encoder layers | 49.2M | 41.6 | 19.8 | 61.9 |
| No feed-forward networks in the transformer | 28.7M | 38.3 | - | - |
| No spatial positional encodings, queries at input only | 41.3M | 32.8 | - | - |
| Loss without GIoU (class + L1) | 41.3M | 35.8 | 13.7 | 57.9 |
| Loss without L1 (class + GIoU) | 41.3M | 39.9 | 19.9 | 57.9 |
Removing the encoder costs 3.9 AP overall and 6.0 on large objects, consistent with the claim that global attention is what helps large objects. Removing the feed-forward networks leaves 10.8M parameters in the transformer and costs 2.3 AP. The 12-layer encoder gains another 1.0 AP for 8M more parameters. The 38.3 in the table is 40.6 minus the reported 2.3; the paper gives no size breakdown for that run, nor for the positional encoding run.
Panoptic segmentation
Panoptic segmentation labels every pixel of an image either with an instance of a countable "thing" class (this pixel belongs to person number 3) or with an amorphous "stuff" class (sky, grass, road). COCO's panoptic annotations have 80 thing and 53 stuff categories. Detectors usually handle stuff with a separate semantic-segmentation branch and then merge the two outputs with heuristics.
DETR treats stuff regions as objects too: it is trained with the same recipe to predict a box for each thing and each stuff region, because the Hungarian matching needs box distances. A mask head is then added and trained for 25 epochs with the rest of DETR frozen. For each decoder output it computes multi-head attention scores against the encoder's output, giving low-resolution heatmaps, and upsamples them with an FPN-style convolutional network to masks at stride 4, trained with the DICE and focal losses. At inference, predictions below 85% confidence are dropped and each pixel goes to the mask with the highest score, which guarantees the masks do not overlap.
The metric is panoptic quality (PQ), which multiplies a segmentation quality term (mean IoU of matched segments) by a recognition quality term (an F1 score over segments). On COCO val, DETR with ResNet-50 reaches 43.4 PQ against 42.4 for PanopticFPN++ with the same backbone, and DETR-R101 reaches 45.1. The difference comes from stuff classes: 36.3 PQ against 32.3. On things DETR is close (48.2 against 49.2) even though its mask AP is 31.1 against 37.7. On the COCO test set DETR scores 46 PQ.
Limitations and what came next
- Small objects. 20.5 AP_S against 26.6 for Faster R-CNN-FPN+, from the stride-32 features. Dilation (DC5) helps by 2 points at twice the computation. A feature pyramid of the kind Faster R-CNN uses is not practical with full self-attention, since the number of tokens grows with the square of the resolution.
- Training time. 500 epochs for the reported numbers, against about 109 for the enhanced Faster R-CNN and 36 for its standard 3× schedule.
- A hard cap of 100 objects. On a synthetic grid of 100 copies of one object, DETR detects all instances up to 50 visible copies, then misses more and more; with all 100 visible it finds about 30.
- Untuned loss weighting. The loss ablation uses the same weights in every run, and the paper notes that other ways of combining the terms may give different results.
Deformable DETR (Zhu et al., 2020) kept the set prediction loss and the encoder-decoder structure but replaced full attention over the feature map with attention to a small number of sampled points per query, over several feature resolutions. That made multi-scale features affordable and, by its own report, improved on DETR, especially on small objects, with 10 times fewer training epochs. Later transformer models for detection and segmentation, Mask2Former (2021) among them, keep the Hungarian set loss and learned object queries.
Implementation notes
The official repository implements the loss in two files: models/matcher.py builds the cost matrix and calls SciPy, and models/detr.py (class SetCriterion) computes the losses from the indices. A condensed version of one image's loss:
# One image with M labeled objects (tgt_cls: M, tgt_box: M x 4).
# DETR repeats this for the output of each of the 6 decoder layers and sums.
p = logits.softmax(-1) # 100 x 92
with torch.no_grad(): # the matching gets no gradient
C = (-1 * p[:, tgt_cls] # 100 x M class term
+ 5 * cdist_l1(boxes, tgt_box) # 100 x M box L1
- 2 * giou(boxes, tgt_box)) # 100 x M, GIoU in [-1, 1]
slots, objs = linear_sum_assignment(C) # scipy, M pairs
target = full(100, NO_OBJECT) # every slot defaults to empty
target[slots] = tgt_cls[objs]
w = ones(92); w[NO_OBJECT] = 0.1 # eos_coef
loss_ce = cross_entropy(logits, target, weight=w) # weighted mean
b, t = boxes[slots], tgt_box[objs] # matched pairs only
loss_l1 = l1(b, t).sum() / num_boxes
loss_giou = (1 - giou_diag(b, t)).sum() / num_boxes
loss = 1 * loss_ce + 5 * loss_l1 + 2 * loss_giou
# backward() reaches the heads, the decoder, the object queries, the encoder
# and the backbone (at a 10x smaller learning rate)- The class head has outputs for COCO because the code uses the largest category id plus one (COCO's 80 classes have ids up to 90, with gaps), and the last index is .
- The matching cost uses the softmax over all 92 outputs, including , and is run on the CPU per image.
- The classification term is a weighted mean over all 100 slots of every image in the batch, not a sum divided by the number of objects as the paper's appendix states; the box terms are divided by the number of objects. The Provenance panel below gives the arithmetic.
- The transformer layers are post-norm (residual addition, then layer norm), with a final layer norm on the decoder output that is also applied to the intermediate outputs used by the auxiliary losses.
- Images in a batch have different sizes, so the code pads them to a common size and passes a mask, used both in attention (padded positions are never attended to) and in the sine encoding (positions are normalized by the unpadded extent).
The attention used throughout is the multi-head attention of the Transformer; the same encoder applied directly to image patches, with no convolutional backbone, is the Vision Transformer, published five months after DETR.
Questions you might still have
Does DETR use non-maximum suppression at all?
No. The paper ran standard NMS on the outputs of every decoder layer as an experiment: it raises AP for the first layer, whose slots cannot see each other yet, the gain shrinks with depth, and on the last layers NMS lowers AP slightly by removing correct boxes. The released model reports its COCO numbers on all 100 raw outputs.
What happens when an image has more than 100 objects?
DETR cannot output more than its 100 slots. In the appendix test, a single object repeated on a 10 by 10 grid is fully detected with up to 50 copies visible, and with all 100 visible the model finds about 30 on average. COCO training images have at most 63 labeled instances and 7 on average, so such crowds are far outside the training data.
Is the Hungarian matching differentiable?
No, and it does not need to be. The matching runs under torch.no_grad and returns integer indices. Given those indices, the class and box losses are ordinary differentiable functions of the network outputs, so gradients flow through the matched pairs and through the no-object targets of the rest. Small changes in the outputs leave the assignment unchanged; when it does change, the loss jumps.
Why does DETR need 500 epochs?
The paper does not identify a cause. Its Faster R-CNN baselines get by with about 109 epochs even after being trained 3 times longer than the standard recipe. Deformable DETR attributes the slow convergence to attention over the whole feature map, which starts out spread almost uniformly over all 850 positions and has to learn to concentrate on a few, and it reaches better accuracy with 10 times fewer epochs by attending to a small set of sampled points instead.
How is DETR different from YOLO?
Both produce all detections in one forward pass. YOLO v1 assigns each object to the grid cell that contains its center, so the assignment is fixed by geometry, and it removes duplicates from neighboring cells with NMS. DETR has no grid: any of its 100 slots can take any object, the matching decides which during training, and duplicate suppression is learned by the decoder's self-attention. The YOLO explainer on this site covers the grid design in detail.
Do the object queries correspond to classes or to regions?
To regions and sizes, loosely. Plotting the boxes each of 20 slots predicts over the COCO validation set shows that every slot has a few preferred areas and box sizes, and all of them can predict an image-wide box. The 24-giraffe test image, more giraffes than any training image holds, is detected in full, which the authors read as evidence that no slot is tied to a class.
Why does the matching cost use probabilities but the loss use log-probabilities?
The loss needs a log-likelihood for a proper classification gradient. In the matching cost the paper uses the probability so the class term lives on the same scale as the box terms (a probability is between 0 and 1, the weighted box terms run up to several units), and reports that it worked better. A log-probability of a near-zero class would dominate the cost and override box geometry.
Can the backbone or transformer be swapped?
Yes; the paper reports ResNet-50 and ResNet-101 backbones and dilated versions of each, and the transformer is the standard encoder-decoder with positions added inside attention. Later work replaced both parts (for example with deformable attention) while keeping the set prediction loss.
Footnotes & further reading
- The paper: Carion, Massa, Synnaeve, Usunier, Kirillov & Zagoruyko, End-to-End Object Detection with Transformers (ECCV 2020). Code and pretrained models: github.com/facebookresearch/detr, files
models/matcher.py,models/detr.py,models/transformer.py,models/position_encoding.pyandmain.py. - The baseline: Ren, He, Girshick & Sun, Faster R-CNN (NeurIPS 2015), and Lin et al., Feature Pyramid Networks for Object Detection (CVPR 2017). The single-pass grid alternative: YOLO.
- The matching: Kuhn, The Hungarian method for the assignment problem (Naval Research Logistics Quarterly, 1955), and the SciPy linear_sum_assignment documentation. Earlier set losses for detection: Stewart, Andriluka & Ng, End-to-end people detection in crowded scenes (CVPR 2016).
- The box loss: Rezatofighi et al., Generalized Intersection over Union (CVPR 2019).
- The follow-up: Zhu et al., Deformable DETR: Deformable Transformers for End-to-End Object Detection (ICLR 2021); Cheng et al., Masked-attention Mask Transformer for Universal Image Segmentation (Mask2Former, CVPR 2022). Panoptic segmentation: Kirillov et al., Panoptic Segmentation (CVPR 2019).
How could this explainer be improved? Found an error, or something unclear? I read every message.