VerifiedarXiv:2005.1287233 min
Vision · Object detection

DETR: End-to-End Object Detection with Transformers

A transformer predicts a fixed set of 100 boxes, trained by one-to-one matching to the real objects, so detection needs no anchors and no duplicate removal.

DETR runs a ResNet and a standard transformer encoder-decoder over an image and outputs 100 (class, box) pairs in one pass. During training the Hungarian algorithm pairs each labeled object with exactly one of those outputs, and every unpaired output is trained to predict "no object".

Explaining the paperEnd-to-End Object Detection with TransformersCarion, Massa, Synnaeve, Usunier, Kirillov, Zagoruyko · ECCV 2020 · arXiv:2005.12872 ↗

A 2020 paper from Facebook AI Research that removed the hand-designed parts of an object detector and still matched a tuned Faster R-CNN on COCO, at 42.0 AP with 41 million parameters.

Nicolas Carion and Francisco Massa (equal contribution), with Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov and Sergey Zagoruyko, posted DETR (DEtection TRansformer) to arXiv in May 2020 and presented it at ECCV 2020. On COCO it scores 42.0 AP, the same as a Faster R-CNN with a feature pyramid network trained with the same augmentations and a longer schedule, at half the GFLOPS: 86 against 180. It is 7.7 points better on large objects and 6.1 points worse on small ones, and it needs a 500-epoch training schedule to get there.

The page covers, in order (equation numbers follow the paper):

  1. what earlier detectors do to produce a list of boxes, and why that list contains duplicates;
  2. detection as set prediction: a fixed number of output slots and a "no object" class;
  3. the bipartite matching of equation (1), its cost, and the Hungarian algorithm that solves it;
  4. the loss of equation (2) and the box loss of equations (9) and (10), L1 plus generalized IoU;
  5. the network: backbone, encoder, decoder, object queries, positional encodings, auxiliary losses;
  6. the training recipe, the results, the ablations, panoptic segmentation, and the limitations;
  7. the official code for the matcher, the loss and the post-processing.

Anchors, proposals and non-maximum suppression

An object detector maps an image to a list of variable length. Each entry is a class (one of COCO's 80 object categories, such as person, car, giraffe), an axis-aligned bounding box, and a confidence score. A COCO training image holds 7 labeled objects on average and at most 63.

Neural networks output tensors of fixed shape, so detectors before DETR turned the list into a dense grid of guesses. Faster R-CNN, the baseline in this paper, places a set of anchor boxes (several sizes and aspect ratios) at every position of a convolutional feature map. A first stage scores each anchor for "is this an object" and regresses an offset from the anchor to a tighter proposal; a second stage classifies the best proposals and refines their boxes again. One-stage detectors such as YOLO skip the proposal stage and predict directly from grid cells or anchors.

Training needs a rule that says which guess should predict which labeled object. The anchor-based rule is many-to-one: every anchor that overlaps a labeled box by more than a threshold, measured in intersection over union (IoU, the shared area divided by the combined area), is trained to predict that box. A person who overlaps twelve anchors gets twelve positive training examples, which helps learning but means the trained network also outputs about twelve boxes around that person at test time. A post-processing step, non-maximum suppression (NMS), sorts the boxes by score and deletes every box that overlaps a higher-scoring box of the same class by more than a fixed IoU, such as 0.5.

Both pieces are hand-tuned. The anchor sizes and ratios, the IoU thresholds for positives and negatives, and the NMS threshold all have to be set per dataset, and the final accuracy depends on them. NMS is not part of the training loss, so the network is never trained against the output it is evaluated on, and when two real objects overlap heavily (two people hugging) NMS can delete one of them. DETR replaces both with a loss that assigns each labeled object to exactly one output.

Detection as set prediction

DETR outputs a fixed number N=100N = 100 of predictions per image, always, in one pass. Each prediction y^i\hat y_i is a probability distribution over the classes plus one extra class written ∅\varnothing ("no object"), and a box b^i∈[0,1]4\hat b_i \in [0,1]^4 given as center xx, center yy, width and height, each as a fraction of the image size. NN is chosen to be much larger than the number of objects in a typical image; on an image with 7 objects, 93 of the 100 outputs should say ∅\varnothing.

The ground truth yy is a set: the order of the 7 labeled objects in the annotation file means nothing. The paper pads it to size NN with ∅\varnothing entries so both sides have 100 elements. Each ground-truth element is a pair yi=(ci,bi)y_i = (c_i, b_i), a class cic_i (possibly ∅\varnothing) and a box bib_i in the same normalized format.

To compute a loss you have to decide which output is compared with which object. Fixing the order does not work. If slot 1 always had to predict the leftmost object, two people standing side by side would swap slots whenever one moved a few pixels past the other, so the target of each slot would jump discontinuously with the image. A loss that is fair to a set has to be permutation-invariant: reordering the 100 predictions must not change it. DETR gets that by choosing, for each image during training, the pairing of predictions with objects that the current network already fits best, and computing the loss under that pairing.

How the Hungarian matching works

A pairing is a permutation σ\sigma of the NN indices: ground-truth element ii is compared with prediction σ(i)\sigma(i). Equation (1) picks the permutation with the lowest total cost:

σ^=arg⁡min⁡σ∈SN  ∑iNLmatch(yi,y^σ(i))\hat{\sigma} = \underset{\sigma \in \mathfrak{S}_N}{\arg\min} \; \sum_{i}^{N} \mathcal{L}_{\text{match}}\big(y_i, \hat{y}_{\sigma(i)}\big)
(1)

where SN\mathfrak{S}_N is the set of all N!N! permutations. The pairwise cost scores how well prediction σ(i)\sigma(i) fits object ii:

Lmatch(yi,y^σ(i))=−1{ci≠∅} p^σ(i)(ci)+1{ci≠∅} Lbox(bi,b^σ(i))\begin{aligned} &\mathcal{L}_{\text{match}}\big(y_i, \hat y_{\sigma(i)}\big) = -\mathbb{1}_{\{c_i \neq \varnothing\}}\, \hat p_{\sigma(i)}(c_i) \\ &\qquad + \mathbb{1}_{\{c_i \neq \varnothing\}}\, \mathcal{L}_{\text{box}}\big(b_i, \hat b_{\sigma(i)}\big) \end{aligned}

The first term rewards a prediction that gives a high probability p^σ(i)(ci)\hat p_{\sigma(i)}(c_i) to the object's class; the second, defined in equation (9) below, penalizes a box that is far from the object's box. Both terms carry the indicator 1{ci≠∅}\mathbb{1}_{\{c_i \neq \varnothing\}}, which is 1 for a real object and 0 for padding. A padding element therefore costs 0 against every prediction, so it does not matter which leftover predictions the padding takes. The problem reduces to assigning the MM real objects to MM distinct predictions out of 100, and the official code solves exactly that: a 100 by MM cost matrix per image, with the padding never built.

A worked example with M=2M = 2 objects and, to keep it small, 4 slots instead of 100. The image holds a dog with box (0.30,0.60,0.36,0.50)(0.30, 0.60, 0.36, 0.50) and a frisbee with box (0.72,0.30,0.12,0.08)(0.72, 0.30, 0.12, 0.08). Slot 1 predicts a box 0.08 away from the dog in summed L1 distance, with GIoU 0.821 (generalized IoU, defined below; 1 means identical boxes) and p^(dog)=0.70\hat p(\text{dog}) = 0.70. With the paper's weights, 5 on L1 and 2 on 1−GIoU1 - \text{GIoU}, its cost against the dog is −0.70 + 5(0.080) + 2(0.179) = 0.057. Slot 2 is a looser second box on the dog with p^(dog)=0.55\hat p(\text{dog}) = 0.55, cost 0.631. Slot 3 sits on the frisbee with p^(frisbee)=0.40\hat p(\text{frisbee}) = 0.40, cost 0.759. Slot 4 predicts a box covering most of the image, and costs 7.7 and 12.0 against the two objects. Of the 12 possible pairings, the cheapest is dog to slot 1 and frisbee to slot 3, total 0.817; the next best, dog to slot 2, totals 1.390. Slots 2 and 4 are assigned ∅\varnothing.

Trying every pairing is out of the question at full size: with 7 objects and 100 slots there are 100 × 99 × ... × 94, about 8 × 10¹³, of them. This is the assignment problem, and the Hungarian method (Kuhn, 1955) solves it in polynomial time; standard implementations take O(n3)O(n^3) steps for an n×nn \times n matrix. It keeps a price for every row and column, adjusted so that the cheapest edges become "tight", and grows the matching one object at a time along augmenting paths, reassigning earlier objects when a later one needs their slot. DETR calls SciPy's linear_sum_assignment, a modified Jonker-Volgenant variant of the same idea that accepts the rectangular 100 by MM matrix directly, on the CPU.

Picture a taxi dispatcher: each passenger needs one taxi, each taxi carries one passenger, and the dispatcher minimizes total driving distance. Letting the first passenger take the nearest taxi, then the next passenger, and so on, can leave a later passenger with a taxi across town. The analogy stops at what happens to idle taxis: in DETR the unmatched predictions are not ignored, they are trained to output ∅\varnothing.

Figure 1 shows the greedy failure on three objects and five slots. In the default layout the dog grabs slot 1, its cheapest option at 2.79, and the cat, whose only good box is also slot 1, is left with slot 3 at 6.34. Matching the dog to slot 2 instead costs 0.09 more and saves the cat 3.94: the Hungarian total is 6.96, the greedy one 10.80.

Figure 1 · the matching of equation (1)
Three labeled objects (amber) and five predicted boxes (teal, numbered). Each matrix cell is the matching cost of one object against one slot, with the slot's fixed class probabilities. The outlined cells are the Hungarian assignment; dashed amber cells and links show where greedy, object by object, differs. Drag a slot (or press 1 to 5 and use the arrow keys); click a cell to read its three terms.

The matching cost uses the probability p^\hat p where the loss below uses its logarithm. The paper chose this so the class term has the same scale as the box terms, which range over several units, and reports better results with it. A log-probability would let a class the network currently rates at 0.001 add 6.9 to a cost and override the box geometry. The official code writes the GIoU term as −2 GIoU-2\,\text{GIoU}, dropping the constant 2 per object from 2(1−GIoU)2(1 - \text{GIoU}); that adds the same amount to every complete assignment and does not change which one is cheapest.

The Hungarian loss

With the assignment σ^\hat\sigma fixed, the loss of equation (2) is computed over all NN pairs:

LHungarian(y,y^)=∑i=1N[−log⁡p^σ^(i)(ci)+1{ci≠∅} Lbox(bi,b^σ^(i))]\begin{aligned} &\mathcal{L}_{\text{Hungarian}}(y, \hat y) = \sum_{i=1}^{N} \Big[ -\log \hat p_{\hat\sigma(i)}(c_i) \\ &\qquad + \mathbb{1}_{\{c_i \neq \varnothing\}}\, \mathcal{L}_{\text{box}}\big(b_i, \hat b_{\hat\sigma(i)}\big) \Big] \end{aligned}
(2)

The class term covers every slot. A slot matched to a real object is trained toward that object's class with the usual negative log-likelihood; a slot matched to padding is trained toward ∅\varnothing. The box term applies only to the matched real objects, because a slot that should predict nothing has no box to be compared with. With 100 slots and about 7 objects, the ∅\varnothing targets outnumber the real ones by more than 10 to 1, so the paper multiplies the log-probability term by 0.1 when ci=∅c_i = \varnothing, the role that sampling a fixed ratio of positive and negative proposals plays in Faster R-CNN.

Back to the worked example. Slot 1 contributes −log 0.70 = 0.357 for the class and 5(0.080) + 2(0.179) = 0.757 for the box. Slot 3 contributes −log 0.40 = 0.916 and 5(0.060) + 2(0.430) = 1.159. Slot 2, the duplicate on the dog, puts probability 0.30 on ∅\varnothing and contributes 0.1 × (−log 0.30) = 0.120, and its gradient raises that 0.30. Slot 2 is penalized even though its box is good: the dog already has a slot, so a second box on it counts as a false positive.

The matching itself receives no gradient. It returns integer indices, computed under torch.no_grad; given those indices the loss is an ordinary differentiable function of the 100 outputs, and backpropagation runs through it into the prediction heads, the transformer and the backbone. The matching is recomputed from scratch at every training step, for every image, so a slot can win an object early in training and lose it later to a slot that has become a better fit.

Figure 2 trains eight free boxes on one toy image under two rules: DETR's one-to-one matching, and the many-to-one rule of anchor-based detectors, here "every box that overlaps an object with IoU at least 0.1 is a positive for the object it overlaps most". The boxes are free parameters rather than network outputs, and the class is reduced to one probability pp of "object"; a box counts as confident when p>0.5p > 0.5.

Figure 2 · one-to-one versus many-to-one targets
step 0
Same eight starting boxes, same three objects; press Play or drag the slider. Under one-to-one matching three slots move onto the objects and the other five are pushed toward "no object" with weight 0.1, so their probability falls more slowly. Under many-to-one, every overlapping box converges on its object with high confidence, and seven confident boxes cover three objects. Labels read slot:probability.

At step 200 the one-to-one run has exactly 3 confident boxes, one per object, with the other five at p=0.10p = 0.10. The many-to-one run has 7: three on the person, three on the dog, one on the kite, all at p=0.99p = 0.99. A real anchor-based detector ends the same way and relies on NMS to delete the four extras. In DETR the slots are outputs of one network, so suppressing duplicates is something the network has to compute: each slot needs to know what the other slots are predicting, which the decoder's self-attention (two sections down) provides.

Box loss: L1 plus generalized IoU

DETR predicts boxes directly, as absolute normalized coordinates, rather than as offsets from an anchor. The obvious loss for that is L1, the summed absolute difference of the four numbers, and L1 grows with box size: a box shifted by 10% of its own size has an L1 error eight times larger if the box is eight times larger, so large objects dominate the gradient while small boxes with the same relative error contribute little. The paper adds the generalized IoU loss of Rezatofighi et al. (2019), which depends only on relative error:

Lbox(bi,b^σ(i))=  λiou Liou(bi,b^σ(i))+λL1∥bi−b^σ(i)∥1\begin{aligned} \mathcal{L}_{\text{box}}\big(b_i, \hat b_{\sigma(i)}\big) = \;&\lambda_{\text{iou}}\, \mathcal{L}_{\text{iou}}\big(b_i, \hat b_{\sigma(i)}\big) \\ &+ \lambda_{\text{L1}} \big\lVert b_i - \hat b_{\sigma(i)} \big\rVert_1 \end{aligned}
(9)

with λiou=2\lambda_{\text{iou}} = 2 and λL1=5\lambda_{\text{L1}} = 5, and the generalized IoU loss

Liou(b,b^)=1−(∣b∩b^∣∣b∪b^∣−∣B(b,b^)∖(b∪b^)∣∣B(b,b^)∣)\mathcal{L}_{\text{iou}}(b, \hat b) = 1 - \left( \frac{|b \cap \hat b|}{|b \cup \hat b|} - \frac{|B(b, \hat b) \setminus (b \cup \hat b)|}{|B(b, \hat b)|} \right)
(10)

Here ∣⋅∣|\cdot| is area and B(b,b^)B(b, \hat b) is the smallest axis-aligned box containing both. The first fraction is plain IoU. The second is the share of the enclosing box that neither box covers. When the boxes overlap well the enclosing box is barely larger than their union and the correction is small. When they do not overlap at all, IoU is 0 whatever the distance, so a pure IoU loss is flat at 1 and gives no gradient, while the second fraction keeps growing as the boxes move apart. GIoU therefore ranges from −1 to 1 and the loss 1−GIoU1 - \text{GIoU} from 0 to 2. Every quantity is a min or max of linear functions of the box coordinates, so the loss is differentiable almost everywhere.

Figure 3 shifts a prediction diagonally by ss box sizes (s=0.3s = 0.3 moves it 30% of its width right and 30% of its height down) for a small box and a large one, and plots each loss term against ss. Watch the two amber L1 curves separate by a factor of 8 while the teal GIoU curve serves both boxes, and watch the grey IoU curve go flat at s=1s = 1.

Figure 3 · L1 depends on box size, GIoU does not
0.30
Drag the shift. The two panels draw the boxes at the same on-screen size; their true sizes differ by a factor of 8 in each side. 5·L1 uses the center offset only (the sizes are correct here); 2(1−GIoU) is identical for both boxes. Past s=1s = 1 the boxes stop overlapping: 1 − IoU stays at 1 while 1 − GIoU keeps rising toward 2.

At s=0.3s = 0.3 the weighted L1 term is 0.165 for the small box and 1.32 for the large one, and 2(1 − GIoU) = 1.56 for both. The ablation in Table 4 of the paper measures what each term contributes. Removing GIoU (class plus L1 only) drops COCO AP from 40.6 to 35.8, and AP on small objects from 19.9 to 13.7. Removing L1 (class plus GIoU) costs only 0.7 AP. Both box losses are summed over the matched pairs and divided by the number of objects in the batch, counted across all GPUs, so they are averages per object.

How DETR works: backbone, encoder, decoder

The network has three parts: a convolutional backbone that turns the image into a grid of feature vectors, a transformer encoder-decoder that turns 100 learned query vectors into 100 object descriptions, and a small head that reads a class and a box off each description. The paper's inference listing in PyTorch is under 50 lines.

Backbone: A ResNet-50 pretrained on ImageNet, with its classification layer removed and its batch-normalization statistics frozen. It downsamples by 32. During training images are resized so the shorter side is between 480 and 800 pixels (at most 1333 on the longer side); a typical 640 by 480 COCO photo evaluated at shorter side 800 becomes 800 by 1066, and the backbone outputs a 2048×25×342048 \times 25 \times 34 feature map: 25 rows, 34 columns, 2048 channels.

Encoder: A 1×1 convolution reduces the 2048 channels to d=256d = 256, and the 25×3425 \times 34 grid is flattened into a sequence of 850 tokens of width 256. Six standard transformer encoder layers (multi-head self-attention with 8 heads, then a feed-forward network of width 2048, each with a residual connection and layer normalization, dropout 0.1) process the sequence. Self-attention lets each of the 850 positions attend to all the others, so after the encoder each feature vector depends on the whole image, not only on its receptive field. The cost is quadratic in the number of tokens: the appendix gives O(d2HW+d(HW)2)O(d^2 HW + d (HW)^2) per layer, with HW=850HW = 850.

Decoder: Six decoder layers take N=100N = 100 vectors of width 256. Each layer runs self-attention among the 100 vectors, then cross-attention from the 100 vectors to the 850 encoder outputs, then a feed-forward network. The original Transformer decoder generates its output one token at a time, each conditioned on the previous ones. DETR's decoder has no causal mask and produces all 100 outputs in one pass, since a set has no order to generate in.

Prediction heads: Each of the 100 output vectors goes through a linear layer that gives class logits (softmax over the classes plus ∅\varnothing) and a 3-layer perceptron with ReLU and hidden width 256 that gives 4 numbers, passed through a sigmoid to land in [0,1][0, 1]. The same heads, with a shared layer norm in front, are applied to the output of every decoder layer during training.

# DETR (ResNet-50) on one 800 x 1066 image, shapes as in the official code
f      = resnet50_conv(image)        # 2048 x 25 x 34   (3 x 800 x 1066 in)
z0     = conv1x1(f)                  # 256 x 25 x 34
src    = z0.flatten(2).T             # 850 x 256        one token per cell
pos    = sine_2d(25, 34)             # 850 x 256        fixed, not learned
memory = encoder(src, pos)           # 850 x 256        6 layers, 8 heads
q_pos  = query_embed.weight          # 100 x 256        learned object queries
tgt    = zeros(100, 256)             # decoder input starts at zero
hs     = decoder(tgt, memory, pos, q_pos)   # 6 x 100 x 256, every layer kept
logits = class_embed(hs)             # 6 x 100 x 92     91 class ids + no-object
boxes  = bbox_mlp(hs).sigmoid()      # 6 x 100 x 4      (cx, cy, w, h) in [0, 1]

With 6 encoder and 6 decoder layers of width 256 the model has 41.3 million parameters: 23.5M in the ResNet-50 and 17.8M in the transformer, about the size of Faster R-CNN with a feature pyramid network (42M).

A stride of 32 is coarse. The feature map has one vector per 32 by 32 pixel cell, and a small object gets about one cell. The DC5 variant ("dilated C5") removes the stride from the backbone's last stage and dilates its convolutions instead, which doubles the resolution: 50 by 67 cells, 3,350 tokens. The paper reports 16 times the encoder self-attention cost and twice the total computation (187 GFLOPS against 86). Figure 4 shows both grids over the 800 by 1066 image.

Figure 4 · the encoder's token grid
40px
The grid is the backbone's feature map over an 800×1066 input; each cell is one encoder token. Drag the object size: the teal cells are the tokens the amber object touches, and the readout converts its side back to the 640×480 original, where COCO's small / medium / large split is defined (area under 32², under 96², above). Switch to DC5 to double the resolution.

At the default 40 pixels (24 in the original photo, a COCO "small" object) the object spans 1.25 cells at stride 32 and 2.5 at stride 16. In Table 1, DC5 raises DETR's AP on small objects from 20.5 to 22.5, so resolution accounts for part of the small-object gap in the results below. Faster R-CNN with a feature pyramid network reads small objects off a map with stride 4, where quadratic self-attention over every position would be far too expensive.

Object queries and positional encodings

Attention by itself does not see order or position: permuting its inputs permutes its outputs and changes nothing else. In the decoder, if the 100 input vectors were identical, all 100 outputs would be identical too. DETR gives each slot a learned 256-dimensional vector, the object query, stored as a 100×256100 \times 256 embedding table trained along with everything else. In the official code the decoder's input is a tensor of zeros, and the object queries are added to the queries and keys of every attention layer of every decoder layer, never to the values. The paper calls them output positional encodings: they say which slot a vector belongs to, while the content of the slot is built up layer by layer from the image.

In the encoder, the 850 tokens have lost their place in the image when the grid was flattened. DETR adds a fixed 2D sine encoding: 128 channels encode the row with sines and cosines at geometrically spaced frequencies, as in the Transformer, and 128 encode the column. Like the object queries, it is added to queries and keys at every attention layer, following equation (7) of the appendix:

[Q;K;V]=[T1′(Xq+Pq);  T2′(Xkv+Pkv);  T3′Xkv][Q; K; V] = \big[T'_1 (X_q + P_q);\; T'_2 (X_{kv} + P_{kv});\; T'_3 X_{kv}\big]
(7)

where XqX_q and XkvX_{kv} are the query and key-value sequences, Pq,PkvP_q, P_{kv} their positional encodings and T1′,T2′,T3′T'_1, T'_2, T'_3 the projection matrices of one head. Adding positions inside every layer rather than once at the input is worth 1.4 AP: Table 3 reports 39.2 AP for sine encodings added once at the input against 40.6 for the default. Removing the spatial encodings entirely costs 7.8 AP (32.8), and removing them from the encoder only costs 1.3 (39.3). Learned spatial encodings in place of the sines give 39.6.

After training, the slots specialize by location and size. The paper plots the boxes that 20 of the 100 slots predict over the 5,000 validation images. Each slot has a few preferred regions and box sizes, and every slot also predicts image-wide boxes, which the authors relate to the distribution of objects in COCO. The paper finds no strong specialization by class. COCO's training set has no image with more than 13 giraffes, and DETR finds all 24 on a synthetic image of 24 giraffes.

Parallel decoding, auxiliary losses, and why NMS is unnecessary

Under the set loss only one slot is matched to each object, so a second slot that predicts the same object is trained toward ∅\varnothing, and avoiding that requires each slot's output to depend on what the other slots output. Decoder self-attention provides that dependence: in every decoder layer each of the 100 vectors attends to the other 99, so a slot whose vector already resembles a confident description of the dog can push the others toward ∅\varnothing. A single decoder layer cannot do this, because its self-attention runs before any slot's vector has passed through cross-attention to the image. The first layer's self-attention is applied to the all-zero input plus object queries, which carry no image content.

The paper measures this. Because a prediction head is trained on every decoder layer's output (below), each layer's output can be evaluated as a detector. AP rises after every layer, by 8.2 AP (9.5 AP50) from the first to the sixth. Running standard NMS on these outputs raises AP for the first layer, whose outputs contain duplicates; the gain shrinks with each layer, and on the last layers NMS lowers AP slightly, because the boxes it deletes are correct detections of separate objects.

For the auxiliary losses, after each of the 6 decoder layers the shared prediction heads produce 100 (class, box) pairs, the Hungarian matching is run on them separately, and the loss of equation (2) is added. The total loss is the sum over the 6 layers. The paper found this helps the model output the correct number of objects of each class. At test time only the last layer's predictions are used.

The decoder's cross-attention maps, visualized in the paper, concentrate on the extremities of each object (heads, feet, the edges of a car), while the encoder's self-attention maps already separate individual instances. The authors' hypothesis is that the encoder does the separation of objects and the decoder only needs the boundary regions to place the box and decide the class.

Training recipe

At inference every slot whose most likely class is ∅\varnothing is reported anyway, with its highest-scoring real class and that class's probability. COCO AP rewards ranking extra low-confidence detections below the confident ones rather than dropping them, and this gains 2 AP over discarding the empty slots. Every image therefore yields exactly 100 scored detections.

# Inference (models/detr.py, PostProcess): no NMS, no threshold for COCO AP
prob = logits[-1].softmax(-1)               # last decoder layer, 100 x 92
scores, labels = prob[:, :-1].max(-1)       # best REAL class, even if the
                                            # slot's top class is no-object
xyxy = cxcywh_to_xyxy(boxes[-1]) * [W, H, W, H]
# 100 scored detections per image go to the COCO evaluator. For a picture,
# keep the slots whose score clears a threshold.

Results against Faster R-CNN

COCO reports average precision (AP), the area under the precision-recall curve averaged over the 80 classes and over ten IoU thresholds from 0.50 to 0.95. AP50 and AP75 use a single threshold each; AP_S, AP_M and AP_L restrict the evaluation to small, medium and large objects.

The baselines are Faster R-CNN models from the Detectron2 library. The authors made them stronger in the same ways DETR is trained (generalized IoU in the box loss, the same random crops, and a 9× schedule of about 109 epochs), which adds 1 to 2 AP; these are marked with "+". Figure 5 compares each DETR model with the enhanced Faster R-CNN on the same backbone, from Table 1.

Figure 5 · Table 1, COCO validation
Pick a backbone. DETR and the enhanced Faster R-CNN with the same backbone; the right column is DETR minus Faster R-CNN. The cost line under each name gives GFLOPS, frames per second on a V100 and parameter count.

With ResNet-50 the two models tie at 42.0 AP. DETR is 7.7 points higher on large objects (61.1 against 53.4) and 6.1 points lower on small ones (20.5 against 26.6). The paper attributes the large-object gain to the encoder's global attention; the small-object deficit matches the stride-32 grid of Figure 4. DETR uses 86 GFLOPS against 180 and runs at 28 frames per second against 26. With ResNet-101, DETR-R101 reaches 43.5 AP against 44.0, and the DC5 models trade speed for accuracy: DETR-DC5-R101 is the best DETR at 44.9 AP but runs at 10 frames per second.

What the ablations show

The ablations use the ResNet-50 model on the 300-epoch schedule (40.6 AP), reporting the median of the last 10 epochs.

ChangeParamsAPAP_SAP_L
Baseline (6 encoder, 6 decoder layers)41.3M40.619.960.2
0 encoder layers33.4M36.716.854.2
3 encoder layers37.4M40.118.558.6
12 encoder layers49.2M41.619.861.9
No feed-forward networks in the transformer28.7M38.3--
No spatial positional encodings, queries at input only41.3M32.8--
Loss without GIoU (class + L1)41.3M35.813.757.9
Loss without L1 (class + GIoU)41.3M39.919.957.9

Removing the encoder costs 3.9 AP overall and 6.0 on large objects, consistent with the claim that global attention is what helps large objects. Removing the feed-forward networks leaves 10.8M parameters in the transformer and costs 2.3 AP. The 12-layer encoder gains another 1.0 AP for 8M more parameters. The 38.3 in the table is 40.6 minus the reported 2.3; the paper gives no size breakdown for that run, nor for the positional encoding run.

Panoptic segmentation

Panoptic segmentation labels every pixel of an image either with an instance of a countable "thing" class (this pixel belongs to person number 3) or with an amorphous "stuff" class (sky, grass, road). COCO's panoptic annotations have 80 thing and 53 stuff categories. Detectors usually handle stuff with a separate semantic-segmentation branch and then merge the two outputs with heuristics.

DETR treats stuff regions as objects too: it is trained with the same recipe to predict a box for each thing and each stuff region, because the Hungarian matching needs box distances. A mask head is then added and trained for 25 epochs with the rest of DETR frozen. For each decoder output it computes multi-head attention scores against the encoder's output, giving low-resolution heatmaps, and upsamples them with an FPN-style convolutional network to masks at stride 4, trained with the DICE and focal losses. At inference, predictions below 85% confidence are dropped and each pixel goes to the mask with the highest score, which guarantees the masks do not overlap.

The metric is panoptic quality (PQ), which multiplies a segmentation quality term (mean IoU of matched segments) by a recognition quality term (an F1 score over segments). On COCO val, DETR with ResNet-50 reaches 43.4 PQ against 42.4 for PanopticFPN++ with the same backbone, and DETR-R101 reaches 45.1. The difference comes from stuff classes: 36.3 PQ against 32.3. On things DETR is close (48.2 against 49.2) even though its mask AP is 31.1 against 37.7. On the COCO test set DETR scores 46 PQ.

Limitations and what came next

Deformable DETR (Zhu et al., 2020) kept the set prediction loss and the encoder-decoder structure but replaced full attention over the feature map with attention to a small number of sampled points per query, over several feature resolutions. That made multi-scale features affordable and, by its own report, improved on DETR, especially on small objects, with 10 times fewer training epochs. Later transformer models for detection and segmentation, Mask2Former (2021) among them, keep the Hungarian set loss and learned object queries.

Implementation notes

The official repository implements the loss in two files: models/matcher.py builds the cost matrix and calls SciPy, and models/detr.py (class SetCriterion) computes the losses from the indices. A condensed version of one image's loss:

# One image with M labeled objects (tgt_cls: M, tgt_box: M x 4).
# DETR repeats this for the output of each of the 6 decoder layers and sums.
p = logits.softmax(-1)                          # 100 x 92
with torch.no_grad():                           # the matching gets no gradient
    C = (-1 * p[:, tgt_cls]                     # 100 x M class term
         + 5 * cdist_l1(boxes, tgt_box)         # 100 x M box L1
         - 2 * giou(boxes, tgt_box))            # 100 x M, GIoU in [-1, 1]
    slots, objs = linear_sum_assignment(C)      # scipy, M pairs

target = full(100, NO_OBJECT)                   # every slot defaults to empty
target[slots] = tgt_cls[objs]
w = ones(92); w[NO_OBJECT] = 0.1                # eos_coef
loss_ce = cross_entropy(logits, target, weight=w)        # weighted mean

b, t = boxes[slots], tgt_box[objs]              # matched pairs only
loss_l1 = l1(b, t).sum() / num_boxes
loss_giou = (1 - giou_diag(b, t)).sum() / num_boxes
loss = 1 * loss_ce + 5 * loss_l1 + 2 * loss_giou
# backward() reaches the heads, the decoder, the object queries, the encoder
# and the backbone (at a 10x smaller learning rate)

The attention used throughout is the multi-head attention of the Transformer; the same encoder applied directly to image patches, with no convolutional backbone, is the Vision Transformer, published five months after DETR.

Provenance Verified against primary literatureHow we verify
Carion, Massa, Synnaeve, Usunier, Kirillov & Zagoruyko (ECCV 2020; arXiv 2005.12872, HTML v3 and PDF v1)Equations (1), (2), (7), (9), (10), the matching cost, N = 100, lambda_L1 = 5, lambda_iou = 2, the 0.1 no-object weight, the training recipe, Tables 1 to 5, the decoder-layer and NMS analysis, the giraffe and 100-instance experiments. v1 and the latest version print the same tables.
github.com/facebookresearch/detr, models/matcher.pyCost matrix C = 5 L1 + 1 (-p) + 2 (-GIoU), with p the softmax over all 92 outputs including no-object. Targets exclude no-object, so the code solves a rectangular 100 x M problem per image with scipy linear_sum_assignment, inside torch.no_grad.
models/detr.py, SetCriterion and PostProcessNo-object class weight eos_coef = 0.1 inside cross_entropy; loss weights 1 / 5 / 2; box losses divided by the number of objects, all-reduced over GPUs; the matcher re-runs for each decoder layer's auxiliary output. PostProcess scores every slot by its best real class.
models/transformer.py, models/position_encoding.pyDecoder input tgt = zeros; object queries and the sine positions are added to queries and keys of every attention layer, never to values; post-norm layers by default; sine encoding uses 128 channels for y and 128 for x, temperature 10000, coordinates normalized to [0, 2 pi].
main.py defaultslr 1e-4, lr_backbone 1e-5, weight_decay 1e-4, clip_max_norm 0.1, epochs 300, lr_drop 200, hidden_dim 256, nheads 8, enc/dec layers 6/6, dim_feedforward 2048, dropout 0.1, num_queries 100.
Rezatofighi et al., Generalized Intersection over Union (CVPR 2019)GIoU = IoU - |C \ (A u B)| / |C| with C the smallest enclosing box; -1 <= GIoU <= 1; the loss 1 - GIoU lies in [0, 2] and still has a gradient when the boxes do not overlap.
Kuhn (1955); SciPy linear_sum_assignment documentationThe assignment problem and the Hungarian method. SciPy's solver is a modified Jonker-Volgenant algorithm and accepts rectangular cost matrices.
Zhu et al., Deformable DETR (ICLR 2021), abstractDETR "suffers from slow convergence and limited feature spatial resolution"; at initialization attention casts "nearly uniform attention weights to all the pixels"; Deformable DETR reaches better accuracy, especially on small objects, "with 10 times less training epochs".
This page, arithmeticToken counts (25 x 34 = 850 at stride 32; 50 x 67 = 3,350 at stride 16, 15.5 times the attention entries), the worked matching example, and the Table 1 differences.
correctionThree points where the paper and its code or its own tables disagree. (1) Appendix A.2 says all losses are normalized by the number of objects in the batch. In the official code only the box losses are; the classification term is PyTorch's weighted-mean cross-entropy over all 100 slots, sum(w_i l_i) / sum(w_i), with w = 1 for a matched slot and 0.1 for a no-object slot. For an image with 7 objects the divisor is 7 + 0.1 x 93 = 16.3, not 7, which changes how strongly the class term weighs against the box terms. (2) Section 4.1 credits DETR with +7.8 AP_L and -5.5 AP_S against Faster R-CNN; Table 1 gives 61.1 vs 53.4 (+7.7) and 20.5 vs 26.6 (-6.1) for the matched-size Faster R-CNN-FPN+, in v1 and in the latest version. This page uses the table. Table 2 also lists the same 41.3M baseline at 23 FPS while Table 1 and Section 4.2 say 28. (3) Equation (2) prints the predicted box as b-hat sub sigma-hat, applied to (i); the intended term is b-hat sub sigma-hat(i), and the "B \ b u b-hat" in equation (10) means B minus the union of the two boxes.

Questions you might still have

?

Does DETR use non-maximum suppression at all?
No. The paper ran standard NMS on the outputs of every decoder layer as an experiment: it raises AP for the first layer, whose slots cannot see each other yet, the gain shrinks with depth, and on the last layers NMS lowers AP slightly by removing correct boxes. The released model reports its COCO numbers on all 100 raw outputs.

?

What happens when an image has more than 100 objects?
DETR cannot output more than its 100 slots. In the appendix test, a single object repeated on a 10 by 10 grid is fully detected with up to 50 copies visible, and with all 100 visible the model finds about 30 on average. COCO training images have at most 63 labeled instances and 7 on average, so such crowds are far outside the training data.

?

Is the Hungarian matching differentiable?
No, and it does not need to be. The matching runs under torch.no_grad and returns integer indices. Given those indices, the class and box losses are ordinary differentiable functions of the network outputs, so gradients flow through the matched pairs and through the no-object targets of the rest. Small changes in the outputs leave the assignment unchanged; when it does change, the loss jumps.

?

Why does DETR need 500 epochs?
The paper does not identify a cause. Its Faster R-CNN baselines get by with about 109 epochs even after being trained 3 times longer than the standard recipe. Deformable DETR attributes the slow convergence to attention over the whole feature map, which starts out spread almost uniformly over all 850 positions and has to learn to concentrate on a few, and it reaches better accuracy with 10 times fewer epochs by attending to a small set of sampled points instead.

?

How is DETR different from YOLO?
Both produce all detections in one forward pass. YOLO v1 assigns each object to the grid cell that contains its center, so the assignment is fixed by geometry, and it removes duplicates from neighboring cells with NMS. DETR has no grid: any of its 100 slots can take any object, the matching decides which during training, and duplicate suppression is learned by the decoder's self-attention. The YOLO explainer on this site covers the grid design in detail.

?

Do the object queries correspond to classes or to regions?
To regions and sizes, loosely. Plotting the boxes each of 20 slots predicts over the COCO validation set shows that every slot has a few preferred areas and box sizes, and all of them can predict an image-wide box. The 24-giraffe test image, more giraffes than any training image holds, is detected in full, which the authors read as evidence that no slot is tied to a class.

?

Why does the matching cost use probabilities but the loss use log-probabilities?
The loss needs a log-likelihood for a proper classification gradient. In the matching cost the paper uses the probability so the class term lives on the same scale as the box terms (a probability is between 0 and 1, the weighted box terms run up to several units), and reports that it worked better. A log-probability of a near-zero class would dominate the cost and override box geometry.

?

Can the backbone or transformer be swapped?
Yes; the paper reports ResNet-50 and ResNet-101 backbones and dilated versions of each, and the transformer is the standard encoder-decoder with positions added inside attention. Later work replaced both parts (for example with deformable attention) while keeping the set prediction loss.

Footnotes & further reading

  1. The paper: Carion, Massa, Synnaeve, Usunier, Kirillov & Zagoruyko, End-to-End Object Detection with Transformers (ECCV 2020). Code and pretrained models: github.com/facebookresearch/detr, files models/matcher.py, models/detr.py, models/transformer.py, models/position_encoding.py and main.py.
  2. The baseline: Ren, He, Girshick & Sun, Faster R-CNN (NeurIPS 2015), and Lin et al., Feature Pyramid Networks for Object Detection (CVPR 2017). The single-pass grid alternative: YOLO.
  3. The matching: Kuhn, The Hungarian method for the assignment problem (Naval Research Logistics Quarterly, 1955), and the SciPy linear_sum_assignment documentation. Earlier set losses for detection: Stewart, Andriluka & Ng, End-to-end people detection in crowded scenes (CVPR 2016).
  4. The box loss: Rezatofighi et al., Generalized Intersection over Union (CVPR 2019).
  5. The follow-up: Zhu et al., Deformable DETR: Deformable Transformers for End-to-End Object Detection (ICLR 2021); Cheng et al., Masked-attention Mask Transformer for Universal Image Segmentation (Mask2Former, CVPR 2022). Panoptic segmentation: Kirillov et al., Panoptic Segmentation (CVPR 2019).