VerifiedarXiv:2301.1259728 min
Multimodal · Vision-language

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

A small trained module turns an image into 32 vectors that a frozen language model reads as its prompt.

BLIP-2 keeps a pretrained image encoder and a pretrained language model fixed and trains only the connector between them, the Q-Former. The Q-Former is trained in two stages: first against the image encoder alone, on three image-text losses, then with its output fed to the language model.

Explaining the paperBLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsLi, Li, Savarese, Hoi · Salesforce Research · ICML 2023 · arXiv:2301.12597 ↗

Salesforce's January 2023 recipe beats Flamingo-80B at zero-shot visual question answering (VQA) with 1/54 of its trainable parameters: 188M against Flamingo's 10.2B.

Two frozen models that do not share an input format

By late 2022 there were very good image encoders and very good language models, trained separately. An image encoder such as EVA-CLIP's ViT-g, a Vision Transformer with about a billion parameters, turns a 224 × 224 image into 257 vectors of 1,408 numbers each: one per 14 × 14 pixel patch (16 × 16 = 256 of them) plus a summary [CLS] token. A language model such as OPT-2.7B reads a sequence of token embeddings, each 2,560 numbers wide, and predicts the next token. Neither has seen the other's vectors. A vision-language model has to connect them, so that a question about a photo gets an answer grounded in the photo.

The expensive way is to train everything together, end to end, on hundreds of millions of image-text pairs. BLIP-2 freezes both big models (their weights never change) for two stated reasons. Compute: the billions of parameters in the two models need no gradients or optimizer state. And catastrophic forgetting: a language model fine-tuned on short image captions tends to lose some of the general text ability it came with. Freezing moves all of the adaptation into the small module between them, and that module has to feed a language model that has never seen anything but word embeddings.

Earlier work had tried this. Frozen (Tsimpoukelli et al., 2021) trained the image encoder so its output worked as a prompt prefix for a frozen 7B language model. Flamingo (2022) kept both models frozen but inserted new cross-attention layers inside the language model and trained those on billions of image-text pairs. Both trained their connector with a single signal: the language model's loss on generating the caption. The BLIP-2 paper argues that this signal alone is a weak way to teach a connector what in an image matters for text, and backs this with an ablation covered below. Its answer has three parts:

  1. a small transformer, the Q-Former, that reads the frozen image features through 32 learned query vectors and outputs exactly 32 vectors;
  2. a first training stage that teaches the Q-Former to pull text-relevant content from the image, using three image-text losses and no language model at all;
  3. a second stage that projects the 32 vectors into the language model's embedding space and trains them as a prompt the frozen model can use.

How the Q-Former works

The Q-Former is a 12-layer transformer initialised from BERT-base (hidden size 768, 12 attention heads). Its input is 32 vectors of size 768 called queries, which are trained parameters of the model like any weight matrix: the same 32 vectors start every forward pass, for every image. In each layer the queries first attend to one another (self-attention); in every other layer (layers 1, 3, 5, 7, 9, 11, counting from 1) they then attend to the frozen image features (cross-attention); then each passes through a feed-forward network. After 12 layers the 32 query positions hold 32 output vectors, which the paper calls ZZ, a 32 × 768 matrix.

Cross-attention moves information from the image into the queries. Each query computes a similarity with each of the 257 image vectors, turns those similarities into weights that sum to 1, and takes the weighted average of (projections of) the image vectors. With XqX_q the 32 query states and FF the frozen features:

CrossAttn(Xq,F)=softmax ⁣((XqWQ)(FWK)⊤dh)FWV,Xq∈R32×768,  F∈R257×1408\mathrm{CrossAttn}(X_q, F) = \mathrm{softmax}\!\left(\frac{(X_q W_Q)(F W_K)^{\top}}{\sqrt{d_h}}\right) F W_V, \qquad X_q \in \mathbb{R}^{32\times 768},\; F \in \mathbb{R}^{257\times 1408}
(1)

WKW_K and WVW_V map width 1,408 down to 768, the softmax runs over the 257 image positions, and the output has one row per query. This is the cross-attention of the original Transformer decoder, run per head (12 heads of size dh=64d_h = 64). Its output shape depends only on the number of queries. A 490-pixel image gives the ViT 35 × 35 + 1 = 1,226 tokens, and the Q-Former still returns 32.

An analogy for the queries: 32 reporters sent to look at a scene, each of whom files a report of fixed length, with an editor (the language model) who reads only the reports. The reporters can compare notes (self-attention among queries) and decide what to look at (cross-attention), and training decides what each one tends to cover. The analogy breaks in two places: the reports are vectors, not words, and nothing assigns each query a topic; any division of labor among them comes from training.

The same 12 layers also process text. When a caption is fed in, its tokens go through the same self-attention weights as the queries. According to the paper the Q-Former is two submodules, an image transformer and a text transformer, that share self-attention layers. The official code is more specific: each layer has one set of self-attention weights, applied to the concatenated sequence of queries and text, but two feed-forward networks, one for query positions and one for text positions (both copied from BERT's at initialisation). Only query positions pass through cross-attention, so text never reads the image directly. With the text side included, the Q-Former has 188M parameters, of which the queries themselves are 32 × 768 = 24,576.

The 32 × 768 output is small on purpose. At 224 pixels it holds 24,576 numbers against 361,856 in the ViT-g output (257 × 1,408), 14.7 times fewer; for the smaller CLIP ViT-L/14 (width 1,024) the ratio is 10.7. The language model gets a short prompt, and the Q-Former has to decide which parts of the image to keep.

Figure 1 · The query bottleneck
224 px6
Drag the image size from 224 to 490 pixels (the pre-training and VQA fine-tuning resolutions) and switch encoders. The amber bars (ViT tokens and numbers) grow with the square of the side; the teal bars stay at 32 tokens and 24,576 numbers. Click a query, or use the query slider, to see an attention map over the patches; those weights are illustrative, not taken from the trained model.

At the far right of the slider (490 pixels on ViT-g) the frozen encoder produces 1,726,208 numbers per image, 70 times what the Q-Former passes on. Which information survives that compression is set by the training losses in the next two sections.

How BLIP-2 is trained, stage 1: three losses, three attention masks

Stage 1 connects the Q-Former to the frozen image encoder and trains on image-caption pairs, with no language model present. It optimises three objectives at once, taken from the earlier BLIP paper by the same group. All three use the same Q-Former weights on the same inputs (32 queries plus the caption). They differ only in which positions may attend to which in self-attention, set by a mask. A mask is a grid with one row per position doing the reading and one column per position being read; a blocked cell gets weight zero in the softmax.

Figure 2 · One network, three masks
wearing
Switch objectives with the tabs; click a row (or use the slider) to list what that position may read. Teal cells are query-to-query, amber text-to-text, violet between the two; crossed cells are blocked. Only query rows have cross-attention into the image, so under ITC and ITG a text row has no path to the image except through the queries.

Under ITC, the queries' outputs cannot depend on the caption and the caption's output cannot depend on the image, so the two embeddings can be compared without either leaking the answer to the other. Under ITG, the text tokens can see the queries but the queries cannot see the text: if they could, a query could copy the caption from the text side and the generation loss would teach nothing about the image. Under ITM, full mixing is allowed because the classifier needs to compare the image and the caption word by word.

The contrastive loss with 32 queries, and how hard negatives are picked

ITC works like CLIP's contrastive loss, with one change forced by the queries. The text side gives one vector: the output at the [CLS] position, projected to 256 dimensions and normalised to length 1, called tt. The image side gives 32 such vectors z1,…,z32z_1, \dots, z_{32}, one per query. The paper scores an image against a caption by the best of the 32 dot products:

s(I,T)=1τ max⁡k=1,…,32  zk⊤ts(I, T) = \frac{1}{\tau}\,\max_{k=1,\dots,32}\; z_k^{\top} t
(2)

Each dot product is a cosine similarity between −1 and 1; τ\tau is a temperature, learned, that starts at 0.07. Dividing by 0.07 multiplies similarities by about 14, so a 0.1 gap in cosine becomes a 1.43 gap in logits, a 4.2 to 1 ratio after exponentiation. Taking the max lets any one query carry the match: a caption about sunglasses can be matched by whichever query ended up covering the face.

For a batch of NN image-caption pairs, the loss scores each image against all NN captions with its own caption as the correct class, and each caption against all NN images:

Litc=−12N∑i=1N[log⁡es(Ii,Ti)∑j=1Nes(Ii,Tj)+log⁡es(Ii,Ti)∑j=1Nes(Ij,Ti)]\mathcal{L}_{\mathrm{itc}} = -\frac{1}{2N}\sum_{i=1}^{N}\left[\log\frac{e^{s(I_i,T_i)}}{\sum_{j=1}^{N} e^{s(I_i,T_j)}} + \log\frac{e^{s(I_i,T_i)}}{\sum_{j=1}^{N} e^{s(I_j,T_i)}}\right]
(3)

Each bracketed term is a softmax cross-entropy whose correct class is the matching pair. The official code gathers embeddings across all GPUs, so NN is the global batch: 1,680 pairs for ViT-g, meaning every image is scored against 1,679 wrong captions per step. It also applies label smoothing of 0.1 (the target puts 0.9 on the true pair and spreads 0.1 over all of them), which the paper does not mention. Because the image encoder is frozen and needs no activations stored for backprop, a GPU fits more pairs than in end-to-end training, and the paper drops the queue of stored past embeddings that BLIP used for extra negatives.

ITM needs negative pairs too, and it reuses the ITC scores to choose them. For each image, the code sets the true caption's score to −10,000, applies a softmax to the row, and samples one caption from that distribution; for each caption it samples one image the same way. Captions with the highest contrastive score against the image are the most likely to be drawn, so ITM spends its effort on pairs that are hard to tell apart:

P(draw Tj as Ii’s negative)=es(Ii,Tj)∑m≠ies(Ii,Tm),j≠iP(\text{draw } T_j \text{ as } I_i\text{'s negative}) = \frac{e^{s(I_i,T_j)}}{\sum_{m\neq i} e^{s(I_i,T_m)}}, \qquad j \neq i
(4)

Each local batch of BB pairs thus becomes 3B3B ITM examples: the BB true pairs, BB captions with a sampled wrong image, and BB images with a sampled wrong caption. The ITM head is a linear layer from 768 to 2 logits applied to each of the 32 query outputs; the 32 pairs of logits are averaged, and the result goes into a two-class cross-entropy.

Figure 3 · Max over queries, softmax over the batch
0.070
Click a cell of the score matrix to see its 32 per-query similarities; the score is the tallest bar. Row i feeds two distributions below: the ITC softmax (whose −log at the true caption is the loss) and the ITM sampling distribution with the true caption removed. Sweep τ from 0.01 to 1. The five pairs and their similarities are made up; the operations are the code's.

Two things to check in the figure. At the initial τ = 0.07, image 1 (the cat) puts about 82% of its ITC probability on its own caption, and nearly all the ITM negative mass on "a dog wearing sunglasses", the caption that differs by one word. At τ = 1 the five captions get between 16% and 26%, the loss rises from 0.19 to 1.34, and the hard-negative sampler flattens: the dog caption drops to 32% and the bowl of ramen rises to 22%, so ITM would spend most of its negatives on pairs that are easy to reject.

Captioning through the queries

ITG trains the Q-Former to write the caption. The first text token is replaced by a [DEC] token to signal decoding, and the loss is the usual next-token cross-entropy over the caption words w1,…,wLw_1, \dots, w_L:

Litg=−∑n=1Llog⁡pθ ⁣(wn∣w<n, Z)\mathcal{L}_{\mathrm{itg}} = -\sum_{n=1}^{L} \log p_\theta\!\left(w_n \mid w_{<n},\, Z\right)
(5)

The conditioning on ZZ is only through self-attention to the 32 query positions (Figure 2, ITG tab). The text tokens cannot reach the image features themselves, so every fact the caption needs from the image has to be pulled into the queries first. This loss pushes the queries to carry caption-relevant detail: which object, what color, what it is doing. The code again uses label smoothing of 0.1, and does not recompute the query states: since queries cannot see text under either the ITC or the ITG mask, their keys and values from the ITC pass are cached and reused.

The three losses are added with equal weight:

Lstage 1=Litc+Litm+Litg\mathcal{L}_{\text{stage 1}} = \mathcal{L}_{\mathrm{itc}} + \mathcal{L}_{\mathrm{itm}} + \mathcal{L}_{\mathrm{itg}}
(6)

The listing below is one step as the code runs it, with shapes, for ViT-g. The frozen ViT runs without gradients; the queries, the Q-Former, the projection heads, the ITM head and the temperature are what the gradient updates.

# One stage-1 step with ViT-g, per GPU. B = local batch, captions padded to 32.
F = ln_vision(vit(images))              # (B, 257, 1408)  frozen ViT, no grad
Q = query_tokens.expand(B, 32, 768)     # the 32 learned queries
Z = qformer(queries=Q, image=F)         # (B, 32, 768)   queries only
z = normalize(vision_proj(Z))           # (B, 32, 256)
h = qformer(text=caption_ids)           # (B, 32, 768)   text only
t = normalize(text_proj(h[:, 0]))       # (B, 256)       the [CLS] output

# ITC: best of 32 query scores, against every caption on every GPU
s_i2t = einsum('bkd,cd->bck', z, all_gather(t)).max(-1) / temp
s_t2i = einsum('bd,ckd->bck', t, all_gather(z)).max(-1) / temp
loss_itc = (xent(s_i2t, pos, smooth=0.1) + xent(s_t2i, pos, smooth=0.1)) / 2

# ITM: one sampled hard negative of each kind, positives masked out first
neg_img = multinomial(softmax(mask_pos(s_t2i)))  # similar image per caption
neg_txt = multinomial(softmax(mask_pos(s_i2t)))  # similar caption per image
# pairs (img, txt) -> 1, (neg_img, txt) -> 0, (img, neg_txt) -> 0
H = qformer(queries=Q, text=pair_txt, image=pair_img, mask='bidirectional')
loss_itm = xent(itm_head(H[:, :32]).mean(1), labels)

# ITG: [DEC] replaces [CLS]; text reads the queries' cached keys/values
loss_itg = qformer_lm(dec_ids, past=Z_cache, mask='causal', smooth=0.1)

loss = loss_itc + loss_itm + loss_itg   # updates Q-Former and queries only

Stage 1 runs 250,000 steps at batch 2,320 (ViT-L) or 1,680 (ViT-g) on 129M images: COCO, Visual Genome, CC3M, CC12M, SBU, and 115M images from LAION-400M. The web captions are noisy, so the paper uses BLIP's CapFilt procedure: a BLIP-large captioner writes 10 synthetic captions per web image, a CLIP ViT-L/14 model ranks them together with the original alt text by image-text similarity, and the top two are kept, one sampled at each step.

Stage 2: 32 soft prompts for a frozen language model

Stage 2 attaches the stage-1 Q-Former (with its frozen image encoder) to a frozen language model. A single fully-connected layer maps each of the 32 output vectors from 768 to the language model's embedding width: 2,560 for OPT-2.7B, 4,096 for OPT-6.7B, 2,048 for FlanT5-XL. The 32 projected vectors are placed before the text embeddings, where the language model treats them like 32 extra input tokens. The paper calls them soft visual prompts. They are "soft" because they are not embeddings of any word in the vocabulary; they are arbitrary points in the same 2,560-dimensional space, found by gradient descent.

With H=ZWp+bpH = Z W_p + b_p, a 32 × dLLMd_{\text{LLM}} matrix, the decoder-only OPT models are trained with the ordinary language-modeling loss on the caption, conditioned on the prefix:

Lstage 2=−∑n=1Llog⁡pLLM ⁣(wn∣H, w<n)\mathcal{L}_{\text{stage 2}} = -\sum_{n=1}^{L} \log p_{\text{LLM}}\!\left(w_n \mid H,\, w_{<n}\right)
(7)

Only the caption tokens are scored; the 32 prefix positions get the ignore label. FlanT5 is an encoder-decoder T5 model, and for it the paper uses a prefix language-modeling loss: the caption is split in two, the first part goes into the encoder after the 32 visual vectors, and the decoder generates the second part. The step for OPT-2.7B:

# One stage-2 step with OPT-2.7B (hidden size 2560)
F = ln_vision(vit(images))              # (B, 257, 1408)  frozen
Z = qformer(queries=Q, image=F)         # (B, 32, 768)    trained
P = opt_proj(Z)                         # (B, 32, 2560)   trained FC layer
E = opt.embed(tokens(caption + "\n"))  # (B, T, 2560)     frozen embedding table
x = concat([P, E], dim=1)               # (B, 32 + T, 2560)
labels = concat([-100] * 32, caption_ids)   # -100: no loss on the prefix slots
loss = opt(inputs_embeds=x, labels=labels).loss
# OPT's weights get no update, but backprop runs through all 32 of its layers
# to reach P, then the FC layer, then the Q-Former.

"Frozen" means no weight updates, not no gradient computation. To train the FC layer, the gradient of the caption loss has to travel back through every layer of OPT to the 32 input positions, so stage 2 pays for the language model's backward pass (though not its weight gradients or optimizer state). The paper keeps the frozen models in FP16 (BFloat16 for FlanT5) and reports no loss in accuracy against 32-bit.

Stage 2 also discards about 81M of the Q-Former's parameters. The language model never sees the Q-Former's text side, so the official code deletes its word embeddings, its output head and its text feed-forward layers. What remains is the query path: 12 self-attention blocks, 12 query feed-forward blocks, 6 cross-attention blocks and the queries. That is 105.2M parameters with ViT-g, and adding the FC layer gives the 107M and 108M trainable counts in the paper's Table 2.

Figure 4 · What is trained in each phase
Step through the two pre-training stages and the three fine-tuning setups. Teal boxes receive weight updates, grey boxes are frozen, dashed boxes are absent. Edge labels are tensor shapes per image at that phase's resolution.

The language model box is grey in every tab where it appears, stage 2 and fine-tuning alike. The ViT box turns teal only in the fine-tuning tabs: for captioning, VQA and retrieval the paper unfreezes the image encoder, which is why those results come from models with 1.1B to 1.2B trainable parameters, not 107M.

At inference the text after the visual prefix is an instruction or question. For zero-shot VQA the paper uses the prompt "Question: {} Answer:" for OPT and "Question: {} Short answer:" for FlanT5, decodes with beam search of width 5, and sets the length penalty to −1 to favor short answers, which the paper says align better with the human annotations:

# Zero-shot VQA at inference (OPT prompt format from Sec. 4.1)
prompt = "Question: what is the cat wearing? Answer:"
x = concat([opt_proj(qformer(queries=Q, image=F)), opt.embed(tokens(prompt))])
answer = opt.generate(inputs_embeds=x, num_beams=5, length_penalty=-1)

Why the first stage matters

The paper claims stage 1 is necessary: a Q-Former trained only with the language model's generation loss bridges the gap badly. Without stage 1, the paper notes, the Q-Former is in the position of Flamingo's Perceiver Resampler, a set of learned queries trained only through the language model's loss. The test runs stage 2 twice, once from the stage-1 Q-Former and once from a Q-Former that skipped stage 1, and measures zero-shot VQAv2 accuracy as stage 2 proceeds.

Figure 5 · Stage 2 with and without stage 1
80k
Zero-shot VQAv2 (validation) accuracy at five stage-2 checkpoints; drag across the plot or use the slider to read a checkpoint, and switch language models with the tabs. Values are read off the paper's Figure 5 to about ±0.5 points.

With FlanT5-XL, skipping stage 1 costs about 19 points at the end of training (44.4 against 63.1). With OPT-6.7B the gap is about 39 points (14.9 against 54.3), and the no-stage-1 curve goes down as training continues, from 29.2 at 16k steps to 14.5 at 48k. The paper calls this catastrophic forgetting. Since OPT's weights are frozen, the degradation has to come from the prefix: one plausible reading is that a Q-Former trained only to make OPT emit captions learns a prompt that steers OPT into captioning regardless of the question that follows. FlanT5, trained on instructions, holds up better against that pull. This mechanism is our reading; the paper reports only the curves.

The paper explains the gain this way: stage 1 hands the language model a prefix whose content already correlates with text. Contrastive training makes the query outputs line up with caption embeddings; matching makes them sensitive to the difference between "cat" and "dog"; generation makes them carry whatever the caption needs. Stage 2 then only has to learn a mapping into the language model's input space, and a single linear layer (plus fine-tuning of the query path) is enough for that.

Table 6 of the paper gives a second, smaller data point, for retrieval: adding ITG to ITC + ITM during COCO fine-tuning raises image-to-text R@1 (recall at 1: the share of queries whose correct match is ranked first) from 84.5 to 85.4 and text-to-image R@1 from 67.2 to 68.3, even though retrieval never generates text.

What the paper reports

The paper leads with zero-shot visual question answering: the model is pre-trained on captions only and then asked questions from VQAv2 with no VQA training. VQA accuracy gives an answer full credit if at least 3 of the 10 human annotators gave the same answer, and partial credit (one third per annotator) otherwise.

Figure 6 · Zero-shot VQA against parameter count
g+T5-XXL
Numbers are the paper's Table 2. Toggle the x-axis between trainable and total parameters (log scale) and switch to OK-VQA; click a point or use the slider to read it. Circles are BLIP-2 variants, squares other methods.

On VQAv2 test-dev the largest BLIP-2 (ViT-g with FlanT5-XXL, 12.1B parameters in total) scores 65.0 against 56.3 for Flamingo-80B. The abstract calls this "8.7%"; it is 8.7 percentage points, or 15.5% relative. The abstract also says BLIP-2 uses "54x fewer trainable parameters": that is Flamingo's 10.2B divided by the full 188M Q-Former. The 65.0 model trains 108M parameters in stage 2, so against that count the ratio is 94. On the trainable axis all six BLIP-2 points sit between 103M and 108M; switch to total parameters and they spread from 3.1B to 12.1B, with accuracy rising along both the image-encoder and the language-model axis.

The paper reads three regularities off Table 2. ViT-g beats ViT-L for both language-model families (FlanT5-XL: 63.0 against 62.3). Within a family, bigger is better (FlanT5-XXL 65.0, FlanT5-XL 63.0). And the instruction-tuned FlanT5 beats OPT at similar size by about 10 points.

On OK-VQA, whose questions need outside knowledge (for example, which country a dish comes from), Flamingo-80B wins, 50.6 to 45.9. The paper's hypothesis is that Flamingo-80B's 70B Chinchilla language model knows more than the 11B FlanT5-XXL. On GQA, a compositional-reasoning benchmark, the best BLIP-2 scores 44.7.

The other tasks are fine-tuned, with the image encoder unfrozen and the language model frozen:

The paper also reports compute. The largest model, ViT-g with FlanT5-XXL, pre-trains on one machine with 16 A100 (40GB) GPUs in under 6 days for stage 1 and under 3 days for stage 2.

Instructions, in-context learning, and failure modes

Because the language model is untouched, it still follows text instructions placed after the visual prefix. The paper's Figure 4 shows ViT-g + FlanT5-XXL holding a conversation about a photo, answering questions that need outside knowledge or common sense, telling a story about an image, and writing text personalised to what it shows. None of this was trained; the pre-training data are single captions. The paper calls these emerging capabilities and attributes them to the language model keeping its ability to follow text prompts once the visual prefix is in place.

Few-shot prompting does not help. Large language models such as GPT-3 improve when the prompt contains solved examples; BLIP-2 given several image-question-answer examples before the real question does no better on VQA. The paper attributes this to the data: every training sample is one image and one caption, so the model never sees a sequence with several images and learns nothing about relating them. Flamingo, which does learn in context, trains on M3W, a closed dataset of web pages with interleaved images and text.

The paper's Figure 6 lists failure cases, with three stated causes: wrong knowledge in the language model, the model following an incorrect reasoning path, and no knowledge of image content newer than its training data. The frozen language model also brings its known risks (offensive output, social bias, leaking private information), which the paper suggests mitigating with instructions or with filtered training data.

Implementation notes

The trainable-parameter counts in Table 2 can be rebuilt from the LAVIS code. A BERT-base self-attention block with its LayerNorm is 2,363,904 parameters; a feed-forward block is 4,723,968; a cross-attention block whose keys and values come from 1,408-wide ViT-g features is 3,346,944 (2,757,120 from 1,024-wide ViT-L). Stage 2 keeps 12 + 12 + 6 of these, plus the 24,576 query parameters, the embedding LayerNorm, the trainable LayerNorm on the ViT output, and the FC layer (768 × 2,560 + 2,560 = 1,968,640 for OPT-2.7B). The total is 107,133,696, the "107M" in Table 2.

Pre-training uses AdamW (the decoupled weight decay variant of Adam) with β₁ = 0.9, β₂ = 0.98 and weight decay 0.05, a cosine learning-rate schedule peaking at 1e-4 after 2,000 linear warm-up steps, and a minimum learning rate of 5e-5 in stage 2. Images are 224 × 224 with random resized crops and horizontal flips. The paper uses the second-to-last layer of each ViT, dropping the final block, which it reports as slightly better; the code builds EVA ViT-g with 39 of its 40 blocks. Captions are truncated to 32 tokens.

The released LAVIS training configs are single-machine examples on COCO and Visual Genome with different batch sizes and a minimum learning rate of 1e-5; they are not the paper's 129M-image runs. For the mechanics (masks, losses, hard-negative sampling, which modules are dropped in stage 2) the code and the paper agree, and the code adds the three details the paper omits: label smoothing of 0.1 on ITC and ITG, a learned temperature starting at 0.07, and separate feed-forward layers for query and text positions.

Provenance Verified against primary literatureHow we verify
Li, Li, Savarese & Hoi (2023), arXiv v3Architecture, both pre-training stages, all result tables (1-6) and fine-tuning hyperparameters (Tables 7-9). The paper has no numbered equations; equations (1)-(6) on this page are written from Sec. 3 and checked against the official code.
LAVIS blip2_qformer.py, forward()ITC score = max over the 32 queries of 256-d normalized dot products, divided by a learned temperature initialised at 0.07; cross-entropy with label smoothing 0.1, averaged over both directions; captions all-gathered across GPUs. ITM negatives drawn with torch.multinomial from the softmaxed ITC row with the positive set to -10000. ITM logit = mean over queries of a 2-way linear head. ITG reuses the query pass's cached keys and values. Total loss is the unweighted sum.
LAVIS Qformer.py, BertLayerQueries and text share the self-attention weights but have separate feed-forward layers (intermediate_query/output_query), both copied from BERT-base at init. Cross-attention exists only in layers where layer_num % 2 == 0 (6 of 12) and only query positions pass through it. The ITG mask is the UniLM-style prefix mask built in get_extended_attention_mask with has_query=True.
LAVIS blip2.py, blip2_opt.py, eva_vit.pyQueries are a (1, 32, 768) parameter drawn from N(0, 0.02^2). A trainable LayerNorm sits on the frozen ViT output. EVA ViT-g is built with depth 39 (the last of 40 blocks removed), width 1408, patch 14. Stage 2 deletes the Q-Former's word and position embeddings, its LM head and every text-side FFN, then adds a Linear(768, d_LLM).
Parameter count, our arithmetic from the codeStage 2: 12 shared self-attention blocks (2.36M each), 12 query FFNs (4.72M each), 6 cross-attention blocks (3.35M each at width 1408) plus the FC layer give 107.1M for ViT-g + OPT-2.7B and 108.3M for FlanT5-XXL, matching Table 2. Stage 1 with the text side added sums to 186.7M; the paper reports 188M.
Alayrac et al. (2022), FlamingoPerceiver Resampler with 64 learned latents and gated cross-attention layers inserted into a frozen Chinchilla LM; 10.2B trainable parameters for Flamingo-80B (as listed in BLIP-2 Table 2).
Antol et al. (2015) / Goyal et al. (2017), VQAVQA accuracy: an answer scores min(#annotators who gave it / 3, 1), averaged over the 10 choose 9 subsets of the 10 human answers.
correctionThree points in the paper need adjusting. (1) Sec. 3.4 lists the AdamW settings as "β1 = 0.9, β1 = 0.98"; the second is β2 = 0.98. (2) "Outperforms Flamingo80B by 8.7%" on zero-shot VQAv2 is 65.0 vs 56.3, a gap of 8.7 percentage points (15.5% relative). (3) "54x fewer trainable parameters" divides Flamingo-80B's 10.2B by the full 188M Q-Former; the model that scores 65.0 trains 108M parameters in stage 2 (Table 2). Table 1 also pairs the 188M figure with NoCaps and Flickr numbers from models whose image encoder was fine-tuned (1.1B and 1.2B trainable, Tables 3 and 5).

Questions you might still have

?

Why 32 queries?
The paper uses 32 throughout and reports no ablation over the count. 32 × 768 is 24,576 numbers per image, 11 to 15 times fewer than the frozen ViT output at 224 pixels, and it costs the language model only 32 extra input positions per image. The count is fixed regardless of image resolution, so fine-tuning at 490 pixels quadruples the ViT tokens without changing the language model input.

?

Is the language model ever updated?
No. It is frozen in stage 2 and in every fine-tuning run in the paper; only the Q-Former, the FC layer and (for fine-tuning) the image encoder change. Gradients still flow backwards through the frozen language model to reach the FC layer, so stage 2 needs memory for its activations.

?

How is this different from LLaVA?
LLaVA, which came three months later, skips the query step: a linear layer maps all 256 CLIP patch tokens into the language model, and its second training stage updates the language model itself on instruction-following conversations. BLIP-2 compresses the image to 32 vectors and never touches the language model. The LLaVA explainer on this site covers its connector and data.

?

How is this different from Flamingo?
Flamingo also freezes a vision model and a language model, and also compresses the image with learned queries (64 latents in its Perceiver Resampler). Flamingo inserts new gated cross-attention layers inside the language model, which is why Flamingo-80B has 10.2B trainable parameters, while BLIP-2 only prepends 32 vectors to the input and adds nothing inside the language model. Flamingo also trains on interleaved image-text web pages, which gives it few-shot in-context learning that BLIP-2 lacks.

?

Can BLIP-2 do image search without a language model?
Yes. The stage-1 model alone is a retrieval model. The paper fine-tunes it on COCO with the same three losses, ranks all candidates by the ITC score, keeps the top 128, and re-ranks those with the ITM head. That gives 97.6% image-to-text R@1 on Flickr30K zero-shot.

?

If the language model is frozen, what is the "catastrophic forgetting" in Figure 5?
The paper uses the term for the OPT curve that falls from about 29 to about 15 as stage-2 training without stage 1 continues. The language model's weights cannot change, so this is not forgetting in the usual sense. What changes is the prefix: trained only to make OPT emit captions, the Q-Former output drifts into a form that pulls OPT towards captioning and away from answering the question that follows. That reading is ours; the paper gives the observation, not the mechanism.

?

Can I run BLIP-2 myself?
Yes. Salesforce released the weights, and Hugging Face transformers loads them with Blip2Processor and Blip2ForConditionalGeneration (for example the Salesforce/blip2-opt-2.7b checkpoint). The official training and evaluation code is in the LAVIS repository. Inference needs the frozen ViT, the Q-Former and the language model in memory; for the OPT-2.7B version that is about 3.8B parameters.

?

Why does FlanT5 beat OPT at the same size?
FlanT5 was instruction-tuned on many tasks phrased as questions; OPT was trained only to continue web text. Zero-shot VQA asks the language model to answer a question in a prompt, which is the format FlanT5 was trained on. At about 3B parameters, ViT-g + FlanT5-XL scores 63.0 on VQAv2 test-dev against 52.3 for ViT-g + OPT-2.7B.

Footnotes & further reading

  1. The paper: Li, Li, Savarese & Hoi, BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (ICML 2023; arXiv v3, June 2023). Official code: salesforce/LAVIS, projects/blip2; the model is in lavis/models/blip2_models/.
  2. The predecessor that introduced the three objectives and CapFilt: Li, Li, Xiong & Hoi, BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation (ICML 2022). The hard-negative sampling comes from Li et al., Align before Fuse (ALBEF) (NeurIPS 2021).
  3. Frozen-model predecessors: Tsimpoukelli et al., Multimodal Few-Shot Learning with Frozen Language Models (NeurIPS 2021), and Alayrac et al., Flamingo: a Visual Language Model for Few-Shot Learning (NeurIPS 2022), covered in the Flamingo explainer.
  4. The frozen components: Fang et al., EVA: Exploring the Limits of Masked Visual Representation Learning at Scale (ViT-g); Zhang et al., OPT; Chung et al., Scaling Instruction-Finetuned Language Models (FlanT5).
  5. The multimodal causal mask follows Dong et al., Unified Language Model Pre-training (UniLM) (NeurIPS 2019). The VQA accuracy metric is defined on the VQA evaluation page.
  6. A simpler connector that came three months later: Liu et al., Visual Instruction Tuning, in the LLaVA explainer.