FLAN: Finetuned Language Models Are Zero-Shot Learners
Finetune a language model on many tasks written as instructions, and it follows instructions for task types it never trained on.
The 2021 Google paper that introduced instruction tuning: a 137B-parameter model finetuned on 62 datasets rephrased through natural-language templates, then tested on whole categories of tasks that were kept out of its training.
Explaining the paperFinetuned Language Models Are Zero-Shot LearnersWith no examples in its prompt, FLAN scored higher than zero-shot 175B GPT-3 on 20 of 25 datasets. The same recipe made every model of 8B parameters or fewer worse on held-out tasks.
Why zero-shot prompting was weak
A language model trained only to predict the next token can still do tasks if you phrase them as text to continue. GPT-3 made this the main way to use a large model: put a task description and a few solved examples (the "shots") in the prompt, add the new input, and read off the continuation. With no examples at all, which the literature calls zero-shot, GPT-3 was much weaker. On natural language inference (NLI: given a premise, is a hypothesis true, false, or undetermined?), its zero-shot average over five datasets was 42.9% against 53.2% few-shot, and chance on the three-way datasets is 33.3%.
The FLAN paper attributes the gap to format: a few-shot prompt shows the model what kind of text comes next. A zero-shot prompt has to look like something from the pretraining data on its own, and an NLI question rarely does, so GPT-3's prompts dress the task up as a document. The same NLI example, prompted three ways (from Appendix E of the paper):
# One NLI example (Appendix E of the paper), three ways of prompting for it
# T5-style: a dataset tag. Only works after finetuning on this dataset.
cb hypothesis: At my age you will probably have learnt one lesson.
premise: It's not certain how many lessons you'll learn by your thirties.
# GPT-3-style: made to look like text the model saw in pretraining.
At my age you will probably have learnt one lesson.
question: It's not certain how many lessons you'll learn by your thirties.
true, false, or neither? answer:
# FLAN-style: an instruction, as you would give a person.
Premise: At my age you will probably have learnt one lesson.
Hypothesis: It's not certain how many lessons you'll learn by your thirties.
Does the premise entail the hypothesis?The third prompt is how you would ask a person. A pretrained model does not treat it as a request. Asked "'The dog runs.' Translate this sentence to French.", the untuned 137B model in this paper continued with "The dog runs after the cat". The paper's idea is to finetune the model on many tasks posed in the third style, so that responding to an instruction becomes part of what it learned, and then ask whether that skill carries over to kinds of tasks it has never been finetuned on.
How instruction tuning works
Instruction tuning is ordinary supervised finetuning on a mixture of tasks where every training input is a natural-language instruction and every target is the answer as text. The paper calls the resulting model FLAN, for Finetuned Language Net. It sits between the two usual ways of using a pretrained model. Pretrain-then-finetune, as in BERT and T5, trains one specialized model per task and needs many labeled examples of that task. Prompting, as in GPT-3, needs no training but depends on how well the prompt imitates pretraining text. Instruction tuning uses labeled data from many tasks once, and at inference expects an instruction for a task it may not have trained on.
Writing a new instruction dataset from scratch is expensive, so the authors converted existing ones. They took 62 public text datasets from TensorFlow Datasets and sorted them into 12 task clusters, where every dataset in a cluster does the same kind of task: NLI (7 datasets), reading comprehension (5), closed-book QA (3), commonsense reasoning (4), coreference (3), reading comprehension with commonsense (2), translation (8), struct-to-text (4), sentiment (4), paraphrase (4), summarization (11), and a Misc. cluster (7) holding CoQA, QuAC, TREC, CoLA, WiC, a math dataset and punctuation fixing.
For each dataset the authors wrote ten templates by hand. A template is a pair of format strings with slots for the dataset's fields, one for the input and one for the target. Up to three of the ten "turn the task around": for a sentiment dataset, one template asks the model to write a movie review with a given sentiment, so the label becomes the input and the review becomes the target. The released code has the templates verbatim. Four of the ten for RTE, an entailment dataset:
# flan/templates.py, "rte" (10 templates; shown: 1, 2, 5 and 10)
("{premise}\n\nBased on the paragraph above can we conclude that "
"\"{hypothesis}\"?\n\n{options_}", "{answer}"),
("{premise}\n\nBased on that paragraph can we conclude that this "
"sentence is true?\n{hypothesis}\n\n{options_}", "{answer}"),
("{premise}\nCan we infer the following?\n{hypothesis}\n\n{options_}",
"{answer}"),
# template 10 turns the task around: the answer becomes the input
("Generate a context and a hypothesis.",
"Context: {premise}\n\nHypothesis: {hypothesis}"),{options_} expands to a list of the allowed answers, covered in its own section below. Filling template 1 with an RTE training example gives one training pair (this example is Table 8 of the paper):
input:
After years of study, the Vatican's doctrinal congregation has sent church
leaders a confidential document concluding that "sex-change" procedures do
not change a person's gender in the eyes of the church.
Based on the paragraph above can we conclude that "Sex-change operations
become more common."?
OPTIONS:
- yes
- no
target:
noThe model reads the input and is trained to produce the target, with the usual next-token cross-entropy loss. All 137B weights are updated; nothing is frozen and no parameters are added. Every example in the mixture is formatted with one of its dataset's ten templates, so the model sees the same task phrased ten different ways over the course of training.
Holding out whole task types
If FLAN trains on RTE and is tested on another entailment dataset, a good score shows that it learned entailment, which tells you little about following new instructions. The paper uses a stricter definition of unseen: a dataset counts as unseen only if no dataset from any cluster it belongs to appeared in instruction tuning. To evaluate on NLI, the model is tuned on everything except the NLI cluster, in any template.
Some clusters overlap in content, so a footnote adds three rules. Paraphrase detection ("do these two sentences mean the same thing?") is close to entailment, so NLI and paraphrase are each removed when the other is evaluated. Reading comprehension with commonsense combines two other clusters, so evaluating it removes reading comprehension and commonsense as well, and evaluating either of those removes it. Summarization and Misc. are never evaluated; they are always part of the training mixture.
Each evaluation cluster therefore needs its own tuned model. The paper evaluates ten clusters, so it trains ten versions of FLAN, each on a different subset of the 62 datasets.
Holding out NLI removes 7 + 4 = 11 datasets and trains on the other 51. Holding out coreference removes only 3, leaving 59.
Classification with options
For generation tasks such as translation or summarization, FLAN's output is the text it generates. Classification needs a rule for turning a language model into a classifier. The standard one, used by GPT-3, is rank classification: compute the model's probability of each allowed answer string given the prompt, and predict the most likely one. For a prompt and answer strings with tokens :
The sum is the log-probability of the whole answer string. The paper points to a weakness of (1): the model may spread its probability for "yes" over many surface forms ("Yes", "yeah", "true", "it does"), and (1) only counts the one string it is given. If yes-meaning answers are phrased many ways and no-meaning answers few, the string "no" can win even when the model puts more total probability on yes.
FLAN's fix is an options suffix: every classification template ends with the token OPTIONS and a list of the allowed answers, as in the RTE example above. The intent is that a model tuned on many prompts that list their answers puts its probability on the listed strings. In the released code the evaluation still scores the options with rank classification; the suffix changes the prompt so that (1) is measuring the right thing.
Without the suffix the prediction flips to "no" at three ways of saying yes, where "yes" gets 0.20 against 0.28 for "no". With the suffix the listed string keeps most of its family's mass and the answer stays correct at every setting. The numbers are made up; the paper does not measure how probability is split. It gives the argument and reports the suffix as part of the method.
The training run
The base model is LaMDA-PT, a 137B-parameter decoder-only Transformer pretrained on web documents (including code), dialog data and Wikipedia: 2.49T tokens with a 32k SentencePiece vocabulary, about 10% of it non-English. The name marks it as the pretrained model behind Google's LaMDA dialog system without LaMDA's dialog finetuning; LaMDA-PT has only language model pretraining.
The 62 datasets differ in size by several orders of magnitude, from 250 training examples for CB to millions of sentence pairs for translation. Sampling examples in proportion to dataset size would let a few big datasets dominate training; sampling datasets uniformly would cycle through CB's examples many times over while touching a small fraction of the big sets. FLAN first caps every dataset at 30,000 training examples, then uses the examples-proportional mixing of T5 with a cap. A dataset with examples is sampled with probability
Below the cap, sampling is proportional to size. Above it, every dataset gets the same weight. FLAN's of 3,000 is ten times smaller than its 30,000-example limit, so most datasets sit at the cap and the mixture is close to uniform over datasets, with only the small ones scaled down.
At the ten datasets with at least 3,000 examples get 8.7% each within this group of 14 and CB gets 0.58%, a ratio of 15. With no cap, CB would get 0.087% and each 30,000-example dataset 13.1%, a ratio of 150.
FLAN trains for 30,000 steps with Adafactor (a memory-light relative of Adam that stores factored second-moment statistics instead of one per weight) at a learning rate of and 8,192 tokens per batch. Inputs are truncated to 1,024 tokens and targets to 256, and several short examples are packed into one sequence with an end-of-sequence token between input and target. That is 30,000 × 8,192 ≈ 246M tokens of finetuning, about one ten-thousandth of the 2.49T-token pretraining corpus, and the paper reports it took about 60 hours on a 128-core TPUv3. Every reported number comes from the final checkpoint at step 30,000.
# Instruction tuning, schematically (one of the c held-out runs)
train_sets = [d for d in DATASETS if cluster(d) not in held_out] # 51-58 sets
for d in train_sets:
d.examples = d.examples[:30_000] # per-dataset cap
d.rate = min(len(d.examples), 3_000) # mixing rate, eq. (1)
for step in range(30_000): # Adafactor, lr 3e-5
batch = []
while tokens(batch) < 8_192:
d = sample(train_sets, weights=[d.rate for d in train_sets])
x, y = d.random_example()
t = d.templates[random_index(10)] # one of 10 templates
batch.append(pack(render(t.input, x), render(t.target, y)))
loss = -log_prob(model, batch, targets_only=True) # assumed; see text
loss.backward(); adafactor.step() # all 137B weights updated
# inputs up to 1024 tokens, targets up to 256; packed with an EOS separatorThe paper and its released code do not say whether the loss also covers the input tokens of each packed pair. The schematic assumes the common setup for input/target pairs in a decoder-only model, which counts only the target tokens.
How FLAN was scored
Each dataset has up to ten templates, and FLAN's score can depend a lot on which one you use. The paper reports two numbers per dataset. The average template score is the mean over all templates, which estimates what a user with a typical phrasing would get. The best-dev template score takes the template that scores highest on a small dev split (usually 200 held-back training examples) and reports its test score, which matches the prompt engineering that GPT-3 results also allow.
The baselines are the untuned LaMDA-PT 137B, zero-shot and few-shot, prompted with GPT-3's prompts because FLAN-style instructions give near-zero scores on generation tasks for an untuned model; GPT-3 175B, using the numbers in its paper; and GLaM 64B/64E, a mixture-of-experts model, also from its paper. Comparing FLAN with LaMDA-PT isolates the effect of instruction tuning: same weights, with and without it. The LaMDA-PT prompts were not re-engineered for that model, and its few-shot runs use the largest in {1, 3, 5, 10} that fits in 1,024 tokens, where GPT-3 picked on a dev set and could use up to 100 examples in its longer context. Generation tasks are decoded greedily, where GPT-3 used beam search.
Where it helps and where it does not
Averaged over templates, zero-shot FLAN scored 56.2 on the five NLI datasets of the paper's Figure 1 (GPT-3: 42.9 zero-shot, 53.2 few-shot), 77.4 on three reading comprehension datasets (63.7 and 72.6), and 56.6 on four closed-book QA datasets (49.8 and 55.7). On each of these clusters the zero-shot tuned model beat GPT-3's few-shot average.
Against GPT-3 zero-shot with best-dev templates, FLAN is higher on 21 of the 28 rows, and on 20 of the 25 that remain once DROP, SQuADv2 and SST-2 are set aside, which is the count in the abstract. Against GPT-3 few-shot it is higher on 12. Against its own untuned base model, averaged over templates, it is higher on 23 of 28. The paper also reports best-dev FLAN above zero-shot GLaM on 13 of 19 datasets and above one-shot GLaM on 11 of 19.
The gains are largest where the task reads naturally as an instruction: NLI, reading comprehension, closed-book QA, and translation into English. Against the untuned base model, template-averaged, closed-book QA rises from 35.9 to 56.6 over its four datasets, reading comprehension from 60.9 to 77.4 over three, and NLI from 47.0 to 56.2 over five. FLAN leads GPT-3 zero-shot on all five NLI datasets, by 8.6 to 37.5 points with best-dev templates, and the paper suggests a reason: an NLI example phrased as a sentence continuation is unnatural text, while "Does <premise> mean that <hypothesis>?" is a normal question.
The losses against GPT-3 zero-shot are HellaSwag (56.7 against 78.9), ReCoRD (72.5 against 90.2), PIQA, WSC273, COPA (a tie at 91.0), and the two † datasets. Most are commonsense and coreference tasks phrased as finishing a sentence, which is already the pretraining objective, so an instruction adds little to the input. On the seven commonsense and coreference datasets, template-averaged FLAN beat untuned LaMDA-PT on four (COPA and PIQA by 0.6 points each, DPR and StoryCloze by more) and lost on HellaSwag, Winogrande and WSC273. On ReCoRD, instruction tuning cost 15 points (72.5 best-dev against LaMDA-PT's 87.8).
Translation is scored in BLEU, the n-gram overlap with reference translations on a 0-100 scale. Into English, FLAN scores 35.9 BLEU on French, 38.9 on German and 37.3 on Romanian with best-dev templates, beating GPT-3 zero-shot on all six directions but trailing GPT-3 few-shot on five of them. Out of English it is weaker (18.9 BLEU into Romanian), which the authors attribute to an English-centric tokenizer and a pretraining corpus that is about 90% English.
Why it works: clusters, scale, instructions
The paper runs three ablations to find which parts of the recipe matter. The first two use one split: NLI, closed-book QA and commonsense reasoning are held out, and up to seven other clusters are available for tuning.
The first ablation adds tuning clusters one at a time, largest first: summarization, translation, reading comprehension, sentiment, data-to-text, coreference, conversational QA. The average over the three held-out clusters rises from 49.9 with one cluster (11 datasets) to 63.5 with seven (39 datasets). Sentiment is the one addition that adds nothing (59.3 to 59.2), and the curve has not flattened at seven, so more clusters might help further.
The second instruction tunes the same split at 422M, 2B, 8B, 68B and 137B parameters, and each tuned model is compared with its untuned counterpart on 13 held-out tasks. At 68B and 137B instruction tuning adds 13 to 15 points. At 8B and below it subtracts about 5. The authors' hypothesis, which the paper does not test, is that the roughly 40 tuning tasks fill a small model's capacity, so it learns those tasks and loses general ability, while a large model has room to also learn how to follow instructions.
The third checks that the gain comes from the instructions and not from multi-task finetuning as such, by finetuning two models on the same data without them. One sees only inputs and outputs ("The dog runs." → "Le chien court."); the other sees the task and dataset name in front of each input ("[Translation: WMT'14 to French] The dog runs."). On four held-out clusters, averaged, FLAN scores 55.2, the no-template model 37.3, and the dataset-name model 46.6 with FLAN's instructions at test time or 47.0 with dataset names, 8 to 18 points below FLAN.
In the model-size tab, step from 8B to 68B: the tuned model goes from 4.8 points below its base to 13.3 above. The paper reads this as instruction following emerging with scale. The data is five model sizes on one split, so where the crossover sits (somewhere between 8B and 68B for this mixture) is not pinned down.
A fourth ablation in Appendix B.1 varies what goes into each cluster. Using four datasets per cluster instead of one raised the held-out average from about 61 to about 70. Using 10 templates per dataset instead of 1 changed it by less than a point, a smaller effect than the authors expected, since the ten templates were meant to keep the model from overfitting to one phrasing.
Few-shot exemplars and prompt tuning
Instruction tuning also combines with the two other ways of adapting a model at inference. For few-shot use, the paper formats solved examples with the same template and concatenates them before the new input:
Here is string concatenation with a delimiter token, is at most 16, and the prompt must stay under 960 tokens (the code reserves 64 of the 1,024 input tokens for separators). The same format is used when tuning FLAN's few-shot variant. Few-shot FLAN beat zero-shot FLAN on all seven evaluated clusters, by 10.2 points on struct-to-text (39.2 to 49.4), 4.6 on NLI, 3.5 on closed-book QA and 0.4 on reading comprehension. The paper attributes the larger gains to exemplars showing the output format. The spread across templates also shrank, so the few-shot model depends less on which instruction wording it gets.
Prompt tuning (Lester et al., 2021) freezes the model and learns a short sequence of input embeddings, a soft prompt, for one task by gradient descent; here the soft prompt is 10 embeddings long. The paper prompt-tunes both FLAN and LaMDA-PT on eight SuperGLUE tasks, each time using a FLAN model that saw no task from that dataset's cluster. With 32 training examples the average goes from 63.8 (LaMDA-PT) to 78.1 (FLAN); with the full training sets, from 79.2 to 87.4 (Table 4, averaged; the prompt-tuning tab of Figure 5 has the per-task numbers).
Limits and what came after
The authors list the limits themselves. Assigning datasets to clusters is partly a judgment call, and the held-out guarantee depends on it. The instructions are single sentences, much shorter than the instructions crowd workers get. The context is 1,024 tokens, too short for most summarization inputs, which is why summarization is trained on and never evaluated. The model is mostly English. It still fails some simple requests; the paper's Figure 22 shows it unable to return the second word of a sentence, and translating a question into Danish when asked to answer it in Danish.
Test contamination was checked after the fact. Following GPT-3, Appendix C splits each test set into examples that share a long n-gram (about 13 words) with the pretraining corpus and clean ones. Many datasets had heavy overlap, but the clean subsets did not score systematically lower than the full sets.
FLAN appeared on arXiv in September 2021. T0 (Sanh et al.), posted a month later, ran a similar experiment on an 11B T5 model with crowd-sourced prompts and also found zero-shot gains on held-out task types. InstructGPT (Ouyang et al., 2022; see the InstructGPT explainer) trained on prompts and demonstrations written by people and added reinforcement learning from human preference rankings. Chung et al. (2022) scaled the FLAN recipe to about 1.8K tasks, added chain-of-thought data, and released the Flan-T5 checkpoints, which is why "Flan" now usually refers to those models. Supervised finetuning on instruction data (usually called SFT) is now a standard stage after pretraining, and the same idea was carried to images in LLaVA.
Questions you might still have
Is FLAN the same model as Flan-T5?
No. FLAN is this 2021 paper's 137B LaMDA-PT model, which Google did not release. Flan-T5 and Flan-PaLM come from a 2022 follow-up (Chung et al., Scaling Instruction-Finetuned Language Models) that reused the name and the recipe with about 1.8K tasks, added chain-of-thought data, and released the T5 checkpoints publicly. The repository for this paper releases the data pipeline and templates only.
How is instruction tuning different from InstructGPT's RLHF?
FLAN is supervised finetuning only: the target text of existing NLP datasets is the label. InstructGPT (covered in the InstructGPT explainer on this site) starts with supervised finetuning on prompts and answers written by people, then trains a reward model on human rankings of outputs and optimizes the policy against it with PPO. FLAN has no human preference signal and no reinforcement learning.
FLAN trained on 62 labeled datasets. In what sense is it zero-shot?
Zero-shot here means the model gets no examples of the evaluation task at test time and saw no dataset of the same task type during tuning. It has seen plenty of labeled data for other task types. The claim is about transfer across task types, which is why the paper holds out whole clusters instead of single datasets.
Why did instruction tuning make the 8B and smaller models worse?
The paper offers a hypothesis, not a tested mechanism: a small model spends its capacity learning the roughly 40 tuning tasks and has none left for following instructions in general, while a larger model learns both. The ablation used one split (NLI, closed-book QA and commonsense held out), and the authors note that cross-validating over splits would have been better. The 2022 Flan-T5 work found instruction tuning helpful for much smaller models with far more tasks, so the crossover depends on the mixture as well as the size.
Why does instruction tuning not help commonsense and coreference tasks?
Those benchmarks are phrased as completing a sentence, which is already what the pretrained model does, so an instruction adds little information to the input. The paper says FLAN beat the untuned LaMDA-PT on only 3 of the 7 commonsense and coreference tasks; its own Table 2 gives 4 of 7 with template-averaged scores, two of them by 0.6 points. On ReCoRD, a cloze-style reading task, FLAN scored 15 points below the untuned model.
Could FLAN have seen the test sets during pretraining?
Partly, yes. Following GPT-3's procedure, Appendix C marks a test example as dirty if any of its n-grams (n around 13) appears in the pretraining corpus, and many datasets had substantial overlap. Scoring only the clean examples did not systematically lower accuracy, so there is no sign that the overlap inflated the results, but datasets with few clean examples give noisy comparisons.
Do more templates per dataset help?
Barely. Appendix B.1 tuned with 1, 4 or 10 templates per dataset. With one dataset per cluster the held-out average went 61.0, 61.6, 61.9; with four datasets per cluster it went 70.1, 70.5, 69.9. Adding datasets per cluster moved the average by about 9 points; adding templates moved it by less than 1.
Footnotes & further reading
- The paper: Wei, Bosma, Zhao, Guu, Yu, Lester, Du, Dai & Le, Finetuned Language Models Are Zero-Shot Learners (ICLR 2022; arXiv September 2021). Data and templates: github.com/google-research/FLAN.
- Prompting, rank classification and the contamination procedure: Brown et al., Language Models are Few-Shot Learners (2020); see the GPT-3 explainer.
- Examples-proportional mixing and packing: Raffel et al., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (2020), Section 3.5.2; see the T5 explainer.
- Prompt tuning: Lester, Al-Rfou & Constant, The Power of Scale for Parameter-Efficient Prompt Tuning (2021).
- T0: Sanh et al., Multitask Prompted Training Enables Zero-Shot Task Generalization (2021).
- Flan-T5 and Flan-PaLM: Chung et al., Scaling Instruction-Finetuned Language Models (2022).
- InstructGPT: Ouyang et al., Training language models to follow instructions with human feedback (2022); see the InstructGPT explainer.
How could this explainer be improved? Found an error, or something unclear? I read every message.