VerifiedarXiv:2109.0165224 min
LLMs · Instruction tuning

FLAN: Finetuned Language Models Are Zero-Shot Learners

Finetune a language model on many tasks written as instructions, and it follows instructions for task types it never trained on.

The 2021 Google paper that introduced instruction tuning: a 137B-parameter model finetuned on 62 datasets rephrased through natural-language templates, then tested on whole categories of tasks that were kept out of its training.

Explaining the paperFinetuned Language Models Are Zero-Shot LearnersWei, Bosma, Zhao, Guu, Yu, Lester, Du, Dai, Le · ICLR 2022 · arXiv:2109.01652 ↗

With no examples in its prompt, FLAN scored higher than zero-shot 175B GPT-3 on 20 of 25 datasets. The same recipe made every model of 8B parameters or fewer worse on held-out tasks.

Why zero-shot prompting was weak

A language model trained only to predict the next token can still do tasks if you phrase them as text to continue. GPT-3 made this the main way to use a large model: put a task description and a few solved examples (the "shots") in the prompt, add the new input, and read off the continuation. With no examples at all, which the literature calls zero-shot, GPT-3 was much weaker. On natural language inference (NLI: given a premise, is a hypothesis true, false, or undetermined?), its zero-shot average over five datasets was 42.9% against 53.2% few-shot, and chance on the three-way datasets is 33.3%.

The FLAN paper attributes the gap to format: a few-shot prompt shows the model what kind of text comes next. A zero-shot prompt has to look like something from the pretraining data on its own, and an NLI question rarely does, so GPT-3's prompts dress the task up as a document. The same NLI example, prompted three ways (from Appendix E of the paper):

# One NLI example (Appendix E of the paper), three ways of prompting for it

# T5-style: a dataset tag. Only works after finetuning on this dataset.
cb hypothesis: At my age you will probably have learnt one lesson.
premise: It's not certain how many lessons you'll learn by your thirties.

# GPT-3-style: made to look like text the model saw in pretraining.
At my age you will probably have learnt one lesson.
question: It's not certain how many lessons you'll learn by your thirties.
true, false, or neither? answer:

# FLAN-style: an instruction, as you would give a person.
Premise: At my age you will probably have learnt one lesson.
Hypothesis: It's not certain how many lessons you'll learn by your thirties.
Does the premise entail the hypothesis?

The third prompt is how you would ask a person. A pretrained model does not treat it as a request. Asked "'The dog runs.' Translate this sentence to French.", the untuned 137B model in this paper continued with "The dog runs after the cat". The paper's idea is to finetune the model on many tasks posed in the third style, so that responding to an instruction becomes part of what it learned, and then ask whether that skill carries over to kinds of tasks it has never been finetuned on.

How instruction tuning works

Instruction tuning is ordinary supervised finetuning on a mixture of tasks where every training input is a natural-language instruction and every target is the answer as text. The paper calls the resulting model FLAN, for Finetuned Language Net. It sits between the two usual ways of using a pretrained model. Pretrain-then-finetune, as in BERT and T5, trains one specialized model per task and needs many labeled examples of that task. Prompting, as in GPT-3, needs no training but depends on how well the prompt imitates pretraining text. Instruction tuning uses labeled data from many tasks once, and at inference expects an instruction for a task it may not have trained on.

Writing a new instruction dataset from scratch is expensive, so the authors converted existing ones. They took 62 public text datasets from TensorFlow Datasets and sorted them into 12 task clusters, where every dataset in a cluster does the same kind of task: NLI (7 datasets), reading comprehension (5), closed-book QA (3), commonsense reasoning (4), coreference (3), reading comprehension with commonsense (2), translation (8), struct-to-text (4), sentiment (4), paraphrase (4), summarization (11), and a Misc. cluster (7) holding CoQA, QuAC, TREC, CoLA, WiC, a math dataset and punctuation fixing.

For each dataset the authors wrote ten templates by hand. A template is a pair of format strings with slots for the dataset's fields, one for the input and one for the target. Up to three of the ten "turn the task around": for a sentiment dataset, one template asks the model to write a movie review with a given sentiment, so the label becomes the input and the review becomes the target. The released code has the templates verbatim. Four of the ten for RTE, an entailment dataset:

# flan/templates.py, "rte" (10 templates; shown: 1, 2, 5 and 10)
("{premise}\n\nBased on the paragraph above can we conclude that "
 "\"{hypothesis}\"?\n\n{options_}", "{answer}"),
("{premise}\n\nBased on that paragraph can we conclude that this "
 "sentence is true?\n{hypothesis}\n\n{options_}", "{answer}"),
("{premise}\nCan we infer the following?\n{hypothesis}\n\n{options_}",
 "{answer}"),
# template 10 turns the task around: the answer becomes the input
("Generate a context and a hypothesis.",
 "Context: {premise}\n\nHypothesis: {hypothesis}"),

{options_} expands to a list of the allowed answers, covered in its own section below. Filling template 1 with an RTE training example gives one training pair (this example is Table 8 of the paper):

input:
After years of study, the Vatican's doctrinal congregation has sent church
leaders a confidential document concluding that "sex-change" procedures do
not change a person's gender in the eyes of the church.

Based on the paragraph above can we conclude that "Sex-change operations
become more common."?

OPTIONS:
- yes
- no

target:
no

The model reads the input and is trained to produce the target, with the usual next-token cross-entropy loss. All 137B weights are updated; nothing is frozen and no parameters are added. Every example in the mixture is formatted with one of its dataset's ten templates, so the model sees the same task phrased ten different ways over the course of training.

Holding out whole task types

If FLAN trains on RTE and is tested on another entailment dataset, a good score shows that it learned entailment, which tells you little about following new instructions. The paper uses a stricter definition of unseen: a dataset counts as unseen only if no dataset from any cluster it belongs to appeared in instruction tuning. To evaluate on NLI, the model is tuned on everything except the NLI cluster, in any template.

Some clusters overlap in content, so a footnote adds three rules. Paraphrase detection ("do these two sentences mean the same thing?") is close to entailment, so NLI and paraphrase are each removed when the other is evaluated. Reading comprehension with commonsense combines two other clusters, so evaluating it removes reading comprehension and commonsense as well, and evaluating either of those removes it. Summarization and Misc. are never evaluated; they are always part of the training mixture.

Each evaluation cluster therefore needs its own tuned model. The paper evaluates ten clusters, so it trains ten versions of FLAN, each on a different subset of the 62 datasets.

Figure 1 · one tuned model per held-out cluster
The 12 clusters of the paper's Figure 3. Click a tile or pick from the menu to hold that cluster out. The amber tile is the evaluation cluster, dashed tiles are removed by the overlap rules, and the teal tiles are the training mixture for that one model. Try NLI, then Reading comp. w/ commonsense.

Holding out NLI removes 7 + 4 = 11 datasets and trains on the other 51. Holding out coreference removes only 3, leaving 59.

Classification with options

For generation tasks such as translation or summarization, FLAN's output is the text it generates. Classification needs a rule for turning a language model into a classifier. The standard one, used by GPT-3, is rank classification: compute the model's probability of each allowed answer string given the prompt, and predict the most likely one. For a prompt xx and answer strings cc with tokens c1,…,cTc_1, \ldots, c_T:

c^  =  arg⁡max⁡c ∈ options  ∑t=1Tlog⁡pθ ⁣(ct∣x, c<t)\hat c \;=\; \arg\max_{c \,\in\, \text{options}} \; \sum_{t=1}^{T} \log p_\theta\!\left(c_t \mid x,\, c_{<t}\right)
(1)

The sum is the log-probability of the whole answer string. The paper points to a weakness of (1): the model may spread its probability for "yes" over many surface forms ("Yes", "yeah", "true", "it does"), and (1) only counts the one string it is given. If yes-meaning answers are phrased many ways and no-meaning answers few, the string "no" can win even when the model puts more total probability on yes.

FLAN's fix is an options suffix: every classification template ends with the token OPTIONS and a list of the allowed answers, as in the RTE example above. The intent is that a model tuned on many prompts that list their answers puts its probability on the listed strings. In the released code the evaluation still scores the options with rank classification; the suffix changes the prompt so that (1) is measuring the right thing.

Figure 2 · splitting probability across surface forms
Toy probabilities for a yes/no question whose correct answer is yes. The model puts 0.60 on yes-meaning answers and 0.40 on no-meaning answers; rank classification compares only the outlined bars for "yes" and "no". Slide the number of ways to say yes from 1 to 8, then switch on the OPTIONS suffix.

Without the suffix the prediction flips to "no" at three ways of saying yes, where "yes" gets 0.20 against 0.28 for "no". With the suffix the listed string keeps most of its family's mass and the answer stays correct at every setting. The numbers are made up; the paper does not measure how probability is split. It gives the argument and reports the suffix as part of the method.

The training run

The base model is LaMDA-PT, a 137B-parameter decoder-only Transformer pretrained on web documents (including code), dialog data and Wikipedia: 2.49T tokens with a 32k SentencePiece vocabulary, about 10% of it non-English. The name marks it as the pretrained model behind Google's LaMDA dialog system without LaMDA's dialog finetuning; LaMDA-PT has only language model pretraining.

The 62 datasets differ in size by several orders of magnitude, from 250 training examples for CB to millions of sentence pairs for translation. Sampling examples in proportion to dataset size would let a few big datasets dominate training; sampling datasets uniformly would cycle through CB's examples many times over while touching a small fraction of the big sets. FLAN first caps every dataset at 30,000 training examples, then uses the examples-proportional mixing of T5 with a capKK. A dataset with eme_m examples is sampled with probability

rm  =  min⁡(em, K)∑nmin⁡(en, K),K=3,000r_m \;=\; \frac{\min(e_m,\, K)}{\sum_n \min(e_n,\, K)}, \qquad K = 3{,}000
(2)

Below the cap, sampling is proportional to size. Above it, every dataset gets the same weight. FLAN's KK of 3,000 is ten times smaller than its 30,000-example limit, so most datasets sit at the cap and the mixture is close to uniform over datasets, with only the small ones scaled down.

Figure 3 · examples-proportional mixing with a cap
Sampling share from (2) for 14 of FLAN's datasets, with the training-example counts the paper lists in Appendix G (shares computed among these 14). Amber bars are datasets below the cap, teal bars are capped, and the white tick is the share with no cap. Slide KK from 1 (every dataset equal) to above 30,000 (plain proportional).

At K=3,000K = 3{,}000 the ten datasets with at least 3,000 examples get 8.7% each within this group of 14 and CB gets 0.58%, a ratio of 15. With no cap, CB would get 0.087% and each 30,000-example dataset 13.1%, a ratio of 150.

FLAN trains for 30,000 steps with Adafactor (a memory-light relative of Adam that stores factored second-moment statistics instead of one per weight) at a learning rate of 3×10−53\times10^{-5} and 8,192 tokens per batch. Inputs are truncated to 1,024 tokens and targets to 256, and several short examples are packed into one sequence with an end-of-sequence token between input and target. That is 30,000 × 8,192 ≈ 246M tokens of finetuning, about one ten-thousandth of the 2.49T-token pretraining corpus, and the paper reports it took about 60 hours on a 128-core TPUv3. Every reported number comes from the final checkpoint at step 30,000.

# Instruction tuning, schematically (one of the c held-out runs)
train_sets = [d for d in DATASETS if cluster(d) not in held_out]  # 51-58 sets
for d in train_sets:
    d.examples = d.examples[:30_000]              # per-dataset cap
    d.rate = min(len(d.examples), 3_000)          # mixing rate, eq. (1)

for step in range(30_000):                        # Adafactor, lr 3e-5
    batch = []
    while tokens(batch) < 8_192:
        d = sample(train_sets, weights=[d.rate for d in train_sets])
        x, y = d.random_example()
        t = d.templates[random_index(10)]         # one of 10 templates
        batch.append(pack(render(t.input, x), render(t.target, y)))
    loss = -log_prob(model, batch, targets_only=True)  # assumed; see text
    loss.backward(); adafactor.step()             # all 137B weights updated
# inputs up to 1024 tokens, targets up to 256; packed with an EOS separator

The paper and its released code do not say whether the loss also covers the input tokens of each packed pair. The schematic assumes the common setup for input/target pairs in a decoder-only model, which counts only the target tokens.

How FLAN was scored

Each dataset has up to ten templates, and FLAN's score can depend a lot on which one you use. The paper reports two numbers per dataset. The average template score is the mean over all templates, which estimates what a user with a typical phrasing would get. The best-dev template score takes the template that scores highest on a small dev split (usually 200 held-back training examples) and reports its test score, which matches the prompt engineering that GPT-3 results also allow.

The baselines are the untuned LaMDA-PT 137B, zero-shot and few-shot, prompted with GPT-3's prompts because FLAN-style instructions give near-zero scores on generation tasks for an untuned model; GPT-3 175B, using the numbers in its paper; and GLaM 64B/64E, a mixture-of-experts model, also from its paper. Comparing FLAN with LaMDA-PT isolates the effect of instruction tuning: same weights, with and without it. The LaMDA-PT prompts were not re-engineered for that model, and its few-shot runs use the largest kk in {1, 3, 5, 10} that fits in 1,024 tokens, where GPT-3 picked kk on a dev set and could use up to 100 examples in its longer context. Generation tasks are decoded greedily, where GPT-3 used beam search.

Where it helps and where it does not

Averaged over templates, zero-shot FLAN scored 56.2 on the five NLI datasets of the paper's Figure 1 (GPT-3: 42.9 zero-shot, 53.2 few-shot), 77.4 on three reading comprehension datasets (63.7 and 72.6), and 56.6 on four closed-book QA datasets (49.8 and 55.7). On each of these clusters the zero-shot tuned model beat GPT-3's few-shot average.

Figure 4 · zero-shot FLAN, dataset by dataset
Tables 1 and 2 of the paper for the 28 datasets that have a GPT-3 number. Each row joins the baseline to zero-shot FLAN; the joining line is teal where FLAN is higher. Switch the baseline and the FLAN template choice, and hover or tap a row for its numbers. † GPT-3's "zero-shot" used exemplars from the same passage; ‡ GPT-3 numbers not from the GPT-3 paper.

Against GPT-3 zero-shot with best-dev templates, FLAN is higher on 21 of the 28 rows, and on 20 of the 25 that remain once DROP, SQuADv2 and SST-2 are set aside, which is the count in the abstract. Against GPT-3 few-shot it is higher on 12. Against its own untuned base model, averaged over templates, it is higher on 23 of 28. The paper also reports best-dev FLAN above zero-shot GLaM on 13 of 19 datasets and above one-shot GLaM on 11 of 19.

The gains are largest where the task reads naturally as an instruction: NLI, reading comprehension, closed-book QA, and translation into English. Against the untuned base model, template-averaged, closed-book QA rises from 35.9 to 56.6 over its four datasets, reading comprehension from 60.9 to 77.4 over three, and NLI from 47.0 to 56.2 over five. FLAN leads GPT-3 zero-shot on all five NLI datasets, by 8.6 to 37.5 points with best-dev templates, and the paper suggests a reason: an NLI example phrased as a sentence continuation is unnatural text, while "Does <premise> mean that <hypothesis>?" is a normal question.

The losses against GPT-3 zero-shot are HellaSwag (56.7 against 78.9), ReCoRD (72.5 against 90.2), PIQA, WSC273, COPA (a tie at 91.0), and the two † datasets. Most are commonsense and coreference tasks phrased as finishing a sentence, which is already the pretraining objective, so an instruction adds little to the input. On the seven commonsense and coreference datasets, template-averaged FLAN beat untuned LaMDA-PT on four (COPA and PIQA by 0.6 points each, DPR and StoryCloze by more) and lost on HellaSwag, Winogrande and WSC273. On ReCoRD, instruction tuning cost 15 points (72.5 best-dev against LaMDA-PT's 87.8).

Translation is scored in BLEU, the n-gram overlap with reference translations on a 0-100 scale. Into English, FLAN scores 35.9 BLEU on French, 38.9 on German and 37.3 on Romanian with best-dev templates, beating GPT-3 zero-shot on all six directions but trailing GPT-3 few-shot on five of them. Out of English it is weaker (18.9 BLEU into Romanian), which the authors attribute to an English-centric tokenizer and a pretraining corpus that is about 90% English.

Why it works: clusters, scale, instructions

The paper runs three ablations to find which parts of the recipe matter. The first two use one split: NLI, closed-book QA and commonsense reasoning are held out, and up to seven other clusters are available for tuning.

The first ablation adds tuning clusters one at a time, largest first: summarization, translation, reading comprehension, sentiment, data-to-text, coreference, conversational QA. The average over the three held-out clusters rises from 49.9 with one cluster (11 datasets) to 63.5 with seven (39 datasets). Sentiment is the one addition that adds nothing (59.3 to 59.2), and the curve has not flattened at seven, so more clusters might help further.

The second instruction tunes the same split at 422M, 2B, 8B, 68B and 137B parameters, and each tuned model is compared with its untuned counterpart on 13 held-out tasks. At 68B and 137B instruction tuning adds 13 to 15 points. At 8B and below it subtracts about 5. The authors' hypothesis, which the paper does not test, is that the roughly 40 tuning tasks fill a small model's capacity, so it learns those tasks and loses general ability, while a large model has room to also learn how to follow instructions.

The third checks that the gain comes from the instructions and not from multi-task finetuning as such, by finetuning two models on the same data without them. One sees only inputs and outputs ("The dog runs." → "Le chien court."); the other sees the task and dataset name in front of each input ("[Translation: WMT'14 to French] The dog runs."). On four held-out clusters, averaged, FLAN scores 55.2, the no-template model 37.3, and the dataset-name model 46.6 with FLAN's instructions at test time or 47.0 with dataset names, 8 to 18 points below FLAN.

Figure 5 · the ablations
Four tabs. Number of clusters: the paper's Figure 6 average. Model size: its Figure 7, with instruction-tuned and untuned models (values measured off the plot). Role of instructions: Table 3. Prompt tuning: Table 4. The slider steps through the points, columns or tasks of the current tab.

In the model-size tab, step from 8B to 68B: the tuned model goes from 4.8 points below its base to 13.3 above. The paper reads this as instruction following emerging with scale. The data is five model sizes on one split, so where the crossover sits (somewhere between 8B and 68B for this mixture) is not pinned down.

A fourth ablation in Appendix B.1 varies what goes into each cluster. Using four datasets per cluster instead of one raised the held-out average from about 61 to about 70. Using 10 templates per dataset instead of 1 changed it by less than a point, a smaller effect than the authors expected, since the ten templates were meant to keep the model from overfitting to one phrasing.

Few-shot exemplars and prompt tuning

Instruction tuning also combines with the two other ways of adapting a model at inference. For few-shot use, the paper formats kk solved examples with the same template and concatenates them before the new input:

instruct(x1)⊕y1⊕instruct(x2)⊕y2⊕⋯⊕instruct(xk)⊕yk⊕instruct(x)\text{instruct}(x_1) \oplus y_1 \oplus \text{instruct}(x_2) \oplus y_2 \oplus \cdots \oplus \text{instruct}(x_k) \oplus y_k \oplus \text{instruct}(x)
(3)

Here ⊕\oplus is string concatenation with a delimiter token, kk is at most 16, and the prompt must stay under 960 tokens (the code reserves 64 of the 1,024 input tokens for separators). The same format is used when tuning FLAN's few-shot variant. Few-shot FLAN beat zero-shot FLAN on all seven evaluated clusters, by 10.2 points on struct-to-text (39.2 to 49.4), 4.6 on NLI, 3.5 on closed-book QA and 0.4 on reading comprehension. The paper attributes the larger gains to exemplars showing the output format. The spread across templates also shrank, so the few-shot model depends less on which instruction wording it gets.

Prompt tuning (Lester et al., 2021) freezes the model and learns a short sequence of input embeddings, a soft prompt, for one task by gradient descent; here the soft prompt is 10 embeddings long. The paper prompt-tunes both FLAN and LaMDA-PT on eight SuperGLUE tasks, each time using a FLAN model that saw no task from that dataset's cluster. With 32 training examples the average goes from 63.8 (LaMDA-PT) to 78.1 (FLAN); with the full training sets, from 79.2 to 87.4 (Table 4, averaged; the prompt-tuning tab of Figure 5 has the per-task numbers).

Limits and what came after

The authors list the limits themselves. Assigning datasets to clusters is partly a judgment call, and the held-out guarantee depends on it. The instructions are single sentences, much shorter than the instructions crowd workers get. The context is 1,024 tokens, too short for most summarization inputs, which is why summarization is trained on and never evaluated. The model is mostly English. It still fails some simple requests; the paper's Figure 22 shows it unable to return the second word of a sentence, and translating a question into Danish when asked to answer it in Danish.

Test contamination was checked after the fact. Following GPT-3, Appendix C splits each test set into examples that share a long n-gram (about 13 words) with the pretraining corpus and clean ones. Many datasets had heavy overlap, but the clean subsets did not score systematically lower than the full sets.

FLAN appeared on arXiv in September 2021. T0 (Sanh et al.), posted a month later, ran a similar experiment on an 11B T5 model with crowd-sourced prompts and also found zero-shot gains on held-out task types. InstructGPT (Ouyang et al., 2022; see the InstructGPT explainer) trained on prompts and demonstrations written by people and added reinforcement learning from human preference rankings. Chung et al. (2022) scaled the FLAN recipe to about 1.8K tasks, added chain-of-thought data, and released the Flan-T5 checkpoints, which is why "Flan" now usually refers to those models. Supervised finetuning on instruction data (usually called SFT) is now a standard stage after pretraining, and the same idea was carried to images in LLaVA.

Provenance Verified against primary literatureHow we verify
Wei et al. (2021), Sections 2.1-2.4 and Figure 3Method as taught here: 62 datasets in 12 clusters, 10 templates per dataset with up to 3 reversed, the cluster hold-out with the footnote-1 overlap rules, the OPTIONS suffix, and the training settings (30k-example cap, mixing cap 3k, 30k steps of 8,192 tokens, Adafactor at 3e-5, lengths 1024/256, packing).
google-research/FLAN, flan/mixtures.py line 25 and tasks.py line 66mixing_rate_3k = seqio.mixing_rate_num_examples(maximum=3000) and NUM_TRAIN_EXAMPLES = 30000. The rate is min(examples, 3000), the capped examples-proportional scheme of Raffel et al. (2020), Section 3.5.2.
flan/preprocessors.py lines 146-152 and 376-428; tasks.py lines 2295-2371format_options writes "OPTIONS:\n- a\n- b". Classification tasks are scored with rank_classification_from_options: the prompt (with the OPTIONS list) is paired with every option string and the highest-scoring option is the prediction. The suffix is added to rank classification, and scoring stays rank classification.
flan/preprocessors.py lines 299-324Templates are assigned round-robin inside each group of 10 consecutive examples, not drawn at random per example as Section 2.1 says. Both give each template a tenth of a dataset's examples.
flan/task_splits.py, cluster tableThe code lists 17 finer clusters (it splits out conversational QA, WiC, CoLA, TREC, math, and text formatting with True Case and Word Segment). The paper merges the small ones into one Misc. cluster of 7 datasets.
Tables 1 and 2, recounted28 datasets have a GPT-3 zero-shot number. Best-dev FLAN is higher on 21. Removing DROP and SQuADv2 (marked † as not comparable) and SST-2 gives 20 of 25, the abstract's count.
Figure 7 of the paperThe paper prints no numbers for the scale ablation. The values in this page's figure were measured from the plot's pixels and are good to about half a point.
Sanh et al. (2021); Chung et al. (2022); Ouyang et al. (2022)T0 (11B T5, arXiv October 2021) is the concurrent work. Flan-T5 and Flan-PaLM come from the 2022 follow-up (1.8K tasks). InstructGPT adds human-preference RL on top of supervised finetuning.
correctionFigure 10 of the paper swaps two bars. Averaging Table 4 over the eight SuperGLUE tasks gives 63.8 for LaMDA-PT and 78.1 for FLAN with 32 prompt-tuning examples, and 79.2 for LaMDA-PT and 87.4 for FLAN with the full training set; the figure prints 79.1 for FLAN at 32 examples and 78.1 for LaMDA-PT on the full set. With the table's numbers, FLAN tuned on 32 examples lands 1.1 points below the untuned model tuned on the full training set, where the figure shows it 1.0 point above. Separately, Section 3 says best-dev FLAN beats few-shot GPT-3 on 10 datasets; Tables 1 and 2 give 12 (ANLI R1-R3, CB, RTE, BoolQ, MultiRC, OBQA, ARC-e, ARC-c, StoryCloze, and WMT'14 En-Fr), and says FLAN beats untuned LaMDA-PT on 3 of the 7 commonsense and coreference tasks, where Table 2 gives 4 with template-averaged scores (COPA, PIQA, StoryCloze, DPR) and 5 with best-dev templates.

Questions you might still have

?

Is FLAN the same model as Flan-T5?
No. FLAN is this 2021 paper's 137B LaMDA-PT model, which Google did not release. Flan-T5 and Flan-PaLM come from a 2022 follow-up (Chung et al., Scaling Instruction-Finetuned Language Models) that reused the name and the recipe with about 1.8K tasks, added chain-of-thought data, and released the T5 checkpoints publicly. The repository for this paper releases the data pipeline and templates only.

?

How is instruction tuning different from InstructGPT's RLHF?
FLAN is supervised finetuning only: the target text of existing NLP datasets is the label. InstructGPT (covered in the InstructGPT explainer on this site) starts with supervised finetuning on prompts and answers written by people, then trains a reward model on human rankings of outputs and optimizes the policy against it with PPO. FLAN has no human preference signal and no reinforcement learning.

?

FLAN trained on 62 labeled datasets. In what sense is it zero-shot?
Zero-shot here means the model gets no examples of the evaluation task at test time and saw no dataset of the same task type during tuning. It has seen plenty of labeled data for other task types. The claim is about transfer across task types, which is why the paper holds out whole clusters instead of single datasets.

?

Why did instruction tuning make the 8B and smaller models worse?
The paper offers a hypothesis, not a tested mechanism: a small model spends its capacity learning the roughly 40 tuning tasks and has none left for following instructions in general, while a larger model learns both. The ablation used one split (NLI, closed-book QA and commonsense held out), and the authors note that cross-validating over splits would have been better. The 2022 Flan-T5 work found instruction tuning helpful for much smaller models with far more tasks, so the crossover depends on the mixture as well as the size.

?

Why does instruction tuning not help commonsense and coreference tasks?
Those benchmarks are phrased as completing a sentence, which is already what the pretrained model does, so an instruction adds little information to the input. The paper says FLAN beat the untuned LaMDA-PT on only 3 of the 7 commonsense and coreference tasks; its own Table 2 gives 4 of 7 with template-averaged scores, two of them by 0.6 points. On ReCoRD, a cloze-style reading task, FLAN scored 15 points below the untuned model.

?

Could FLAN have seen the test sets during pretraining?
Partly, yes. Following GPT-3's procedure, Appendix C marks a test example as dirty if any of its n-grams (n around 13) appears in the pretraining corpus, and many datasets had substantial overlap. Scoring only the clean examples did not systematically lower accuracy, so there is no sign that the overlap inflated the results, but datasets with few clean examples give noisy comparisons.

?

Do more templates per dataset help?
Barely. Appendix B.1 tuned with 1, 4 or 10 templates per dataset. With one dataset per cluster the held-out average went 61.0, 61.6, 61.9; with four datasets per cluster it went 70.1, 70.5, 69.9. Adding datasets per cluster moved the average by about 9 points; adding templates moved it by less than 1.

Footnotes & further reading

  1. The paper: Wei, Bosma, Zhao, Guu, Yu, Lester, Du, Dai & Le, Finetuned Language Models Are Zero-Shot Learners (ICLR 2022; arXiv September 2021). Data and templates: github.com/google-research/FLAN.
  2. Prompting, rank classification and the contamination procedure: Brown et al., Language Models are Few-Shot Learners (2020); see the GPT-3 explainer.
  3. Examples-proportional mixing and packing: Raffel et al., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (2020), Section 3.5.2; see the T5 explainer.
  4. Prompt tuning: Lester, Al-Rfou & Constant, The Power of Scale for Parameter-Efficient Prompt Tuning (2021).
  5. T0: Sanh et al., Multitask Prompted Training Enables Zero-Shot Task Generalization (2021).
  6. Flan-T5 and Flan-PaLM: Chung et al., Scaling Instruction-Finetuned Language Models (2022).
  7. InstructGPT: Ouyang et al., Training language models to follow instructions with human feedback (2022); see the InstructGPT explainer.