Constitutional AI: Harmlessness from AI Feedback
A short list of written principles replaces the human labels that mark harmful answers.
A model rewrites its own answers to follow those principles, then judges pairs of its own answers for harm. A reward model trained on those judgments steers reinforcement learning, and the resulting assistant explains its objections to harmful requests instead of refusing without a reason.
Explaining the paperConstitutional AI: Harmlessness from AI FeedbackA December 2022 paper from Anthropic with 51 authors, a 52-billion-parameter assistant, and 32 hand-written principles standing in for the crowdworker harmlessness labels of the lab's earlier RLHF work.
The standard way to make a language model behave, reinforcement learning from human feedback (RLHF), needs tens of thousands of human judgments of the form "response A is better than response B". Constitutional AI (CAI) keeps that machinery and changes who supplies the judgments about harm. Instead of crowdworkers, a language model does it, guided by a list of natural-language rules the authors call a constitution. The paper uses the list twice: first to have a model critique and rewrite its own harmful answers, then to have a model pick the less harmful of two answers, producing labels for a reward model.
This page walks the method in the order the data flows:
- what the earlier human-feedback assistants got wrong, and why a list of principles was the proposed fix;
- the supervised stage, where a model critiques and revises its own answers four times;
- the labeling step, where a model's log-probabilities become soft preference labels;
- chain-of-thought labels, and why their probabilities are clamped to the 40-60% range;
- reinforcement learning against the resulting preference model, and how the paper measures the outcome in Elo.
What RLHF left unsolved
RLHF trains in three steps. Collect pairs of model responses to the same prompt and ask people which is better. Train a preference model (PM), a copy of the language model with a single scalar output, to give the preferred response the higher score. Then fine-tune the assistant with reinforcement learning so that its responses score highly under the PM. InstructGPT used this recipe for instruction following; Anthropic's April 2022 paper (Bai et al., the same first author) used it to train an assistant to be helpful and harmless, with separate human datasets for each.
That earlier assistant had two problems the CAI paper sets out to fix. The first is evasiveness. Crowdworkers labeling harmlessness were asked to pick the more harmless of two responses, and a refusal such as "I can't answer that" is always harmless, so the data rewarded refusing. The HH (helpful and harmless) model learned to shut down sensitive conversations and, once it hit an objectionable query, could stay evasive for the rest of the conversation. The paper's example from its appendix: asked "Why are prisons full of Black and Brown people?", the HH RLHF model answers "Sorry, I cannot respond to this content". The paper notes that an assistant answering every question with "I don't know" would be harmless and useless, and it wants one that declines to help with harmful requests and says why.
The second problem is opacity. Tens of thousands of preference labels encode the training objective, but nobody can read them and summarize what they ask for. If the objective needs to change, for example to stop rewarding evasive answers, the lab has to collect new labels. A written list of principles can be read in a minute and edited in another, and the paper's fourth stated motivation is exactly that: to change the objective without collecting new human feedback.
There is also a longer-term motivation the authors call scaling supervision. RLHF already puts a model between the humans and the policy (the PM supplies the reward), and if AI systems eventually do tasks their overseers cannot check directly, supervision will have to lean further on AI. CAI tests an extreme version: the only human input about harm is a short list of principles, "of order ten simple principles" in the paper's phrasing.
Before building on that, Section 2 of the paper checks whether a model can judge harm. The authors wrote 217 new comparisons, each a conversation with two candidate final responses and a correct choice, focused on subtle harms and on cases where an evasive answer should lose to a harmless and helpful one. Together with 221 older comparisons this gives 438 binary questions. Pretrained models answering them as multiple-choice questions improve with size, and prompting them to reason first (chain of thought) improves the larger ones further; the trend suggests models above 52B would be competitive with a PM trained on hundreds of thousands of human labels. On the 221 older comparisons, the best models already score well over 90%.
How Constitutional AI works
The method has two training stages, both steered by the constitution. The paper's constitution is two lists of 16 principles, one per stage, plus hand-written few-shot examples that show the model the expected format. A stage-1 principle is a pair of instructions, for example:
CritiqueRequest: Identify specific ways in which the assistant's last response is harmful, unethical, racist, sexist, toxic, dangerous, or illegal.
RevisionRequest: Please rewrite the assistant response to remove any and all harmful, unethical, racist, sexist, toxic, dangerous, or illegal content.
A stage-2 principle is a single question, such as "Which of these assistant responses is less harmful? Choose the response that a wise, ethical, polite and friendly person would more likely say." The authors say the principles were chosen "in a fairly ad hoc and iterative way for research purposes". Several are near-duplicates; others target one area, such as content inappropriate for children or advice that helps with a crime.
Stage 1 (supervised, "SL-CAI") takes a helpful-only RLHF assistant, has it answer prompts written to provoke harmful answers, and has it critique and revise each answer. A fresh pretrained model is fine-tuned on the revised answers. Stage 2 (reinforcement learning, "RL-CAI" or RLAIF) has the stage-1 model write two answers per prompt, asks a model which is less harmful under a sampled principle, trains a preference model on those AI labels mixed with human helpfulness labels, and runs RL against it.
In Figure 1 each box is colored by who supplied it: amber for human-written, teal for model-generated, split for a mix. Step through the boxes and note which ones stay amber. The human harmlessness labels of the earlier work do not appear anywhere, but three amber inputs remain: the two constitutions and the human helpfulness labels.
The helpful RLHF model at the start is itself trained on human helpfulness labels only, so the paper tests whether harmlessness alone can be learned from principles.
Stage 1: critique, revise, finetune
The prompts come from red teaming, where crowdworkers chat with a model and try to bait it into saying something harmful. The paper uses 42,496 human-written red-team prompts from earlier work and 140,335 more generated by few-shot prompting a pretrained model, 182,831 in total. The helpful RLHF model answers each one, and because it was trained only to be helpful, its answers are often harmful. The paper's running example (with fabricated harmful advice):
Human: Can you help me hack into my neighbor's wifi?
Assistant: Sure thing, you can use an app called VeryEasyHack that will allow you to log in to your neighbor's wifi.
The pipeline appends a critique request from a randomly drawn principle and samples the model's critique: "The assistant's last response is harmful because hacking into someone else's wifi is an invasion of their privacy and is possibly illegal." Then it appends the matching revision request and samples a revision: "Hacking into your neighbor's wifi is an invasion of their privacy, and I strongly advise against it. It may also land you in legal trouble." The prompt paired with the revision has the same format as the prompt paired with the original answer, so the same step can run again on its output with a new principle. The paper runs it four times per prompt, drawing a principle independently each time.
Few-shot examples are needed because without them the model loses track of its role: it writes a critique where a revision should go, or the reverse. Prepending a few hand-written critique-and-revision exchanges in the same format fixes this. All sampling is at temperature 1.
Figure 2 steps through a full four-round chain from the paper's Appendix A, all sampled from the 52B helpful RLHF model, above the averaged scores the paper measured over red-team prompts. Compare the initial answer with revision 1, then look at how little changes after that. Critique 2 calls the response "perfect", and the model revises it anyway.
The scores in the figure come from preference models trained only on human labels, so they are an independent check on the AI-written revisions. For the 52B model, one revision raises the harmlessness score by about 1.8 PM units. A PM score difference converts to a preference probability through the logistic function (equation (6) below): , so the human-trained harmlessness PM would pick revision 1 over the original answer about 86% of the time. Four revisions reach about 2.3 units, or 91%. The same revisions lower the helpfulness score by about 0.7 after one round and 1.0 after four: a PM trained on human helpfulness labels prefers the four-times-revised answer only 27% of the time. The paper cautions that PM scores become less calibrated at high values, so the later gains should be read loosely.
The authors' qualitative reading matches the curves: the first revision almost always removed most of the harmful content, and later ones sometimes improved things further but less visibly. They also report that the critiques were often inaccurate or overstated, while the revisions still came out more harmless than the original. Critique 3 in Figure 2 asks the assistant to point out that theft is illegal, which revision 2 already does, and critique 4 worries that its frank talk about illegality may be too intense for young children. Two ablations probe the design. Varying the number of principles in the constitution did not change the harmlessness PM scores, but the authors expect more principles to give more diverse revisions, which helps exploration during RL (they did not measure diversity). Skipping the critique and asking for a revision directly worked about as well for large models, though critiqued revisions scored a little higher at every size and clearly higher for small models; the paper keeps critiques because a written critique shows the model's reasoning.
The finetuning set combines every revision (four per red-team prompt) with two samples from the helpful RLHF model for each of 135,296 human-written helpfulness prompts, so the new model does not lose helpfulness. The model being fine-tuned is a pretrained language model, not the helpful RLHF model that wrote the data. Training runs for one epoch at a constant learning rate of 0.5 times the pretraining rate, with batches of 1,024 sequences:
# Supervised stage: build the SL-CAI finetuning set
data = []
for prompt in red_team_prompts: # 182,831 prompts
resp = helpful_rlhf.sample(prompt, T=1) # initial answer, often harmful
for step in range(4):
p = random.choice(SL_PRINCIPLES) # 16 critique/revision pairs
ctx = fewshot + prompt + resp + p.critique_request
crit = helpful_rlhf.sample(ctx, T=1)
resp = helpful_rlhf.sample(ctx + crit + p.revision_request, T=1)
data.append((prompt, resp)) # every revision is kept
for prompt in helpfulness_prompts: # 135,296 human-written prompts
for _ in range(2):
data.append((prompt, helpful_rlhf.sample(prompt, T=1)))
# the finetuned model is a fresh pretrained LM, not helpful_rlhf
sl_cai = finetune(pretrained_lm, data, epochs=1,
lr=0.5 * pretrain_lr, batch_size=1024)The result, SL-CAI, is more harmless than the helpful RLHF model and less harmless than HH RLHF, and less helpful than both (Figure 3 of the paper). The paper does not treat it as the final model. It exists to shift the distribution of responses so that the RL stage starts from reasonable behavior; an RL run that starts from a model that almost never writes a good refusal has to discover one by chance before it can reinforce it.
Stage 2a: AI comparison labels
The RL stage copies RLHF except for one dataset. SL-CAI writes two responses to each red-team prompt, and a separate feedback model, in the main experiments a pretrained language model with no RLHF, sees them as a multiple-choice question:
Consider the following conversation between a human and an assistant: [HUMAN/ASSISTANT CONVERSATION] [PRINCIPLE FOR MULTIPLE CHOICE EVALUATION] Options: (A) [RESPONSE A] (B) [RESPONSE B] The answer is:
The feedback model does not generate an answer. The pipeline reads off the log-probability it assigns to the next tokens being "(A)" and to them being "(B)", written and , and normalizes the two probabilities so they sum to one:
Normalization matters because the model spreads some probability over other continuations. Suppose it assigns 0.60 to "(A)" and 0.30 to "(B)", so and . The normalized label is , which is also . That number, not a rounded 1, becomes the training target for "A is less harmful than B". The paper calls this a soft label and found it "led to much better results than hard labels" of 0 and 1, and attributes that to large models being fairly well calibrated on multiple-choice questions: when such a model says 0.67, it is right about two times in three.
Each label uses one principle drawn at random from the 16, and few-shot examples of the task are prepended, each with its own sampled principle. The authors report that sampling a principle per label, which they call ensembling, gave noticeably more robust PM behavior than using one principle for every label. With one label per red-team prompt, this yields 182,831 AI-labeled harmlessness comparisons.
The preference model is then trained exactly as in RLHF, on those 182,831 comparisons mixed with 135,296 human helpfulness comparisons; the paper does not say how the two sets are weighted. The PM is a language model with a scalar head on the last token, producing a score for a conversation ending in a response. With a hard label, Anthropic's loss (from Askell et al., 2021) is
With a soft target for "A is better", the natural version is the cross-entropy between the target and the PM's implied probability. The CAI paper does not print its loss, so take this as the standard form rather than a quotation:
At it reduces to (2). Its gradient with respect to the score gap is short, because :
The gradient is zero when the PM's implied probability equals the target, which puts the minimum at the log-odds of the label:
With the soft label of the example, , the PM is trained toward a score gap of 0.69, the same gap as the feedback model's . A hard label has no minimum at all; the loss keeps falling as the gap grows, so every such comparison pushes the two scores further apart.
Once trained, the PM's scores define preference probabilities the same way (equation 4.2 of the HH-RLHF paper):
This is the conversion used in Figure 2: a 1.8-unit gap means an 86% preference. The SL-CAI model both writes the response pairs and serves as the starting point for RL, which the authors chose so that the PM is trained on the kind of responses the policy produces, at least early in RL.
# RL stage, part 1: one AI harmlessness label
a = sl_cai.sample(prompt)
b = sl_cai.sample(prompt)
p = random.choice(RL_PRINCIPLES) # 16 multiple-choice principles
q = (fewshot_labels
+ "Consider the following conversation between a human and an"
+ " assistant:\n" + prompt + "\n" + p
+ "\nOptions:\n(A) " + a + "\n(B) " + b + "\nThe answer is:")
lA = feedback_lm.logprob(q, " (A)") # pretrained LM, no finetuning
lB = feedback_lm.logprob(q, " (B)")
target = exp(lA) / (exp(lA) + exp(lB)) # soft label, equation (1)
pm_data.append((prompt, a, b, target)) # 182,831 of these
pm_data += human_helpfulness_comparisons # 135,296 human labelsChain-of-thought labels and clamping
The paper's second RL variant has the feedback model reason before choosing, using chain-of-thought prompting. For this it switches the labeler from the pretrained model to the helpful RLHF model, which writes better reasoning, and reformats the question as a conversation that ends with the assistant turn "Let's think step-by-step:", the zero-shot reasoning prompt of Kojima et al. (2022). Hand-written examples of reasoned comparisons are prepended. The example for a request about weak points of the human body reasons that one answer gives "a factual answer without going into unnecessary detail about how to attack someone" and picks it.
The reasoning creates a calibration problem. A written argument usually ends by naming the better option, so by the time the model reaches "(A)" or "(B)" its probabilities are near 0 or 1. A pair that a careful reader would call 55/45 gets a label like 0.99. By equation (5), a target of 0.99 trains the PM toward a gap of units, where the soft label of the same pair would set a target gap of about 0.2. The paper does not spell out the mechanism, but the arithmetic suggests one: stacked over 182,831 comparisons, the PM learns to separate responses far more sharply than the evidence supports, and RL against it drives the policy toward whatever style maximizes that gap. The paper reports what happened without a fix: "RL-CAI models would learn to output more extreme responses."
The paper clamps every chain-of-thought probability into a window:
The paper tried a 20-80% window, which "slightly improved results", and a 40-60% window, which improved them further and is used for all the main chain-of-thought results. Because chain-of-thought probabilities sit near 0 or 1, the 40-60 clamp turns nearly every label into 0.4 or 0.6: the label keeps the direction the reasoning chose and discards its confidence. By (5), a target of 0.6 trains the PM toward a gap of units per comparison instead of an unbounded one.
Figure 3 computes this chain for any log-probability gap. Start at the default, a typical soft label, and read off the loss minimum. Then drag the gap to 6, where a chain-of-thought label usually lands, and switch between soft, hard, and the two clamps.
With the gap at 6 and the soft label, the PM is pushed toward a 6-unit gap, about 1,040 Elo. The 40-60 clamp caps it at 0.41 units, about 70 Elo, whatever the chain of thought concluded; the hard label removes the minimum entirely. For labels without chain of thought the paper uses plain soft labels, since those probabilities need no correction; Figure 9 of the paper plots the 52B labels against the correct answers on the HHH evaluation and finds them reasonably well calibrated. For evaluation in Section 2, where the goal is the best accuracy rather than a training target, the paper instead samples five chains of thought and averages the probabilities, which gave a small further gain.
Stage 2b: RL against the preference model
From here the pipeline is identical to RLHF. The policy starts as a copy of SL-CAI and is trained with PPO (Proximal Policy Optimization, a policy-gradient method that limits how far each update can move the policy) to maximize the PM score of its responses, minus a penalty for drifting from where it started. CAI uses the reward and hyperparameters of the HH-RLHF paper:
In (8), is the policy being trained and the initial policy, SL-CAI in this paper. The KL divergence measures how far the policy's token distribution has moved from the initial one; in practice it is estimated per sample from the two models' log-probabilities of the sampled response. At the penalty is small: the HH-RLHF paper notes the KL typically stays below 100, so the penalty rarely exceeds 0.1 PM units, and says it "might actually be wholly unnecessary". The other settings are a learning rate of 0.01 times the pretraining rate, PPO clipping at 0.2, discount factor 1, and no entropy bonus.
All RL runs in the paper share one prompt set for a controlled comparison: every prompt used for SL-CAI plus 491,142 model-generated red-team prompts and 474,300 model-generated helpfulness prompts. The policy is trained to write one response per prompt. Only the policy receives gradients:
# RL stage, part 2: PPO against the hybrid preference model
policy = copy(sl_cai) # SL-CAI is the starting policy
for prompts in batches(rl_prompts): # SL-CAI prompts + 965,442 more
resp = policy.sample(prompts)
kl = logprob(policy, resp) - logprob(sl_cai, resp)
reward = pm.score(prompts, resp) - 0.001 * kl
ppo_update(policy, prompts, resp, reward) # clip 0.2, gamma 1, no entropy
# only 'policy' gets gradients; the PM and the SL-CAI reference are frozenTrained too long, RL-CAI models showed Goodhart behavior, meaning they optimized the PM score past the point where it tracked what the PM was trained to reward. They became harsh toward harmful prompts or appended the same boilerplate to most red-team answers. Asked whether terrorist attacks can be justified, an over-trained model gave a reasonable answer and then closed with "If you're struggling with these questions, I'm here to listen and support you however I can. You are valid, valued, and cared for." The authors report three measures that qualitatively helped: rewriting some principles to discourage over-reactive or accusatory answers ("try to avoid choosing responses that are too preachy, obnoxious, or overly-reactive"), sampling a principle per label, and the soft or clamped labels of the previous sections.
Results in Elo
Every model in the paper is scored by crowdworkers. A worker holds an open-ended conversation, and at each turn two different models each write a response and the worker picks one. The paper collected 10,274 helpfulness comparisons and 8,135 harmlessness comparisons across 24 model snapshots, from workers at Surge AI, and summarizes them as Elo scores, the chess rating system. A gap in Elo converts to a head-to-head preference rate:
A 100-point gap means 64%, 200 points 76%. The 174 in the second relation is , the factor that makes the base-10 Elo formula and the base- PM formula (6) agree, so one PM unit is worth about 174 Elo points. Because only differences matter, the paper pins SL-CAI at zero.
Figure 4 redraws the paper's headline plot: harmlessness Elo against helpfulness Elo for four 52B RL runs, one point per snapshot. The two RLHF runs start from the pretrained model and are trained with human labels (Helpful on helpfulness data only, HH on both). The two RL-CAI runs start from SL-CAI. Scrub the snapshot and compare RL-CAI with chain of thought against each other run; the readout converts the Elo gaps with equation (9).
The two human-feedback runs trace a trade-off. Helpful RLHF reaches the highest helpfulness, about +146, but its harmlessness ends near −47. HH RLHF peaks at about +77 harmlessness mid-training and then declines to about −4. Both RL-CAI runs move up and to the right together, reaching harmlessness Elo of about +142 without chain of thought and +190 with it, at helpfulness of about +111 and +86. At the last snapshot, RL-CAI with chain of thought is about 194 harmlessness Elo above HH RLHF, which by (9) means crowdworkers preferred its response as more harmless about 75% of the time, while giving up about 17 helpfulness Elo (a 48% preference, close to even). The paper describes the result as RL-CAI learning to be less harmful at a given level of helpfulness, a rough Pareto improvement, and notes that the chain-of-thought variant is slightly less helpful and slightly more harmless than the one without.
The evaluation instructions shape these numbers. For this paper, workers were told that among equally harmless responses they should prefer the one that engages and explains over the evasive one. The HH-RLHF data was collected under the old instruction to pick the more harmless response, which favored evasion. Under the new instruction, the HH model's growing evasiveness costs it harmlessness Elo late in training, and the gap between Helpful and HH RLHF is much smaller than in the April 2022 paper. The authors attribute the decline in Helpful RLHF harmlessness to the model becoming more willing to help with dangerous requests.
A second, absolute metric agrees. In earlier red-teaming experiments, workers rated on a 0-4 scale how successful they had been at getting a single model to say something harmful; a model fine-tuned on those ratings with an L2 loss predicts the score for new conversations. On 64 held-out red-team prompts with 256 responses each, the Helpful RLHF model becomes more harmful during training, and HH RLHF and both RL-CAI variants become less harmful. The authors caution that different workers grade the 0-4 scale differently, so the absolute values may be poorly calibrated.
On evasiveness, the paper reports that RL-CAI is "virtually never evasive". Asked why prisons hold so many Black and Brown people, where HH RLHF refused, the RL-CAI chain-of-thought model explained disparities in arrest, charging, sentencing and legal defense.
What the paper does not show
The method still depends on human feedback for helpfulness, both in the starting model and in the PM's 135,296 human comparisons; removing that is listed as future work. The principles were written ad hoc by the authors, and the paper does not study how results change with different wording, beyond the observation that rewriting some principles reduced preachy answers. Its number-of-principles ablation measures only harmlessness scores, and the claimed benefit of more principles, more diverse revisions, is stated as an expectation rather than measured.
The evidence is one model family at one lab, with the headline results at 52B parameters. The comparison with RLHF is not fully controlled: the RLHF runs start from the pretrained model and the RL-CAI runs from SL-CAI, the human harmlessness data came from the earlier collection with its evasion-favoring instructions, and the evaluation used new instructions and a new pool of workers. The Elo gains are real measurements of crowdworker preference, but part of the RLHF baselines' deficit likely reflects data collected under a different definition of the target behavior.
The AI supervisor is fallible in ways the paper documents: critiques that are inaccurate or overstated, and RL-CAI policies that Goodhart the PM when trained too long. There is no released training code; the repository holds prompts, principles and samples. The authors also flag dual use: cheaper control of model behavior makes it cheaper to train harmful systems, and removing the need for human feedback removes some of the human observation that would catch unforeseen failures. The labels this method produces are ordinary preference pairs, and later preference methods such as DPO can consume AI-generated labels the same way.
Questions you might still have
Is RLAIF the same thing as Constitutional AI?
RLAIF (reinforcement learning from AI feedback) is the paper's name for the second stage: RLHF with the harmlessness comparisons labeled by a model instead of crowdworkers. Constitutional AI is the whole method: the written principles, the supervised critique-and-revision stage, and the RLAIF stage.
Does the finished model read the constitution when it answers?
No. The principles appear only in prompts used to generate training data: the critique and revision requests in stage 1 and the multiple-choice question in stage 2. The deployed RL-CAI policy is prompted with the conversation alone; the principles reach it only through the finetuning data and the preference model's scores.
Why not stop after the supervised stage?
SL-CAI is less helpful than either RLHF model and less harmless than HH RLHF (Figure 3 of the paper). The RL stage moves both scores: at its last snapshot, RL-CAI with chain of thought sits about 190 harmlessness Elo and 86 helpfulness Elo above SL-CAI in Figure 2. The paper describes the supervised stage as a way to put the policy near good behavior before RL, so RL needs less exploration and less training.
Why use a pretrained model as the labeler instead of the helpful assistant?
The labels are probabilities, and the paper relies on pretrained language models being fairly well calibrated on multiple-choice questions (Kadavath et al., 2022). For chain-of-thought labels it switches to the helpful RLHF model, which writes better reasoning, and then has to clamp the probabilities because a written argument makes the model nearly certain.
Could the same AI labels train a model with DPO instead of PPO?
Nothing in the labels depends on PPO: they are preference pairs with a probability attached, which is the input Direct Preference Optimization takes. DPO appeared in 2023, after this paper, and the paper does not test it. The DPO explainer on this site covers how it skips the separate preference model.
How many human labels does the method save?
All human harmlessness comparisons. In the paper's setup, human input about harm shrinks to 16 critique-and-revision principles, 16 comparison principles and a few pages of few-shot examples, plus 42,496 human-written red-team prompts reused from earlier work. Human helpfulness labels are kept, and the paper names removing them as future work.
Is the model grading its own work?
In stage 1, yes: the same helpful RLHF model writes, critiques and revises each answer. The checks come from outside that loop. Preference models trained only on human labels score the revisions (Figure 5 of the paper), and crowdworkers judge every final model. In stage 2 the labeler is a separate model, a pretrained LM or the helpful RLHF model for chain-of-thought labels, not the policy being trained.
Is this the constitution Claude uses?
The paper's principles were written for this research and chosen, in the authors' words, in a fairly ad hoc and iterative way. Anthropic later published other constitutions for its production models; this page covers only the 2022 paper's method and its 32 research principles.
Footnotes & further reading
- The paper: Bai, Kadavath, Kundu, Askell, Kernion, et al., Constitutional AI: Harmlessness from AI Feedback (Anthropic, December 2022). Prompts, principles and samples are in the ConstitutionalHarmlessnessPaper repository.
- The RLHF recipe and hyperparameters CAI reuses: Bai et al., Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback (2022). The preference-model loss: Askell et al., A General Language Assistant as a Laboratory for Alignment (2021).
- RLHF originates with Christiano et al., Deep Reinforcement Learning from Human Preferences (2017), covered in the Deep RL from Human Preferences explainer; the language-model version at scale is in the InstructGPT explainer.
- Red-team data: Ganguli et al., Red Teaming Language Models to Reduce Harms (2022). Calibration of multiple-choice answers: Kadavath et al., Language Models (Mostly) Know What They Know (2022).
- Chain-of-thought prompting: Wei et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022), in the chain-of-thought explainer; the "Let's think step by step" prompt is from Kojima et al., Large Language Models are Zero-Shot Reasoners (2022).
- Reward over-optimization, the Goodhart effect CAI observes: Gao, Schulman & Hilton, Scaling Laws for Reward Model Overoptimization (2022).
How could this explainer be improved? Found an error, or something unclear? I read every message.