VerifiedarXiv:1706.0374138 min
Reinforcement learning · Alignment

Deep Reinforcement Learning from Human Preferences

An agent learns a task from a person picking the better of two short clips of its behavior, with no hand-written reward.

A network learns to predict which clip the person will pick, and the agent is trained to maximize that network's output. About 900 comparisons taught a simulated robot to backflip.

Explaining the paperDeep Reinforcement Learning from Human PreferencesChristiano, Leike, Brown, Martic, Legg, Amodei · OpenAI · DeepMind · NeurIPS 2017 · arXiv:1706.03741 ↗

On the Atari runs a contractor watches two 1.7-second clips of the agent side by side and clicks the one that plays better, which takes 3 to 5 seconds. Over a 50-million-step training run that happens about 5,500 times.

Reinforcement learning trains an agent from a reward: a number the environment returns after every action, which the agent learns to make large over time. A game supplies one in its score and a physics simulator in the distance walked. Most tasks people want done supply nothing. The paper's examples are a robot that cleans a table or scrambles an egg. A reward for either would have to be computed from the robot's sensor readings, and a simple formula that approximates the goal tends to produce behavior that scores well on the formula without doing what the person wanted.

Christiano, Leike, Brown, Martic, Legg and Amodei, from OpenAI and DeepMind, replace the hand-written reward with a learned one. A person watches pairs of short clips of the agent and says which is better. A neural network, the reward model, is trained to predict those answers. The agent is trained by ordinary RL to maximize the reward model's output, and all three run at the same time, so the person keeps judging the agent's newest behavior. The method is the origin of RLHF (reinforcement learning from human feedback), and its reward model reappears, scoring text instead of clips, in InstructGPT.

The agent never sees the true reward. On eight simulated robotics tasks, 700 human comparisons nearly match RL trained on the true reward. On seven Atari games, 5,500 comparisons give substantial learning on most games and match or beat RL on a few. The abstract states the human's share as feedback on less than 1% of the agent's interactions with the environment.

The sections below take the method in order: why the person compares instead of scoring; how the Bradley-Terry model turns comparisons into a reward for every timestep; what one comparison does to the reward model; what comparisons cannot determine; how the experiments keep the true reward hidden; how the three processes share a run; why the policy optimizer and the online labeling were chosen; how pairs are picked for the human; and what the experiments measured.

Goals you can recognize but cannot write down

At each timestep the agent receives an observation oto_t (an Atari screen, or a robot's joint angles and velocities) and chooses an action ata_t. A policy π\pi maps observations to actions. In standard RL the environment also returns a reward rtr_t, and training adjusts π\pi to maximize the discounted sum of future rewards. Everything the agent learns about the goal arrives through that one number per step.

The paper sets a narrower target than writing better rewards: solve tasks where a person can only recognize the desired behavior, not necessarily demonstrate it. Demonstrations would allow inverse reinforcement learning (recover a reward that explains the demonstrations) or imitation learning (copy them). Both need someone who can perform the task, which rules them out for robots with very non-human bodies. For the simulated Hopper, a one-legged robot, nobody has written a formula whose maximum is a backflip that lands upright, and nobody can demonstrate one. A person watching the Hopper can still say which of two clips looks more like a backflip.

Using that judgment directly as the reward would mean a person scoring every timestep, and deep RL agents need hundreds or thousands of hours of experience. The introduction states the requirement: the amount of feedback has to fall by several orders of magnitude. The paper's approach, learning a reward model from a small number of judgments and optimizing it, had been tried before (Wilson et al. 2012; Akrour et al. 2012 and 2014), on low-dimensional problems or with rewards linear in hand-coded features. This paper runs it with deep networks on raw pixels and on robot locomotion.

Why ask for comparisons instead of scores

The person is shown two clips of the agent, each 1 to 2 seconds long, and answers one of four ways: the left clip is better, the right clip is better, they are equally good, or they cannot be compared.

The authors found it much easier for people to give consistent comparisons than consistent absolute scores, especially on the robot tasks and on the new behaviors with no reward. A rater scoring a half-trained walker out of ten has to hold a fixed scale in mind across hundreds of clips; asked which of two walkers moves forward better, the rater only has to get the sign of the difference right. An optometrist uses the same property when asking "better with lens one, or lens two?": the patient cannot state their prescription but can pick the sharper lens. The analogy stops there, because an optometrist only needs the winning lens, while this method needs a number for every timestep, and the next section shows how comparisons produce one.

An ablation tested the choice with a synthetic oracle in place of the person. Instead of comparisons, the reward model was fit by mean squared error to the true total reward of each clip (the "target" variant). On the continuous-control tasks, predicting comparisons worked much better. The paper's explanation is that reward scale varies substantially across states there, which makes the regression hard, and a comparison smooths that out. Atari rewards are clipped to their sign, which removes the scale problem; there the two variants performed significantly differently but neither consistently won.

The unit being compared is a trajectory segment, or clip: a run of kk observation-action pairs,

σ=((o0,a0), (o1,a1), …, (ok−1,ak−1)).\sigma = \big((o_0,a_0),\,(o_1,a_1),\,\dots,\,(o_{k-1},a_{k-1})\big).

Write σ1≻σ2\sigma^1 \succ \sigma^2 for "the person prefers clip 1". On Atari a clip is 25 timesteps, 1.7 seconds at 15 frames per second after frame skipping. On the robots it is 1.5 seconds, which is 15 to 60 timesteps depending on the task. A single frame usually cannot show whether a Pong paddle is about to return the ball, and on the robots an ablation using single states instead of clips performed much worse: matching the clip results would have needed significantly more comparisons. Longer clips were more useful per clip and less useful per frame. On very short clips raters spent time just working out what was happening, and beyond that the evaluation time grew roughly linearly with length, so the authors used the shortest clip length at which it had become linear.

Two clips in a pair usually start from different states, since the method cannot reset a simulator to a chosen state the way Wilson et al. assumed. Each answer is stored as a triple (σ1,σ2,μ)(\sigma^1, \sigma^2, \mu) in a database D\mathcal{D}, where μ\mu is a distribution over the two clips. A clear preference puts all of μ\mu on the chosen clip. "Equally good" sets μ=(0.5,0.5)\mu = (0.5, 0.5), and that pair still trains the model, pulling the two clips' scores toward each other. "Cannot compare" drops the pair: it never enters D\mathcal{D}.

How the reward model turns comparisons into a reward

The reward model r^\hat r is a neural network that scores one timestep, r^(o,a)\hat r(o, a). On Atari its input is the last four frames rather than one, so that motion is visible. A clip's score is the sum of its per-step scores,

si=∑t=0k−1r^(oti, ati),i∈{1,2}.s_i = \sum_{t=0}^{k-1} \hat r\big(o^i_t,\, a^i_t\big), \qquad i \in \{1, 2\}.

The model then needs a rule that turns two clip scores into a probability that the person picks clip 1. The paper uses the Bradley-Terry model (1952), the standard model for paired comparisons:

P^[σ1≻σ2]=exp⁡∑tr^(ot1,at1)exp⁡∑tr^(ot1,at1)+exp⁡∑tr^(ot2,at2)\hat P\big[\sigma^1 \succ \sigma^2\big] = \frac{\exp \sum_t \hat r(o^1_t,a^1_t)}{\exp \sum_t \hat r(o^1_t,a^1_t) + \exp \sum_t \hat r(o^2_t,a^2_t)}
(1)

Dividing the numerator and denominator by es1e^{s_1} turns (1) into the logistic sigmoid of the difference between the two clip scores:

P^[σ1≻σ2]=11+e−(s1−s2)=σ(s1−s2).\hat P\big[\sigma^1 \succ \sigma^2\big] = \frac{1}{1 + e^{-(s_1 - s_2)}} = \sigma(s_1 - s_2).

The paper compares this to the Elo rating system in chess, where the difference between two players' ratings sets the probability that one beats the other. Here each clip's summed reward plays the part of a rating and the person's choice is the game result. A gap of 0 gives a probability of 0.5, a gap of +2 gives σ(2)=0.88\sigma(2) = 0.88, and a gap of +4 gives 0.98.

The sum in (1) has no discount factor. A footnote in the paper reads this as modeling the person as indifferent to when, within the 1 to 2 seconds, the good moments happen. The discount factors that do appear in the paper (γ=0.99\gamma = 0.99 on Atari, 0.9950.995 on the robots) belong only to the policy optimizer.

Fitting r^\hat r is now supervised learning. The model predicts a probability for each stored pair, and the loss is the cross-entropy between that prediction and the person's answer μ\mu, summed over the database:

loss(r^)=− ⁣ ⁣∑(σ1,σ2,μ)∈D[μ(1)log⁡P^[σ1≻σ2]+μ(2)log⁡P^[σ2≻σ1]]\mathrm{loss}(\hat r) = -\!\!\sum_{(\sigma^1,\sigma^2,\mu)\in\mathcal{D}} \Big[\mu(1)\log \hat P\big[\sigma^1 \succ \sigma^2\big] + \mu(2)\log \hat P\big[\sigma^2 \succ \sigma^1\big]\Big]
(2)

The paper leaves this loss unnumbered; this page calls it (2). It is the negative log-likelihood of the Bradley-Terry model, so minimizing it is maximum-likelihood estimation of the reward that best explains the answers. For a reader who knows logistic regression: the input feature is the score gap s1−s2s_1 - s_2, the label is the person's choice, and the feature is itself computed by the network being trained.

The paper changes (1) in one way before using it. A pure sigmoid approaches 0 or 1 at large gaps, which claims the person will certainly agree with a confident model. Real raters misclick or lose attention at some constant rate that does not shrink on easy pairs, so the paper assumes a 10% chance that the person answers uniformly at random. Half of those random answers pick clip 1, which adds 0.1×0.5=0.050.1 \times 0.5 = 0.05:

P^[σ1≻σ2]=0.9 σ(s1−s2)+0.05  ∈  [0.05, 0.95].\hat P\big[\sigma^1 \succ \sigma^2\big] = 0.9\,\sigma(s_1 - s_2) + 0.05 \;\in\; [0.05,\ 0.95].

The bright curve in Figure 1 is this version, and the dim curve behind it is the plain sigmoid. Drag the gap toward either end and the bright curve levels off at 0.05 and 0.95 instead of reaching 0 and 1.

Figure 1 · the preference model
P = 77%
The chance a person prefers clip 1 is a sigmoid of the gap between the two clips' summed rewards. Drag the gap. The bright curve is the version the loss uses, 0.9 σ+0.050.9\,\sigma + 0.05, which levels off at the amber lines at 0.05 and 0.95 because the paper assumes a 10% chance of a random answer.

The floor limits how much one comparison can contribute. The predicted probability of any answer is at least 0.05, so the loss from a single comparison is at most −log⁡0.05≈3.0-\log 0.05 \approx 3.0. Without the floor, a label that contradicts a confident model produces a loss that grows without limit as the gap grows, and its gradient stays at full strength. With the floor, the gradient stays finite and fades toward zero at large gaps: at a gap of 6 against the label it is 0.042, against 0.998 for the plain sigmoid, so a contradicted label at a large gap moves the parameters about as little as one of the expected 10% random answers would. This is separate from "equally good", which is a real answer with μ=(0.5,0.5)\mu = (0.5, 0.5).

What one comparison teaches the reward model

Take one stored pair and follow the update. Clip 1 has six timesteps whose current predicted rewards sum to s1=0.90s_1 = 0.90; clip 2 has six that sum to s2=1.40s_2 = 1.40. The person preferred clip 1, so μ=(1,0)\mu = (1, 0). The model currently leans the wrong way:

s1−s2=−0.50,σ(−0.50)=0.378,P^=0.9×0.378+0.05=0.390,loss=−log⁡0.390=0.942.s_1 - s_2 = -0.50,\quad \sigma(-0.50) = 0.378,\quad \hat P = 0.9 \times 0.378 + 0.05 = 0.390,\quad \text{loss} = -\log 0.390 = 0.942.

Every timestep's output enters the loss only through its clip's sum, so the chain rule gives every timestep in clip 1 the same derivative, and every timestep in clip 2 the negative of it. Writing Δ=s1−s2\Delta = s_1 - s_2:

∂ loss∂r^(ot1,at1)=∂ loss∂Δ=−0.9 σ(Δ) (1−σ(Δ))P^=−0.9×0.378×0.6220.390=−0.543for every t,\frac{\partial\, \mathrm{loss}}{\partial \hat r(o^1_t,a^1_t)} = \frac{\partial\, \mathrm{loss}}{\partial \Delta} = -\frac{0.9\,\sigma(\Delta)\,\big(1-\sigma(\Delta)\big)}{\hat P} = -\frac{0.9 \times 0.378 \times 0.622}{0.390} = -0.543 \quad \text{for every } t,

and +0.543+0.543 for every timestep of clip 2. A gradient step raises all six outputs of clip 1 by the same amount and lowers all six of clip 2 by the same amount. The comparison carries no information about which moment in clip 1 made it better. Figure 2 shows that update with the numbers above as its starting state.

Figure 2 · one labeled pair, one gradient step
+0.00
Per-timestep rewards for clip 1 and clip 2, with the predicted preference, the loss and the per-step gradient above them. Press the gradient step a few times: every bar in a clip moves by the same amount, and the loss stops falling near −log⁡0.95=0.051-\log 0.95 = 0.051. Pick other verdicts, then move the offset cc, which adds the same constant to every bar; with clip 2 cut to 4 steps the same offset changes the prediction.

Repeated steps on this one pair push the gap up until P^\hat P nears its 0.95 ceiling, where the loss is 0.051 and the gradient is close to zero. With "equally good" selected, the gradient instead pushes the two sums toward each other and vanishes when s1=s2s_1 = s_2.

Which timesteps deserve the credit is learned across many comparisons. r^\hat r is one network applied to every timestep, so a situation that keeps appearing in preferred clips (the Hopper partway through a rotation, an Enduro car drawing level with another) has its score raised by many comparisons, while situations that appear equally in winners and losers have their increases and decreases cancel.

The Atari reward network takes the same 84×84×484 \times 84 \times 4 stacked frames as the policy. It has four convolutional layers (7×77\times7, 5×55\times5, 3×33\times3, 3×33\times3 kernels with strides 3, 2, 1, 1, and 16 filters each, leaky ReLU with slope 0.01), then a 64-unit fully connected layer and one scalar output. Every convolutional layer uses batch normalization and dropout, and an adaptive ℓ2\ell_2 penalty is added, because the network fits a few thousand labels. On the robots the reward model is a two-layer network with 64 hidden units per layer. A pair of 25-step Atari clips is 50 forward passes, two sums and one loss:

# one reward-model update on one labeled pair (Atari shapes)
seg1, seg2, mu = sample(D)       # each seg: 25 inputs of shape (84, 84, 4)
r1 = r_hat(seg1)                 # shape (25,): one scalar per timestep
r2 = r_hat(seg2)                 # shape (25,)
s1, s2 = r1.sum(), r2.sum()      # plain sum, no discount inside a clip
p = 0.9 * sigmoid(s1 - s2) + 0.05            # P[seg1 preferred]
loss = -(mu[0] * log(p) + mu[1] * log(1 - p))
loss.backward()                  # gradient reaches r_hat's weights only

What comparisons cannot pin down

Only the gap s1−s2s_1 - s_2 enters (1). Add a constant cc to the reward at every timestep and each clip's score rises by kckc. When both clips have the same length kk, the gap does not change, so no prediction and no loss changes:

r^→r^+c    ⟹    (s1+kc)−(s2+kc)=s1−s2    ⟹    P^ unchanged.\hat r \to \hat r + c \;\;\Longrightarrow\;\; (s_1 + kc) - (s_2 + kc) = s_1 - s_2 \;\;\Longrightarrow\;\; \hat P \text{ unchanged}.

The offset slider in Figure 2 shows this. With both clips at six steps, moving cc from −1-1 to +1+1 shifts both sums by 12 and leaves the gap, P^\hat P and the loss where they were. Cut clip 2 to four steps and the gap becomes s1−s2+2cs_1 - s_2 + 2c. The invariance needs equal-length clips, and every pair in the paper is two segments of the same length kk, so the comparisons fix the reward only up to an additive constant. The paper says the position of the rewards is underdetermined by the learning problem.

The scale is not determined by the task either. Multiplying every reward by 2 doubles every gap. That changes the predicted probabilities, so minimizing the loss does fix a scale: the one at which the predicted probabilities match how consistently the raters answered, through the fixed 10% noise and the unit-temperature sigmoid. The ordering of clips is the same at any positive scale, and multiplying a reward by a positive constant does not change which policy maximizes it.

The paper removes both freedoms by normalizing the rewards that r^\hat r hands to the policy to zero mean and a fixed standard deviation: 0.05 on Atari and 1 on the robots. The Atari value was picked so the A2C hyperparameters tuned for the game's clipped real rewards could be reused unchanged; the paper notes it could equally have adjusted the learning rate and entropy bonus. The same reasoning explains the ensemble, described in a later section, which averages its three predictors only after normalizing each one separately: two predictors that agree on every comparison can still sit at different offsets and scales, and averaging them raw would mix those arbitrary choices.

Comparisons do determine the shape of the reward: which behaviors score higher than which, and by how much relative to the raters' noise. Figure 3 fits a reward over 22 behaviors from comparisons alone. Each comparison picks two behaviors, reports the one with the higher hidden reward (with occasional errors), and takes one gradient step on the Bradley-Terry loss. Scrub from 0 to 620 comparisons and the learned curve settles onto the peaks and valleys of the true one; both curves are drawn normalized to zero mean and unit standard deviation because their level and scale are never determined.

Figure 3 · fitting a reward from comparisons
110 cmp
A hidden true reward over a range of behaviors, recovered by fitting the Bradley-Terry loss to pairwise comparisons. Scrub the number of comparisons or press Play, and the learned reward converges to the true shape. Both curves are normalized to zero mean and unit standard deviation, because comparisons fix the shape and not the level or scale.

Keeping the real reward out of reach

To test learning from preferences alone, the experiments remove every channel through which the task's real reward could leak into the agent's inputs. Atari screens show the score, so the score area is replaced by a constant black background in all seven games. On BeamRider the enemy ship count is also blanked, and on Enduro the speedometer.

Episode endings leak information too. A Gym robot's episode ends when it falls over, and an agent can learn "avoid whatever ends the episode" without any reward. The paper removes all variable-length episodes. On the robots, termination conditions become a penalty that the agent has to learn. On Atari, the agent is never sent life-loss or game-over signals; the environment still resets, but the agent sees one continuous episode. (With the synthetic oracle, game over is replaced by a penalty in every game except Pong.) The Gym robot rewards also penalize large joint torques, which a person watching a clip cannot see, so the paper removes that term from the true reward.

Fixed-length episodes also make the unknown constant from the previous section harmless, though the paper gives the information leak as its reason, not this. In one continuing episode discounted by γ\gamma, a constant cc added to every step adds c/(1−γ)c/(1-\gamma) to the value of every state, which changes no decision. If episodes could end early, a positive cc would pay the agent for staying alive and a negative one would pay it for dying quickly, and the arbitrary offset would change the behavior.

How the RLHF training loop runs: three processes at once

The method keeps a policy π\pi and a reward model r^\hat r, both neural networks, and updates them with three processes. The policy acts in the environment and is trained by RL on the rewards r^\hat r assigns. Pairs of clips from its recent behavior are sent to a person for comparison. The reward model is refit to all comparisons so far and its new parameters go back to the policy. The processes run asynchronously: clips flow from the first to the second, comparisons from the second to the third, and reward-model parameters from the third back to the first.

# three processes, each on its own clock
policy:   run A2C (16 workers) or TRPO on rewards from the latest r_hat;
          the game score is blacked out and never used
labeler:  sample clip pairs from recent rollouts, show the chosen pair,
          append (seg1, seg2, mu) to D
reward:   fit 3 predictors on the newest 3,000 labels, normalize each,
          average them, publish the new r_hat

Running them asynchronously keeps the simulator busy while the person thinks. A synchronous loop would stop the agent at every query. A person answers one in 3 to 5 seconds, and since a 50-million-step Atari run took about a day, the agent covers roughly 1,700 to 2,900 steps in that time. On the paper's Atari hardware the reward model processes about one label per 10 RL timesteps, looping over a buffer that holds only the newest 3,000 labels, so that new labels from new parts of the state space are not swamped by old ones.

Labels are not spread evenly. Before RL starts, the person compares clips from the untrained, randomly initialized policy: 500 comparisons on Atari, and 25% of the total on the robots. On Atari the reward model is then pretrained on these for 200 epochs, to make it less likely that the policy learns an irreversibly bad behavior from an untrained reward. After that, the label rate falls inversely with training time TT (on Atari it is cut in steps every 5 million frames to follow this curve):

label rate  ∝  5⋅106T+5⋅106    (Atari),2⋅106T+2⋅106    (robots).\text{label rate} \;\propto\; \frac{5 \cdot 10^{6}}{T + 5 \cdot 10^{6}} \;\;\text{(Atari)}, \qquad \frac{2 \cdot 10^{6}}{T + 2 \cdot 10^{6}} \;\;\text{(robots)}.

Early labels shape a reward model that is still mostly guesswork; later labels only need to track the states the improving agent reaches. With real contractors the schedule was approximate, since they labeled at uneven rates. Figure 4 lays this schedule over one Atari run of 50 million steps and 5,500 labels. Scrub through the run and read the bottom row: labels per environment step, and the share of environment steps a person actually watched, counting 2 clips of 25 timesteps per label.

Figure 4 · the loop and its label budget
25.0M
Top: the policy, the human and the reward model run concurrently; pick a process to read what it does. Bottom: the label rate over a 50M-step Atari run, reconstructed from the paper's schedule (500 labels up front, then a rate cut every 5M steps, 5,500 in total). Scrub or press Play to move through the run.

At the end of the run, 5,500 labels over 50 million steps is one label per 9,000 steps, or 0.011%. The frames a person watched are 5,500×2×25=275,0005{,}500 \times 2 \times 25 = 275{,}000 timesteps, 0.55% of the run. Both are under 1%. The paper does not show which count its "less than 1%" refers to, and early in the run the watched share is higher: in this reconstruction it stays above 1% until about 20 million steps, because the first 500 labels arrive before any training.

The Discussion states a separate figure: learning a reward model by supervised learning reduces the interaction complexity, the amount of human interaction needed, by roughly 3 orders of magnitude compared with having a person supply the reward directly.

Optimizing a reward that keeps changing

Once r^\hat r assigns a reward to every step, the policy side is a standard RL problem, with one difference: r^\hat r is non-stationary. It is refit continually on new comparisons, so the reward for the same behavior changes during training. The environment's physics do not change; only the learned reward does. The paper prefers methods that are robust to changes in the reward, and chooses policy gradient methods, citing their successful use in that setting by Ho and Ermon (2016). The paper gives no further reason. A likely factor is that value-based methods like DQN learn from a replay memory of transitions whose rewards were recorded when they were collected, and those rewards go stale when r^\hat r changes, while the methods chosen here learn from fresh rollouts scored by the current r^\hat r.

On Atari the optimizer is A2C, the synchronous form of A3C (Mnih et al. 2016). An actor network picks actions and a critic network estimates how much better than expected each action turned out (the advantage); the actor is pushed toward actions with positive advantage. Settings are the standard ones: 16 parallel workers, updates every 5 steps, γ=0.99\gamma = 0.99, an entropy bonus of 0.01, and Adam with a learning rate of 0.0007 decayed linearly toward zero at 80 million timesteps, though runs stopped at 50 million.

On the robots it is TRPO (Schulman et al. 2015), which takes the largest policy update that stays within a fixed KL-divergence distance of the current policy, with γ=0.995\gamma = 0.995. The entropy bonus was the only hyperparameter the authors adjusted. An entropy bonus adds reward for keeping the action distribution spread out. TRPO relies on its trust region to keep exploring, which the paper says can lead to too little exploration when the reward is changing, so they add an entropy bonus of 0.01 on every task (0.001 on Swimmer) to encourage the extra exploration.

Provenance Verified against primary literatureHow we verify
Bradley & Terry (1952)The paired-comparison model behind Eq (1). The paper also calls it the specialization of the Luce-Shepard choice rule to trajectory segments, and compares it to Elo ratings in chess.
Mnih et al. (2016)A3C. The Atari agents use its synchronous form, A2C, with standard settings: 16 workers, 5-step updates, gamma = 0.99, entropy bonus 0.01.
Schulman et al. (2015)TRPO, used on the eight MuJoCo tasks with gamma = 0.995. The only hyperparameter the authors adjusted was the entropy bonus.
Ho & Ermon (2016)The precedent the paper cites for running policy gradient methods against a reward that changes during training.
Amodei et al. (2016)Concrete Problems in AI Safety. Cited for the offline-training failure; the terms reward hacking and Goodhart are later vocabulary, not the paper's.
Wilson et al. (2012); Akrour et al. (2012, 2014)Earlier preference-based RL. The query format follows Wilson et al. (without resetting to chosen states) and the overall approach follows Akrour et al., in low-dimensional domains.
Daniel et al. (2014)The uncertainty-based query selection that the ensemble-disagreement heuristic is modeled on.
This page, Figure 4The label-budget curve reconstructs one Atari run from App. A.2: 500 initial comparisons, the rate cut every 5M steps in proportion to 5e6/(T + 5e6), 5,500 labels in total. The paper does not plot this curve.
correctionTwo corrections. First, the popular account that the first RLHF paper used PPO is wrong: PPO (arXiv:1707.06347, 20 July 2017) appeared about five weeks after this paper (12 June 2017) and is never mentioned in it. The Atari agents were trained with A2C and the robots with TRPO; PPO entered the recipe with later work (Ziegler et al. 2019, Stiennon et al. 2020, InstructGPT). Second, the Atari results text and Figure 3 compare 5,500 human queries to "350, 700, or 1400 synthetic queries". That phrase is copied from the MuJoCo paragraph: the same section says synthetic labels match RL on BeamRider and Pong "with only 3,300 such labels", and the Atari ablations use 5,500 synthetic labels. The MuJoCo counts of 350, 700 and 1400 are correct.

PPO is absent because it did not exist yet. The pipeline that later work uses with PPO (comparisons, a Bradley-Terry reward model, RL on that model) is complete in this paper, with A2C or TRPO in the optimizer slot.

Why the feedback has to stay online

Suppose all comparisons are collected at the start, the reward model is fit once and frozen, and only then does RL begin. The ablation called "no online queries" does this, and it performs poorly. The frozen model was fit on clips of early behavior. As the policy improves it reaches states that never appeared in a labeled clip; the paper attributes the failure to this change in the distribution of visited states. In those states the model's output is whatever its training happened to extrapolate, and an RL optimizer moves toward any state where that output is high, whether or not the person would agree.

Figure 5 shows the mechanism on a one-dimensional behavior axis. The amber curve is the true reward, which peaks at the goal. The teal curve is the reward model the agent climbs. In offline mode the model was fit on the left part of the axis and keeps rising past the goal, so as the agent climbs, its predicted reward keeps going up while the true reward it achieves falls. In online mode new comparisons correct the model wherever the agent is, and the agent stops at the goal.

Figure 5 · online versus offline
0%
The agent climbs the learned reward; the true reward is drawn for comparison. Offline, the frozen model keeps rising past the goal, so the agent's predicted score climbs while its true score falls. Online, the model is corrected wherever the agent goes, and it stops at the goal. Toggle the mode and scrub the run.

The paper's example is Pong. Trained offline, the agent sometimes learns to avoid losing points but not to score them, which produces extremely long volleys that repeat the same sequence of events indefinitely. The frozen reward model had learned that conceding is bad but had not learned that scoring is good in the states the agent ended up in. With online labeling, the long volleys would show up in the clips the person sees, and the labels on them would correct the model. The paper concludes that human feedback needs to be intertwined with RL learning rather than provided statically.

The paper's description is that the predictor captures only part of the true reward, and maximizing that partial reward produces bizarre behavior that is undesirable under the true reward. It cites Concrete Problems in AI Safety (Amodei et al. 2016) for this failure. Later work calls it reward hacking, or an instance of Goodhart's law (a measure that is optimized stops tracking what it measured); those names are not used in this paper.

Choosing which pairs to ask about

A comparison whose answer the model already predicts confidently changes little. The paper picks queries by an approximation of the reward model's uncertainty, following Daniel et al. (2014). It trains an ensemble of three predictors, each on ∣D∣|\mathcal{D}| comparisons resampled from D\mathcal{D} with replacement, so each sees a different mix of the data. It samples 10 times more candidate clip pairs than it will show, has every predictor give its probability that the first clip wins, and shows the person the pairs where the three probabilities vary most:

# pick n queries for the human (ensemble of 3 predictors)
cands = [sample_pair(recent_clips) for _ in range(10 * n)]
def spread(pair):
    probs = [prob_first_wins(m, *pair) for m in ensemble]  # 3 numbers
    return variance(probs)
ask = sorted(cands, key=spread, reverse=True)[:n]

Where many labels exist, the three predictors are pinned to the same answers; where few exist, they spread apart. In Figure 6 the band between the three curves is their disagreement, and the next query goes to its widest point. Click the plot to add a label anywhere, or press Play to let the rule choose.

Figure 6 · query selection by disagreement
uncertainty sampling
Three reward predictors agree where labels exist and spread into a wide disagreement band where they do not. The next query goes to the widest part of the band. Click to add a label and the band collapses there.

The ensemble has a second use: the reward handed to the policy is the average of the three normalized predictors, which varies less than any single one. Each predictor also holds out a fraction 1/e1/e (about 37%) of its data for validation, and the ℓ2\ell_2 coefficient is adjusted to keep validation loss between 1.1 and 1.5 times training loss, so the model is allowed to fit its labels somewhat more closely than unseen ones but not to memorize them.

The paper calls the variance rule a crude approximation. In the ablations it is not a reliable gain: on some tasks choosing queries uniformly at random did better. The quantity the authors would prefer is the expected value of information of a query, which they leave to future work.

Results

Each task is run three ways: with real human comparisons from contractors who got a one- or two-sentence description of the task; with a synthetic oracle that prefers whichever clip has the higher true reward (and answers "equal" when both clips have zero reward, common in sparse Atari games); and with ordinary RL on the true reward as the baseline. The goal stated in the paper is to come close to the baseline without access to the reward.

Robots. The eight MuJoCo tasks include a simple cartpole ("pendulum") for comparison with prior work. With 700 human labels the agent nearly matches RL on all eight. Learned rewards give less stable, higher-variance training with a comparable mean. Real human labels were between half as efficient as synthetic labels and equally efficient, depending on the task. With 1,400 synthetic labels the method does slightly better than RL on the true reward, perhaps because the learned reward is better shaped: it gives positive reward to all behaviors that are typically followed by high reward. On Ant, human labels beat synthetic ones because raters were asked to prefer clips where the robot was "standing upright", which turned out to be useful shaping; the true reward has an upright bonus too, but a simpler one that helped less.

Atari. The seven games are the ones from the original DQN paper. With 5,500 human labels the method shows substantial learning on most of them and matches or exceeds RL on some. With synthetic labels, BeamRider and Pong match or come close to RL with only 3,300 labels; Seaquest and Qbert approach RL but learn more slowly; SpaceInvaders and Breakout never match RL, though the agent often passes the first level of SpaceInvaders and reaches a Breakout score of 20, or 50 with enough labels. Real human labels often perform about as well as synthetic labels with 40% fewer, that is, 5,500 human labels doing the work of about 3,300 synthetic ones. The paper suggests labeling errors, disagreement between contractors, and uneven labeling rates as causes. With human labels, Qbert never gets past the first level, possibly because short Qbert clips are confusing. Enduro goes the other way: A3C struggles to pass cars by random exploration, but human raters reward any progress toward passing a car, which shapes the reward, and the agent outperforms A3C, comparable to DQN's result.

New behaviors. Three behaviors with no reward function at all were trained on the authors' own feedback, with the same settings. The Hopper does a sequence of backflips, landing upright and repeating, from 900 queries in under an hour. The Half-Cheetah moves forward standing on one leg, from 800 queries in under an hour. An Enduro car stays almost exactly level with other cars for much of the episode, from about 1,300 queries and 4 million frames, though changes in the background confuse it.

Ablations. Besides the comparison-versus-regression and clip-versus-single-state results above, the paper removes one piece at a time: random queries instead of disagreement-based ones, one predictor instead of an ensemble, offline labels only, and dropout without ℓ2\ell_2. Offline labeling performed poorly; disagreement-based selection was not consistently better than random.

Cost. For Atari the paper estimates compute at about $25, a 16-CPU machine with one K80 GPU (about $700 a month) running for about a day, and labor at about $36, since 5,000 labels take roughly 5 hours at US minimum wage. With compute and human feedback costing about the same, the authors conclude that further cuts in the number of labels give diminishing returns.

From this paper to RLHF for language models

Later work kept the reward model of (1) and (2) and changed the rest. Ziegler et al. (2019) fine-tuned language models from human preferences, using a softmax over four candidate completions; Stiennon et al. (2020) used pairwise comparisons of summaries; InstructGPT (2022) trained an instruction-following model the same way. In each, a reward model scores a whole completion with one number, and a sigmoid of the difference between two completions' scores predicts which one a labeler prefers. Those recipes add a pretrained and supervised starting model, PPO as the optimizer, and a KL penalty that keeps the policy near a reference model, which limits how far it can move into regions where the reward model is unreliable, the failure shown in Figure 5. Direct Preference Optimization (2023) optimizes the same Bradley-Terry objective directly on the policy, with no separate reward model and no RL loop.

Questions you might still have

?

Did this paper use PPO?
No. It used A2C, the synchronous form of A3C, on Atari and TRPO on the MuJoCo robots. PPO appeared about five weeks later, in July 2017, and became the RLHF optimizer with later work such as Ziegler et al. (2019), Stiennon et al. (2020) and InstructGPT. The PPO explainer covers that algorithm.

?

If the agent never sees the true reward, how do the authors know it worked?
The true reward is withheld from the agent and used only by the researchers to score it. Each task is run three ways: with real human comparisons, with a synthetic oracle that prefers whichever clip has the higher true reward, and with ordinary RL on the true reward. On MuJoCo, 700 human labels nearly match the RL baseline; for the backflip and other new behaviors there is no reward, so the evidence is the videos.

?

What does a single comparison teach the reward model?
Only that one clip's summed reward should rise relative to the other's. The gradient on every timestep of a clip is the same number, so one comparison cannot say which moment made the clip better. The reward model learns that from many comparisons, because the same network scores similar situations in many different clips.

?

Why ask for comparisons instead of numeric scores?
The authors found people give consistent comparisons far more easily than consistent absolute scores, especially on the robot tasks and the new behaviors. In an ablation that regressed true clip rewards instead of comparisons, comparisons worked much better on continuous control, likely because reward scale varies a lot there; on Atari, with rewards clipped to their sign, neither consistently won.

?

Why must the reward model keep training while the agent learns?
A reward model fit once is only reliable on the behavior it was fit on. As the policy improves it visits new states where the frozen model is wrong, and maximizing that partial reward led to bizarre behavior, such as a Pong agent that stops losing points but never scores. Labeling clips of the current behavior keeps the reward model accurate where the agent actually is.

?

How much human time and money did a run take?
Contractors answered a query in 3 to 5 seconds, so the real-feedback experiments took between 30 minutes and 5 hours of human time. For Atari the paper estimates about $25 of compute (a 16-CPU, one-K80 machine for about a day) and about $36 of labor for 5,000 labels at US minimum wage.

?

How does this connect to ChatGPT-style RLHF?
The reward model carries over: a Bradley-Terry probability over a pair, fit by cross-entropy, now scoring whole text completions. The language-model recipe adds a pretrained and supervised starting model, PPO as the optimizer, and a KL penalty that keeps the policy near a reference model. DPO later optimized the same Bradley-Terry objective without a separate reward model or RL loop. The InstructGPT, PPO and DPO explainers follow that line.

Footnotes & further reading

  1. The paper: Christiano, Leike, Brown, Martic, Legg, Amodei, Deep Reinforcement Learning from Human Preferences (arXiv v1 12 June 2017; NeurIPS 2017). Christiano and Amodei were at OpenAI; Leike, Martic and Legg at DeepMind; Brown is listed without an institution. Environment changes, architectures and the label schedules are in Appendix A.
  2. Bradley and Terry, Rank Analysis of Incomplete Block Designs (1952), the two-item case of the Luce choice rule. The paper's Elo comparison is to the logistic-of-a-rating-difference form; Elo's own original curve was Gaussian.
  3. Policy optimizers: A2C is the synchronous form of A3C from Mnih et al., Asynchronous Methods for Deep Reinforcement Learning (2016), explained on the A3C page. The robots used TRPO (Schulman et al. 2015). PPO (Schulman et al., arXiv:1707.06347) appeared about five weeks after this paper and is not used in it. The precedent for policy gradients on a changing reward is Ho and Ermon, Generative Adversarial Imitation Learning (2016).
  4. The offline-training failure is framed with Amodei et al., Concrete Problems in AI Safety (2016), which the paper also cites in its introduction on misaligned objectives.
  5. On the Atari label counts: the results paragraph and the Figure 3 caption compare 5,500 human queries to "350, 700, or 1400 synthetic queries", which repeats the MuJoCo sentence word for word. The same paragraph says synthetic labels match RL on BeamRider and Pong "with only 3,300 such labels", and the Atari ablations average 3 runs with 5,500 synthetic labels. The MuJoCo counts of 350, 700 and 1400 are correct.
  6. The language-model line: Ziegler et al., Fine-Tuning Language Models from Human Preferences (2019); Stiennon et al., Learning to Summarize from Human Feedback (2020); InstructGPT (Ouyang et al. 2022); and DPO (Rafailov et al. 2023). They reuse the Bradley-Terry reward-model skeleton for whole text completions and add the pretrained start, PPO and the KL-to-reference penalty that this paper did not need.