VerifiedarXiv:1509.0297133 min
Reinforcement learning · Continuous control

DDPG: Continuous Control with Deep Reinforcement Learning

A network outputs one action per state and is trained along the gradient of a learned Q-function, so deep Q-learning works for continuous actions such as joint torques.

DDPG trains two networks from a replay buffer: a critic that estimates the return of a state and action, and an actor that maps a state to an action. The actor is updated by backpropagating the critic's output through the action into the actor's weights.

Explaining the paperContinuous control with deep reinforcement learningLillicrap, Hunt, Pritzel, Heess, Erez, Tassa, Silver, Wierstra · ICLR 2016 · arXiv:1509.02971 ↗

A 2015 DeepMind paper that took the replay buffer and target network from DQN, combined them with the deterministic policy gradient, and trained one algorithm with one set of hyperparameters on more than 20 simulated physics tasks, from state vectors and from pixels.

Timothy Lillicrap and Jonathan Hunt (equal contribution), with Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver and Daan Wierstra, posted the paper in September 2015 and presented it at ICLR 2016. They named the method Deep DPG, or DDPG. On 26 tasks in the MuJoCo physics simulator, scored so that a random policy gets 0 and a planner with access to the simulator's equations gets 1, the average DDPG run scores 0.72 from joint angles and 0.50 from 64 by 64 pixel images; the best of five runs beats the planner on 10 tasks. With the same buffer and batch normalization but no target networks, the average drops to 0.22.

The page covers, in order (equation numbers follow the paper):

  1. why Q-learning needs a finite set of actions, and why discretizing a robot's actions fails;
  2. the critic, equations (1) to (5), and why it can learn from old data;
  3. the actor update of equation (6), the deterministic policy gradient, and what Silver et al. proved about it;
  4. the three stabilizers: replay buffer, soft target networks, batch normalization;
  5. exploration with Ornstein-Uhlenbeck noise, equation (7);
  6. Algorithm 1 as code with tensor shapes, the networks and hyperparameters;
  7. the results in Table 1, the limitations, and what TD3 and SAC changed.

Why DQN cannot output a torque

In reinforcement learning an agent sees a state sts_t, takes an action ata_t, receives a reward rtr_t and a new state, and tries to maximize the discounted sum of future rewards. The Deep Q-Network (DQN) learned Atari games from pixels by training a network to output Q(s,a)Q(s, a), the expected discounted return after taking action aa in state ss, for each of the 4 to 18 joystick actions. To act, it picks the action with the largest output. To train, it needs the same maximum at the next state, inside its target r+γmax⁡a′Q(s′,a′)r + \gamma \max_{a'} Q(s', a'). Both take one forward pass and a max over at most 18 numbers.

A robot's action is a vector of real numbers, usually joint torques. The cheetah task in this paper has 6 action dimensions, so an action is a point in [−1,1]6[-1, 1]^6, and max⁡aQ(s,a)\max_a Q(s, a) is now a 6-dimensional continuous optimization over a neural network's input. It would have to be solved at every environment step to act, and at every next-state in every minibatch to train. Gradient ascent on the action for a few dozen iterations per query would multiply the cost of training by that number of iterations, and it still finds only a local maximum.

Discretizing the actions runs into the size of the grid. The paper's example is a 7-joint arm with each torque restricted to {−k,0,k}\{-k, 0, k\}: 3⁷ = 2,187 joint actions, each needing its own output unit, and three levels per joint is too coarse for fine control. With 11 levels per joint the cheetah would have 11611^6, about 1.77 million, outputs, and the paper's quadruped (hyq) has 12 action dimensions. A grid also discards the fact that a torque of 0.51 behaves almost like 0.50; each grid point's value is learned separately.

DDPG keeps the action continuous and trains a second network, the actor μ(s)\mu(s), to output the maximizing action directly. Evaluating μ\mu is one forward pass, so the max inside the target becomes Q(s′,μ(s′))Q(s', \mu(s')). The rest of the paper is about training μ\mu and keeping both networks stable.

The critic: learning Q for a deterministic policy

The return from time tt is RtR_t, the sum of rewards from tt to the end of the episode at TT, each discounted by γ\gamma per step of delay:

Rt=∑i=tTγ i−t r(si,ai),γ∈[0,1]R_t = \sum_{i=t}^{T} \gamma^{\,i-t}\, r(s_i, a_i), \qquad \gamma \in [0, 1]

The paper uses γ=0.99\gamma = 0.99, so a reward 100 steps away counts 0.99¹⁰⁰ ≈ 0.37 times as much as an immediate one. The goal is a policy that maximizes JJ, the expected return from the start state. The action-value function of a policy π\pi is the expected return after taking ata_t in sts_t and following π\pi afterward:

Qπ(st,at)=Eri≥t, si>t∼E,  ai>t∼π[Rt∣st,at]Q^{\pi}(s_t, a_t) = \mathbb{E}_{r_{i \ge t},\, s_{i > t} \sim E,\; a_{i > t} \sim \pi}\big[ R_t \mid s_t, a_t \big]
(1)

The subscript EE denotes the environment, which may be random. Splitting off the first reward gives the Bellman equation, a recursion between the value at one step and the value at the next:

Qπ(st,at)=Ert, st+1∼E[r(st,at)+γ Eat+1∼π[Qπ(st+1,at+1)]]Q^{\pi}(s_t, a_t) = \mathbb{E}_{r_t,\, s_{t+1} \sim E}\Big[ r(s_t, a_t) + \gamma\, \mathbb{E}_{a_{t+1} \sim \pi}\big[ Q^{\pi}(s_{t+1}, a_{t+1}) \big] \Big]
(2)

For a deterministic policy, a function μ\mu from states to actions, the inner expectation over at+1a_{t+1} has a single term:

Qμ(st,at)=Ert, st+1∼E[r(st,at)+γ Qμ(st+1,μ(st+1))]Q^{\mu}(s_t, a_t) = \mathbb{E}_{r_t,\, s_{t+1} \sim E}\big[ r(s_t, a_t) + \gamma\, Q^{\mu}(s_{t+1}, \mu(s_{t+1})) \big]
(3)

The expectation in (3) is over the environment only. A recorded transition (st,at,rt,st+1)(s_t, a_t, r_t, s_{t+1}) supplies everything (3) needs, whichever policy chose ata_t: the next action comes from evaluating μ\mu at st+1s_{t+1}, not from the data. So QμQ^{\mu} can be learned off-policy, from transitions generated by a different behavior policy β\beta, such as μ\mu plus noise, or an older version of μ\mu.

With a neural network Q(s,a∣θQ)Q(s, a \mid \theta^Q), the critic is trained by regression onto the right-hand side of (3):

L(θQ)=Est∼ρβ, at∼β, rt∼E[(Q(st,at∣θQ)−yt)2]L(\theta^Q) = \mathbb{E}_{s_t \sim \rho^{\beta},\, a_t \sim \beta,\, r_t \sim E}\Big[ \big( Q(s_t, a_t \mid \theta^Q) - y_t \big)^2 \Big]
(4)
yt=r(st,at)+γ Q(st+1,μ(st+1)∣θQ)y_t = r(s_t, a_t) + \gamma\, Q(s_{t+1}, \mu(s_{t+1}) \mid \theta^Q)
(5)

ρβ\rho^{\beta} is the distribution of states the behavior policy visits. The target yty_t depends on θQ\theta^Q too, and the gradient through it is ignored: the target is computed, treated as a constant, and only Q(st,at)Q(s_t, a_t) is pushed toward it. This is the semi-gradient update of Q-learning. Because the same weights produce the target, every step on θQ\theta^Q also moves the targets that the next step regresses onto, which can make the regression oscillate or diverge; the target networks below address that.

If μ(s)\mu(s) is replaced by arg⁡max⁡aQ(s,a)\arg\max_a Q(s, a), (4) and (5) are Q-learning. DDPG uses the actor's output instead, and needs a way to make that output approach the argmax.

How DDPG works: the actor update

The actor μ(s∣θμ)\mu(s \mid \theta^{\mu}) is a network that takes a state and outputs an action through a final tanh⁡\tanh, so each dimension lies in [−1,1][-1, 1]. Its objective is the critic's value of the actions it chooses, averaged over states. Because QQ is a differentiable network and a=μ(s)a = \mu(s) is its input, the chain rule gives the gradient of that objective with respect to the actor's weights:

∇θμJ≈Est∼ρβ[∇θμQ(s,a∣θQ)∣s=st, a=μ(st∣θμ)]=Est∼ρβ[∇aQ(s,a∣θQ)∣s=st, a=μ(st) ∇θμμ(s∣θμ)∣s=st]\begin{aligned} \nabla_{\theta^{\mu}} J &\approx \mathbb{E}_{s_t \sim \rho^{\beta}}\Big[ \nabla_{\theta^{\mu}} Q(s, a \mid \theta^Q)\big|_{s = s_t,\, a = \mu(s_t \mid \theta^{\mu})} \Big] \\ &= \mathbb{E}_{s_t \sim \rho^{\beta}}\Big[ \nabla_a Q(s, a \mid \theta^Q)\big|_{s = s_t,\, a = \mu(s_t)}\, \nabla_{\theta^{\mu}} \mu(s \mid \theta^{\mu})\big|_{s = s_t} \Big] \end{aligned}
(6)

Read the second line right to left. ∇θμμ\nabla_{\theta^{\mu}} \mu is the Jacobian of the action with respect to the actor's weights: how each action dimension moves when each weight moves. ∇aQ\nabla_a Q is a vector with one entry per action dimension: how the critic's estimated return changes when that dimension of the action is nudged, at the action the actor currently outputs. Their product tells each weight how to move so the action shifts in the direction the critic rates as better. The critic's weights are not changed by this update.

A worked example with a linear actor and a one-dimensional action. Let the state be s=(0.5,−1.0)s = (0.5, -1.0) and the actor μ(s)=w⊤s+b\mu(s) = w^{\top} s + b with w=(0.4,0.2)w = (0.4, 0.2) and b=0.1b = 0.1, so the action is 0.2 − 0.2 + 0.1 = 0.1. Suppose the critic, evaluated at that state and action, has slope 2.0 in the action: a slightly larger action is estimated to give a higher return. The gradient of μ\mu with respect to ww is ss and with respect to bb is 1, so (6) multiplies the slope 2.0 into each:

∇wJ=2.0⋅(0.5,−1.0)=(1.0,−2.0),∇bJ=2.0⋅1=2.0\nabla_w J = 2.0 \cdot (0.5, -1.0) = (1.0, -2.0), \qquad \nabla_b J = 2.0 \cdot 1 = 2.0

A gradient-ascent step of size 0.01 makes w=(0.41,0.18)w = (0.41, 0.18) and b=0.12b = 0.12, and the new action at this state is 0.205 − 0.18 + 0.12 = 0.145, up by 0.045, which is 0.01 × 2.0 × (0.25 + 1 + 1): step size, times slope, times the squared length of the features (s,1)(s, 1). The reward never appears in this update; it reaches the actor only through what the critic has learned.

Picture the critic as a terrain map over actions, one map per state, with height equal to estimated return. The actor stands at one point on each map and steps uphill along the local slope. Over many states and steps it becomes a function that outputs, for any state, an action near a peak of that state's map: a learned, amortized argmax that costs one forward pass. The analogy breaks because the map is a learned network, not the terrain, and the actor climbs the map's errors as readily as its true peaks.

Figure 1 runs equation (6) on a toy with one state dimension and one action dimension, with a fixed critic so that only the actor moves. Look at where the white arrows point at the start and how long they are compared with the end, and compare the two starting points.

Figure 1 · the actor update of equation (6)
step 0
0.30
The amber field is a fixed critic Q(s,a)Q(s, a): brighter means higher estimated return. It has a high ridge and a lower ridge below it; the dashed line is the true argmax over actions at each state. The teal curve is the actor μ(s)\mu(s), a tanh of 11 Gaussian features. Each step samples 16 states, reads ∂Q/∂a\partial Q / \partial a at the actor's action (white arrows), and takes one step on the actor's weights. Press Play or scrub; drag on the field or use the state slider to see the slice Q(s,⋅)Q(s, \cdot) at one state. Switch the starting point to see the actor settle on the lower ridge.

From μ=0\mu = 0, the average value JJ rises from 0.243 to 1.036 by step 60 and 1.040 by step 120, against a maximum of 1.041, and the arrows shrink to nothing as the actor reaches the ridge, because the slope of QQ in the action direction is zero at a peak. From μ=−0.85\mu = -0.85, the nearest uphill direction at every state leads to the lower ridge, and the actor stops there with J=0.766J = 0.766. The update follows the local slope at the current action and never evaluates actions far from it. Exploration noise in DDPG, and a critic that keeps changing, sometimes move the actor out of such a basin; nothing guarantees it.

What the deterministic policy gradient theorem covers

The update in (6) is the deterministic policy gradient (DPG) of Silver et al. (2014), three of whose authors (Heess, Silver, Wierstra) are also on this paper. A reasonable worry is that changing the policy also changes which states the agent visits, and (6) ignores that: it averages over a fixed set of states. Their Theorem 1 shows that for the true action-value function and the states the policy itself visits, ignoring it is exact:

∇θJ(μθ)=Es∼ρμ[∇θμθ(s) ∇aQμ(s,a)∣a=μθ(s)]\nabla_{\theta} J(\mu_{\theta}) = \mathbb{E}_{s \sim \rho^{\mu}}\Big[ \nabla_{\theta} \mu_{\theta}(s)\, \nabla_a Q^{\mu}(s, a)\big|_{a = \mu_{\theta}(s)} \Big]

In this theorem ρμ\rho^{\mu} is the discounted distribution of states visited by μ\mu itself. DDPG differs from that statement in three ways. It averages over replay states drawn from ρβ\rho^{\beta}, the behavior policy's distribution, which Silver et al. introduce as an approximation that drops a term; hence the ≈\approx in (6). It uses a learned critic in place of QμQ^{\mu}, and the DPG paper guarantees the gradient is exact only for a "compatible" critic of a specific linear form, which a neural network is not. And, as the paper's footnote says, it ignores the discount in the state distribution, as most policy gradient implementations do. So (6) is the gradient of a surrogate, the critic's average value over replay states, and it improves the true return to the extent the critic is accurate near the actor's actions.

The stochastic policy gradient that A3C, TRPO and PPO build on has the form E[∇θlog⁡πθ(a∣s) Q(s,a)]\mathbb{E}[\nabla_{\theta} \log \pi_{\theta}(a \mid s)\, Q(s, a)]: it samples actions, scores them by their returns, and makes high-scoring ones more likely. It needs no derivative of QQ, but it learns about the direction of improvement only through sampled actions, and the estimate gets noisier as the action dimension grows. The deterministic version replaces that sampling with the critic's derivative; the DPG paper's Theorem 2 shows it is the limit of the stochastic gradient as the policy's variance goes to zero. The expectation in (6) is over states only, while the stochastic gradient also averages over actions; the DPG authors write that the stochastic version "may require more samples, especially if the action space has many dimensions."

DPG with neural networks had been tried before. NFQCA (Hafner and Riedmiller, 2011) used the same update rules with batch training to keep it stable; the paper notes that a minibatch version of NFQCA is the original DPG, which is what it compares against. That baseline is unstable on hard tasks, and the next three sections are the changes that fix it.

Replay buffer and soft target networks

Two of DDPG's three stabilizers come from DQN. The first is the replay buffer: every transition (st,at,rt,st+1)(s_t, a_t, r_t, s_{t+1}) goes into a first-in, first-out store of the last 10610^6 transitions, and each update trains both networks on a minibatch of 64 sampled uniformly from it. Consecutive steps of a physical system are nearly identical; training on them in order would fit the networks to whatever the robot is doing this second. Sampling from a million stored steps makes the minibatch closer to the independent samples that SGD assumes. A buffer is possible only because the critic learns off-policy, as (3) showed.

The second is the target network. In (5) the target yty_t is computed with the same critic being trained, so every update that raises Q(st,at)Q(s_t, a_t) also tends to raise Q(st+1,⋅)Q(s_{t+1}, \cdot) through shared weights, which raises the next target, and estimates can chase themselves upward or oscillate. DDPG computes targets with separate copies of both networks, Q′Q' and μ′\mu':

yi=ri+γ Q′(si+1, μ′(si+1∣θμ′)∣θQ′)y_i = r_i + \gamma\, Q'\big(s_{i+1},\, \mu'(s_{i+1} \mid \theta^{\mu'}) \mid \theta^{Q'}\big)

and moves the copies a fraction τ\tau of the way toward the online networks after every update:

θ′←τ θ+(1−τ) θ′,τ=0.001\theta' \leftarrow \tau\, \theta + (1 - \tau)\, \theta', \qquad \tau = 0.001

DQN's Nature version froze its target network and replaced it with a full copy every 10,000 updates. The soft update makes the target weights an exponential moving average of the online weights instead. If the online weights jump to new values and stay there, the target closes a fraction τ\tau of the remaining gap per update, so the gap left after nn updates is (1−τ)n(1-\tau)^n of the original. With τ=0.001\tau = 0.001 half the gap is closed after

n1/2=ln⁡2−ln⁡(1−τ)=0.69310.0010005≈693 updates,n_{1/2} = \frac{\ln 2}{-\ln(1 - \tau)} = \frac{0.6931}{0.0010005} \approx 693 \text{ updates,}

63% after 1,000, and 99% after about 4,600. The paper reports that both target copies were needed: "having both a target μ' and Q' was required to have stable targets yiy_i in order to consistently train the critic without divergence."

Figure 2 applies both rules to a synthetic critic output. Watch how far the amber target trails the teal online value at the paper's τ\tau, and what happens at the two ends of the slider.

Figure 2 · soft and hard target updates
0.0010
Teal: the online critic's output for one fixed next-state over 6,000 updates, a synthetic trace that rises as the critic learns and jitters from minibatch noise. Amber: the target network's output for the same input. Drag τ\tau (log scale; it starts at the paper's 0.001); at τ=1\tau = 1 the target is the online network, which is DPG without a target network. The second tab replaces the moving average with a frozen copy refreshed every 1/τ1/\tau updates.

At τ=0.001\tau = 0.001 the target's update-to-update jitter is about 0.3% of the online value's, and it trails a rising trend by about 1,000 updates ((1−τ)/τ(1-\tau)/\tau). The paper names this lag as a cost, since value information propagates through the bootstrap only as fast as the target moves. At τ=10−4\tau = 10^{-4} the target barely leaves its starting value within the plotted 6,000 updates. The hard copy holds every target fixed for 1,000 updates, an average delay of 500, and then changes all of them at once.

In Table 1 the variant labeled cntrl is the original DPG with the replay buffer and batch normalization but no target networks. Its average score over the 26 tasks is 0.22, against 0.72 with target networks; it is below DDPG on 25 of the 26 tasks and below the random policy on 7. The paper's Figure 2 caption puts it as "Target networks are crucial."

Batch normalization for mixed units

The third stabilizer addresses the inputs. A low-dimensional observation mixes quantities in different units: the cart's position in meters, the pole's angle in radians, angular velocities that can be ten times larger. Across tasks the ranges differ again. A network whose first layer sees one input in the hundreds and another near 0.01 learns slowly on the small one, and a learning rate tuned for one task's scales is wrong for another's, which defeats one hyperparameter setting for all tasks.

Instead of scaling each task's features by hand, DDPG applies batch normalization to the state input and to every layer of the actor, and to the critic's layers before the action enters. Each feature is normalized over the minibatch to zero mean and unit variance and then scaled and shifted by learned parameters; running averages of the mean and variance are used when the network acts in the environment. With it, the paper reports learning across tasks with different units without checking their ranges. The re-tuned DDPG baseline in the TD3 paper (2018) dropped batch normalization.

Exploration with Ornstein-Uhlenbeck noise

A deterministic actor takes the same action every time it sees a state, so on its own it never tries anything else and the critic never learns what other actions are worth. Because DDPG is off-policy, exploration can be added to the behavior policy without changing the learning rule. The paper adds noise to the actor's output:

μ′(st)=μ(st∣θtμ)+N\mu'(s_t) = \mu(s_t \mid \theta^{\mu}_t) + \mathcal{N}
(7)

The prime here marks the exploration policy, not the target actor of the previous section; the paper uses the same symbol for both. The noise N\mathcal{N} is an Ornstein-Uhlenbeck (OU) process, a random walk pulled back toward zero, with rate θ=0.15\theta = 0.15 (unrelated to the network weights θμ\theta^{\mu}) and scale σ=0.2\sigma = 0.2. Stepped once per action (the paper does not state its time step), each action dimension gets

nt+1=nt−θ nt+σ εt,εt∼N(0,1)n_{t+1} = n_t - \theta\, n_t + \sigma\, \varepsilon_t, \qquad \varepsilon_t \sim \mathcal{N}(0, 1)

so each step keeps 85% of the previous noise and adds a fresh Gaussian of standard deviation 0.2. In steady state the noise has standard deviation

σθ(2−θ)=0.20.2775≈0.38\frac{\sigma}{\sqrt{\theta(2 - \theta)}} = \frac{0.2}{\sqrt{0.2775}} \approx 0.38

on actions in [−1,1][-1, 1], and its correlation with itself halves every 4.3 steps. A push in one direction therefore tends to last several steps. On a body with inertia, consecutive pushes in the same direction add up to real movement, while independent pushes of the same size mostly cancel before the body responds.

Figure 3 · correlated and independent exploration noise
0.15
Top: 200 steps of OU noise and of independent Gaussian noise (gray) with the same steady-state standard deviation. Bottom: each noise used as the force on a toy cart with friction, and the root-mean-square final position over 300 episodes. Drag θ\theta (log scale; it starts at the paper's 0.15). At θ=1\theta = 1 the OU update becomes nt+1=σεtn_{t+1} = \sigma \varepsilon_t and the two noises are the same process; toward θ=0.01\theta = 0.01 OU becomes a slow random walk that leaves the action range.

At θ=0.15\theta = 0.15 the OU-driven cart ends about three times farther from where it started than the cart driven by independent noise of the same size, so it visits more of the state space per episode. The unstated time step changes the noise by orders of magnitude: public implementations with the same θ\theta and σ\sigma step the process with dt=1dt = 1 (rllab) or dt=0.01dt = 0.01 (OpenAI baselines), which gives correlation times of 6.7 and 667 agent steps. For Torcs, the driving game, the paper changed the noise because of the different time scale, without giving the values.

The full DDPG algorithm

Algorithm 1 in the paper interleaves acting and learning: one environment step, then one minibatch update of both networks, then a soft update of both targets. With the minibatch of NN transitions, the critic minimizes

L=1N∑i(yi−Q(si,ai∣θQ))2L = \frac{1}{N} \sum_i \big( y_i - Q(s_i, a_i \mid \theta^Q) \big)^2

and the actor steps along the sampled version of (6):

∇θμJ≈1N∑i∇aQ(s,a∣θQ)∣s=si, a=μ(si) ∇θμμ(s∣θμ)∣si\nabla_{\theta^{\mu}} J \approx \frac{1}{N} \sum_i \nabla_a Q(s, a \mid \theta^Q)\big|_{s = s_i,\, a = \mu(s_i)}\, \nabla_{\theta^{\mu}} \mu(s \mid \theta^{\mu})\big|_{s_i}

The paper reuses NN for three things: the action dimension (a∈RNa \in \mathbb{R}^N), the minibatch size, and, as N\mathcal{N}, the noise process. Below, the action dimension is 6 and the minibatch is 64. In an autodiff framework the actor gradient is written as a loss: the negative mean of Q(si,μ(si))Q(s_i, \mu(s_i)), backpropagated through the critic into the actor, with only the actor's optimizer stepping.

# One DDPG update on cheetah: state 17, action 6, minibatch N = 64.
s, a, r, s2 = buffer.sample(64)          # 64x17, 64x6, 64, 64x17

# critic: regress Q(s, a) onto a target built from the TARGET networks
with torch.no_grad():                    # y is a constant for this step
    a2 = actor_targ(s2)                  # 64x6, in [-1, 1] from tanh
    y = r + 0.99 * critic_targ(s2, a2)   # 64
q = critic(s, a)                         # 64, a = the action actually taken
loss_q = ((y - q) ** 2).mean() + weight_decay_1e-2(critic)
critic_opt.zero_grad(); loss_q.backward(); critic_opt.step()   # lr 1e-3

# actor: ascend Q(s, actor(s)); the gradient passes THROUGH the critic
loss_mu = -critic(s, actor(s)).mean()    # actor's own action, not a
actor_opt.zero_grad(); loss_mu.backward(); actor_opt.step()    # lr 1e-4
# loss_mu.backward() also fills critic grads; only actor_opt steps

# targets trail the online networks
for p_t, p in zip(all_target_params, all_online_params):
    p_t.mul_(1 - 0.001).add_(0.001 * p)  # tau = 0.001

Three different action sources appear in that update. The critic is evaluated at aia_i, the stored action that was taken, with noise, possibly many updates ago. The target uses μ′(si+1)\mu'(s_{i+1}), the target actor's choice at the next state. The actor loss uses μ(si)\mu(s_i), the current actor's choice at the stored state, which is never executed. The acting loop around the update:

# Acting in the environment (one episode)
noise = OUProcess(theta=0.15, sigma=0.2, size=6); noise.reset()
s = env.reset()
for t in range(T):
    actor.eval()                         # batch norm uses running stats
    a = actor(s) + noise.sample()        # equation (7); usually clipped
    s2, r, done = env.step(a)            # to the valid action range
    buffer.add(s, a, r, s2)              # buffer keeps the last 1e6
    actor.train()
    ddpg_update()                        # one minibatch update per step
    s = s2

Each transition is sampled into a minibatch about 64 times on average over its stay in a full buffer (a million updates, 64 draws each from a million entries), so every environment step contributes to about 64 updates. An on-policy method such as A3C uses a transition in one update, and PPO in a few epochs over one batch, before discarding it.

Networks and hyperparameters

Every MuJoCo task used the same settings, and Torcs changed only the exploration noise. Both networks have two hidden layers of 400 and 300 ReLU units for low-dimensional input. The actor ends in a tanh layer with one unit per action dimension. The critic takes the state at its first layer and concatenates the action to the 400 features before the second layer. For the cheetah (17 observation dimensions, 6 actions), that is

matching the paper's "≈ 130,000 parameters", not counting batch normalization's learned scales and shifts. The final layers start uniform in [−3×10−3, 3×10−3][-3 \times 10^{-3},\, 3 \times 10^{-3}] so that initial actions and Q values are near zero; other layers start uniform in [−1/f, 1/f][-1/\sqrt{f},\, 1/\sqrt{f}] with ff the fan-in. Optimization uses Adam with learning rate 10−410^{-4} for the actor and 10−310^{-3} for the critic, L2 weight decay of 10−210^{-2} on the critic only, γ=0.99\gamma = 0.99, τ=0.001\tau = 0.001, minibatch 64, buffer 10610^6.

From pixels, each observation is three consecutive renders of 64 by 64 RGB, stacked into 9 channels and scaled to [0,1][0, 1]. The agent's action is repeated for the three simulator steps, so differences between the frames carry velocity. Three convolutional layers of 32 filters each with no pooling feed two fully connected layers of 200 units, about 430,000 parameters, with actions entering at the fully connected layers. The minibatch is 16 and the final-layer initialization range is ten times smaller.

Results on 26 MuJoCo tasks and Torcs

The tasks are MuJoCo simulations: cart-pole swing-up and balance with up to three linked poles, 1- to 7-joint reaching arms, grippers that pick up and move a block, an arm with a hockey stick that hits a puck (canada), a one-legged hopper, a 2D cheetah and walker, and the 12-joint hyq quadruped. Actions are joint torques in every domain except cheetah. Rewards are dense: distance to a goal, progress forward, and a small action cost in every task.

Each score in Table 1 is normalized per task between two reference policies:

score=R−RrandomRiLQG−Rrandom\text{score} = \frac{R - R_{\text{random}}}{R_{\text{iLQG}} - R_{\text{random}}}

RrandomR_{\text{random}} is the mean return of actions drawn uniformly from the valid range. RiLQGR_{\text{iLQG}} is a model-predictive controller that, at every step, runs one iteration of iterative LQG trajectory optimization over a 250 to 600 ms horizon from the true simulator state, using the simulator's dynamics and their derivatives. A score above 1 means DDPG beat a planner that knows the physics; a negative score means worse than random. Every score is after at most 2.5 million steps, over 5 runs.

Figure 4 · Table 1, normalized score per task
Each row is one task, sorted by DDPG's score from state. Teal: DDPG from low-dimensional state. Violet: DDPG from pixels. Amber: DPG with replay buffer and batch normalization but no target networks (the paper's cntrl). Switch between the average and the best of 5 runs; hover or tap a row for its three values. Torcs is reported in raw reward and is not plotted.

Averaged over the 26 tasks, the mean run scores 0.72 from state and 0.50 from pixels. The best run scores 1.08 and 0.93. The best of five beats the planner on 10 tasks from state and 8 from pixels, and on the 7-joint block-lifting task (blockworld3da) the best pixel run scores 2.225. Two average scores from state exceed 1: blockworld1 (1.156) and hardCheetah (1.311), the cheetah variant with its stabilizing joint stiffness removed. The gap between best and average is large on several tasks, 0.303 against 1.735 on canada, so outcomes vary a lot between runs with the same settings.

On 5 tasks the average pixel run beats the average state run (blockworld3da, cart, movingGripper, reacherSingle, walker2d). The paper notes that on some simpler tasks learning from pixels is as fast as from state, and suggests that the action repeats simplify the problem or that convolutional features separate the states well; it does not test either.

On Torcs, a racing simulator with acceleration, braking and steering as actions, the reward is speed along the track with a penalty of 1 per collision. Some runs learned to complete laps, best returns 1,840 from state and 1,876 from pixels, while others failed, which pulls the averages to −393 and −402.

The paper also compared the critic's estimates with the returns actually obtained on test episodes. On pendulum and cart-pole they line up without a systematic bias; on harder tasks the estimates are less accurate, and the policies still perform well. Nearly all tasks were solved within 2.5 million steps, which the paper calls 20 times fewer than DQN needs on Atari. That ratio compares agent steps with DQN's 50 million emulator frames; DQN chose an action every 4th frame, so in decisions the ratio is 5.

Limitations and what came next

The paper names one limitation: like most model-free methods, DDPG needs many training episodes, up to 2.5 million steps of simulated experience per task. It also notes that nonlinear function approximation removes any convergence guarantee; the stability is empirical.

In 2018 Fujimoto, van Hoof and Meger measured DDPG's value estimates on the Gym versions of Hopper and Walker2d and found them above the true returns. The actor is trained to find actions where the critic is high, so it moves toward actions where the critic's error happens to be positive, and those overestimates enter the next targets. They proposed TD3 (Twin Delayed DDPG). TD3 takes the minimum of two critics in the target, updates the actor and targets every second critic update, and adds clipped noise to the target action so that the target does not depend on a narrow peak of QQ. Their re-tuned DDPG baseline also dropped batch normalization, weight decay and OU noise.

Soft Actor-Critic (2018) kept the off-policy actor-critic structure and the soft target update, and replaced the deterministic actor with a stochastic one trained to maximize return plus entropy. Its actor update still backpropagates the critic through the action, by sampling the action as a differentiable function of noise, which this paper's related-work section anticipated: "The techniques we described here for scaling DPG are also applicable to stochastic policies by using the reparametrization trick".

Implementation notes

DeepMind did not release code for the paper, so details the paper leaves out were filled in differently by each reimplementation. The ones that change behavior:

Provenance Verified against primary literatureHow we verify
Lillicrap, Hunt, Pritzel, Heess, Erez, Tassa, Silver & Wierstra (ICLR 2016; arXiv 1509.02971v6)Equations (1) to (7), Algorithm 1, Table 1 (all 27 rows, copied into Figure 4), Table 2 dimensions, Figure 2 and 3 captions, and Supplementary Section 7 hyperparameters: Adam 1e-4 actor / 1e-3 critic, L2 1e-2 on Q, gamma 0.99, tau 0.001, 400-300 units, actions entering Q at the second hidden layer, final-layer init U[-3e-3, 3e-3], batch 64 (16 from pixels), buffer 1e6, OU theta 0.15 sigma 0.2.
Silver, Lever, Heess, Degris, Wierstra & Riedmiller, Deterministic Policy Gradient Algorithms (ICML 2014)Theorem 1, equation (9): the gradient is an expectation over the on-policy discounted state distribution rho-mu with the true Q-mu. Equation (15): the off-policy form over rho-beta, introduced as an approximation that drops a term depending on the gradient of Q-mu with respect to theta. Theorem 2: the stochastic policy gradient tends to the deterministic one as the policy variance goes to 0. Theorem 3: compatible function approximation.
Mnih et al., Human-level control through deep reinforcement learning (Nature 2015)The target network is cloned from Q every C updates and held fixed between; 50 million training frames; frame skip k = 4, so 12.5 million agent decisions.
Mnih et al., Playing Atari with Deep Reinforcement Learning (2013)Algorithm 1 bootstraps from the same network it updates; replay memory, no target network.
Ioffe & Szegedy, Batch Normalization (ICML 2015)Each activation is normalized over the minibatch to zero mean and unit variance, then scaled and shifted by learned gamma and beta; inference uses population statistics.
Fujimoto, van Hoof & Meger, Addressing Function Approximation Error in Actor-Critic Methods (ICML 2018)Figure 1: DDPG value estimates exceed the true returns on Hopper-v1 and Walker2d-v1. TD3 adds clipped double Q-learning, delayed policy updates and target policy smoothing. Table 3: their re-tuned DDPG drops batch normalization and weight decay and replaces OU noise with Gaussian N(0, 0.1).
rllab ou_strategy.py; OpenAI baselines ddpg/noise.pyNo official DDPG code was released. rllab steps OU noise with an implicit time step of 1 (x += theta(mu - x) + sigma eps); baselines uses dt = 0.01 with sigma sqrt(dt). Same theta and sigma, very different noise: correlation time 6.7 agent steps against 667.
This page, arithmetic and toysParameter counts for cheetah (129,306 actor, 129,601 critic, excluding batch-norm parameters), the soft-update half-life ln 2 / -ln(1 - tau) = 693 updates, the OU stationary std 0.38 and half-life 4.3 steps under dt = 1, the Table 1 column means, and the toy critic, trace and cart in Figures 1 to 3.
correctionFive points where the paper says something other than what its sources support. (1) Section 2 writes the deterministic policy as mu: S <- A; a policy maps states to actions, mu: S -> A. (2) Section 3 says batch normalization normalizes each dimension "to have unit mean and variance"; Ioffe and Szegedy normalize to zero mean and unit variance, then apply a learned scale and shift. (3) Section 3 says DDPG's target network is "similar to the target network used in (Mnih et al., 2013)". The 2013 DQN paper has no target network; it bootstraps from the network being trained. The target network appears in the 2015 Nature paper. (4) The supplement gives the pixel networks' final-layer initialization as [3 x 10^-4, 3 x 10^-4], an interval of one point; the intended range is [-3 x 10^-4, 3 x 10^-4], matching the low-dimensional [-3 x 10^-3, 3 x 10^-3]. (5) The paper states that Silver et al. (2014) "proved that this is the policy gradient". Their theorem covers the on-policy state distribution and the true action-value function. The off-policy form that DDPG uses, an average over replay states, is introduced in that paper as an approximation, and with a learned critic the gradient is guaranteed exact only for a compatible critic of a specific linear form, which a neural network is not. A softer note: the conclusion's "a factor of 20 fewer steps than DQN" compares 2.5 million agent steps with 50 million Atari frames; DQN acted on every fourth frame, so in agent decisions the ratio is 5.

Questions you might still have

?

Is DDPG on-policy or off-policy?
Off-policy. The critic's target uses the actor's choice at the next state, not the action the behavior policy took there, so transitions collected by any earlier version of the actor, with any noise, are valid training data, so DDPG can keep a million transitions in a replay buffer and reuse each one many times. On-policy methods such as A3C, TRPO and PPO have to collect fresh data after each policy change.

?

How is DDPG related to DQN?
DDPG keeps DQN's replay buffer, its target network (made soft), and its squared-error regression of Q onto a bootstrapped target. It replaces the max over actions in DQN's target, which needs a finite action set, with a second network that outputs the action, trained to make Q large. With a perfectly trained actor, mu(s) equals argmax Q(s, a) and the target is the same as DQN's.

?

Why does the actor learn from the critic's gradient instead of from rewards?
Because the gradient of Q with respect to the action says, for every action dimension, which direction raises the expected return and by how much, from a single state with no sampling over actions. A likelihood-ratio policy gradient (REINFORCE) has to infer that direction from the returns of many sampled actions, which gives much noisier estimates. In exchange, the actor improves only as fast as the critic's gradient is correct.

?

Does DDPG find the best action in each state?
Not necessarily. The actor follows the local slope of Q in the action direction. If Q has two peaks over the action at some state and the actor starts near the lower one, gradient ascent climbs the lower one and stays there; Figure 1 on this page shows that case. Exploration noise and a changing critic help it escape in practice, with no guarantee.

?

Is the Ornstein-Uhlenbeck noise necessary?
No. It was chosen for tasks with inertia, where correlated pushes move the body farther than independent ones. Later work, including the TD3 paper's re-tuned DDPG baseline, used independent Gaussian noise with standard deviation 0.1 and reported strong results. The paper also never states the time step of its OU process, and public implementations disagree by a factor of 100.

?

Why do the actions enter the critic at the second layer?
The paper states the choice without an ablation. One reading is that the first layer, batch-normalized, builds features of the state alone, and the action is combined with those features rather than with raw positions and velocities. The supplement also notes that batch normalization is applied only to the critic's layers before the action input, so the action itself is never normalized.

?

What goes wrong with DDPG in practice?
The critic overestimates: the actor is trained to find actions where Q is high, and errors in Q that happen to be high are what it finds. TD3 (2018) measured this on Hopper and Walker2d and fixed it with the minimum of two critics, slower actor updates and noise on the target action. Results also vary a lot between runs: in the paper's own Table 1 the best of 5 runs on the canada task scores 1.735 and the average 0.303. Soft Actor-Critic, with a stochastic policy and an entropy bonus, is a common replacement.

?

How are terminal states handled?
Algorithm 1 in the paper does not mention them, although the locomotion tasks end early when the body falls. Implementations multiply the bootstrap term by (1 - done), so that at the last transition of an episode the target is the reward alone; an episode cut off by a time limit is usually treated as not done.

Footnotes & further reading

  1. The paper: Lillicrap, Hunt, Pritzel, Heess, Erez, Tassa, Silver & Wierstra, Continuous control with deep reinforcement learning (ICLR 2016). Videos of learned policies were linked from the paper at goo.gl/J4PIAz.
  2. The theory: Silver, Lever, Heess, Degris, Wierstra & Riedmiller, Deterministic Policy Gradient Algorithms (ICML 2014). The earlier neural version: Hafner & Riedmiller, Reinforcement learning in feedback control (Machine Learning, 2011).
  3. The stabilizers: Mnih et al., Playing Atari with Deep Reinforcement Learning (2013; explained in DQN) and Human-level control through deep reinforcement learning (Nature 2015); Ioffe & Szegedy, Batch Normalization (ICML 2015).
  4. The noise: Uhlenbeck & Ornstein, On the theory of the Brownian motion (Physical Review, 1930). Implementations compared: rllab ou_strategy.py and OpenAI baselines ddpg/noise.py.
  5. What came next: Fujimoto, van Hoof & Meger, Addressing Function Approximation Error in Actor-Critic Methods (TD3, ICML 2018); Haarnoja et al., Soft Actor-Critic (ICML 2018).