VerifiedarXiv:2301.0410434 min
Reinforcement learning · World models

DreamerV3: Mastering Diverse Domains through World Models

An agent learns a model of its environment, practices inside it, and one fixed set of hyperparameters works on over 150 tasks.

DreamerV3 trains three networks at once: a world model that predicts what happens next, a critic that estimates how good a state is, and an actor that picks actions. The actor and critic learn only from trajectories the world model imagines. A handful of normalizations keep every loss on the same scale, so the configuration that plays Atari also controls simulated robots and collects diamonds in Minecraft.

Explaining the paperMastering Diverse Domains through World ModelsHafner, Pasukonis, Ba, Lillicrap · arXiv 2023; Nature 2025 · arXiv:2301.04104 ↗

The first agent to collect a diamond in Minecraft from scratch, with no human demonstrations, on one GPU in nine days, using the same settings it uses for Atari.

The problem: one set of hyperparameters for every task

Reinforcement learning (RL) trains an agent by trial and error. At each step the agent receives an observation xtx_t (a camera image, or a vector of joint angles), chooses an action ata_t, and the environment returns a reward rtr_t, a number that is high when something good happened. The agent's goal is to maximize the return, the sum of future rewards with each one discounted by how far away it is:

Rt≐∑τ=0∞γτ rt+τ,γ=0.997R_t \doteq \sum_{\tau=0}^{\infty} \gamma^{\tau}\, r_{t+\tau}, \qquad \gamma = 0.997

With γ=0.997\gamma = 0.997 a reward 333 steps away still counts for 0.997333≈0.370.997^{333} \approx 0.37 of its face value, so the agent cares about roughly the next 333 steps (the paper calls 1/(1−γ)1/(1-\gamma) the discount horizon). A policy, called the actor, maps the current state to a distribution over actions. A value function, called the critic, estimates the expected return from a state under the current policy. Actor-critic methods train both: the critic learns to predict returns, and the actor shifts probability toward actions that turned out better than the critic expected.

The algorithms that work well are specialized. DQN and its descendants handle discrete actions, SAC handles continuous control, MuZero plans with a learned model on board games and Atari. Each needs its hyperparameters retuned when the domain changes: the entropy bonus that makes an agent explore under sparse rewards makes it dither under dense ones, a squared-error loss that suits rewards of size 1 can diverge on rewards of size 1000, and a representation-loss weight chosen for Atari's static backgrounds is wrong for 3D scenes. PPO is the common fallback because it is comparatively robust, but it needs a lot of experience and usually loses to the specialists.

DreamerV3, by Danijar Hafner, Jurgis Pasukonis, Jimmy Ba and Timothy Lillicrap, sets out to remove the retuning. The paper evaluates one configuration on 8 domains and over 150 tasks: Atari, ProcGen, DMLab, Atari100k, continuous control from joint angles and from pixels, BSuite, and Minecraft. It is the third version of Dreamer. Version 1 (2019) handled continuous control and version 2 (2020) Atari; version 3 keeps their structure and adds the robustness techniques this page spends most of its length on.

How DreamerV3 works

The agent has three neural networks and a replay buffer, a store of everything it has experienced (up to 5 million steps). Three things happen concurrently:

  1. The actor plays in the real environment, and every step goes into the replay buffer.
  2. The world model trains on sequences sampled from the buffer to predict, from its compact internal state, the next internal state, the reward, whether the episode continues, and the observation itself.
  3. From each replayed state the world model and actor imagine 15 steps ahead. The critic learns to predict returns on those imagined trajectories, and the actor learns to pick actions with high predicted return.

Training a policy on imagined experience is an old idea (Sutton's Dyna, 1991, and the 2018 World Models paper of Ha and Schmidhuber both do it). Think of a flight simulator built from the pilot's own flight logs: practice inside it is cheap, but the simulator is only as good as the logs, and a pilot who trains too long in it learns the simulator's bugs. Dreamer keeps imagined trajectories short (15 steps) and keeps refitting the model to fresh experience for the same reason.

Here is a full update with the tensor shapes of the default 200M-parameter model; the sections below go through it line by line:

# One DreamerV3 update. 200M model, Table 4 defaults, shapes per batch.
batch = replay.sample(B=16, T=64)          # images, actions, rewards, flags
# 1. World model: run the RSSM over the replayed sequence (eq. 1)
h, z, post, prior = rssm.observe(batch)    # h: 16x64x8192, z: 16x64x32x64
L_pred = (decoder.nll(batch.x, h, z)       # images: MSE; vectors: symlog MSE
          + reward_head.twohot_ce(batch.r, h, z)
          + cont_head.bce(batch.cont, h, z))
L_dyn = max(1, KL(sg(post) || prior))      # moves the dynamics predictor
L_rep = max(1, KL(post || sg(prior)))      # moves the encoder
L_model = 1.0 * L_pred + 1.0 * L_dyn + 0.1 * L_rep
# 2. Imagination: 15 steps from each of the 16*64 = 1024 replayed states
traj = rssm.imagine(sg(h, z), actor, steps=15)   # 1024 x 16 model states
r, c, v = reward_head(traj), cont_head(traj), critic(traj)
R = lambda_return(r, c, v, lam=0.95)       # eq. 5, backward over 15 steps
# 3. Actor and critic, trained on imagined states only
S = ema(percentile(R, 95) - percentile(R, 5), decay=0.99)
adv = sg((R - v[:, :-1]) / max(1, S))
L_actor = -(adv * actor.logp(traj.a) + 3e-4 * actor.entropy(traj))
L_critic = (critic.twohot_ce(sg(R))        # eq. 5
            + critic.twohot_ce(sg(critic_ema(traj))))   # pull toward EMA
L_critic_replay = 0.3 * critic_loss_on_replayed_states(batch)
# 4. One optimizer step on everything: LaProp, lr 4e-5, AGC at 0.3
step(L_model + L_actor + L_critic + L_critic_replay)

sg is stop-gradient: the value passes forward, but no gradient flows back through it. In the listing it has three jobs: it keeps the actor and critic losses from changing the world model, it splits the world model's KL term into two losses that each move one network, and it turns the advantages and the critic's targets into constants, so gradients reach the actor and critic only through their own outputs.

The world model: a recurrent state-space model

The world model is a recurrent state-space model (RSSM), introduced by Hafner and colleagues in 2018 for the PlaNet agent. Its state at time tt has two parts. The recurrent state hth_t is a deterministic vector that carries memory forward, like the hidden state of any recurrent network. The stochastic state ztz_t is a sample from a distribution, and holds what the model learned from the current observation. The model state is the pair st={ht,zt}s_t = \{h_t, z_t\}. Six components connect them:

Sequence model:ht=fϕ(ht−1,zt−1,at−1)Encoder:zt∼qϕ(zt∣ht,xt)Dynamics predictor:z^t∼pϕ(z^t∣ht)Reward predictor:r^t∼pϕ(r^t∣ht,zt)Continue predictor:c^t∼pϕ(c^t∣ht,zt)Decoder:x^t∼pϕ(x^t∣ht,zt)\begin{aligned} &\text{Sequence model:} && h_t = f_\phi(h_{t-1}, z_{t-1}, a_{t-1}) \\ &\text{Encoder:} && z_t \sim q_\phi(z_t \mid h_t, x_t) \\ &\text{Dynamics predictor:} && \hat z_t \sim p_\phi(\hat z_t \mid h_t) \\ &\text{Reward predictor:} && \hat r_t \sim p_\phi(\hat r_t \mid h_t, z_t) \\ &\text{Continue predictor:} && \hat c_t \sim p_\phi(\hat c_t \mid h_t, z_t) \\ &\text{Decoder:} && \hat x_t \sim p_\phi(\hat x_t \mid h_t, z_t) \end{aligned}
(1)

All six share the parameters ϕ\phi. The sequence model is a GRU, a gated recurrent cell, and it advances hh one step using the previous stochastic state and action. The encoder looks at the current observation xtx_t (and hth_t) and outputs a distribution over ztz_t. The dynamics predictor outputs a distribution over the same ztz_t from hth_t alone, without the observation: it is the model's guess of what it is about to see. The other three heads read the full model state and predict the reward, the continue flag ct∈{0,1}c_t \in \{0, 1\} (0 when the episode ends), and the observation.

The shapes at the default 200M size, for one 64×64 RGB frame:

The stochastic state is categorical because DreamerV2 found that categorical latents beat Gaussian ones on Atari. Sampling a class is not differentiable, so Dreamer uses straight-through gradients: the forward pass uses the one-hot sample, and the backward pass treats it as if it were the vector of class probabilities. In the code this is one line, sample + probs - sg(probs), whose value equals sample and whose gradient is that of probs.

Equation (1) produces ztz_t in two ways. While the world model trains on replayed data, it observes: each ztz_t comes from the encoder, which sees the real frame. During imagination there are no frames, so each ztz_t comes from the dynamics predictor, and the actions come from the actor instead of the replay buffer. In Figure 1, drag the observed-steps slider to move the boundary, and click a column to see what that step is computed from.

Figure 1 · observing and imagining
3 of 8
4
The RSSM unrolled over time. Observed steps take ztz_t from the encoder, which sees the frame xtx_t; imagined steps take it from the dynamics predictor, which sees only hth_t, and their frames (dashed) are predictions. Click a column or use the second slider to inspect a step.

Imagination never touches pixels: from a start state, one imagined step is one GRU update, one dynamics-predictor sample, and the reward and continue heads. The decoder is not run during behavior learning. That makes imagining 16,384 states per update (1024 start states times 16) cheap enough to do at every training step, and it is why the dynamics predictor must produce latents that look like the encoder's: the actor was trained on imagined latents but acts on encoded ones.

Training the world model: prediction, dynamics and representation losses

Given a replayed batch of 16 sequences of 64 steps, the world model minimizes three losses, with weights βpred=1\beta_{\text{pred}} = 1, βdyn=1\beta_{\text{dyn}} = 1 and βrep=0.1\beta_{\text{rep}} = 0.1:

L(ϕ)≐Eqϕ[∑t=1T(βpred Lpred(ϕ)+βdyn Ldyn(ϕ)+βrep Lrep(ϕ))]\mathcal{L}(\phi) \doteq \mathrm{E}_{q_\phi}\Big[\sum_{t=1}^{T}\big(\beta_{\text{pred}}\,\mathcal{L}_{\text{pred}}(\phi) + \beta_{\text{dyn}}\,\mathcal{L}_{\text{dyn}}(\phi) + \beta_{\text{rep}}\,\mathcal{L}_{\text{rep}}(\phi)\big)\Big]
(2)
Lpred(ϕ)≐−ln⁡pϕ(xt∣zt,ht)−ln⁡pϕ(rt∣zt,ht)−ln⁡pϕ(ct∣zt,ht)Ldyn(ϕ)≐max⁡(1, KL[sg⁡(qϕ(zt∣ht,xt)) ∥ pϕ(zt∣ht)])Lrep(ϕ)≐max⁡(1, KL[qϕ(zt∣ht,xt) ∥ sg⁡(pϕ(zt∣ht))])\begin{aligned} \mathcal{L}_{\text{pred}}(\phi) &\doteq -\ln p_\phi(x_t \mid z_t, h_t) - \ln p_\phi(r_t \mid z_t, h_t) - \ln p_\phi(c_t \mid z_t, h_t) \\ \mathcal{L}_{\text{dyn}}(\phi) &\doteq \max\big(1,\ \mathrm{KL}\big[\operatorname{sg}(q_\phi(z_t \mid h_t, x_t)) \,\big\|\, p_\phi(z_t \mid h_t)\big]\big) \\ \mathcal{L}_{\text{rep}}(\phi) &\doteq \max\big(1,\ \mathrm{KL}\big[q_\phi(z_t \mid h_t, x_t) \,\big\|\, \operatorname{sg}(p_\phi(z_t \mid h_t))\big]\big) \end{aligned}
(3)

The prediction loss trains the heads to reconstruct the observation, predict the reward, and predict whether the episode continues. Images use a squared error, vector observations a squared error on symlog-transformed targets, the reward head the twohot loss of a later section, and the continue head a logistic (binary cross-entropy) loss.

The other two terms compare the encoder's distribution qq (which saw the frame) with the dynamics predictor's pp (which did not) using the KL divergence, a measure in nats of how far one distribution is from another; it is zero only when they are equal. This is the structure of a variational autoencoder, with the dynamics predictor playing the role of a learned prior. A single KL term would move both networks toward each other with equal force. DreamerV3 writes it twice with the stop-gradient in different places, so each copy moves only one network:

This is KL balancing, from DreamerV2 (which used a 0.8/0.2 split). The weights are unequal because a better predictor is always useful, while a more predictable representation is useful only up to a point: past it the encoder makes itself predictable by throwing information away: in the limit, an encoder that outputs the same zz for every frame is perfectly predictable and useless.

Free bits keep the encoder away from that limit. The max⁡(1,⋅)\max(1, \cdot) clips each KL term at 1 nat (about 1.44 bits). While the KL is below 1 nat, the loss is the constant 1 and its gradient is zero, so neither network is pushed further; the prediction loss alone shapes the representation. The idea is from Kingma and colleagues (2016). In the official code the KL is first summed over the 32 latents of a time step and then clipped, so the free allowance is 1 nat for the whole state per step, not per latent.

Figure 2 is a toy version of these three terms. There is one latent with 8 classes and four possible observations x1x_1 to x4x_4 that the predictor cannot tell apart, as when the next frame depends on something random. A fixed decoder maps class jj to observation xjx_j for the first four classes; the last four decode to nothing in particular. The best the prior can do is the average of the four posteriors, so the KL that remains equals the information the latent carries about which observation it saw. At the paper's βrep=0.1\beta_{\text{rep}} = 0.1 each posterior settles on its own class and the latent keeps 1.38 nats, nearly all of ln⁡4=1.39\ln 4 = 1.39. Raise βrep\beta_{\text{rep}} to 3 and turn free bits off, and the posteriors collapse onto the prior: 0.006 nats, and the decoder is right 26% of the time, chance level for four observations. Turn free bits back on at the same βrep\beta_{\text{rep}} and the KL stops near 1 nat, keeping 0.96 nats and a 78% decode rate.

Figure 2 · KL balancing and free bits (toy)
800
0.10
Four observations, one 8-class latent. Prior bars are the dynamics predictor's guess; posterior bars are the encoder's output for each observation. The shaded first four classes decode to the right observation. The trace shows the KL (white) and the information kept (teal) against the 1-nat line. Play to watch training; the slider sets βrep\beta_{\text{rep}}, with βdyn=1\beta_{\text{dyn}} = 1 fixed.

Earlier Dreamers had to choose the representation weight per domain: a strong pull toward predictability for 3D environments full of irrelevant detail, a weak one for games with static backgrounds where single pixels matter. The paper reports that a small βrep\beta_{\text{rep}} combined with free bits works for both, so DreamerV3 uses 0.1 everywhere.

The paper also saw occasional spikes in the KL losses, consistent with reports from deep VAEs. A categorical that puts probability near zero on a class makes the log ratio inside the KL explode. DreamerV3 mixes every categorical (encoder, dynamics predictor and actor) as 99% network output and 1% uniform, called unimix. No class can then fall below 0.01/64≈1.6×10−40.01/64 \approx 1.6 \times 10^{-4}, which caps the KL of one 64-class latent at ln⁡6400≈8.76\ln 6400 \approx 8.76 nats.

Learning in imagination: the critic and λ-returns

The actor and critic never see an observation. They read model states st={ht,zt}s_t = \{h_t, z_t\}, and since hth_t carries the history, a single model state is meant to contain everything needed to predict the future (the paper calls the representation Markovian). Both are 3-layer MLPs:

Actor:at∼πθ(at∣st)Critic:vψ(Rt∣st)\text{Actor:}\quad a_t \sim \pi_\theta(a_t \mid s_t) \qquad\qquad \text{Critic:}\quad v_\psi(R_t \mid s_t)
(4)

The critic outputs a distribution over returns, not a single number; its prediction vtv_t is the mean of that distribution. Starting from every state of the replayed batch, the world model and actor imagine 15 steps, giving 16 model states per start including the start itself (the paper's T=16T = 16, the configuration's imagination horizon of 15). The reward, continue and critic heads label every imagined state.

The return of an imagined state depends on rewards up to 333 steps away, and the trajectory stops after 15. The critic fills the gap: after the last imagined step, its own estimate stands in for the rest. Dreamer blends the imagined rewards and the critic's estimates with the λ-return, written here with the indices of the official code:

Rtλ≐rt+1+γ ct+1((1−λ) vt+1+λ Rt+1λ),RTλ≐vTR^\lambda_t \doteq r_{t+1} + \gamma\, c_{t+1}\big((1-\lambda)\, v_{t+1} + \lambda\, R^\lambda_{t+1}\big), \qquad R^\lambda_{T} \doteq v_{T}
(5)

Read the recursion from the end. The return of the last state is the critic's value. Each earlier state takes the reward for the next step, then adds the discounted future, which is a mix: a fraction 1−λ1-\lambda of the critic's value of the next state, and a fraction λ\lambda of the next state's λ-return. At λ=0\lambda = 0 the target is one predicted reward plus the critic's value of the next state (one-step temporal-difference learning). At λ=1\lambda = 1 it is the sum of all imagined rewards plus the discounted critic value at step 15. DreamerV3 uses λ=0.95\lambda = 0.95, which unrolls into a weighted average of all the nn-step returns for n=1,…,15n = 1, \dots, 15, with weight (1−λ)λn−1(1-\lambda)\lambda^{n-1} on the nn-step one and the leftover λ14≈0.49\lambda^{14} \approx 0.49 on the full 15-step one. The continue flag ct+1c_{t+1} zeroes everything past a predicted episode end.

A three-step example with γ=0.997\gamma = 0.997, λ=0.95\lambda = 0.95, no episode end, rewards r1=r2=0r_1 = r_2 = 0 and r3=1r_3 = 1, and critic values v1=0.5v_1 = 0.5, v2=0.6v_2 = 0.6, v3=0.2v_3 = 0.2:

R3=v3=0.2R2=1+0.997 (0.05⋅0.2+0.95⋅0.2)=1.199R1=0+0.997 (0.05⋅0.6+0.95⋅1.199)=1.166R0=0+0.997 (0.05⋅0.5+0.95⋅1.166)=1.129\begin{aligned} R_3 &= v_3 = 0.2 \\ R_2 &= 1 + 0.997\,(0.05 \cdot 0.2 + 0.95 \cdot 0.2) = 1.199 \\ R_1 &= 0 + 0.997\,(0.05 \cdot 0.6 + 0.95 \cdot 1.199) = 1.166 \\ R_0 &= 0 + 0.997\,(0.05 \cdot 0.5 + 0.95 \cdot 1.166) = 1.129 \end{aligned}

The reward at step 3 reaches the start state almost undiminished, while the critic's values (0.5 and 0.6, too low) contribute only 5% each. In code the recursion is a backward loop:

def lambda_return(r, c, v, lam=0.95):
    # r, c, v: predictions for imagined states 0..15 (index 0 = start state)
    # c already carries the discount: the continue head is trained on
    # (1 - terminal) * 0.997
    R = [v[15]]                                    # R_15 = v_15
    for t in reversed(range(15)):                  # t = 14, ..., 0
        R.insert(0, r[t + 1] + c[t + 1] * ((1 - lam) * v[t + 1] + lam * R[0]))
    return R[:15]                                  # targets for states 0..14

Figure 3 runs the same recursion over a full 15-step trajectory with two rewards and a critic whose values are roughly right but smoothed across steps. Slide λ to 0 and each target is one reward plus the critic's next value, so the critic's errors pass straight into the targets. Slide it to 1 and the targets follow the rewards exactly up to the last step, at the cost of depending on 15 imagined steps of the world model.

Figure 3 · λ-returns over an imagined trajectory
0.95
0
λ-return targets, the critic values they are built from, and the imagined rewards (made-up numbers). The readout expands equation (5) for the selected step; click a column or use the second slider to move it.

The critic is trained by maximum likelihood on these targets:

L(ψ)≐−∑t=1Tln⁡pψ(Rtλ∣st)\mathcal{L}(\psi) \doteq -\sum_{t=1}^{T} \ln p_\psi(R^\lambda_t \mid s_t)

Three details keep it stable:

The output weights of the reward head and the critic are initialized to zero. A randomly initialized head predicts large random rewards and values at the start of training, and the actor would be trained toward them; with zero output weights every prediction starts at exactly 0.

The actor: why returns are normalized and advantages are not

The actor is trained with REINFORCE, the basic policy-gradient estimator: increase the log-probability of each imagined action in proportion to its advantage, how much better the λ-return turned out than the critic expected. An entropy bonus keeps the policy from becoming deterministic too early, so it keeps trying actions. DreamerV3 uses REINFORCE for discrete and continuous actions alike; a continuous action is drawn from a normal distribution whose mean is squashed into [−1, 1] by a tanh and whose standard deviation is kept between 0.1 and 1. The loss, with the sign of the entropy term as the official code has it:

L(θ)≐−∑t=1T[sg⁡ ⁣(Rtλ−vψ(st)max⁡(1,S))ln⁡πθ(at∣st)+η H[πθ(at∣st)]]\mathcal{L}(\theta) \doteq -\sum_{t=1}^{T}\Big[\operatorname{sg}\!\Big(\frac{R^\lambda_t - v_\psi(s_t)}{\max(1, S)}\Big)\ln \pi_\theta(a_t \mid s_t) + \eta\,\mathrm{H}\big[\pi_\theta(a_t \mid s_t)\big]\Big]
(6)

Minimizing it raises the probability of actions with positive advantage, lowers it for negative advantage, and raises the entropy H\mathrm{H}. The entropy scale is fixed at η=3×10−4\eta = 3 \times 10^{-4} for every domain, and that only works if the advantage term has a comparable size everywhere. Raw returns do not: Atari scores run into the hundreds of thousands, a robot's reward per step is at most 1, and BSuite deliberately rescales its rewards to test exactly this. DreamerV3 divides the advantage by max⁡(1,S)\max(1, S), where SS is the spread of recent returns:

S≐EMA⁡(Per⁡(Rtλ,95)−Per⁡(Rtλ,5), 0.99)S \doteq \operatorname{EMA}\big(\operatorname{Per}(R^\lambda_t, 95) - \operatorname{Per}(R^\lambda_t, 5),\ 0.99\big)
(7)

Per⁡(R,95)\operatorname{Per}(R, 95) is the 95th percentile of the λ-returns in the batch, so SS is the width of the middle 90% of returns, smoothed across batches with decay 0.99.

The scale is measured on the returns, which reflect how much reward is reachable, and then applied to the advantages. Subtracting a constant from every return would not change the actor's gradient, so dividing by a range is enough; no mean is subtracted.

The floor at 1 is for sparse rewards. Large returns are scaled down into roughly [0,1][0, 1]; small ones are left alone. If the agent has found no reward yet, every return is zero plus the critic's noise, with a standard deviation of perhaps 0.01. Dividing by that, as standard deviation normalization does, multiplies the noise by 100 and gives the actor a large gradient in a random direction, which drowns out the entropy bonus and stops exploration. With the floor the noise stays at its own size and the entropy term keeps the policy random until a real reward shows up.

The range uses percentiles instead of the minimum and maximum. In randomized environments a few episodes can be worth far more than the rest (a level that happens to be easy). Dividing by the full range would let one such episode shrink everyone else's advantages. The 5th-to-95th percentile range ignores the top and bottom 5%.

In Figure 4, pick a batch and a normalizer and watch the bottom axis, which shows the spread of the normalized returns on a log scale next to η\eta. With dense rewards every normalizer brings the spread to within a factor of four of 1. With "sparse, nothing found" the standard-deviation normalizer amplifies pure noise by a factor of about 100 while max⁡(1,S)\max(1, S) leaves it alone. With "rare big episodes" the full-range normalizer divides by about 200 because of 2% of the batch, while SS stays near 8. The reward-scale slider shows that above a spread of 1, DreamerV3's normalized spread stays at exactly 1.

Figure 4 · return normalization
×1.00
Raw returns of one synthetic batch with the 5th and 95th percentiles dashed, and the normalized spread on a log axis, next to the entropy scale η=3×10−4\eta = 3 \times 10^{-4}. The readout says whether the denominator shrinks the returns or amplifies them.

The paper compares this against the usual alternatives and reports that none had stable hyperparameters across domains: advantage normalization (PPO's default) puts a fixed weight on the return term whether or not reward is within reach; normalizing rewards or returns by their standard deviation fails under sparse rewards as above; and constrained entropy targets (as in SAC's automatic temperature) are robust but explore slowly under sparse rewards and converge lower under dense ones.

Predicting numbers of unknown size: symlog and twohot

Three heads predict raw numbers: the decoder for vector observations, the reward predictor, and the critic. With fixed hyperparameters, the scale of their targets, anywhere from 0.01 to 100,000, is not known in advance. A squared-error loss on a target of 100,000 has a gradient of size 100,000 at initialization, large enough to swamp every other loss on the shared network. Absolute or Huber losses keep the gradient bounded but learn large values slowly. Normalizing targets by running statistics makes the regression target drift as the statistics change. DreamerV3 handles the vector observations with a transform and the stochastic targets with a classification loss.

The transform is symlog, a logarithm that works on both sides of zero:

symlog⁡(x)≐sign⁡(x)ln⁡(∣x∣+1),symexp⁡(x)≐sign⁡(x)(exp⁡(∣x∣)−1)\operatorname{symlog}(x) \doteq \operatorname{sign}(x)\ln(|x| + 1), \qquad \operatorname{symexp}(x) \doteq \operatorname{sign}(x)\big(\exp(|x|) - 1\big)
(9)

Near zero, ln⁡(∣x∣+1)≈∣x∣\ln(|x|+1) \approx |x|, so small targets are left nearly unchanged; far from zero it grows like a logarithm, so 10610^6 becomes 13.8 and −50-50 becomes −3.93-3.93. A plain log would fail on negative and zero targets. The network predicts in symlog space and the prediction is read out through the inverse:

L(θ)≐12(f(x,θ)−symlog⁡(y))2,y^≐symexp⁡(f(x,θ))\mathcal{L}(\theta) \doteq \tfrac{1}{2}\big(f(x, \theta) - \operatorname{symlog}(y)\big)^2, \qquad \hat y \doteq \operatorname{symexp}\big(f(x, \theta)\big)
(8)

Dreamer applies symlog to vector observations twice: to the encoder's inputs, so an input of 300 enters the network as 5.71, and to the decoder's targets, so its reconstruction gradients stay small next to the representation loss.

Rewards and returns are stochastic: the same state can lead to a reward of 0 or of 10. A single-number output can only represent the mean, and a return distribution can have two modes whose mean is a value that never occurs. For these the network instead outputs a softmax over 255 fixed bins, and its prediction is the expected bin value:

y^≐softmax⁡(f(x))⊤B,B≐symexp⁡([−20,…,+20])\hat y \doteq \operatorname{softmax}(f(x))^{\top} B, \qquad B \doteq \operatorname{symexp}\big([-20, \dots, +20]\big)
(10)

The bins are evenly spaced in symlog units (0.157 apart) and so exponentially spaced in real units: 0.17 apart near zero, about 1 apart near 5, about 16 apart near 100, and reaching ±4.85×108\pm 4.85 \times 10^{8} at the ends. The training target is the twohot encoding of yy, a probability vector that is zero everywhere except at the two bins that bracket yy, weighted so that their average is exactly yy:

twohot⁡(x)i≐{∣bk+1−x∣ / ∣bk+1−bk∣if i=k∣bk−x∣ / ∣bk+1−bk∣if i=k+10elsek≐∑j=1∣B∣δ(bj<x)\operatorname{twohot}(x)_i \doteq \begin{cases} |b_{k+1} - x| \,/\, |b_{k+1} - b_k| & \text{if } i = k \\ |b_k - x| \,/\, |b_{k+1} - b_k| & \text{if } i = k+1 \\ 0 & \text{else} \end{cases} \qquad k \doteq \sum_{j=1}^{|B|} \delta(b_j < x)
(12)

For y=5y = 5 the bracketing bins (numbered from 0) are 138 at 4.654 and 139 at 5.618. The weights are (5.618−5)/0.964=0.641(5.618 - 5)/0.964 = 0.641 and 0.3590.359, and 0.641⋅4.654+0.359⋅5.618=5.0000.641 \cdot 4.654 + 0.359 \cdot 5.618 = 5.000. The loss is the cross-entropy between that vector and the network's softmax:

L(θ)≐−twohot⁡(y)⊤log⁡softmax⁡(f(x,θ))\mathcal{L}(\theta) \doteq -\operatorname{twohot}(y)^{\top} \log \operatorname{softmax}\big(f(x, \theta)\big)
(11)

The bin positions enter only through the target weights. The gradient of a softmax cross-entropy with respect to the logits is the predicted probabilities minus the target probabilities, whose length is at most 2\sqrt{2} no matter how large yy is. A target of a million produces the same size of gradient as a target of one. Because the expected value can fall anywhere between bins, the head can still output any continuous value in the range. Figure 5 shows the two bins and their weights for any target, and compares the gradient at an untrained output for the three losses: drag the target toward a million and watch the squared-error bar grow by six orders of magnitude while the twohot bar stays below 1.

Figure 5 · symlog bins and the twohot target
5.000
Top: all 255 bins on the symlog axis and the target. Middle: a zoom on the neighboring bins in real units, with the twohot weights. Bottom: gradient size at an untrained output for each loss, on a log scale.

Earlier agents handled reward scale in ways DreamerV3 avoids: DQN clipped Atari rewards to [−1,1][-1, 1], which throws away the difference between small and large rewards, and PopArt rescales the network's output layer when it sees a new extreme value. In the code, the expected value is summed in a fixed order (positive and negative bins separately, small to large) so that a uniform softmax over symmetric bins gives exactly 0 rather than a rounding error of the size of the largest bin.

What DreamerV3 achieved

Every number below uses the same hyperparameters. The comparison methods are mostly the best published specialist for each benchmark, plus PPO with one configuration tuned for all domains (the authors check it against the original PPO on ProcGen, where it scores 42.80 against the published 41.16).

Benchmark (budget)MetricDreamerV3Best comparedPPO
Atari, 57 games (200M frames)gamer median830%MuZero 693%180%
ProcGen, 16 games (50M)normalized mean66.01PPG 64.8942.80
DMLab, 30 tasks (100M)human mean capped71.4%IMPALA at 1B: 66.3%35.9%
Atari100k, 26 games (400K)gamer mean125%IRIS 105%11%
Proprio control, 18 tasks (500K)task mean871DMPO 80194
Visual control, 20 tasks (1M)task mean861DrQ-v2 77094
BSuite, 23 environmentstask mean66%Boot DQN 60%49%
Minecraft Diamond (100M)episode return9.1IMPALA 7.15.1

"Gamer median 830%" means that on the median Atari game, Dreamer's score minus a random policy's is 8.3 times a professional human tester's score minus random. The DMLab comparison gives the baseline ten times as much data: Dreamer at 100M steps beats IMPALA at 1B, though IMPALA and R2D2+ at 10B steps (85.1% and 85.4%) remain higher. Each Dreamer agent ran on one Nvidia A100, and one 200M-frame Atari run took 7.7 GPU-days.

In the Minecraft Diamond task, every episode starts in a new randomly generated world and lasts until the player dies or 36,000 steps pass (30 minutes at 20 actions per second). The reward is +1 the first time in an episode the agent obtains each of 12 items on the way to a diamond: log, plank, stick, crafting table, wooden pickaxe, cobblestone, stone pickaxe, iron ore, furnace, iron ingot, iron pickaxe, diamond. The agent sees a 64×64 image plus its inventory, and has to learn a chain of crafting steps from those sparse rewards with no demonstrations. Dreamer's average return of 9.1 at 100M steps means it typically gets about nine items deep, around the furnace. IMPALA and Rainbow, whose learning rates and entropy scales the authors tuned for this task, reach 7.1 and 6.3, and PPO 5.1. All 10 Dreamer runs collected at least one diamond during training and none of the baselines did. The previous diamond result, OpenAI's VPT, needed recorded human gameplay and 720 GPUs for 9 days; Dreamer used 1 GPU for 9 days.

Removing the robustness techniques one at a time on 14 tasks lowers the average score in every case; the largest drop comes from removing KL balancing and free bits, then return normalization, then the symexp twohot loss. And stopping the reconstruction gradients hurts far more than stopping the reward and value gradients into the world model, so Dreamer's representations come mostly from predicting observations, and the reward and value gradients add little. Scaling the model from 12M to 400M parameters raises the final score and also reduces the number of environment steps needed, and more gradient steps per environment step (the replay ratio) also help monotonically, so more compute buys better results without a new hyperparameter search.

What the results do and do not show

"Fixed hyperparameters" covers every value in Table 4, but two settings still change between benchmarks: the replay ratio (from 32 on Atari to 1024 on BSuite) and the model size (12M parameters on the control suites, 200M elsewhere). The paper picks them to fit each benchmark's step budget rather than tuning them for score, and its scaling experiments on Crafter and DMLab show larger values helping monotonically.

The diamond result is a discovery rate, not a success rate. At the 100M-step budget, 0.4% of Dreamer's episodes end with a diamond; the 100% figure counts runs that found at least one diamond at any point in training. The environment also follows the block-breaking setting of earlier work, in which blocks break much faster than in the normal game (the released environment uses a speed multiplier of 100), because the stochastic policy would otherwise need to hold the attack key for many consecutive steps. Crafting is a single abstract action, as in the MineRL competition.

Not every table favors Dreamer. On Atari100k its gamer median (49%) is below TWM's (51%) even though its mean is higher, and EfficientZero scores 190% with a different evaluation protocol. On BSuite's exploration category Dreamer scores 0.01 against Bootstrapped DQN's 0.68: entropy-based exploration does not solve Deep Sea, a task built to require deliberate exploration. Prioritized replay improved Dreamer in the authors' experiments, but they kept uniform replay for simplicity, so the reported numbers are not the best the method can do.

Implementation notes

Everything is trained concurrently from one loss sum with one optimizer. The optimizer is LaProp, a variant of Adam that divides the gradient by its running RMS first and applies momentum afterwards, which the authors found allows a tiny ϵ=10−20\epsilon = 10^{-20} without Adam's occasional instabilities. Gradients are clipped per tensor with adaptive gradient clipping (AGC): a tensor's gradient is scaled down if its norm exceeds 30% of the norm of the weights it belongs to, so the threshold does not depend on the scale of any loss. The learning rate is 4×10−54 \times 10^{-5} for all networks, with batches of 16 sequences of 64 steps. There is no learning-rate schedule, weight decay or dropout. Every layer uses RMSNorm (a layer normalization without the mean subtraction) and SiLU activations.

The replay ratio counts how many replayed time steps are trained on per environment step. On Atari it is 32 with an action repeat of 4: a batch covers 16×64=102416 \times 64 = 1024 steps, so one gradient step happens every 1024/32 = 32 agent steps, or 128 frames, and 200M frames make about 1.56M gradient steps. The replay buffer mixes the newest trajectories (an online queue) with uniformly sampled old ones, and stores the latent states computed during data collection so replayed sequences can start from a sensible hh.

Model sizes are set by one number, the hidden size dd of the MLPs. The GRU gets 8d8d units in 8 blocks, each latent gets d/16d/16 classes, and the convolutions' base channel count is d/16d/16; the number of layers and of latents (32) stays fixed, as do all hyperparameters.

Parameters12M25M50M100M200M400M
Hidden size d25638451276810241536
Recurrent units 8d2048307240966144819212288
Classes per latent d/16162432486496

The 12M column's 2048 recurrent units follow the 8d8d rule and the code; the paper's Table 3 prints 1024 there. The official implementation (JAX) is at github.com/danijar/dreamerv3, and the continue head there is trained on (1−terminal)×0.997(1 - \text{terminal}) \times 0.997, so the discount rides inside the predicted continue flag.

Provenance Verified against primary literatureHow we verify
Hafner et al. (2023), arXiv v2Equations (1)-(12), Tables 2-4 and all benchmark numbers on this page are from v2 (April 2024). The Nature version (Hafner et al., 2025, "Mastering diverse control tasks through world models") prints the same equations (5) and (6).
Official code, github.com/danijar/dreamerv3Read for every constant the paper leaves implicit. agent.py imag_loss: policy loss = -(logpi * adv + 3e-4 * entropy), adv = (R - v) / max(1, P95 - P5). agent.py lambda_return: uses rew[t+1] and val[t+1]. rssm.py loss: KL(sg(post) || prior) and KL(post || sg(prior)), each summed over the 32 latents and then clipped at 1 nat. configs.yaml: 255 bins, imag_length 15, horizon 333, lr 4e-5, agc 0.3, unimix 0.01, slowvalue rate 0.02.
Hafner et al. (2020), DreamerV2KL balancing (weight 0.8 on the prior, 0.2 on the posterior), categorical latents with straight-through gradients, and the lambda-return with v(z_{t+1}) (Eq. 4) that DreamerV3 inherits. Its actor loss also subtracts the entropy term (Eq. 6).
Kingma et al. (2016), Appendix C.8Free bits: a floor on the KL per group of latents, below which it is not penalized. DreamerV3 applies one floor of 1 nat to the KL summed over all 32 latents of a time step.
Hafner et al. (2023), Tables 6-13Atari gamer median 830% vs MuZero 693%; ProcGen 66.01 vs PPG 64.89; DMLab 71.4% at 100M vs IMPALA 66.3% at 1B (R2D2+ 85.4% at 10B); Atari100k mean 125% vs IRIS 105% (median 49% vs TWM 51%); BSuite 66% vs Boot DQN 60%, exploration category 0.01 vs 0.68.
Hafner et al. (2023), Figure 9 and Table 5Minecraft return 9.1 at 100M steps (IMPALA 7.1, Rainbow 6.3, PPO 5.1). All 10 Dreamer runs collect at least one diamond during training; at the 100M budget 0.4% of episodes end with a diamond. The released environment sets a block-breaking speed multiplier of 100 (embodied/envs/minecraft_flat.py).
correctionThree slips in the paper, all checked against the official code. (1) Equation (6) adds the entropy bonus, +ηH[π], to a loss that is minimized, which would push the policy toward determinism. The code subtracts it (agent.py: -(logpi * adv + actent * ent)), as do DreamerV2 Eq. 6 and DreamerV3 v1 Eq. 11; this page writes −ηH. (2) Equation (5) bootstraps R^λ_t from v_t, the critic's value of the same state whose return it defines. v1 (Eq. 7) and DreamerV2 (Eq. 4) use the next state, and the code's lambda_return uses rew[t+1] and val[t+1]; this page writes r_{t+1}, c_{t+1} and v_{t+1}. Both slips also appear in the 2025 Nature version. (3) Table 3 lists 1024 recurrent units for the 12M model, but the rule in the same table (8 × hidden size 256) and the code (size12m: deter 2048) give 2048.

Questions you might still have

?

Does Dreamer plan ahead when it plays?
No. To act, it runs the encoder and the recurrent model on the current observation and samples an action from the actor network, one forward pass with no lookahead search. The world model is used only during training, to generate the imagined trajectories the actor and critic learn from. MuZero, by contrast, runs a tree search with its learned model at every move.

?

Is it really one set of hyperparameters?
Every value in Table 4 (learning rate, loss scales, entropy scale, discount, horizon, unimix, free bits) is the same on all benchmarks. Two things do change per benchmark, chosen to fit the compute budget: the replay ratio (32 on Atari and Minecraft, 512 on the control suites, 1024 on BSuite) and the model size (200M parameters, or 12M on the two control suites). The paper argues both only trade compute for data efficiency, and Figure 6 shows that larger values help monotonically rather than needing a search.

?

Why does the actor-critic ignore the real rewards?
The actor is trained only on imagined trajectories, so it sees predicted rewards. The real rewards still matter: they train the reward predictor, and the critic also gets a second loss on replayed real trajectories (scale 0.3) whose lambda-returns use the recorded rewards, bootstrapped from the imagination returns at the start states.

?

What happens when the world model is wrong?
The actor exploits whatever the model predicts, so errors compound with the number of imagined steps. DreamerV3 limits imagination to 15 steps and lets the critic stand in for everything beyond, and the world model keeps training on fresh data the actor collects, which includes the states where its predictions were wrong.

?

How is this different from the 2018 World Models paper?
Ha and Schmidhuber trained a VAE, then an MDN-RNN, then a tiny linear controller with an evolution strategy, in separate stages. Dreamer trains the world model, critic and actor concurrently from one replay buffer, with gradient-based actor-critic learning on imagined trajectories, and the world model is a single recurrent state-space model rather than a separate VAE and RNN.

?

Why not normalize advantages, as PPO does?
Normalizing advantages divides by their standard deviation every batch, which puts the same weight on the return term whether rewards are near or not. When no reward has been found, the advantages are mostly critic noise, and dividing by a tiny standard deviation blows that noise up until it outweighs the entropy bonus, so the policy stops exploring. DreamerV3 divides returns by their 5-95 percentile range only when the range exceeds 1.

?

What do the 1% unimix and the 255 bins buy?
Unimix mixes every categorical with 1% uniform, so no class ever has probability below 0.01/64 = 0.00016 and the KL of one latent is capped at ln(6400) = 8.76 nats. The 255 bins span about plus or minus 485 million in real units while staying 0.157 apart in symlog units, so the reward and value heads cover any scale with one output layer.

Footnotes & further reading

  1. The paper: Hafner, Pasukonis, Ba & Lillicrap, Mastering Diverse Domains through World Models (arXiv 2023, v2 April 2024), published as Mastering diverse control tasks through world models (Nature 640, 2025). Code: github.com/danijar/dreamerv3.
  2. Earlier Dreamers: Hafner et al., Dream to Control (ICLR 2020) and Mastering Atari with Discrete World Models (ICLR 2021), which introduced categorical latents and KL balancing. The RSSM is from PlaNet (ICML 2019).
  3. Free bits: Kingma et al., Improved Variational Inference with Inverse Autoregressive Flow (NeurIPS 2016), Appendix C.8. The VAE background is in the VAE explainer.
  4. Twohot targets for value regression appear in MuZero (see the MuZero explainer); the symlog transform belongs to the bi-symmetric log family of Webber, A bi-symmetric log transformation for wide-range data (2012).
  5. Optimizer pieces: Ziyin et al., LaProp (2020), and Brock et al., High-Performance Large-Scale Image Recognition Without Normalization (ICML 2021), which introduced adaptive gradient clipping.
  6. Minecraft baselines: Baker et al., Video PreTraining (VPT) (2022), and Kanitscheider et al., Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft (2021), whose block-breaking setting Dreamer follows.