DreamerV3: Mastering Diverse Domains through World Models
An agent learns a model of its environment, practices inside it, and one fixed set of hyperparameters works on over 150 tasks.
DreamerV3 trains three networks at once: a world model that predicts what happens next, a critic that estimates how good a state is, and an actor that picks actions. The actor and critic learn only from trajectories the world model imagines. A handful of normalizations keep every loss on the same scale, so the configuration that plays Atari also controls simulated robots and collects diamonds in Minecraft.
Explaining the paperMastering Diverse Domains through World ModelsThe first agent to collect a diamond in Minecraft from scratch, with no human demonstrations, on one GPU in nine days, using the same settings it uses for Atari.
The problem: one set of hyperparameters for every task
Reinforcement learning (RL) trains an agent by trial and error. At each step the agent receives an observation (a camera image, or a vector of joint angles), chooses an action , and the environment returns a reward , a number that is high when something good happened. The agent's goal is to maximize the return, the sum of future rewards with each one discounted by how far away it is:
With a reward 333 steps away still counts for of its face value, so the agent cares about roughly the next 333 steps (the paper calls the discount horizon). A policy, called the actor, maps the current state to a distribution over actions. A value function, called the critic, estimates the expected return from a state under the current policy. Actor-critic methods train both: the critic learns to predict returns, and the actor shifts probability toward actions that turned out better than the critic expected.
The algorithms that work well are specialized. DQN and its descendants handle discrete actions, SAC handles continuous control, MuZero plans with a learned model on board games and Atari. Each needs its hyperparameters retuned when the domain changes: the entropy bonus that makes an agent explore under sparse rewards makes it dither under dense ones, a squared-error loss that suits rewards of size 1 can diverge on rewards of size 1000, and a representation-loss weight chosen for Atari's static backgrounds is wrong for 3D scenes. PPO is the common fallback because it is comparatively robust, but it needs a lot of experience and usually loses to the specialists.
DreamerV3, by Danijar Hafner, Jurgis Pasukonis, Jimmy Ba and Timothy Lillicrap, sets out to remove the retuning. The paper evaluates one configuration on 8 domains and over 150 tasks: Atari, ProcGen, DMLab, Atari100k, continuous control from joint angles and from pixels, BSuite, and Minecraft. It is the third version of Dreamer. Version 1 (2019) handled continuous control and version 2 (2020) Atari; version 3 keeps their structure and adds the robustness techniques this page spends most of its length on.
How DreamerV3 works
The agent has three neural networks and a replay buffer, a store of everything it has experienced (up to 5 million steps). Three things happen concurrently:
- The actor plays in the real environment, and every step goes into the replay buffer.
- The world model trains on sequences sampled from the buffer to predict, from its compact internal state, the next internal state, the reward, whether the episode continues, and the observation itself.
- From each replayed state the world model and actor imagine 15 steps ahead. The critic learns to predict returns on those imagined trajectories, and the actor learns to pick actions with high predicted return.
Training a policy on imagined experience is an old idea (Sutton's Dyna, 1991, and the 2018 World Models paper of Ha and Schmidhuber both do it). Think of a flight simulator built from the pilot's own flight logs: practice inside it is cheap, but the simulator is only as good as the logs, and a pilot who trains too long in it learns the simulator's bugs. Dreamer keeps imagined trajectories short (15 steps) and keeps refitting the model to fresh experience for the same reason.
Here is a full update with the tensor shapes of the default 200M-parameter model; the sections below go through it line by line:
# One DreamerV3 update. 200M model, Table 4 defaults, shapes per batch.
batch = replay.sample(B=16, T=64) # images, actions, rewards, flags
# 1. World model: run the RSSM over the replayed sequence (eq. 1)
h, z, post, prior = rssm.observe(batch) # h: 16x64x8192, z: 16x64x32x64
L_pred = (decoder.nll(batch.x, h, z) # images: MSE; vectors: symlog MSE
+ reward_head.twohot_ce(batch.r, h, z)
+ cont_head.bce(batch.cont, h, z))
L_dyn = max(1, KL(sg(post) || prior)) # moves the dynamics predictor
L_rep = max(1, KL(post || sg(prior))) # moves the encoder
L_model = 1.0 * L_pred + 1.0 * L_dyn + 0.1 * L_rep
# 2. Imagination: 15 steps from each of the 16*64 = 1024 replayed states
traj = rssm.imagine(sg(h, z), actor, steps=15) # 1024 x 16 model states
r, c, v = reward_head(traj), cont_head(traj), critic(traj)
R = lambda_return(r, c, v, lam=0.95) # eq. 5, backward over 15 steps
# 3. Actor and critic, trained on imagined states only
S = ema(percentile(R, 95) - percentile(R, 5), decay=0.99)
adv = sg((R - v[:, :-1]) / max(1, S))
L_actor = -(adv * actor.logp(traj.a) + 3e-4 * actor.entropy(traj))
L_critic = (critic.twohot_ce(sg(R)) # eq. 5
+ critic.twohot_ce(sg(critic_ema(traj)))) # pull toward EMA
L_critic_replay = 0.3 * critic_loss_on_replayed_states(batch)
# 4. One optimizer step on everything: LaProp, lr 4e-5, AGC at 0.3
step(L_model + L_actor + L_critic + L_critic_replay)sg is stop-gradient: the value passes forward, but no gradient flows back through it. In the listing it has three jobs: it keeps the actor and critic losses from changing the world model, it splits the world model's KL term into two losses that each move one network, and it turns the advantages and the critic's targets into constants, so gradients reach the actor and critic only through their own outputs.
The world model: a recurrent state-space model
The world model is a recurrent state-space model (RSSM), introduced by Hafner and colleagues in 2018 for the PlaNet agent. Its state at time has two parts. The recurrent state is a deterministic vector that carries memory forward, like the hidden state of any recurrent network. The stochastic state is a sample from a distribution, and holds what the model learned from the current observation. The model state is the pair . Six components connect them:
All six share the parameters . The sequence model is a GRU, a gated recurrent cell, and it advances one step using the previous stochastic state and action. The encoder looks at the current observation (and ) and outputs a distribution over . The dynamics predictor outputs a distribution over the same from alone, without the observation: it is the model's guess of what it is about to see. The other three heads read the full model state and predict the reward, the continue flag (0 when the episode ends), and the observation.
The shapes at the default 200M size, for one 64×64 RGB frame:
- : 8192 floats. The GRU's weights are block-diagonal with 8 blocks of 1024 units, so the parameter count grows like instead of , an eighth.
- : 32 categorical variables with 64 classes each. A sample is 32 one-hot vectors, flattened to 2048 numbers of which exactly 32 are 1.
- The encoder is a stack of stride-2 convolutions that shrinks the image from 64×64 to 4×4 before flattening; the decoder mirrors it with transposed convolutions. Vector observations go through 3-layer MLPs instead.
- The actor, critic, reward and continue heads all read the 8192 + 2048 = 10,240 numbers of the model state.
The stochastic state is categorical because DreamerV2 found that categorical latents beat Gaussian ones on Atari. Sampling a class is not differentiable, so Dreamer uses straight-through gradients: the forward pass uses the one-hot sample, and the backward pass treats it as if it were the vector of class probabilities. In the code this is one line, sample + probs - sg(probs), whose value equals sample and whose gradient is that of probs.
Equation (1) produces in two ways. While the world model trains on replayed data, it observes: each comes from the encoder, which sees the real frame. During imagination there are no frames, so each comes from the dynamics predictor, and the actions come from the actor instead of the replay buffer. In Figure 1, drag the observed-steps slider to move the boundary, and click a column to see what that step is computed from.
Imagination never touches pixels: from a start state, one imagined step is one GRU update, one dynamics-predictor sample, and the reward and continue heads. The decoder is not run during behavior learning. That makes imagining 16,384 states per update (1024 start states times 16) cheap enough to do at every training step, and it is why the dynamics predictor must produce latents that look like the encoder's: the actor was trained on imagined latents but acts on encoded ones.
Training the world model: prediction, dynamics and representation losses
Given a replayed batch of 16 sequences of 64 steps, the world model minimizes three losses, with weights , and :
The prediction loss trains the heads to reconstruct the observation, predict the reward, and predict whether the episode continues. Images use a squared error, vector observations a squared error on symlog-transformed targets, the reward head the twohot loss of a later section, and the continue head a logistic (binary cross-entropy) loss.
The other two terms compare the encoder's distribution (which saw the frame) with the dynamics predictor's (which did not) using the KL divergence, a measure in nats of how far one distribution is from another; it is zero only when they are equal. This is the structure of a variational autoencoder, with the dynamics predictor playing the role of a learned prior. A single KL term would move both networks toward each other with equal force. DreamerV3 writes it twice with the stop-gradient in different places, so each copy moves only one network:
- stops the gradient into , so it trains only the dynamics predictor, to predict the encoder's output better. Weight 1.
- stops the gradient into , so it trains only the encoder, to make its representations easier to predict. Weight 0.1.
This is KL balancing, from DreamerV2 (which used a 0.8/0.2 split). The weights are unequal because a better predictor is always useful, while a more predictable representation is useful only up to a point: past it the encoder makes itself predictable by throwing information away: in the limit, an encoder that outputs the same for every frame is perfectly predictable and useless.
Free bits keep the encoder away from that limit. The clips each KL term at 1 nat (about 1.44 bits). While the KL is below 1 nat, the loss is the constant 1 and its gradient is zero, so neither network is pushed further; the prediction loss alone shapes the representation. The idea is from Kingma and colleagues (2016). In the official code the KL is first summed over the 32 latents of a time step and then clipped, so the free allowance is 1 nat for the whole state per step, not per latent.
Figure 2 is a toy version of these three terms. There is one latent with 8 classes and four possible observations to that the predictor cannot tell apart, as when the next frame depends on something random. A fixed decoder maps class to observation for the first four classes; the last four decode to nothing in particular. The best the prior can do is the average of the four posteriors, so the KL that remains equals the information the latent carries about which observation it saw. At the paper's each posterior settles on its own class and the latent keeps 1.38 nats, nearly all of . Raise to 3 and turn free bits off, and the posteriors collapse onto the prior: 0.006 nats, and the decoder is right 26% of the time, chance level for four observations. Turn free bits back on at the same and the KL stops near 1 nat, keeping 0.96 nats and a 78% decode rate.
Earlier Dreamers had to choose the representation weight per domain: a strong pull toward predictability for 3D environments full of irrelevant detail, a weak one for games with static backgrounds where single pixels matter. The paper reports that a small combined with free bits works for both, so DreamerV3 uses 0.1 everywhere.
The paper also saw occasional spikes in the KL losses, consistent with reports from deep VAEs. A categorical that puts probability near zero on a class makes the log ratio inside the KL explode. DreamerV3 mixes every categorical (encoder, dynamics predictor and actor) as 99% network output and 1% uniform, called unimix. No class can then fall below , which caps the KL of one 64-class latent at nats.
Learning in imagination: the critic and λ-returns
The actor and critic never see an observation. They read model states , and since carries the history, a single model state is meant to contain everything needed to predict the future (the paper calls the representation Markovian). Both are 3-layer MLPs:
The critic outputs a distribution over returns, not a single number; its prediction is the mean of that distribution. Starting from every state of the replayed batch, the world model and actor imagine 15 steps, giving 16 model states per start including the start itself (the paper's , the configuration's imagination horizon of 15). The reward, continue and critic heads label every imagined state.
The return of an imagined state depends on rewards up to 333 steps away, and the trajectory stops after 15. The critic fills the gap: after the last imagined step, its own estimate stands in for the rest. Dreamer blends the imagined rewards and the critic's estimates with the λ-return, written here with the indices of the official code:
Read the recursion from the end. The return of the last state is the critic's value. Each earlier state takes the reward for the next step, then adds the discounted future, which is a mix: a fraction of the critic's value of the next state, and a fraction of the next state's λ-return. At the target is one predicted reward plus the critic's value of the next state (one-step temporal-difference learning). At it is the sum of all imagined rewards plus the discounted critic value at step 15. DreamerV3 uses , which unrolls into a weighted average of all the -step returns for , with weight on the -step one and the leftover on the full 15-step one. The continue flag zeroes everything past a predicted episode end.
A three-step example with , , no episode end, rewards and , and critic values , , :
The reward at step 3 reaches the start state almost undiminished, while the critic's values (0.5 and 0.6, too low) contribute only 5% each. In code the recursion is a backward loop:
def lambda_return(r, c, v, lam=0.95):
# r, c, v: predictions for imagined states 0..15 (index 0 = start state)
# c already carries the discount: the continue head is trained on
# (1 - terminal) * 0.997
R = [v[15]] # R_15 = v_15
for t in reversed(range(15)): # t = 14, ..., 0
R.insert(0, r[t + 1] + c[t + 1] * ((1 - lam) * v[t + 1] + lam * R[0]))
return R[:15] # targets for states 0..14Figure 3 runs the same recursion over a full 15-step trajectory with two rewards and a critic whose values are roughly right but smoothed across steps. Slide λ to 0 and each target is one reward plus the critic's next value, so the critic's errors pass straight into the targets. Slide it to 1 and the targets follow the rewards exactly up to the last step, at the cost of depending on 15 imagined steps of the world model.
The critic is trained by maximum likelihood on these targets:
Three details keep it stable:
- Its output is a softmax over 255 bins, trained with the twohot loss of the symlog section, because the return distribution can have several modes and its scale differs by orders of magnitude between environments.
- The critic regresses targets built from its own predictions, so an error in its prediction feeds into its next target. A second loss term pulls its output toward that of an exponential moving average (EMA) of its own weights, with decay 0.98. This plays the role of DQN's target network but still lets the returns be computed with the current critic.
- Imagined rewards can be wrong where rewards are hard to predict, so the same loss is also applied, with weight 0.3, to real replayed trajectories: their λ-returns use the recorded rewards and bootstrap from the imagination returns at the start states.
The output weights of the reward head and the critic are initialized to zero. A randomly initialized head predicts large random rewards and values at the start of training, and the actor would be trained toward them; with zero output weights every prediction starts at exactly 0.
The actor: why returns are normalized and advantages are not
The actor is trained with REINFORCE, the basic policy-gradient estimator: increase the log-probability of each imagined action in proportion to its advantage, how much better the λ-return turned out than the critic expected. An entropy bonus keeps the policy from becoming deterministic too early, so it keeps trying actions. DreamerV3 uses REINFORCE for discrete and continuous actions alike; a continuous action is drawn from a normal distribution whose mean is squashed into [−1, 1] by a tanh and whose standard deviation is kept between 0.1 and 1. The loss, with the sign of the entropy term as the official code has it:
Minimizing it raises the probability of actions with positive advantage, lowers it for negative advantage, and raises the entropy . The entropy scale is fixed at for every domain, and that only works if the advantage term has a comparable size everywhere. Raw returns do not: Atari scores run into the hundreds of thousands, a robot's reward per step is at most 1, and BSuite deliberately rescales its rewards to test exactly this. DreamerV3 divides the advantage by , where is the spread of recent returns:
is the 95th percentile of the λ-returns in the batch, so is the width of the middle 90% of returns, smoothed across batches with decay 0.99.
The scale is measured on the returns, which reflect how much reward is reachable, and then applied to the advantages. Subtracting a constant from every return would not change the actor's gradient, so dividing by a range is enough; no mean is subtracted.
The floor at 1 is for sparse rewards. Large returns are scaled down into roughly ; small ones are left alone. If the agent has found no reward yet, every return is zero plus the critic's noise, with a standard deviation of perhaps 0.01. Dividing by that, as standard deviation normalization does, multiplies the noise by 100 and gives the actor a large gradient in a random direction, which drowns out the entropy bonus and stops exploration. With the floor the noise stays at its own size and the entropy term keeps the policy random until a real reward shows up.
The range uses percentiles instead of the minimum and maximum. In randomized environments a few episodes can be worth far more than the rest (a level that happens to be easy). Dividing by the full range would let one such episode shrink everyone else's advantages. The 5th-to-95th percentile range ignores the top and bottom 5%.
In Figure 4, pick a batch and a normalizer and watch the bottom axis, which shows the spread of the normalized returns on a log scale next to . With dense rewards every normalizer brings the spread to within a factor of four of 1. With "sparse, nothing found" the standard-deviation normalizer amplifies pure noise by a factor of about 100 while leaves it alone. With "rare big episodes" the full-range normalizer divides by about 200 because of 2% of the batch, while stays near 8. The reward-scale slider shows that above a spread of 1, DreamerV3's normalized spread stays at exactly 1.
The paper compares this against the usual alternatives and reports that none had stable hyperparameters across domains: advantage normalization (PPO's default) puts a fixed weight on the return term whether or not reward is within reach; normalizing rewards or returns by their standard deviation fails under sparse rewards as above; and constrained entropy targets (as in SAC's automatic temperature) are robust but explore slowly under sparse rewards and converge lower under dense ones.
Predicting numbers of unknown size: symlog and twohot
Three heads predict raw numbers: the decoder for vector observations, the reward predictor, and the critic. With fixed hyperparameters, the scale of their targets, anywhere from 0.01 to 100,000, is not known in advance. A squared-error loss on a target of 100,000 has a gradient of size 100,000 at initialization, large enough to swamp every other loss on the shared network. Absolute or Huber losses keep the gradient bounded but learn large values slowly. Normalizing targets by running statistics makes the regression target drift as the statistics change. DreamerV3 handles the vector observations with a transform and the stochastic targets with a classification loss.
The transform is symlog, a logarithm that works on both sides of zero:
Near zero, , so small targets are left nearly unchanged; far from zero it grows like a logarithm, so becomes 13.8 and becomes . A plain log would fail on negative and zero targets. The network predicts in symlog space and the prediction is read out through the inverse:
Dreamer applies symlog to vector observations twice: to the encoder's inputs, so an input of 300 enters the network as 5.71, and to the decoder's targets, so its reconstruction gradients stay small next to the representation loss.
Rewards and returns are stochastic: the same state can lead to a reward of 0 or of 10. A single-number output can only represent the mean, and a return distribution can have two modes whose mean is a value that never occurs. For these the network instead outputs a softmax over 255 fixed bins, and its prediction is the expected bin value:
The bins are evenly spaced in symlog units (0.157 apart) and so exponentially spaced in real units: 0.17 apart near zero, about 1 apart near 5, about 16 apart near 100, and reaching at the ends. The training target is the twohot encoding of , a probability vector that is zero everywhere except at the two bins that bracket , weighted so that their average is exactly :
For the bracketing bins (numbered from 0) are 138 at 4.654 and 139 at 5.618. The weights are and , and . The loss is the cross-entropy between that vector and the network's softmax:
The bin positions enter only through the target weights. The gradient of a softmax cross-entropy with respect to the logits is the predicted probabilities minus the target probabilities, whose length is at most no matter how large is. A target of a million produces the same size of gradient as a target of one. Because the expected value can fall anywhere between bins, the head can still output any continuous value in the range. Figure 5 shows the two bins and their weights for any target, and compares the gradient at an untrained output for the three losses: drag the target toward a million and watch the squared-error bar grow by six orders of magnitude while the twohot bar stays below 1.
Earlier agents handled reward scale in ways DreamerV3 avoids: DQN clipped Atari rewards to , which throws away the difference between small and large rewards, and PopArt rescales the network's output layer when it sees a new extreme value. In the code, the expected value is summed in a fixed order (positive and negative bins separately, small to large) so that a uniform softmax over symmetric bins gives exactly 0 rather than a rounding error of the size of the largest bin.
What DreamerV3 achieved
Every number below uses the same hyperparameters. The comparison methods are mostly the best published specialist for each benchmark, plus PPO with one configuration tuned for all domains (the authors check it against the original PPO on ProcGen, where it scores 42.80 against the published 41.16).
| Benchmark (budget) | Metric | DreamerV3 | Best compared | PPO |
|---|---|---|---|---|
| Atari, 57 games (200M frames) | gamer median | 830% | MuZero 693% | 180% |
| ProcGen, 16 games (50M) | normalized mean | 66.01 | PPG 64.89 | 42.80 |
| DMLab, 30 tasks (100M) | human mean capped | 71.4% | IMPALA at 1B: 66.3% | 35.9% |
| Atari100k, 26 games (400K) | gamer mean | 125% | IRIS 105% | 11% |
| Proprio control, 18 tasks (500K) | task mean | 871 | DMPO 801 | 94 |
| Visual control, 20 tasks (1M) | task mean | 861 | DrQ-v2 770 | 94 |
| BSuite, 23 environments | task mean | 66% | Boot DQN 60% | 49% |
| Minecraft Diamond (100M) | episode return | 9.1 | IMPALA 7.1 | 5.1 |
"Gamer median 830%" means that on the median Atari game, Dreamer's score minus a random policy's is 8.3 times a professional human tester's score minus random. The DMLab comparison gives the baseline ten times as much data: Dreamer at 100M steps beats IMPALA at 1B, though IMPALA and R2D2+ at 10B steps (85.1% and 85.4%) remain higher. Each Dreamer agent ran on one Nvidia A100, and one 200M-frame Atari run took 7.7 GPU-days.
In the Minecraft Diamond task, every episode starts in a new randomly generated world and lasts until the player dies or 36,000 steps pass (30 minutes at 20 actions per second). The reward is +1 the first time in an episode the agent obtains each of 12 items on the way to a diamond: log, plank, stick, crafting table, wooden pickaxe, cobblestone, stone pickaxe, iron ore, furnace, iron ingot, iron pickaxe, diamond. The agent sees a 64×64 image plus its inventory, and has to learn a chain of crafting steps from those sparse rewards with no demonstrations. Dreamer's average return of 9.1 at 100M steps means it typically gets about nine items deep, around the furnace. IMPALA and Rainbow, whose learning rates and entropy scales the authors tuned for this task, reach 7.1 and 6.3, and PPO 5.1. All 10 Dreamer runs collected at least one diamond during training and none of the baselines did. The previous diamond result, OpenAI's VPT, needed recorded human gameplay and 720 GPUs for 9 days; Dreamer used 1 GPU for 9 days.
Removing the robustness techniques one at a time on 14 tasks lowers the average score in every case; the largest drop comes from removing KL balancing and free bits, then return normalization, then the symexp twohot loss. And stopping the reconstruction gradients hurts far more than stopping the reward and value gradients into the world model, so Dreamer's representations come mostly from predicting observations, and the reward and value gradients add little. Scaling the model from 12M to 400M parameters raises the final score and also reduces the number of environment steps needed, and more gradient steps per environment step (the replay ratio) also help monotonically, so more compute buys better results without a new hyperparameter search.
What the results do and do not show
"Fixed hyperparameters" covers every value in Table 4, but two settings still change between benchmarks: the replay ratio (from 32 on Atari to 1024 on BSuite) and the model size (12M parameters on the control suites, 200M elsewhere). The paper picks them to fit each benchmark's step budget rather than tuning them for score, and its scaling experiments on Crafter and DMLab show larger values helping monotonically.
The diamond result is a discovery rate, not a success rate. At the 100M-step budget, 0.4% of Dreamer's episodes end with a diamond; the 100% figure counts runs that found at least one diamond at any point in training. The environment also follows the block-breaking setting of earlier work, in which blocks break much faster than in the normal game (the released environment uses a speed multiplier of 100), because the stochastic policy would otherwise need to hold the attack key for many consecutive steps. Crafting is a single abstract action, as in the MineRL competition.
Not every table favors Dreamer. On Atari100k its gamer median (49%) is below TWM's (51%) even though its mean is higher, and EfficientZero scores 190% with a different evaluation protocol. On BSuite's exploration category Dreamer scores 0.01 against Bootstrapped DQN's 0.68: entropy-based exploration does not solve Deep Sea, a task built to require deliberate exploration. Prioritized replay improved Dreamer in the authors' experiments, but they kept uniform replay for simplicity, so the reported numbers are not the best the method can do.
Implementation notes
Everything is trained concurrently from one loss sum with one optimizer. The optimizer is LaProp, a variant of Adam that divides the gradient by its running RMS first and applies momentum afterwards, which the authors found allows a tiny without Adam's occasional instabilities. Gradients are clipped per tensor with adaptive gradient clipping (AGC): a tensor's gradient is scaled down if its norm exceeds 30% of the norm of the weights it belongs to, so the threshold does not depend on the scale of any loss. The learning rate is for all networks, with batches of 16 sequences of 64 steps. There is no learning-rate schedule, weight decay or dropout. Every layer uses RMSNorm (a layer normalization without the mean subtraction) and SiLU activations.
The replay ratio counts how many replayed time steps are trained on per environment step. On Atari it is 32 with an action repeat of 4: a batch covers steps, so one gradient step happens every 1024/32 = 32 agent steps, or 128 frames, and 200M frames make about 1.56M gradient steps. The replay buffer mixes the newest trajectories (an online queue) with uniformly sampled old ones, and stores the latent states computed during data collection so replayed sequences can start from a sensible .
Model sizes are set by one number, the hidden size of the MLPs. The GRU gets units in 8 blocks, each latent gets classes, and the convolutions' base channel count is ; the number of layers and of latents (32) stays fixed, as do all hyperparameters.
| Parameters | 12M | 25M | 50M | 100M | 200M | 400M |
|---|---|---|---|---|---|---|
| Hidden size d | 256 | 384 | 512 | 768 | 1024 | 1536 |
| Recurrent units 8d | 2048 | 3072 | 4096 | 6144 | 8192 | 12288 |
| Classes per latent d/16 | 16 | 24 | 32 | 48 | 64 | 96 |
The 12M column's 2048 recurrent units follow the rule and the code; the paper's Table 3 prints 1024 there. The official implementation (JAX) is at github.com/danijar/dreamerv3, and the continue head there is trained on , so the discount rides inside the predicted continue flag.
Questions you might still have
Does Dreamer plan ahead when it plays?
No. To act, it runs the encoder and the recurrent model on the current observation and samples an action from the actor network, one forward pass with no lookahead search. The world model is used only during training, to generate the imagined trajectories the actor and critic learn from. MuZero, by contrast, runs a tree search with its learned model at every move.
Is it really one set of hyperparameters?
Every value in Table 4 (learning rate, loss scales, entropy scale, discount, horizon, unimix, free bits) is the same on all benchmarks. Two things do change per benchmark, chosen to fit the compute budget: the replay ratio (32 on Atari and Minecraft, 512 on the control suites, 1024 on BSuite) and the model size (200M parameters, or 12M on the two control suites). The paper argues both only trade compute for data efficiency, and Figure 6 shows that larger values help monotonically rather than needing a search.
Why does the actor-critic ignore the real rewards?
The actor is trained only on imagined trajectories, so it sees predicted rewards. The real rewards still matter: they train the reward predictor, and the critic also gets a second loss on replayed real trajectories (scale 0.3) whose lambda-returns use the recorded rewards, bootstrapped from the imagination returns at the start states.
What happens when the world model is wrong?
The actor exploits whatever the model predicts, so errors compound with the number of imagined steps. DreamerV3 limits imagination to 15 steps and lets the critic stand in for everything beyond, and the world model keeps training on fresh data the actor collects, which includes the states where its predictions were wrong.
How is this different from the 2018 World Models paper?
Ha and Schmidhuber trained a VAE, then an MDN-RNN, then a tiny linear controller with an evolution strategy, in separate stages. Dreamer trains the world model, critic and actor concurrently from one replay buffer, with gradient-based actor-critic learning on imagined trajectories, and the world model is a single recurrent state-space model rather than a separate VAE and RNN.
Why not normalize advantages, as PPO does?
Normalizing advantages divides by their standard deviation every batch, which puts the same weight on the return term whether rewards are near or not. When no reward has been found, the advantages are mostly critic noise, and dividing by a tiny standard deviation blows that noise up until it outweighs the entropy bonus, so the policy stops exploring. DreamerV3 divides returns by their 5-95 percentile range only when the range exceeds 1.
What do the 1% unimix and the 255 bins buy?
Unimix mixes every categorical with 1% uniform, so no class ever has probability below 0.01/64 = 0.00016 and the KL of one latent is capped at ln(6400) = 8.76 nats. The 255 bins span about plus or minus 485 million in real units while staying 0.157 apart in symlog units, so the reward and value heads cover any scale with one output layer.
Footnotes & further reading
- The paper: Hafner, Pasukonis, Ba & Lillicrap, Mastering Diverse Domains through World Models (arXiv 2023, v2 April 2024), published as Mastering diverse control tasks through world models (Nature 640, 2025). Code: github.com/danijar/dreamerv3.
- Earlier Dreamers: Hafner et al., Dream to Control (ICLR 2020) and Mastering Atari with Discrete World Models (ICLR 2021), which introduced categorical latents and KL balancing. The RSSM is from PlaNet (ICML 2019).
- Free bits: Kingma et al., Improved Variational Inference with Inverse Autoregressive Flow (NeurIPS 2016), Appendix C.8. The VAE background is in the VAE explainer.
- Twohot targets for value regression appear in MuZero (see the MuZero explainer); the symlog transform belongs to the bi-symmetric log family of Webber, A bi-symmetric log transformation for wide-range data (2012).
- Optimizer pieces: Ziyin et al., LaProp (2020), and Brock et al., High-Performance Large-Scale Image Recognition Without Normalization (ICML 2021), which introduced adaptive gradient clipping.
- Minecraft baselines: Baker et al., Video PreTraining (VPT) (2022), and Kanitscheider et al., Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft (2021), whose block-breaking setting Dreamer follows.
How could this explainer be improved? Found an error, or something unclear? I read every message.