Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training
Reward an agent only for the prediction error its model can still reduce.
Prediction error splits in two: the part that more training removes, and the noise that no amount of training removes. A small second network learns how much of that noise each transition carries, and subtracting its estimate keeps a curious agent away from transitions it can never learn to predict.
Explaining the paperCuriosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model TrainingIn the paper's grid world, an agent paid for raw prediction error spent more than 98% of its 35,000 steps on a wall of random pixels, and its world model finished where a model that predicts 0.5 for every pixel would.
Model-based reinforcement learning first learns a world model, a network that predicts what the environment does next, and then plans or trains a policy against that model instead of the real environment. A model is much cheaper to query than an environment is to act in: MuZero runs 800 simulated continuations per move in Go and chess without touching the real board. The approach goes back to Sutton's Dyna (1990), and the name most people know it by comes from Ha and Schmidhuber's World Models (2018).
A planner chains the model's predictions, so the model is only useful if it is accurate. A model that is slightly wrong about one step is more wrong about two, and how fast the error grows depends on whether the dynamics stretch small differences or shrink them.5 A central practical question in model-based RL is therefore how to get an accurate model out of a limited number of environment steps.
That depends on which transitions the model is trained on. An agent that wanders at random spends much of its budget on transitions the model already predicts well. Schmidhuber's proposal, in a February 1990 technical report and a 1991 conference paper, was to pay the agent an intrinsic reward, one the environment does not provide, for visiting transitions where the model's predictions are bad.
The simplest version of that reward fails in a way common enough to have a name: the noisy-TV problem. Put a television showing static in the environment and the reward in front of it never runs out. Static cannot be predicted, so the error there stays high, the reward stays high, and the agent stays. Curiosity-Critic subtracts the unpredictable part of the error before paying for it.
Turning "subtract the part training cannot remove" into a number an agent can compute at every step takes three ideas: rewrite a sum over the agent's whole history by changing the order its terms are added in, identify the single term per transition that survives the rewrite, and train a second network to estimate that term while the agent explores. The paper then tests the resulting reward against eight others in a 30×30 grid world.
Why raw prediction error traps an agent
Every reward on this page is one formula with one term changed, so the notation comes first. The agent is in state , takes action , and the environment moves it to . The world model has weights , and is its prediction of . The prediction error is the distance between the prediction and what happened:
The bars are the Euclidean length of the residual vector, not its square. The world model trains on the squared length, and the gap between the two matters when the noise floor is computed below. Schmidhuber's first curiosity reward pays this error directly as the reward. The paper calls it Curiosity V1, and uses the name for the whole family of rewards that pay raw error under any choice of error measure.
Paying for error assumes that high error marks something the model can learn. Schmidhuber said so in a later retrospective: "it was implicitly and optimistically assumed that the predictor will indeed improve whenever its error is high."2 A coin about to be flipped breaks the assumption. Its outcome has high prediction error on every visit and at every point in training, so an agent that finds one is paid well for staying next to it, and the payment never shrinks.
Figure 1 is the paper's test environment, replayed from the authors' released runs. The left half of the 30×30 grid is learnable: each cell shows the same fixed pattern of 200 black-and-white pixels on every visit, so a model can memorize it and drive its error toward zero. The right half redraws all 200 pixels at random on every visit, so training cannot help there. With Curiosity V1 selected, watch which half the lit cells stay in and what the error readout does over 35,000 steps. The other eight rules are the alternatives the rest of the page works through; come back to each as it is introduced.
In seed 1, Curiosity V1 never enters the learnable half, and its world model's mean error on the 450 learnable cells stays near 7.3 for the whole run. Across five seeds it spends 1.7% of its steps on learnable cells. The policy is maximizing what it is paid: on a noisy cell the reward stays near 7.07 on every visit, while on a learnable cell it falls as soon as the model starts to learn the pattern.
Rewarding improvement instead
Schmidhuber described this failure in an April 1991 technical report: "in non-deterministic environments the controller will focus on parts of the environmental dynamics which are inherently unpredictable. This is because the adaptive model usually will produce incorrect predictions for the uncertain parts of the environment. Therefore the controller will receive reinforcement although it cannot be expected that the world model will improve."3
His fix pays for improvement in the model instead of error. Take one gradient step on the transition just observed, measure the error on the same transition again, and pay the difference:
Both terms use the same observed , so the environment is sampled once and the second error costs one extra forward pass. The paper calls this Curiosity V2. On a learnable transition the step removes some error that stays removed. On a coin flip the step moves the prediction toward this particular outcome, the next flip is as likely to go the other way, and the model makes no lasting progress. The single-visit difference is still slightly positive, because the step was taken on the same sample it is scored on. In our re-run of the released code (Provenance panel), V2 paid 0.172 per visit on learnable cells and 0.014 on noisy cells over the last 5,000 steps.
As an estimate of how much is left to learn at a transition, (2) has two weaknesses, and the paper names both. It is one sample, and a difference of two noisy numbers is noisier than either. And the subtracted term, , is the error of the current model. It equals the error of a fully trained model only once the model is nearly trained, so early in training, when the reward has the most influence on what gets learned, it is furthest from the quantity it stands in for.
Schmidhuber's 1991 system had already dealt with the first weakness, which the paper does not mention. It did not subtract one post-update sample. It trained a second network to predict the model's expected change and paid the agent according to that prediction. His later description: "although noise was unpredictable and led to wildly varying target signals for the predictor, in the long run these signals did not change the adaptive predictor parameters much, and the predictor of predictor changes was able to learn this."2 Equation (2) is the one-step form the paper uses as its comparison point, and the paper argues against it as if the 1991 work had stopped there.
Scoring an update against the whole history
Equations (1) and (2) depend only on the transition the agent is standing on. A gradient step taken on that transition also changes the model's predictions on every other transition, some for the better and some for the worse, and the goal is a world model that is accurate on all of them.
The paper writes the side effects into the reward. Let be the history: every transition visited up to time . At each step, pay the total change in error that this step's update produced across the entire history:
Each term in the sum compares the model's error on one old transition before and after this step's update. Summed over the history, the reward credits the current step with its effect on everything the model has seen, not only on the transition it was trained on. An update that improves one cell and degrades fifty others scores badly under (3) and well under (1) and (2).
Noisy transitions in the history contribute about zero to (3) in expectation, so (3) mostly counts improvement on learnable transitions. The paper states this in its Section 3 and writes it as an approximation in Appendix A. The mechanism: suppose the model already predicts the average outcome of a noisy transition, and an update made elsewhere nudges that prediction a little. Whether the error on one stored noisy sample rises or falls depends on which side of the prediction the sample fell. For fair coin flips the samples fall on each side equally often and at the same distance, so the first-order changes cancel. Two conditions limit this. The noise has to be balanced around the prediction (Figure 4 shows what happens when it is not), and the second-order term is positive, so unrelated updates make noisy transitions very slightly worse rather than leaving them exactly unchanged. The argument also excludes the transition the agent is standing on, whose error the update always reduces a little.
Evaluating (3) is expensive. At step it runs the world model over every stored transition twice, before and after the update: on the order of forward passes at that step, and on the order of over a run of steps (the paper writes the running total as compute and memory). At the paper's 35,000 steps that is about 1.2 billion extra forward passes, spent producing one scalar reward per step.
Neither the objective nor its cost is new. Schmidhuber's formal theory of curiosity specifies that "both the old and the new compressor have to be tested on the same data, namely, the history so far", and lists among the improvements still to be made computing learning progress "without frequent expensive compressor performance evaluations on the entire history so far".4 The next section is the paper's way around that cost.
How the history sum telescopes
Instead of the reward at one step, look at what (3) adds up to over a run. An RL agent maximizes the discounted sum of its rewards, with a discount setting how much a later reward is worth now:
With , (4) is a double sum over pairs with : a triangle with one row per training step and one column per visited transition. Adding it up row by row is the expensive order, since each row is one reward from (3) and is as wide as the history. Adding it up column by column gives the same total, and each column collapses to two terms.
Follow column down the triangle. Row contributes and . Row contributes and . The minus term from one row and the plus term from the next are the same number, the error on transition under weights , so they cancel, and the same happens between every pair of adjacent rows. Only the first plus term and the last minus term remain:
A column with made-up numbers: transition 0 is visited at step 0, the run lasts three steps, and its error under through is 5.0, 4.2, 3.9 and 3.7. The three rows contribute 0.8, 0.3 and 0.2, which sum to 1.3, and 5.0 minus 3.7 is 1.3. Nothing was approximated. A transition's total contribution to the cumulative improvement is its error at the moment it was visited minus its error at the end of training.
Figure 2 draws the triangle. Switch between the two summation orders and pick a column to see the pairs cancel. At the triangle has 66 entries, and summing by column leaves 11 two-term differences, one per transition.
Equation (6) has one term per visited transition, so it can be paid out one step at a time. The per-step reward it defines is
For below 1 the adjacent terms carry weights and , so each cancelling pair leaves a remainder proportional to . The paper's Equation (5) collects those remainders into a history term:
For between 0 and 1 the coefficient is negative, and the errors it multiplies are norms and never negative, so the history term is at most zero and the per-step term is an upper bound on . At the history term vanishes and the per-step term is (6). The paper argues for on its own terms too: what matters is the model's accuracy when training stops, not at some earlier point.
The limit has an order. For near 1 the history term is roughly , with the average error per transition, so it is small only when . At that needs well below . So means fixing the horizon and taking the discount to exactly 1; a discount of 0.99 over 35,000 steps is far from the limit.
The irreducible noise floor
Equation (7) is exact and per-step but still cannot be computed during the run, because is the model at the end of training. The paper makes two substitutions. It replaces the end-of-training model by its limit , which is accurate for a long run. Then it replaces the error on the one observed outcome by the average error over all the outcomes the transition can produce, a stable quantity that a network can learn to predict:
The subtracted term is the paper's asymptotic error baseline, and is the environment's distribution over next states. The world model trains on squared error, and the prediction that minimizes expected squared error is the mean outcome, so a converged model predicts . Substituting that prediction into the error gives
The right side depends only on the environment: how far a transition's outcomes scatter around their own mean. A transition with a fixed outcome has a baseline of zero. A coin-flip transition has a baseline set by how far its flips scatter. The current error minus this baseline is the error the model can still remove.
A worked case: a transition emits 200 bits, of which are redrawn at random on every visit and are fixed. A converged model predicts 0.5 on each random bit and the true value on each fixed bit, so it is off by exactly 0.5 on coordinates and by 0 on the rest, and the floor is . That value holds on every draw, not just on average, because a bit that lands on 0 and a bit that lands on 1 are both 0.5 away from 0.5. At the floor is , which is ≈ 7.071: the paper's noisy half. At it is 0: the learnable half.
In Figure 3, drag from 0 to 200. Both ends are the paper's environment; the values between are partially learnable transitions, which is the case the reward has to grade. The teal band between the error curve and the floor is what Curiosity-Critic pays. At the curve starts on the floor and stays there, so there is no reward on the first visit or any later one.
Two caveats on (9). The paper calls the quantity the mean absolute deviation. The two names agree for a single number; for a 200-dimensional vector, (9) is the mean Euclidean distance to the mean, a different number from the per-coordinate absolute deviation. Second, (9) is a floor only for a model trained the way this one is. The reward's metric (1) is the unsquared distance, which the geometric median of the outcomes minimizes; the training loss is the squared distance, which the mean minimizes. Where the median and the mean differ, some other prediction scores lower on (1) than the converged model does, so (9) is the floor for an MSE-trained model and not the lowest error any model could reach.
Figure 4 shows both caveats on one pixel that is 1 with probability . For a prediction , the expected unsquared error is , a straight line, and the expected squared error is a parabola with its minimum at . MSE training lands at , where the unsquared error is ; the lowest unsquared error any prediction reaches is , at the median. At those are 0.32 and 0.20.
At the two floors agree at 0.5 and the teal line is flat, so nudging the prediction away from 0.5 does not change a fair-coin pixel's expected error at first order. That flatness is the cancellation the noisy terms of (3) rely on. For any other the line has slope , and an update that nudges the prediction changes a noisy transition's error in proportion to the nudge. The paper's grid uses fair coins, so neither caveat is active in its experiment.
An appendix bounds the floor from above with Jensen's inequality:
So the floor is at most the square root of the summed per-coordinate variance, with equality only when the distance comes out the same on every draw. Equality holds exactly in the grid world, where is 7.071, and does not hold in general, so the floor is neither the standard deviation nor the variance.
How Curiosity-Critic learns the floor
The baseline in (9) needs the environment's outcome distribution and a converged model, and the agent has neither. It does have the post-update error at almost no cost: it ran the model before the step to get , took the step, and one more forward pass on the same input gives . On a noisy transition is near the floor from early on; on a learnable one it falls toward zero as the model learns.
Subtracting that single number is Curiosity V2. Curiosity-Critic instead trains a small network , the critic, to predict from the same input the world model sees, and subtracts the critic's prediction:
The subscript on means the critic is queried after its own update on this step's , so its estimate already includes the sample just taken. One environment step of the released code:
# one environment step of Curiosity-Critic, in the order curiosity_experiment.py
# runs it (batch size one, no replay buffer)
obs = env.observe(s) # the 200 bits the current cell shows
e_before = norm(world(s) - obs) # L2 error before the update
world.step(s, obs) # one Adam step on the MSE loss
e_after = norm(world(s) - obs) # same input, same obs, updated weights
critic.step(s, e_after) # the critic regresses e_after (MSE)
r = max(0.0, e_before - critic(s)) # Eq. (10) with phi_{t+1}, clipped at 0
a = policy.act(s) # eps-greedy over the 4 neighbours' V
V[s] += 0.05 * (r / run_std(r) - V[s]) # discount 0 in the released config
s = step(s, a) # movement is deterministicA second learned network, trained on the first network's output, choosing where to train the first network, looks circular. It works because the two networks learn problems of very different size. The world model outputs a 200-dimensional vector per cell. The critic outputs one number per cell, the error that will remain there after a step, and in this grid it gets close to the right value long before the world model does (Figure 6 measures how long before).
The paper gives two reasons the critic should beat a raw . Variance: is a regression fit over many visits, so it returns roughly the mean of where (2) subtracts one draw of it. Generalization, which is also what separates the neural critic from the tabular ablation that keeps a separate average per cell: the critic's input is the same one-hot row concatenated with one-hot column that the world model gets, so a visit to one cell moves the estimate for every cell in the same row or column. A cell the agent has never stood on already has a baseline, inherited from its row and column.
Figure 5 lines up the four baselines on one cell. Each bar is the current error, split into the part subtracted and the part paid as reward. On the noisy cell, Curiosity V1 subtracts nothing and pays 7.07 on every visit; the other three all subtract roughly the floor and pay roughly nothing. On the learnable cell the oracle subtracts its true floor of zero and pays the whole error, like V1, while V2 and the critic subtract the post-update error and pay only the part one step removes.
Our re-run of the released code measures how much each of the paper's two reasons contributes here. The variance argument is weak in this environment: over the last 5,000 steps V2's reward on noisy cells was 0.014 ± 0.009 per visit, small and steady, while the critic's was 0.067 ± 0.101 (clipped to exactly zero on 47% of those visits). The critic's reward varies more because its estimate on any one cell moves whenever the agent trains it on a different cell in the same row or column. The difference shows up early instead. An untrained critic outputs small numbers, so for the first few hundred steps it pays close to the full error everywhere (2.37 per visit on noisy cells over the first 1,000 steps of the re-run), and the stored seed-1 trajectory visits 425 learnable and 434 noisy cells by step 2,000. Curiosity V2's stored seed-1 run visits 304 noisy cells but only 73 learnable ones in its first 5,000 steps.
The paper's claim that the critic converges before the world model checks out in the released traces. Averaged over five seeds, the critic's mean estimate over the noisy cells first comes within 5% of 7.071 at step 400, about 1% of the run, when the world model's error on the learnable cells is still 7.006, essentially untrained. Over the learnable cells the critic's estimate rises for the first 600 steps, because early in training the post-update error there is large, and then falls with the model.
The oscillation on the noisy cells has the same cause as the critic's reward spread. The logged value is the critic's prediction averaged over all 450 noisy cells, most of which the agent is not standing on. While the agent works the learnable half, the critic trains on low targets, and the shared row weights pull the noisy-cell predictions down with them. Over the last 5,000 steps the five-seed average is 6.551, 7.35% under the floor. A visit pulls the visited cell back up: in the re-run, at the noisy cells the agent actually stood on late in the run, averaged 7.084 and the critic 7.085.
All four rewards in Figure 5 are one formula with different baselines. A zero baseline, as if the environment were deterministic, gives Curiosity V1. The single post-update error gives V2. The critic's prediction gives Curiosity-Critic, and the true floor gives the oracle the paper uses as a reference. The framing holds for any error measure, so feature-space variants such as ICM belong to the V1 family. The pieces had precedents: Schmidhuber's 1991 confidence network was trained to output the model's expected error, and in his reported experiments it predicted "the absolute value of the difference between M's (one-dimensional) output and the current target value", the one-dimensional case of (9).3
The word "asymptotic" is exact on noisy cells and loose on learnable ones. The critic regresses , the error of the current model after one step. On a noisy cell that number has no trend, so it already sits near the asymptotic floor. On a learnable cell it is still falling at the end of the run, so the critic tracks the model's current error, not its eventual zero: the critic's mean estimate over learnable cells peaks at 7.14 around step 600 and ends at 1.18, still falling.
That gap shows up in where the agent spends its time. On a learnable cell the oracle subtracts zero and pays the whole error. The critic subtracts its estimate of the post-update error, most of the current error, and pays the remainder, so it separates learnable from noisy cells less sharply: in the re-run's last 5,000 steps it paid 0.368 per visit on learnable cells against 0.067 on noisy ones. The critic agent spends 70.9% of its last 5,000 steps in the learnable half; the oracle spends 95.3%.
The experiment: one grid, two halves
The test environment is small on purpose. In a 30×30 grid with the noise built in by hand, a difference between methods can only come from the reward: every method trains the same world model with the same optimizer, starting from the same warm-up weights.
Columns 0 to 14, all 30 rows, are the learnable half: 450 cells, each emitting a fixed 200-bit pattern. Columns 15 to 29 are the noisy half: each visit draws 200 fresh fair coin flips. The agent starts at cell (15, 15), in the centre row and on the noisy side of the border, and moves one cell up, down, left or right per step. Movement is deterministic; all the randomness is in what the cells show.
The world model does less than (1) suggests, as an appendix says. It takes a 60-dimensional input, a one-hot code for the row concatenated with a one-hot code for the column, and outputs that cell's 200 pixels through one hidden layer of 1,024 ReLU units. The action is not an input. Movement is deterministic, so predicting the next state would be trivial, and the model predicts the current cell's observation instead. For the question the experiment asks, whether a reward separates learnable transitions from unlearnable ones, that is a reasonable simplification. It also means the experiment never tests transition modelling, action conditioning, or the compounding rollout error from the introduction. The neural critic is a 60 → 128 → 1 network with a ReLU and its output clamped at zero; the tabular ablation keeps a 30×30 table of exponential moving averages of with decay 0.9.
Two visits from our re-run of the released code, one in each half, with the quantities the loop computes:
# two visits logged in our re-run of the released code (seed 1)
# step 8,264: learnable cell, row 24, column 14
s = onehot(row=24) ++ onehot(col=14) # (60,) two ones, 58 zeros
obs = patterns[24, 14] # (200,) fixed bits, 88 of them 1
e_before = 5.616 # the model has learned part of this pattern
e_after = 5.413 # one Adam step removed 0.203 (Curiosity V2 would pay 0.203)
critic = 4.152 # phi's estimate of e_after here, after its own step
reward = 1.464 # 5.616 - 4.152
# step 30,000: noisy cell, row 22, column 24
obs = bernoulli(0.5, size=200) # redrawn on every visit
e_before = 7.114 # close to sqrt(50) = 7.071 whatever the draw
e_after = 7.110 # the step fit this draw by 0.004
critic = 7.431 # the estimate is above e_before on this visit
reward = 0.0 # 7.114 - 7.431 < 0, clipped to zeroThe policy is a 30×30 table of values, one per cell, not a network. After each step the value of the current cell moves 5% of the way toward the normalized reward received there (step size 0.05), and the agent moves to the neighbour with the highest value, or to a random neighbour 30% of the time. The discount is 0 in the released configuration; the paper does not state it. Values therefore never propagate across the grid, and the agent compares the four cells next to it and looks no further. This is the policy's discount and has nothing to do with the of the derivation, which is the discount on the curiosity objective (4), so the zero here neither contradicts the derivation nor tests it.
The world model and the critic each take one Adam step per environment step at learning rate 0.001, and the world model is first warmed up for 100 random steps, from identical weights for every method within a seed. Rewards are clipped at zero before use, which only affects the rewards that subtract a baseline and can go negative. They are then divided by a running estimate of their own standard deviation (an exponential moving average with decay 0.95), so a method with large rewards does not effectively get a larger step size. The division has no mean subtraction, so each method's rewards are rescaled by a different factor that drifts during the run. The paper says the shared value-table initialization of 3.0 "eliminates initial optimism asymmetries" across methods; because of the rescaling, 3.0 means a different amount of optimism for each method.
Results: which reward builds the best world model
Nine methods, five seeds each, 35,000 steps per run. The metric is the world model's mean L2 error over all 450 learnable cells, queried directly every 100 steps, so it does not depend on which cells the agent happened to visit. Figure 7 ranks the methods by final error and by where they spent their last 5,000 steps.
Among the methods not given the floor, the neural critic finishes lowest at 1.858 ± 0.080, with the tabular critic next at 1.912 ± 0.070; the oracle, given the true floor, reaches 1.736 ± 0.063. Every method without a critic finishes worse than all three. The ± is the population standard deviation over the five per-seed results. The neural critic is also the first non-oracle method below error 3.0, at step 13,400, against 21,700 for the best method without a critic, Random Network Distillation (RND, a novelty bonus described below) fed the cell's row and column, and 34,400 for Curiosity V2. It reaches 2.5 at step 16,000 and 2.0 at 26,700. The oracle is slower to 3.0 and 2.5 (13,900 and 18,500) and first to 2.0, at 24,400.
The released traces add two things the paper's comparison leaves out. An undirected random walk reaches 3.0 at step 16,800 and 2.5 at 24,500, faster than RND on the cell address (21,700 and 28,800) and faster than the tabular critic (25,300 and 29,500); the paper explains the walk's strength by the open geometry, since a random walk from the border diffuses into the learnable half with high probability. And five seeds do not separate the neural critic from the tabular one: an exact permutation test over the per-seed finals gives p = 0.35. What five seeds do support is that every critic-based reward beats every reward without a critic, and that the oracle beats the learned critic.
Curiosity V1 finishes at 7.114 ± 0.147 and does not improve. The first version of the paper called that indistinguishable from an untrained model. An untrained network here outputs roughly zero, and every fixed pattern has 88 ones out of 200, so an untrained model scores about ≈ 9.38. The value 7.114 is close to ≈ 7.071, the score of a model that outputs 0.5 on every pixel. The noisy half can teach one thing, the average of a fair coin, and a model that predicts 0.5 everywhere is off by 0.5 on every bit of any binary target, learnable or not.
Visitation count does not fit the paper's account of it. Its bonus is , larger for cells visited less, with every count starting at 1, which should spread visits evenly. It ends at 5.588 ± 0.794, worse than the random walk. The first version of the paper attributed this to the agent continuing "to allocate substantial time to stochastic cells throughout training"; the current version says it plateaus "despite spreading coverage broadly regardless of cell learnability". The released trajectories contradict both. Visitation count puts 62.2% of its 35,000 steps on learnable cells, more than the neural critic's 56.4% and the random walk's 51.9%. It does not spread broadly either: it reaches 333 of the 450 learnable cells on average, so about 117 cells, a quarter of the region the model is graded on, never get a single training step.
Across the nine methods, the number of learnable cells ever reached correlates with final error at −0.94, while the share of steps spent in the learnable half correlates at only −0.64. Random, V2, RND on the cell address and all three critics reach essentially all 450 cells. Visitation count's seeds show the same pattern: their late-run learnable shares are 95.1, 1.2, 100.0, 97.3 and 0.0%, and seed 3, which spends 99.8% of its whole run on learnable cells, still finishes at 4.578. It reaches 393 cells and visits them very unevenly, a median of 30 visits per cell and a maximum of 623. In Figure 8, switch the horizontal axis between the two measures and watch where visitation count lands on each.
The tabular critic is the other counterexample to reading time spent as the cause. It spends 26.5% of its steps in the learnable half, about half the random walk's 51.9%, reaches all 450 learnable cells in every seed, and finishes at 1.912 against the random walk's 2.348.
A plausible reason a count bonus fails to cover this grid is the policy it was given. The bonus is largest at the least-visited cells, but with a discount of 0 the agent compares only the four cells next to it and cannot travel toward a distant unvisited region. Count-based exploration in its original form, Strehl and Littman's MBIE-EB (2008), put the bonus inside a value function re-solved by value iteration, and that planning step is what carried the agent toward unvisited states. Without it, a one-step greedy agent lowers the bonus only in the neighbourhood it is already in. Prediction-error rewards may survive the same policy better because the model is bad in contiguous patches, which the agent can detect from an adjacent cell; the paper does not test this.
The two RND variants are a matched pair. Random Network Distillation (Burda et al., 2019, originally used with a PPO agent) scores how new an input is by how badly a trained predictor network matches a fixed random network on it, and the mismatch shrinks as the same input recurs. Fed the cell's row and column, it behaves like a soft visit count and is the best method without a critic, at 2.220 ± 0.109. Fed the 200 noisy pixels, every visit to the noisy half is a new input from a space of , the mismatch never shrinks, and it fails like V1 at 6.842 ± 0.212. The algorithm and the reward equation are the same; changing the input moved the final error by a factor of 3.08.
What the grid world cannot show
The derivation of the reward is general and the experiment is not. The paper says its evidence comes from one controlled environment and lists MuJoCo, Atari and VizDoom as next steps. There is no transition model, no action conditioning and no multi-step rollout, so no test of the compounding-error problem the introduction opens with. Nothing has been run on Atari, the standard benchmark for exploration rewards (see DQN), or on continuous control.
The environment also favours the neural critic. "This cell is noisy" is the predicate column ≥ 15, a function of one block of the critic's one-hot input, the same down all 30 rows. A shared-weight network over that encoding picks up most of the rule from a handful of visits, which is a large part of why its estimate is calibrated by step 400, and a real environment's learnability boundary need not have that shape. The paper's description of the grid also overstates the structure in it. It says the fixed patterns are cyclic shifts of one random base vector "such that neighbouring cells have strongly correlated observations". In the released code adjacent cells disagree on 46% of their pixels and correlate at about 0.07, and the shift wraps around, so the 450 cells carry only 200 distinct patterns. What lets the world model generalize across cells is the shared row and column weights of the one-hot encoding, not any resemblance between patterns.
The baseline the paper names most prominently is one it does not run. It presents ICM, the inverse-dynamics feature-space method, as an instance of Curiosity V1 that its framework covers. ICM attacks the noisy TV from another side: it measures error in a feature space trained only to predict which action was taken, so parts of the observation the agent cannot influence are filtered out before the error is computed. The noise in this grid does not depend on the action, which is the case ICM was designed for, and Table 1 has no ICM row.
The closest published relative goes uncited. Mavor-Parker and colleagues (ICML 2022) predict the mean and variance of the next state with separate heads and reduce the intrinsic reward where the predicted variance is high, which is prediction error minus a learned estimate of the noise. Curiosity-Critic reaches a similar reward from a different direction, by asking what the cumulative objective collapses to, and regresses a single scalar error instead of a per-dimension variance.
On the evidence it collected, the paper shows that a learned scalar estimate of how wrong a model will stay is cheap to train next to the model, converges within the first 1% of training in this grid, and gives a reward that beats every reward in the comparison without one. On this grid, swapping the baseline from zero to the learned critic took the final error from 7.114 to 1.858, within 7% of the oracle's 1.736.
Questions you might still have
If the critic is trained on the model’s current error, how can it estimate the error at convergence?
In general it cannot, and on learnable transitions it does not: there it tracks the model’s present post-update error, which is still falling at the end of the run (the critic’s mean estimate on learnable cells ends at 1.18). The two quantities coincide on unlearnable transitions, where the error has no trend, and those are the transitions the reward most needs to get right.
Does the reward really drop to zero on a noisy transition?
Nearly. In our re-run of the released code, the reward on the noisy cells the agent visited averaged 0.067 per visit over the last 5,000 steps, and was clipped to exactly zero on 47% of those visits, against 0.368 on learnable cells. The logged critic estimate averaged over all noisy cells is lower, 6.551 against the 7.071 floor, because the critic’s shared row weights pull the cells the agent is not on toward the low targets of the learnable half; a visit pulls the visited cell back up.
Why does an undirected random walk do so well here?
Geometry, as the paper notes. The grid has no obstacles and the learnable half starts one cell from the agent, so a random walk diffuses into it with high probability and reaches error 3.0 at step 16,800, ahead of every method without a critic and of the tabular critic. In a maze that needs long directed travel, an undirected walk would rarely arrive at all, so the strong showing is specific to this environment.
How is this different from Random Network Distillation?
RND scores how novel an input looks, not how learnable it is. Given the deterministic cell address, novelty decays with visits, so it behaves like a soft visit count and finishes at 2.220. Given the 200 noisy pixels, every visit is a new input out of 2^200, the bonus never decays, and it fails like raw prediction error at 6.842. Curiosity-Critic asks whether there is anything left to learn at a transition, not whether the transition has been seen.
What did Schmidhuber’s 1991 system actually compute?
Not the one-line difference the paper uses for comparison. The April 1991 report trained a confidence network to output the world model’s expected error (in its experiments, the absolute difference between output and target); the curiosity reward came from changes in that learned estimate. The retrospective in his later survey describes a network that learned to predict the predictor’s changes, so that unpredictable noise, which does not change the predictor’s parameters much in the long run, earned little reward.
How does this compare with ensemble disagreement?
Disagreement methods train several forward models and pay for how much they disagree. On a purely noisy transition every model converges to the same average, so the disagreement goes to zero even though each model’s error stays high, which handles the noisy TV without estimating a floor. The cost is several models instead of one model plus a small critic. The paper sets ensembles aside and does not compare against them.
Would this transfer to a real environment?
Untested. The reward does not depend on how error is measured, so it carries over to feature-space or distributional world models by changing one function. What the grid world does not test is transition modelling, action conditioning, planning, or a learnability boundary more complicated than one column index, which is the boundary the critic’s one-hot input was best placed to learn.
Footnotes & further reading
- The paper: Bhaskara and Wang, Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training (ICML 2026 Workshop on Epistemic Intelligence in Machine Learning). Code and result files. Every number on this page that the paper does not print was recomputed from those files; the per-visit reward statistics come from our re-run of the same code.
- Schmidhuber, Driven by Compression Progress, sections 3.1 and 3.2, which retell the 1990 prediction-error reward and the 1991 improvement reward in his own words. The published survey version is Formal Theory of Creativity, Fun, and Intrinsic Motivation (1990–2010), IEEE TAMD 2010.
- Schmidhuber, Adaptive Confidence and Adaptive Curiosity, technical report FKI-149-91 (30 April 1991), the long version of Curious Model-Building Control Systems (IJCNN Singapore, 1991). The labels “1991a” and “1991b” are bibliography artefacts that different papers, Schmidhuber's own included, assign in opposite orders, so these papers are best named by title.
- Both quotes are from the arXiv version of the formal theory in footnote 2: the same-data requirement in Appendix A.5, and the full-history evaluation cost in its closing list of planned improvements. The 2010 IEEE survey restates the theory.
- Asadi, Misra and Littman, Lipschitz Continuity in Model-based Reinforcement Learning (2018) bound the -step error of a -accurate model by , which grows exponentially only when the model's Lipschitz constant exceeds 1, is exactly at 1, and stays bounded below it. Janner et al., When to Trust Your Model (2019) get a bound linear in the rollout length. Curiosity-Critic states that model errors compound without a citation and does not claim exponential growth.
- The noisy-TV lineage: Pathak et al., Curiosity-driven Exploration by Self-supervised Prediction (ICM, 2017) raises the white-noise screen and credits Schmidhuber for it; Burda et al., Large-Scale Study of Curiosity-Driven Learning (2018) coins the name and takes it literally, adding a TV to a maze along with an action that changes the channel.
- Exploration by Random Network Distillation (Burda et al., 2019). The RND here is a stripped-down version. Published RND whitens and clips its inputs before both networks, trains the predictor on a quarter of the collected experience (with the same Adam learning rate, 1e-4, as the policy), and normalizes by the spread of discounted returns rather than of raw rewards. It also uses two value heads to mix an episodic extrinsic stream with a non-episodic intrinsic one, which this grid world does not need, since it has no extrinsic reward and no episodes.
- The closest published relative, uncited by the paper: Mavor-Parker, Young, Barry and Griffin, How to Stay Curious while Avoiding Noisy TVs using Aleatoric Uncertainty Estimation (ICML 2022), which predicts the mean and variance of the next state separately and reduces the intrinsic reward where the predicted variance is high.
- The count-based lineage: Strehl and Littman's MBIE-EB (2008) puts a bonus inside a Bellman equation re-solved by value iteration, with chosen so the bonus is a valid optimism guarantee; Bellemare et al., Unifying Count-Based Exploration and Intrinsic Motivation (2016) generalize counts to a density model. The baseline in this paper keeps the shape and drops the action index, the constant, and the planner.
- Ensemble alternatives the paper sets aside: Pathak et al., Self-Supervised Exploration via Disagreement (2019), where several forward models converge to the same average on a purely noisy transition so their variance vanishes even though each one's error stays high; and Houthooft et al., VIME (2016), which pays information gain about the dynamics parameters.
How could this explainer be improved? Found an error, or something unclear? I read every message.