DDIM: Denoising Diffusion Implicit Models
Get 1,000-step diffusion quality from 20 to 100 steps, with the same trained network.
DDIM changes only the sampling loop of a diffusion model. It defines a family of samplers that all fit a network trained as a standard diffusion model (DDPM), and the deterministic member of that family can skip most of the 1,000 noise levels.
Explaining the paperDenoising Diffusion Implicit ModelsBecause the DDIM sampler is deterministic, the starting noise works as a latent code: the same code gives the same image at any step count, two codes can be blended, and a real image can be encoded and reconstructed.
Why DDPM sampling is slow
A denoising diffusion probabilistic model (DDPM) generates an image by starting from pure Gaussian noise and removing a little of it at a time. A standard DDPM uses noise levels, and every level costs one forward pass of a large U-Net, the image-to-image network that predicts the noise. The passes cannot run in parallel, because each one needs the output of the one before. The DDIM paper measures the cost on one Nvidia 2080 Ti: about 20 hours to sample 50,000 images of pixels, against under a minute for a GAN, and nearly 1,000 hours for 50,000 images at .
Taking 50 large steps instead of 1,000 small ones would cut the cost 20×, but a DDPM sampler resists this because it was derived as the reverse of one particular noising process: a Markov chain in which each step adds a small amount of Gaussian noise to the previous step's output. The reverse of that chain is a 1,000-step procedure, and the DDPM derivation gives no reason why a 50-step version should work. DDIM's framework does justify skipping, for DDPM's sampler as well, but the results section shows the DDPM sampler still degrades fast when the steps are few.
DDIM keeps the trained network and derives a new sampler. The argument goes through four ideas, one section each after a short recap: the training loss only ever looks at one noise level at a time; so many different noising processes are compatible with the same trained network; one of them gives a deterministic sampler; and a deterministic sampler can take large steps and run backwards as an encoder.
What the trained network computes
DDIM reuses DDPM's training unchanged, so the notation comes first. The forward process turns an image into noise over steps. For a 32×32 RGB image, is a tensor of shape 3 × 32 × 32, 3,072 numbers scaled to . The forward process is built so that you can jump straight to any noise level in closed form:
is a number between 0 and 1 that falls as grows. It sets how much of the image survives: scales the image down and scales the noise up, and because a unit-variance image stays at unit variance. With the CIFAR-10 schedule (per-step noise rising linearly from to 0.02), , , and . By convention , no noise at all.
The paper's is cumulative. It is the product of all the per-step signal factors up to step , the quantity the DDPM paper and most tutorials call (alpha-bar). DDPM's per-step appears here as the ratio . The paper states this only in Appendix C.2, so formulas copied from a DDPM tutorial will look wrong next to it until the two notations are lined up. This page uses the cumulative throughout.
The network takes a noisy image and its noise level and returns a tensor of the same shape: its estimate of the noise that was mixed in. Training picks a training image, a level and a noise draw, builds with (1), and penalizes the squared error:
is a vector of positive weights, one per noise level. DDPM trains with every weight equal to 1, which the paper writes . The subscript is the all-ones vector; the loss is a squared L2 error, not an L1 loss. The same objective trains the noise-conditional score networks of Song and Ermon, where it appears as denoising score matching, which is why DDPMs and score-based models can share one network.
Predicting the noise is equivalent to predicting the image, because (1) can be solved for . The network's guess of is
Take one pixel. Say , the noisy pixel is , and the network predicts for it. Then (3) gives
The loss only sees one noise level at a time
Look at what each term of (2) uses: one training image, one noise level, one noise draw. It builds directly from with (1) and scores the network on undoing that single jump. No term involves and together, so the loss says nothing about how the noisy versions of one image at neighbouring levels are related. In probability terms, the loss depends only on the marginals , the distribution of the noisy image at each level on its own, and not on the joint distribution of the whole sequence .
An exam graded question by question against an answer key cannot tell in what order the student answered. In the same way, any noising process whose marginal at every level is (1) produces exactly the same loss, so the network trained for DDPM's process is also the trained network for any of them.
DDPM's Markov chain draws fresh noise at every step: is a scaled-down plus a small new Gaussian, a random walk whose per-step noise is sized so that its spread matches (1) at every level. A second process draws a single noise tensor once and builds every level from (1) with that same . Each on its own has distribution (1), so the marginals match, but the sequence is a smooth curve fixed by , where DDPM's is a jagged walk. The second process is the member of the family in the next section, the one behind DDIM. To reverse it, given and , recover from (1) and put it back at the lower level, with no new random draw.
A family of forward processes
The paper connects these two extremes with a family of processes indexed by a vector of non-negative numbers. Each member is specified by its reverse conditional, which says where is when both and are known:
Read the mean in two parts. is scaled for level . The fraction is the unit noise that is in , recovered by solving (1); the mean reuses it with weight . On top of the mean, the step adds fresh Gaussian noise with variance .
The weight is chosen so that the noise variances add up to what (1) requires at level . The reused noise has unit variance, so it contributes ; the fresh noise contributes ; and independent variances add:
The sum is the noise variance of the marginal (1) at level , whatever is. The paper proves the general statement (Lemma 1): every gives a process with marginals exactly (1) at every level. The square root needs , a condition the paper leaves implicit. At the step adds nothing new and is a fixed function of and , the reuse-one-noise process from the previous section. At the particular given in (6) below, the family reproduces DDPM's Markov chain.
Figure 1 draws sample paths from the family for one clean pixel. The amber band is the marginal (1), which the family is built to keep fixed. Drag , a single scale on all the that (6) makes precise, and watch the paths change while the band stays put:
The paths at different are different processes, so a natural worry is that each needs its own trained network. The paper's Theorem 1 answers it. Each member of the family has its own variational training objective (the evidence lower bound a model would be trained on if that member were its forward process), and for every there exist positive weights and a constant with
The weights depend on , so this is a reweighted version of (2), not itself. The step to "the DDPM network works for every σ" uses one more fact: if the network had separate weights for each noise level, each term of (2) could be minimized on its own, and the minimizer would not depend on . A real network shares one set of weights across all levels, so that step is an idealization. The experiments show that the -trained network samples well across the family.
How one DDIM sampling step works
At sampling time is unknown. The sampler substitutes the network's guess from (3) into the reverse conditional (4), and the recovered unit noise becomes the network's prediction . That gives one update rule for every member of the family:
The labels are the paper's. The first term places the predicted image at the next level's scale. The second adds back the predicted noise, scaled for the next level; it points from the guess toward , and it is deterministic, because is a fixed output of the network. Only the third term, a fresh draw times , is random.
Set at every step and nothing random is left after the initial draw of . The paper calls the resulting model the denoising diffusion implicit model. "Implicit" is used in the sense of an implicit probabilistic model, one that produces samples by pushing a latent variable (here ) through a fixed procedure (here the sampling loop), as a GAN does, instead of defining a density step by step.
Continue the one-pixel example, stepping from to . The noise budget at the new level is . With all of it goes to the reused noise:
This puts the pixel exactly where it would sit at level if were the true image and the true noise. With the DDPM value of from (6) below, and the direction coefficient drops to , so the mean is and the pixel lands at for a fresh standard normal . Figure 2 does the same split in two dimensions, with an exact denoiser for a four-cluster toy dataset:
To compare deterministic and random sampling with one knob, the paper scales a fixed shape of by a scalar :
is DDIM. is DDPM: with this the forward process becomes Markovian again and (5) is DDPM's own sampler. Values in between are valid samplers that use the same network. In the worked step, can go up to 1.449 before uses the whole budget and the direction term vanishes; that is the right end of the slider in Figure 2.
The paper also tests a third setting, , the larger variance used in Ho et al.'s code for their CIFAR-10 samples. It is not a point on the line. Appendix D.3 gives its update: the same first two terms as , with noise in place of . In the worked step , so the total noise variance is , 13% more than the 0.55 that (1) calls for. Between adjacent levels of the 1,000-step CIFAR-10 schedule the excess is at most ; across a large skip it is large, and the paper attributes 's poor few-step results to this extra noise.
Skipping steps
The marginals argument also allows skipping. The derivation of (4) and (5) never used the fact that is adjacent to ; it only needs the marginal (1) at the two levels involved. So pick an increasing subsequence of , define the process on those levels alone, and sample along in reverse, replacing and in (5) with and . The cost is network calls instead of , and the network is the one trained on all 1,000 levels.
The paper tries two spacings, each of the form (linear) or (quadratic), with set so the last level is close to . Quadratic spacing puts more steps at low noise, where fine detail is resolved. It used quadratic for CIFAR-10 and linear for the other datasets, reporting that each gave slightly better FID than the alternative on its dataset. Figure 3 plots both against the CIFAR-10 schedule's noise variance:
Under this schedule the noise variance passes 0.92 by , so with linear spacing about half of the visited levels sit where the image is almost all noise. At , linear spacing puts 4 levels at and quadratic puts 8.
Whether skipping works well depends on . A large jump makes a cruder estimate, because the network has to guess from a much noisier input. The deterministic update then carries that estimate's own predicted noise forward, while the stochastic one replaces part of it with fresh noise that the remaining few steps must remove. The full sampler is one loop for every setting:
# One sampling run. Same trained eps_theta for every setting;
# only tau (which levels) and eta (how much fresh noise) change.
def sample(eps_theta, alpha, tau, eta): # alpha[t]: cumulative, alpha[0] = 1
x = randn(3, 32, 32) # x_T: pure Gaussian noise
for i in reversed(range(len(tau))): # walk reversed(tau): noise -> data
t = tau[i] # current level
s = tau[i - 1] if i > 0 else 0 # next, less noisy level
a_t, a_s = alpha[t], alpha[s]
eps = eps_theta(x, t) # one network call
x0 = (x - sqrt(1 - a_t) * eps) / sqrt(a_t) # eq (3)
sig = eta * sqrt((1 - a_s) / (1 - a_t)) * sqrt(1 - a_t / a_s) # eq (6)
dir = sqrt(1 - a_s - sig**2) * eps # deterministic
x = sqrt(a_s) * x0 + dir + sig * randn(3, 32, 32) # eq (5)
return x # a sample x_0Each pass through the loop is one network call, so costs 50 forward passes. With eta = 0 the randn inside the loop is multiplied by zero and the only randomness is the first line.
How much faster, at what quality
The paper trains one model per dataset at with and changes only and . Quality is measured with FID (Fréchet Inception Distance), which compares statistics of generated and real images in the feature space of an image classifier; lower is better. Figure 4 plots the paper's Table 1 for DDIM, DDPM and :
On CIFAR-10 at 10 steps, DDIM scores 13.36, DDPM 41.07 and 367.43. At 100 steps DDIM reaches 4.16, close to its own 1,000-step 4.04. On CelebA the 100-step DDPM (13.93) matches the 20-step DDIM (13.73), so DDIM gets there with a fifth of the network calls. The two intermediate settings in Table 1 fall in between at every step count: at 10 steps on CIFAR-10, gives 14.04 and gives 16.66.
The paper's headline is that DDIM matches 1,000-step quality in 20 to 100 steps, which it states as a 10× to 50× speedup in wall-clock time. Sampling time is linear in the number of steps, so a 50-step run takes about 1/20 of the 1,000-step time; the 20 hours for 50,000 CIFAR-10 images become roughly one. The paper also writes "10× to 100×" in two places, as the range of step reductions it tests (1,000 steps down to 100 or 10), not as a matched-quality figure.
At the full 1,000 steps the order flips: scores 3.17 on CIFAR-10 against DDIM's 4.04, and 3.26 against 3.51 on CelebA. With every step taken, the added noise helps a little; DDIM's advantage is confined to small step counts, the regime that sets sampling cost.
The starting noise as a latent code
With the sample is a deterministic function of . The paper finds that this function barely depends on the step count: for a fixed , images generated with 10, 20, 50, 100 and 1,000 steps share their high-level features, such as pose, layout and colors, and the 20-step image is often already very close to the 1,000-step one.
Figure 5 shows the same effect in a toy. Eight fixed starting points are each decoded by the exact denoiser of an eight-cluster dataset. A ring marks where each lands with a 100-step run, a dot where it lands with steps. Under DDIM the dots sit on their rings from about 4 steps up. Switch to DDPM and the dots move off their rings at every , including 100, because the reference run draws its own fresh noise:
Since determines the image, it works like the latent code of a GAN or a variational autoencoder. The paper interpolates in it: draw two codes and , blend them, decode each blend with 50 DDIM steps, and the images change smoothly from one sample to the other. The blend is spherical linear interpolation (slerp), not a straight average:
The paper does not say why it chose slerp, but the geometry of high-dimensional Gaussians gives a reason. A 3,072-dimensional standard normal vector almost always has length close to , give or take about 2. Two independent codes are nearly orthogonal, so their straight-line midpoint has length about , a length the network never saw at . Slerp moves along the arc between the codes and keeps the length near 55.4. Figure 6 plots the length along both blends:
A DDPM cannot do this with one code, because its output also depends on the fresh noise drawn at every step: the same gives a different image on each run. The paper notes that interpolating all noise draws might work instead. Running the deterministic sampler backwards gives the remaining latent-code operation, finding the code of a real image, covered in the next section.
DDIM as an ODE solver
The deterministic update can be rewritten as one step of Euler's method, the simplest numerical method for an ordinary differential equation (ODE): from the current point, move a step along the current velocity. Change variables to and . By (1), : the clean image plus times unit noise, with growing from 0 at the clean end to 157 at on the CIFAR-10 schedule. (This is a different quantity from the sampler's noise scale in (4) to (6); the overloaded symbol is the paper's.) The DDIM update (5) with becomes
an Euler step with step size for the ODE
The network's noise prediction is the velocity and plays the role of time. Since , the network input is just , the kind of noisy image it was trained on. Check it on the one-pixel example, where , and . The Euler step lands on the DDIM result from before:
An ODE can be integrated in either direction. Run the same update from the clean end toward the noisy end and it encodes an image into its ; run it back and it decodes. Each direction makes discretization error, so the round trip lands near the start, not on it, and the gap shrinks as steps are added. Figure 7 runs both directions on a one-dimensional toy with three data clusters, where the denoiser is exact and only step size causes error:
In the default view (6 steps, ), the code comes out at 0.77 instead of the 1,000-step value 1.04, and decoding it lands at 0.80, between two clusters and 0.50 away from the start. At 20 steps the miss is 0.012, and at 100 steps 0.004. The paper measures the same thing on the CIFAR-10 test set, encoding and decoding with the same and reporting per-dimension squared error with pixels scaled to :
| steps S (encode and decode) | squared error per pixel |
|---|---|
| 10 | 0.014 |
| 20 | 0.0065 |
| 50 | 0.0023 |
| 100 | 0.0009 |
| 200 | 0.0004 |
| 500 | 0.0001 |
| 1000 | 0.0001 |
An error of 0.0009 is a root-mean-square pixel error of 0.03, about 8 levels out of 255. A DDPM has no such round trip, because decoding draws new noise at every step. The encoder is the sampling loop run forwards:
# Encoding: the same update with eta = 0, run from clean to noisy.
def encode(eps_theta, alpha, tau, x): # x: a real image x_0
levels = [0] + list(tau) # 0 -> tau_1 -> ... -> tau_S
for t, s in zip(levels[:-1], levels[1:]): # s > t: noisier
eps = eps_theta(x, t)
x0 = (x - sqrt(1 - alpha[t]) * eps) / sqrt(alpha[t])
x = sqrt(alpha[s]) * x0 + sqrt(1 - alpha[s]) * eps
return x # x_T, the image's codeEquation (7) also connects DDIM to score-based SDEs, which describe diffusion in continuous time and come with a deterministic "probability flow" ODE. Proposition 1 of the paper shows that, with the optimal network, (7) is the probability flow ODE of the variance-exploding SDE, the continuous-time form of the NCSN noise process. That ODE is usually written in with a factor ; (7) is written in with no ½, and the change of variables absorbs the ½. The two samplers still differ: DDIM takes Euler steps in , the score-SDE sampler takes them in , and with few steps those land in different places. The view of sampling as solving an ODE with a learned velocity also links DDIM to Neural ODEs, which Section 5.4 cites alongside normalizing flows when describing the round trip.
Questions you might still have
Do I need to retrain my diffusion model to use DDIM?
No. Every DDIM result in the paper uses one network per dataset, trained with the ordinary DDPM loss at T = 1000. Theorem 1 shows that each member of the σ family has a training objective equal to a weighted DDPM loss plus a constant. If the network had separate parameters for each noise level, the loss weighting γ would not matter and the DDPM-trained network would be optimal for every σ. Real networks share parameters across levels, so the transfer is an idealization that the experiments confirm.
Is DDIM just DDPM with fewer steps?
No. Both can skip steps, since the skipping argument only uses the marginals. The difference is the update: DDPM (η = 1) adds fresh noise at every step, DDIM (η = 0) adds none. With 10 steps on CIFAR-10 the same network scores FID 13.36 under DDIM and 41.07 under DDPM.
Why is it called "implicit"?
In the sense of an implicit probabilistic model: samples come from a latent variable through a fixed procedure, as in a GAN, rather than from a density you can write down step by step. With σ = 0 the latent is the starting noise x_T and the procedure is the deterministic sampling loop. The update formula itself is explicit.
If DDIM is deterministic, where does sample diversity come from?
From x_T. Each sample starts from a fresh Gaussian draw of the full image size (3,072 numbers for a 32×32 RGB image), and the deterministic loop maps different draws to different images. Determinism means the same draw always gives the same image.
Is DDIM always better than DDPM?
In the η family, yes, in every column of Table 1: FID rises with η at every step count, and at 1,000 steps CIFAR-10 goes 4.04, 4.09, 4.29, 4.73 for η = 0, 0.2, 0.5, 1. The exception is σ̂, the larger-variance DDPM from Ho et al.'s code, which beats DDIM at 1,000 steps (3.17 vs 4.04 on CIFAR-10, 3.26 vs 3.51 on CelebA) and is far worse at 10 steps (367.43).
Can DDPM interpolate between images too?
Not by blending one x_T, because a DDPM sample also depends on the fresh noise drawn at every step, so the same x_T gives different images. The paper notes in a footnote that interpolating all T noise draws, as done for NCSN, might work. DDIM needs only the one code.
How does DDIM relate to the probability flow ODE of score SDEs?
With the optimal network they describe the same ODE: DDIM's update is an Euler step of the variance-exploding probability flow ODE written in terms of σ. They differ as samplers, because DDIM takes each Euler step in σ between the chosen levels while the score-SDE sampler steps in t, and with few steps the two land in different places. The Score SDEs explainer covers the continuous-time view.
Footnotes & further reading
- The paper: Song, Meng, Ermon, Denoising Diffusion Implicit Models (Stanford, ICLR 2021). Official code.
- The model and objective DDIM reuses: Ho, Jain, Abbeel, Denoising Diffusion Probabilistic Models (our explainer: DDPM). The α/ᾱ notation map is in DDIM's Appendix C.2.
- The score, SDE and probability flow ODE view: Song et al., Score-Based Generative Modeling through Stochastic Differential Equations (our explainer: Score SDEs), and the denoising score matching objective in NCSN.
- Residual networks read as Euler integration of an ODE: Chen et al., Neural Ordinary Differential Equations (our explainer: Neural ODEs).
- Figure 4 and the reconstruction table are transcribed from the paper's Tables 1 and 2. Figures 2, 5 and 7 use exact denoisers for small Gaussian-mixture datasets, so their errors come from step size alone; the CIFAR-10 schedule values (α₁₀₀ = 0.897, α₅₀₀ = 0.079, σ at t = 1000 ≈ 157) are computed from the linear schedule in the official config.
- For how DDPM, score matching and SDEs fit together, see the diffusion tutorial explainer; for another deterministic generator trained by regression, see flow matching; and for a model whose paper samples with DDIM, see latent diffusion.
How could this explainer be improved? Found an error, or something unclear? I read every message.