Diffusion & Video at Inference Time
Page 12 explains how diffusion works. This page is about the bill. A published sampler spends a hundred forward passes on one image, and the same loop over five seconds of video runs for half an hour.
Diffusion quality comes from a loop, and the loop is the whole bill. Fifty steps, doubled by guidance, is a hundred forward passes of a large network per image. Every technique on this page is a way of cutting that number.
Four levers, ordered by what they cost you. A better ODE solver is free. Feature caching is free and has a ceiling you can derive. Guidance distillation costs a training run and removes the factor of two. Step distillation costs a training run and some diversity.
Video is not images times frames. Joint attention over a clip costs a factor equal to the number of latent frames more than attending to each frame alone, so temporal compression in the autoencoder is worth more than any sampler change.
One sentence for an interview: diffusion serving is compute-bound where LLM serving is bandwidth-bound, which inverts nearly every instinct from page 31.
01The bill, in forward passes
One number governs this entire page. Count it before anything else.
A text-to-image pipeline has three parts and the costs are not close. The text encoder runs once. The decoder runs once, at the end. The denoiser runs once per step, and twice per step when guidance is on.
That last line is the budget. Page 12 makes the point in passing; this page starts from it.
The unit is the number of function evaluations, or NFE: how many times the large network is run to produce one output. NFE is what you are billed for, and it is the only quantity the rest of this page moves.
Two thousand passes down to one is a factor of two thousand. Nobody gets all of it, and the rows are not interchangeable, but the range is why this is a subject and not a setting.
02What one step buys
A step is one increment of a numerical integration. Once you see it that way, the question of where quality falls off has a precise answer.
Sampling is solving the reverse ODE that page 12 derives. The network supplies the derivative; the sampler integrates it. Step count is not a quality dial. It is a discretisation, and fewer steps means a coarser one.
Derivation Why a second-order solver needs roughly the square root of the steps
Take the sampling ODE $\dot{x} = v(x, t)$ integrated from $t=0$ to $t=1$ in $N$ equal steps of size $h = 1/N$. Assume $v$ is smooth enough for a Taylor expansion, which is the usual assumption and the one that fails near $t=1$.
-
$$x(t+h) = x(t) + h\,\dot{x}(t) + \tfrac{1}{2}h^{2}\ddot{x}(t) + O(h^{3})$$Why: plain Taylor. An Euler step keeps the first two terms and discards the rest, so it is wrong by $O(h^{2})$ on that single step.
-
$$\text{global error} \;\approx\; N \cdot O(h^{2}) \;=\; O(h) \;=\; O(1/N)$$Why: you take $N$ steps and each contributes its own error. One power of $h$ is eaten by the step count, which is the standard result that local order two gives global order one.
-
$$\text{second-order method:}\quad N \cdot O(h^{3}) = O(h^{2}) = O(1/N^{2})$$Why: a method that also cancels the $\ddot{x}$ term is locally wrong by $O(h^{3})$, so the same accounting leaves $O(1/N^{2})$.
-
$$\frac{C_{1}}{N_{1}} = \frac{C_{2}}{N_{2}^{2}} \quad\Longrightarrow\quad N_{2} = \sqrt{N_{1} C_{2} / C_{1}}$$Why this is the payoff: equate the two error budgets and the second-order step count is the square root of the first. Ignore the constants and $N_1 = 1000$ becomes $N_2 \approx 32$.
DDIM versus ancestral sampling
The original DDPM sampler injects fresh noise at every step. This is what makes it ancestral: the trajectory is stochastic, so skipping steps does not approximate the same path more coarsely, it walks a different path.
DDIM removes the injected noise. The reverse process becomes deterministic, the sample is a function of the initial noise alone, and skipping steps is now ordinary numerical error and not a different random walk. Its authors report samples 10 to 50 times faster in wall-clock terms than DDPM.
Once the path is fixed, a coarse integration of it is an approximation of a specific answer. You can talk about error, measure it against a fine reference, and reduce it with a better solver.
With noise injected each step there is no fixed answer to approximate. So every fast sampler on the list is deterministic, and why the stochastic ones are still preferred when you want variety more than speed.
Where quality falls off
Truncation error accumulates along the trajectory and not on the output. Under-stepping therefore does not give you a noisy version of the right image. It lands you somewhere else, with a clean version of a slightly different one.
Watch for composition drifting, fine texture thinning and prompt adherence weakening while nothing looks broken. Cutting steps until an image looks bad is the wrong test, because the image stops being the right image well before it stops looking plausible.
The practical floor for training-free samplers sits around ten evaluations. Below that, no amount of solver cleverness helps, because the error is no longer dominated by discretisation. Getting lower means changing the weights, which is section 04.
03Guidance costs double, and nobody says so
Every quoted step count for a guided model is half the real forward-pass count, which is the most common accounting error in the subject.
Classifier-free guidance, from Ho and Salimans, evaluates the model twice per step, once with the prompt and once without, then extrapolates away from the unconditional prediction:
Two evaluations of $\epsilon_\theta$ per step. The guidance scale $w$ is free; the second evaluation is not.
So 50 steps with guidance is 100 forward passes. You can stack the conditional and unconditional inputs into one batch of two, which saves kernel launches and nothing else. The arithmetic is the same, and section 09 explains why that matters here and would not for an LLM.
Folding the factor of two into the weights
Meng and colleagues train a student conditioned on the guidance scale to reproduce the combined output of both branches in one pass. The scale survives as an input; the second forward pass does not.
Guidance distillation is standard in current open checkpoints. The FLUX.1 [dev] model card says the model was "trained using guidance distillation". HunyuanVideo distils "the combined output for unconditional and conditional inputs into a single student model", and reports roughly 1.9 times acceleration from that alone.
A guidance-distilled model no longer has an unconditional branch to extrapolate from. Negative prompts, which are just a different conditional in that slot, stop working the way they did, and the guidance scale becomes a learned input with its own trained range and no longer a free multiplier.
People tune it anyway, well outside the range it was distilled over, and then report that the model is insensitive to guidance. The model is answering a question about a number it was taught to interpret.
The cheaper alternative, if you keep the two branches
Kynkäänniemi and colleagues found that guidance is "clearly harmful toward the beginning of the chain (high noise levels), largely unnecessary toward the end (low noise levels), and only beneficial in the middle". Restricting it to an interval improved ImageNet-512 FID from 1.81 to 1.40 in their setting.
That buys two things. Steps outside the interval need one pass instead of two, so a limited interval is cheaper and, in their measurements, better. It is a rare change that costs nothing.
That covers the budget in forward passes, why a better solver cuts it for free, and why guidance doubles it, which is everything you can do without touching the weights.
The next three sections each change something. Distillation changes the weights, rectified flow changes the training objective, and caching changes what the network is asked to compute.
04Distillation: buying steps with a training run
Below roughly ten evaluations you cannot integrate your way out. You train a new model to take bigger jumps, and you pay for it in something other than compute.
Four families, in the order they appeared. Each one is a different answer to the same question: what should a student be trained to match?
- Progressive distillation (Salimans and Ho, 2022). It trains a student to take one step that lands where two teacher steps land, then make that student the teacher and repeat. Each round halves the step count. They take a model from 8192 steps to 4, reporting FID 3.0 on CIFAR-10 at 4 steps.
- Consistency models (Song and colleagues, 2023). It trains so that every point on a trajectory maps directly to that trajectory's endpoint. If that property holds, one evaluation is enough. They report one-step FID of 3.55 on CIFAR-10 and 6.20 on ImageNet 64.
- Latent consistency models (Luo and colleagues, 2023). This distils the same property into an existing latent diffusion model instead of training from scratch. Their result is 2 to 4 step generation at 768 by 768, from about 32 A100 hours of distillation.
- Adversarial distillation (Sauer and colleagues, 2023, shipped as SDXL-Turbo). A discriminator pushes the student toward the image manifold as well as toward the teacher's answer. One to four steps, matching SDXL in four by their evaluation.
The later work mixes them. SDXL-Lightning describes itself as combining "progressive and adversarial distillation to achieve a balance between quality and mode coverage", which names the tradeoff in its own abstract. DMD2 adds a GAN loss to distribution matching and reports one-step FID of 1.28 on ImageNet 64 and 8.35 zero-shot on COCO.
What a 4-step or 1-step model gives up, concretely
Sharpness is not what suffers, which is the surprise. The DMD2 authors put their 4-step student at FID 19.32 and their 1-step student at 19.01 against an SDXL teacher at 19.36 with 50 steps and guidance scale 6. On a distribution metric the student is level with its teacher.
Diversity is what moves. The same paper states that "our distilled generator experiences a slight degradation in image diversity compared to the teacher models". FID over thousands of samples is not the metric that would show it.
Gandikota and Bau give the mechanism. They find that "distilled models commit to their final image structure almost immediately at the first timestep, while base models distribute structural decisions across many steps", and that distilled models "possess variational directions needed for diversity yet fail to activate them".
"Every seed gives me the same picture." The seed still changes texture and small details. It has stopped changing composition, because composition was decided in the first evaluation and there is only one.
This matters for a gallery product where the user regenerates until something surprises them, and far less for an interactive canvas where the user is steering. Decide which one you are building before you pick a step count.
05Rectified flow: straighten the path, and the steps come free
Integrating a curve needs steps, while a straight line needs one at any resolution. Rectified flow is built on this.
Liu, Gong and Liu proposed learning an ODE whose paths are "as straight as possible" between noise and data. Their claim is that recursively rectifying yields "nearly straight flows that give high quality results even with a single Euler discretization step".
The part worth understanding is what does the straightening, because it is not the interpolant. Training on a linear interpolant still gives a curved flow, because the marginal velocity field averages over every data point a given noise sample might become.
Reflow fixes the coupling. You run the learned ODE to see where each noise sample lands and retrain on those pairs, so each noise sample has one destination, the paths stop crossing, and the field they induce is straight by construction.
Try it Curved path, straight path, and the Euler error at one step
import numpy as np
# A 1-D flow from noise (t=0) to a 3-point data set (t=1), on the
# linear interpolant x_t = (1-t)z + t*x0 that rectified flow uses.
data = np.array([-2.0, 0.5, 3.0])
def v_indep(x, t): # marginal velocity, z drawn independently
s = max(1.0 - t, 1e-6)
lp = -0.5 * ((x[:, None] - t * data[None, :]) / s) ** 2
w = np.exp(lp - lp.max(1, keepdims=True))
m = (w * data[None, :]).sum(1) / w.sum(1) # E[x0 | x_t = x]
return (m - x) / s
def euler(x0, n, v):
x = x0.copy()
for i in range(n):
x = x + v(x, i / n) / n
return x
grid = np.linspace(-6, 6, 4001) # solve the ODE on a grid once
endp = euler(grid, 2000, v_indep) # z -> x1(z), the coupling
def v_reflow(x, t): # same endpoints, re-paired: straight by
g = (1 - t) * grid + t * endp # construction
o = np.argsort(g)
z = np.interp(x, g[o], grid[o])
return np.interp(z, grid, endp) - z
rng = np.random.default_rng(0)
z = rng.normal(size=6000)
ref = euler(z, 2000, v_indep)
print("Euler steps mean |error| vs the 2000-step solution")
print(" independent coupling after reflow")
for n in (1, 2, 4, 8, 16):
a = np.abs(euler(z, n, v_indep) - ref).mean()
b = np.abs(euler(z, n, v_reflow) - ref).mean()
print(f"{n:>7} {a:>8.4f} {b:>8.4f}")
for name, v in (("independent", v_indep), ("reflowed", v_reflow)):
x, L = z.copy(), 0.0 # arc length of the path
for i in range(2000):
s = v(x, i / 2000) / 2000
L += np.abs(s).mean()
x = x + s
print(f"{name:>12}: arc length {L:.4f}, endpoint distance "
f"{np.abs(x - z).mean():.4f}")
1.1292 against an endpoint distance of 1.1292. The left column halves its error each time the steps double, which is section 02's $O(1/N)$. The interpolant is linear in both columns. Only the pairing differs.Scaling Rectified Flow Transformers, the Stable Diffusion 3 paper, is what settled the argument in practice. Esser and colleagues report that rectified flow "surpasses conventional diffusion methods for high-resolution text-to-image" at scale, and the formulation is now what current open image and video models train under. HunyuanVideo states plainly that it uses "Flow Matching for model training".
06Feature caching, and the ceiling you can derive
Adjacent denoising steps compute nearly the same thing. Reusing what they share is the cheapest win here, and the size of the win is not a mystery.
DeepCache is the canonical version for UNets. It exploits what its authors call "the inherent temporal redundancy observed in the sequential denoising steps", reusing high-level features across adjacent steps while updating the low-level ones cheaply. It needs no retraining.
They report 2.3 times on Stable Diffusion v1.5 for a 0.05 drop in CLIP score, and 4.1 times on LDM-4-G for a 0.22 rise in FID on ImageNet.
The cost model is one line. Run the full network on one step in every $N$, and a cheap partial path costing a fraction $r$ of it on the other $N-1$:
The interval you choose is a knob. The reuse fraction $r$ is a ceiling, and the knob cannot get past it.
Try it The cache reuse model, its ceiling, and why it stops helping
# Uniform feature caching: run the full network on 1 step in every N,
# reuse cached features on the other N-1 at a fraction r of the cost.
# Cost per step = (1 + (N-1)*r) / N, so the speedup is:
speedup = lambda N, r: N / (1 + (N - 1) * r)
print(" r=0.10 r=0.20 r=0.29 r=0.40")
for N in (2, 3, 5, 10, 20, 10**7):
row = " ".join(f"{speedup(N, r):6.2f}"
for r in (.10, .20, .29, .40))
print(f"{('N -> inf' if N > 100 else 'N = %d' % N):>8} {row}")
print(" the last row is the 1/r ceiling: the interval cannot")
print(" buy you more than the cheap path already costs")
# Invert it. DeepCache report 2.30x on Stable Diffusion v1.5 at N=5.
N, s = 5, 2.30
r = (N / s - 1) / (N - 1)
print(f"\n{s}x at N={N} implies a cheap path costing r = {r:.3f}")
print(f"which caps that configuration at 1/r = {1/r:.2f}x")
# Caching composes with step reduction, and that is the catch: the
# interval can never exceed the number of steps you have left.
print("\nsteps cfg cache NFE full-pass equivalents")
for steps, cfg, on in ((50, 2, 0), (50, 2, 1), (20, 1, 0),
(4, 1, 0), (4, 1, 1)):
frac = (1 + (steps - 1) * r) / steps if on else 1.0
print(f"{steps:>5} {cfg:>4}x {'on ' if on else 'off':>5}"
f" {steps*cfg:>5} {steps*cfg*frac:>8.1f}")
N -> inf row is the point. Push the interval as far as you like and the speedup stops at 1/r, because the cheap path is still paid on nearly every step. Inverting DeepCache's 2.30x at N=5 gives r = 0.293, capping that configuration at 3.41x. The last table is the awkward part: at 4 steps there is barely an interval to have.That limit is the useful part. A reuse scheme is characterised by how cheap its cheap path is, and how far apart you space the full evaluations matters less, and a single measurement at any interval pins down both.
Where the artefacts appear
DeepCache's own reading of their results: below an interval of 5 "there is only a slight reduction in the quality", while at an interval of 20 the degradation is "non-negligible". What changes at that point is specific and worth quoting, because it tells you what to look for.
They describe "subtle details such as the color of clothing and the shape of the cat" being modified. This is drift and not corruption: the image stays coherent and stops being the image the uncached sampler would have produced, which is exactly the failure a spot check does not catch.
Uniform intervals are the wrong shape
TeaCache starts from the observation that "differences among model outputs are not uniform across timesteps". Some neighbouring steps change the output a great deal and some barely at all, so spacing full evaluations evenly spends effort in the wrong places.
It estimates how much the output will change using the timestep embedding, which is available before the expensive part runs, and skips adaptively. Its authors report up to 4.41 times on Open-Sora-Plan for a 0.07 percent drop in VBench score.
Both are harvesting the same redundancy: the fact that consecutive steps agree. A 4-step model has already spent it. There are only three steps to skip, the interval cannot exceed the step count, and the remaining steps disagree with each other far more than 50 steps of the base model did.
Expect the two techniques to be substitutes, and measure the combination instead of assuming the speedups multiply.
07Video is not images times frames
The arithmetic in this section is the most important on the page and it is four lines long.
A modern video model is a diffusion transformer over a three-dimensional latent. A causal autoencoder compresses the clip in space and in time, the latent is cut into patches, and all of those patches become one flat sequence that attends to itself.
HunyuanVideo publishes every number needed to price that. Its autoencoder uses a temporal factor of 4, a spatial factor of 8 and 16 latent channels. The transformer is 13B parameters, 3072 wide, with 20 dual-stream and 40 single-stream blocks, and its documented 720p configuration is 129 frames at 1280 by 720 in 50 steps.
Try it Price a five-second 720p clip from the published constants
# HunyuanVideo's published 720p configuration, from its paper and repo:
W, H, F, STEPS = 1280, 720, 129, 50
CS, CT, CH = 8, 4, 16 # VAE: 8x spatial, 4x temporal, 16 chans
PS, PT = 2, 1 # patch size on the latent, spatial/temp
D, BLOCKS = 24 * 128, 20 + 40 # width 3072, 60 attention blocks
PARAMS = 13e9
def tokens(w, h, f, ct=CT):
lf = (f - 1) // ct + 1 if ct > 1 else f
return lf * (h // CS // PS) * (w // CS // PT // PS), lf
N, LF = tokens(W, H, F)
print(f"latent grid: {LF} x {H//CS} x {W//CS}, patched -> {N:,} tokens")
print(f"pixels per latent value: {CT*CS*CS*3/CH:.0f}x compression")
attn = 4 * N**2 * D # QK^T and AV, summed over heads
print(f"attention matmuls: {attn/1e12:,.1f} TFLOP per block, "
f"{attn*BLOCKS/1e15:,.2f} PFLOP per step")
clip = attn * BLOCKS * STEPS
print(f"{STEPS} steps -> {clip/1e15:,.0f} PFLOP of attention alone")
print(f"the repo reports 1904 s on one GPU, so {clip/1904/1e12:,.0f}"
" TFLOP/s sustained on that term")
# Why video is not images times frames.
one, _ = tokens(W, H, 1, ct=1)
print(f"\none 720p frame is {one:,} tokens; {LF} of them held jointly")
print(f"is {N//one}x the tokens and {(N/one)**2:,.0f}x the attention,")
print(f"so {(N/one)**2/(N/one):,.0f}x more attention work per frame")
flat, _ = tokens(W, H, F, ct=1) # drop temporal compression
print(f"\nno temporal compression: {flat:,} tokens, "
f"{(flat/N)**2:.1f}x the attention cost")
act = N * D * 2
print(f"\nweights in bf16: {PARAMS*2/1e9:.0f} GB; one full-width "
f"activation: {act/1e6:.0f} MB")
print(f"Q, K, V and output at once: {4*act/1e9:.1f} GB, before any"
" of the MLP")
118,800 tokens for five seconds, longer than many LLM context windows, and it is the output and not the prompt. 1,089x the attention of one frame for 33 latent frames, so 33x more work per frame than generating them separately. And 15.3x, which is what temporal compression is worth before the transformer runs at all.The rule that makes video expensive
Split $N$ tokens into $k$ per-frame groups and attend within each. That costs $k \cdot (N/k)^2 = N^2/k$. Attend over all of them jointly and it costs $N^2$. The ratio is exactly $k$:
Here $k$ is the number of latent frames and not output frames, which is what the autoencoder is for.
So temporal coherence is expensive. It costs a factor equal to the latent frame count, on the dominant term, on every step, and it is the price of making a video a video and not a flipbook.
It also explains where the engineering effort has gone. Temporal compression divides $k$ and attention is quadratic, so the 4 times compression in the autoencoder is worth about 15 times on attention. No sampler change is worth 15 times.
Frame count scaling and the memory wall
Doubling the clip length doubles $k$, doubles the token count, and quadruples attention. A ten-second clip is not two five-second clips, and a capacity plan built at five seconds does not survive the product asking for ten.
Memory follows the same shape. The weights are a fixed 26 GB in bf16, but one full-width activation at 118,800 tokens is 730 MB, and the attention inputs and output alone are 2.9 GB before the MLP. HunyuanVideo's repository asks for 60 GB at this configuration.
One caveat on that 520 PFLOP figure: it counts the two attention matmuls and nothing else, so it is a floor on the cost of a step and not an estimate of it. The projections, the MLPs and the normalisations all sit on top.
Sliding Tile Attention measured where the time goes: generating "a 5-second 720P video, attention alone takes 800 out of 945 seconds of total inference time". At that ratio, halving the step count cannot help you as much as making attention sub-quadratic.
Which is what the recent work does. Sparse VideoGen separates spatial from temporal heads and reports 2.28 and 2.33 times end to end on CogVideoX-v1.5 and HunyuanVideo. Radial Attention exploits spatiotemporal energy decay for an $O(n \log n)$ scheme, reporting up to 3.7 times on inference.
Prefix caching works for an LLM because the sequence is the prompt: it arrives fixed, it is often shared, and its keys and values can be computed once and kept.
In a diffusion transformer the sequence is the output. Every token changes at every step, so there is no prefix, no KV cache, and nothing to share between two users asking for different videos. The redundancy that does exist is across timesteps, which is section 06, and it is a much smaller prize.
08Build this
The claim you should least take on trust is that quality degrades gracefully with step count. Measure your own frontier, on your own checkpoint, in an afternoon.
Take one checkpoint and one fixed set of 32 prompts. You are not looking for the best images. You are looking for the step count where the sampler stops reproducing its own 100-step answer, which is a different and much sharper question.
- Generate a reference set at 100 steps with a fixed seed per prompt. This is the ground truth for everything that follows, and everything is compared against it.
- Sweep step counts over 2, 4, 8, 16, 32 and 50 for DDIM and for a second-order solver, same seeds throughout. Record LPIPS against the reference and CLIP score against the prompt.
- Plot both. LPIPS will move long before CLIP does, so note the step count where each one starts to bend, since they differ.
- Enable feature caching at intervals 2, 3, 5 and 10 at your chosen step count. Fit $N/(1+(N-1)r)$ to the measured speedups and read off $r$.
- Now measure diversity in place of quality. Generate 32 samples of one prompt with 32 different seeds, from the base model and from a distilled few-step model. Embed all 64 and compare mean pairwise distance within each group.
You now have every lever: solvers, guidance, distillation, straightened paths, caching, and the token arithmetic that makes video its own problem.
What remains is putting it on a machine. Section 09 is the part that surprises people who arrive from LLM serving, and it is one division.
09Serving diffusion, and why it is not LLM serving
Page 31 taught you that batching is nearly free and bandwidth is the enemy. Both halves of that invert here, and the reason is one line of arithmetic.
Repeat page 31's roofline count for one denoiser pass. A model with $P$ parameters in bf16 moves $2P$ bytes of weights. It does about two floating-point operations per parameter per token, over $T$ tokens per sample and $B$ samples:
Same derivation as decode, one extra factor. For LLM decode $T = 1$ and the intensity is $B$. Here $T$ is the whole latent.
A 1024 pixel image with an 8 times autoencoder and a patch size of 2 is 4,096 tokens. The five-second clip above is 118,800. An H100 SXM lists 989 bf16 TFLOP/s against 3.35 TB/s, putting its ridge point near 295 operations per byte.
So a diffusion step is past the ridge at batch 1, by a factor of roughly 14 for that image and 400 for that clip. Argus, a text-to-image serving system, states the consequence directly: these models are "highly compute-bound".
Four things follow, and each contradicts an instinct from page 31.
- Batching buys much less. The arithmetic units were already saturated at batch 1, so a larger batch mostly adds latency. There is no equivalent of the decode batching cliff, because you are not waiting on memory.
- There is no KV cache to run out of. No per-request state grows with time. Capacity is weights plus activations, and activations scale with resolution and clip length and not with how long a user has been talking.
- Scaling out means splitting one sample. More GPUs cannot serve more requests per second at the same latency if one request already fills the device. So video inference reaches for sequence parallelism across GPUs before it reaches for more concurrency.
- Quantisation pays differently. Weight-only quantisation is the standard LLM win because it cuts bytes moved, and bytes moved is not the constraint here. You need low-precision arithmetic, meaning FP8 activations as well as weights, or you get nothing.
Requests in a diffusion batch all run the same number of steps and finish together. The ragged-completion problem that makes continuous batching essential for an LLM barely exists, because generation length is a server setting and is not something the model decides.
The scheduling problem that replaces it is admission: one request can occupy the whole device for minutes, and there is no chunking trick that makes a denoising step interruptible.
The hardware conclusion is to choose on arithmetic throughput per pound and not bandwidth per pound. A card that wins on LLM decode because of its memory system can lose here, and the ranking you carry over from a serving benchmark will be the wrong ranking.
10What breaks
Almost none of these are crashes. They are quality regressions that no dashboard is watching for.
- Caching drift passes the spot check. Details move and the image stays plausible, so nobody notices in review. The only way to see it is to diff against an uncached reference at a fixed seed.
- The step count was tuned on one sampler and copied to another. Twenty steps of a second-order solver and twenty of DDIM are not the same amount of integration. The number is not portable.
- A distilled model is benchmarked at 4 steps against a base model at 50. That comparison measures the distillation and not the checkpoint, and it is how a worse model wins an evaluation.
- Diversity collapse arrives as a support ticket. FID over thousands of samples is a distribution distance and will not show it, while per-prompt pairwise distance will.
- Guidance scale stopped doing anything. The checkpoint is guidance-distilled, the scale is now a trained input with a range, and it is being driven outside it.
- The video memory plan was made at five seconds. Attention is quadratic in clip length. Doubling the duration roughly quadruples the dominant term and the budget was linear.
- Batch size was raised to improve throughput and only raised latency. See section 09: the arithmetic units were busy already.
- A negative prompt quietly stopped applying. It was a second conditional in the unconditional slot, and the model no longer evaluates that slot.
11Where you meet this in the wild
The same budget, four different things worth spending it on.
Throughput per GPU is the unit economics and nobody is watching a progress bar. Every free lever gets pulled and the paid ones get a quality review.
Guidance-distilled checkpoint, a second-order solver in the 20s, caching on, FP8 arithmetic.
The user is drawing and the image updates as they move. A latency budget under a couple of hundred milliseconds leaves room for one to four evaluations and nothing else.
Use a distilled few-step model. The diversity loss costs nothing here, because the user is the source of variation.
Minutes per clip on one GPU. The step count is already low and attention is most of the remaining bill, so the work is in the attention kernel and in splitting one sample across devices.
Sparse or tiled attention, sequence parallelism, the most aggressive autoencoder you can tolerate.
More concurrency does not help, because one request already owns the device.
On a phone or a laptop the constraint is memory and sustained power, and peak throughput matters less. Attack token count before step count.
A deeply compressing autoencoder plus a small step count. SANA's 32 times compression is the shape of the answer.
12Interview questions
BeginnerWhy is a diffusion model expensive to run compared with a GAN?
Because quality comes from an iterative loop rather than a single pass. A GAN generator produces its output in one forward evaluation. A diffusion sampler integrates a reverse process, and each increment of that integration is a full forward pass of a large network.
The other two components barely count. The text encoder runs once per request and the decoder runs once at the end, so a 50-step generation is essentially 50 network evaluations, doubled to 100 if classifier-free guidance is on.
That is the whole cost model, and it is why every acceleration technique for diffusion is ultimately about reducing the number of function evaluations or making each one cheaper.
BeginnerWhat does classifier-free guidance cost at inference, and how do people avoid it?
It doubles the forward passes. Each step evaluates the model once with the prompt and once without, then extrapolates away from the unconditional prediction, so a 50-step guided generation is 100 passes rather than 50.
The usual fix is guidance distillation: train a student that takes the guidance scale as an input and reproduces the combined output in a single pass. FLUX.1 [dev] ships this way, and HunyuanVideo reports about 1.9 times acceleration from it.
A cheaper partial fix is to apply guidance only over an interval of noise levels, which removes the second pass outside that interval and, in the paper that proposed it, also improved sample quality.
IntermediateYou need to get from 50 steps to 10. What do you try, and in what order?
Start with things that do not touch the weights. Swap the sampler for a higher-order ODE solver, since that changes the step floor by roughly a square root and costs nothing. Then check whether guidance can be limited to an interval, which removes a second pass from the steps outside it.
Next, feature caching. Measure the speedup at two intervals, fit the reuse model, and you get both the achievable range and the ceiling from the same two measurements.
Only then reach for distillation, because it costs a training run and changes the model's behaviour. Guidance distillation first, since it removes a factor of two with little quality argument. Step distillation last, because it is the one that costs diversity.
IntermediateWhat does a 4-step distilled model give up, and how would you measure it?
Mostly diversity, not fidelity. The DMD2 authors put their 4-step and 1-step students at FID 19.32 and 19.01 against an SDXL teacher at 19.36, so on a distribution metric the student is level, and they separately report a degradation in image diversity.
The mechanism is that distilled models fix their image structure in the first evaluation, where a base model spreads structural decisions over many steps. With one evaluation there is nowhere for the seed to change composition, so it changes only surface detail.
FID will not show this, because it compares distributions over thousands of samples. Measure it directly: generate many samples of a single prompt with different seeds, embed them, and compare mean pairwise distance against the base model on the same prompts.
IntermediateWhy do straighter probability paths need fewer sampling steps?
Because Euler integration of a straight line is exact at any step size. All the discretisation error comes from curvature, so a perfectly straight path can be traversed in one step and a curved one cannot.
The subtlety is that a linear interpolant does not by itself give a straight flow. The marginal velocity field averages over every data point a given noise sample might become, and that average bends. Straightness is a property of the coupling between noise and data, not of the schedule.
Reflow fixes this by running the learned ODE, recording where each noise sample actually lands, and retraining on those pairs. With one destination per starting point the paths stop crossing and the induced field is straight.
DeepDerive the speedup from feature caching and state its limit.
Let the full network cost 1 and the cached path cost a fraction r of it, and run the full network once every N steps. Average cost per step is (1 + (N-1)r)/N, so the speedup is N divided by that, which is N/(1 + (N-1)r).
Take N to infinity and the speedup tends to 1/r. The interval is a knob but the reuse fraction is a ceiling, so a scheme is characterised by how cheap its cheap path is rather than by how far apart the full evaluations are spaced.
This is usable in reverse. DeepCache report 2.30 times on Stable Diffusion v1.5 at interval 5; inverting gives r of about 0.293 and a ceiling near 3.4 times, which tells you not to spend a week tuning the interval. It also predicts that caching will not compose with a 4-step model, because the interval cannot exceed the step count.
DeepWhy is generating a five-second video not the same as generating 120 images?
Because the whole clip is one attention sequence. Split N tokens into k per-frame groups and attend within each and you pay k times (N/k) squared, which is N squared over k. Attend jointly and you pay N squared. The ratio is exactly k, the number of latent frames.
Concretely, HunyuanVideo compresses 129 frames of 1280 by 720 by 8 in space and 4 in time, patchifies by 2 spatially, and gets 33 by 45 by 80, which is 118,800 tokens. That is 33 times the tokens of one frame and 1,089 times the attention, so 33 times more attention work per frame of output.
Two consequences. Temporal compression in the autoencoder is worth more than any sampler change, since dropping it here would cost about 15 times on attention. And doubling the clip length roughly quadruples the dominant term, so duration budgets are not linear.
DeepHow does serving a diffusion model differ from serving an LLM?
The roofline position is opposite. One denoiser pass moves 2P bytes of weights and does roughly 2PTB operations over T tokens per sample and B samples, so arithmetic intensity is TB. For LLM decode T is 1 and intensity is the batch size; for a 1024 pixel image T is about 4,096 and for a five-second clip it is over 100,000.
Against an H100's ridge point near 295 operations per byte, a diffusion step is compute-bound at batch 1. Batching therefore buys much less than it does during decode, and raising it mostly adds latency.
The rest follows. There is no KV cache, so capacity is weights plus activations rather than context length. Scaling out means sequence parallelism over one sample rather than more concurrent requests. Weight-only quantisation gains little, because bytes moved was never the constraint. And hardware should be chosen on arithmetic throughput rather than on memory bandwidth.
13Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.
Other techniques for this problem
A scoped slice of the full Technique Map: every technique this page covers, grouped by what it solves.