←Home KnowML
TL;DR

Every request runs in two phases. Prefill reads the whole prompt in one parallel forward pass and saturates GPU compute. Decode then emits tokens one at a time, and each step re-reads every weight to produce a single token, so it is starved of work and limited by memory bandwidth.

That asymmetry sets everything else. Prefill decides time to first token; decode decides inter-token latency. Batching barely helps prefill and is the only thing that helps decode. The whole discipline is scheduling work so neither phase stalls the other, which is what continuous batching, chunked prefill and prefix caching each attack from a different angle.

One sentence for an interview: serving an LLM is a scheduling problem wrapped around a memory-bandwidth problem, and the batch size is the dial that connects them.

01The two phases, and why everything follows

Read this section properly and the rest of the page is mostly consequences.

A request arrives as a prompt. The server turns it into tokens, then runs the model twice in two very different ways.

  • Prefill processes the entire prompt in a single forward pass. Every token is available at once, so they all go through the model in parallel. It ends by producing the first output token.
  • Decode produces every token after the first, one per forward pass. Each pass takes the single token just generated and appends to the KV cache built during prefill.

The two look similar in code. They behave nothing alike on hardware.

Prefill vs decode — same weights, opposite hardware profile
PREFILL · one pass all 6 prompt tokens enter together the whole model read once from memory token 1 compute-bound 6 tokens of work per weight load DECODE · one pass per token one token in, one token out, repeatedly just the last token the whole model read again, every single step token 2 ↻ then token 3, 4, 5… bandwidth-bound 1 token of work per weight load
The indigo box is the same weights in both halves, and it is read in full on every pass. Prefill amortises that read across all 6 prompt tokens at once, so the arithmetic units stay busy. Decode pays the identical read to produce one token, then pays it again — which is why it is starved of work, and why adding more concurrent requests is nearly free during decode and nearly useless during prefill.

Prefill has thousands of tokens to multiply against each loaded weight, so the GPU's arithmetic units stay busy. Decode has one token per request. It reads the entire model out of memory to compute a single token's worth of arithmetic, then does it again.

The Sarathi-Serve authors put it plainly: prefill iterations "saturate GPU compute", while decode iterations have "low compute utilization" because each one processes only a single token per request.

The sentence that makes it click

Decode is not slow because the model is big. Decode is slow because you are moving a large model through the memory bus to do a tiny amount of arithmetic with it. Adding more requests to the same pass costs almost nothing, because the weights are already in flight.

That last observation is the entire economics of LLM serving. Batching is close to free during decode, and nearly useless during prefill.

02The request lifecycle, end to end

Follow one request through a modern server. Each stage below exists because of something in section 01.

  1. Tokenize. The prompt becomes token IDs. Cheap, but it decides your real input length, which is what you are billed and rate-limited on.
  2. Admission. The scheduler decides whether there is room. Room means KV cache blocks as well as weights, and the cache is usually what runs out first.
  3. Prefill. It runs one forward pass over the whole prompt, populates the KV cache for every prompt token, and emits token one.
  4. Decode loop. One forward pass per token. Each pass reads the whole model and appends one entry per layer to the cache.
  5. Eviction. The request finishes or is cancelled, and its cache blocks return to the pool for the next request.

Two things about this list are worth pausing on.

First, the KV cache is per request and grows with every token. It is not a fixed overhead you pay once. On long conversations it routinely exceeds the size of the model weights, which is why page 23 spends so long on managing it.

Second, steps 3 and 4 compete. A server running both is constantly choosing whether to admit a new request's prefill or advance the decodes already in flight. That choice is the scheduler, and section 04 is about what happens when it chooses badly.

03The four numbers that define a serving system

Almost every argument about inference is an argument about which of these four you are optimising, and which you are quietly sacrificing.

  • Time to first token (TTFT). This is how long the user waits before anything appears. Prefill dominates it, so it grows with prompt length, and it decides whether a chat interface feels alive.
  • Inter-token latency (ITL), also called time per output token. The gap between successive tokens once streaming starts. Dominated by decode, and roughly flat per token. It sets the perceived reading speed.
  • Throughput. This is the total tokens per second across all concurrent requests, and it decides your cost per million tokens.
  • Goodput. This is throughput counting only the requests that met their latency target. A server can post excellent throughput while failing every user, and goodput exposes that.

The first three are in direct tension. Raising concurrency raises throughput, because decode batches get fuller. It also raises TTFT and ITL, because each request waits behind more work. There is no setting that maximises all of them.

Why "how fast is it" is not a question

A single tokens-per-second figure is meaningless without saying whether it was measured at batch size 1 or 256. At batch 1 you get the best possible ITL and terrible throughput. At batch 256 you get the reverse. Benchmarks that quote one number have chosen a point on that curve and not told you which.

The practical version: pick a latency budget first, then find the highest throughput that stays inside it. Goodput measures exactly this, which is why the DistServe authors built their system around optimising it and not raw throughput.

04The math: why batching is the only lever during decode

One short derivation explains why every serving system is built around keeping the batch full.

Take a model with $P$ parameters stored in FP16, so two bytes each. Consider one decode step with a batch of $B$ requests. Ignore the KV cache for a moment and count only the weights.

Bytes moved. Every weight is read once, whatever the batch size:

$$\text{bytes} = 2P$$

Arithmetic performed. A forward pass costs roughly two floating-point operations per parameter per token, and there are $B$ tokens in flight, one per request:

$$\text{FLOPs} = 2PB$$

The ratio of those two is the arithmetic intensity, the quantity page 23 uses to place an operation on the roofline:

$$\text{intensity} = \frac{2PB}{2P} = B$$

The model size cancels, and decode's arithmetic intensity equals the batch size.

This small result carries the whole section. A 7B model and a 70B model sit at the same point on the roofline during decode. Only concurrency moves you along it.

A modern datacentre accelerator has a ridge point in the hundreds of FLOPs per byte, so at batch 1 you are running at a small fraction of the hardware's arithmetic capability, and no amount of kernel optimisation changes that. The only way out is more requests in the same pass. Compute your own ridge point with the project on page 23 and do not trust a number from a blog post.

Prefill sits at the opposite end. A prompt of $L$ tokens gives intensity of roughly $LB$, and $L$ is in the thousands, so prefill is compute-bound almost immediately. So one long prompt can occupy the GPU completely.

05The scheduler: continuous batching and chunked prefill

Two ideas, each fixing a way the naive scheduler wastes the machine.

Continuous batching

The obvious way to batch is to collect requests, run them together, and return when all are done. It fails badly here, because generation lengths differ wildly. One request wanting 2,000 tokens holds the whole batch hostage while nine requests that finished at 50 tokens sit idle in it.

Continuous batching, introduced as iteration-level scheduling in the Orca paper, schedules per step and not per request. After every decode step the scheduler drops finished requests and admits waiting ones. The batch works like a revolving door and is not a fixed cohort.

It is the largest throughput win in modern serving, and it is why a naive model.generate() loop is not a serving system.

Chunked prefill

Continuous batching leaves one problem unsolved. Prefill and decode still contend for the same GPU, and prefill is the bigger, greedier job.

When a request arrives with a 30,000-token prompt, its prefill occupies the device for a long single step. Every request already streaming stalls for that entire duration. Users see their output freeze mid-sentence because somebody else pasted a document. This is head-of-line blocking, and it shows up as an ITL spike unrelated to the affected users' own requests.

Chunked prefill, from the Sarathi-Serve paper, splits a long prefill into fixed-size pieces and interleaves them with ongoing decodes. Each scheduler step carries one chunk of prefill plus all the decodes, so no single step is long enough to starve anyone.

The tradeoff it makes explicit

Chunking does not make prefill faster. It makes prefill interruptible, trading a little total throughput for a large reduction in worst-case latency. Sarathi-Serve is titled "Taming Throughput-Latency Tradeoff" for exactly this reason: the technique does not remove the tradeoff, it gives you a knob to set it deliberately.

A related idea takes the separation further. Prefill/decode disaggregation runs the two phases on different GPUs entirely, so neither can interfere with the other, at the cost of shipping the KV cache between them. DistServe is the reference design, and it optimises goodput and not raw throughput.

06Prompt caching: the cheapest win available

Most production traffic re-sends the same prefix thousands of times. Prefill computes it from scratch every time unless you stop it.

A system prompt, a few-shot block, a long document being asked about repeatedly, an agent loop replaying its own history each turn. In all of these the first several thousand tokens are byte-identical across requests. Prefill is deterministic, so their KV cache entries are identical too.

Prefix caching keeps those entries and reuses them. A cache hit skips the prefill work for the shared prefix entirely, which cuts TTFT directly and frees compute for everyone else.

SGLang generalised this as RadixAttention, storing cached prefixes in a radix tree so requests share the longest prefix they have in common and do not need an exact match. It is the main reason SGLang does well on agent and few-shot workloads, where every request in a batch shares a long head and differs only in the tail.

Two properties are worth remembering, because they explain the API design you see from model providers:

  • It only works on a prefix. Cache reuse stops at the first differing token, because every entry after that point depended on it. Putting a timestamp at the top of your system prompt invalidates everything below it.
  • It is a memory-for-compute trade. Cached prefixes occupy the same pool the active requests need, so caching too aggressively reduces how many requests you can run at once.

The practical consequence is a prompt-ordering rule: put the stable content first and the variable content last. Static instructions, then retrieved documents, then the user's turn. Reversing that order can cost every cache hit you would otherwise have got.

07Build this

The throughput-latency frontier is the one thing on this page you should not take on trust. It takes an afternoon to draw yours.

Project Draw the latency wall, then break it on purpose ~4 hours · vLLM + one GPU

Serve a small model, drive it with increasing concurrency, and plot TTFT and ITL against throughput. You are not benchmarking the model. You are finding the point where your server stops trading gracefully and starts falling over.

  1. Start a server with a 1B to 8B model. A free Colab T4 is enough, because the shape of the curve matters more than the absolute numbers.
  2. Write a load generator that fires $N$ concurrent requests with fixed prompt and output lengths, and records per-request TTFT and per-token arrival times. Do not use a benchmarking harness. Recording the timestamps yourself is what makes the metrics stop being abstract.
  3. Sweep $N$ over powers of two, from 1 up to whatever the server accepts. Plot throughput on one axis and TTFT and ITL on the other.
  4. Now cause head-of-line blocking by holding a steady stream of short requests and injecting one request with a very long prompt. Watch the ITL of the short requests while the long prefill runs.
  5. Restart the server with chunked prefill enabled and repeat step 4 unchanged.
You'll know it worked when the sweep produces a hockey stick: throughput climbs while latency barely moves, then bends sharply upward at a concurrency the hardware picked, not you. That knee is your real capacity, and it sits well below the point where the server starts refusing requests.
What the injection teaches. Without chunking, the long prefill shows up as a single fat spike in everyone else's inter-token latency, at a moment none of those users did anything unusual. With chunking the same work spreads across many steps and the spike flattens into a small, wide bump. That before-and-after pair is section 05's argument as a measurement rather than a claim, and it is the clearest demonstration of why interruptibility is worth paying throughput for.

08The engines, and how to choose

There are four serious options, and they target different workloads and do not compete as implementations of the same thing.

vLLM

The default for GPU serving. Originated PagedAttention, has the broadest model coverage, and is the easiest to get to good throughput without tuning. Its paper reports 2 to 4 times the throughput of the prior state of the art at comparable latency.

Reach for it when you are serving on GPUs and have no reason to do otherwise.

SGLang

Built around RadixAttention and structured generation. Its advantage shows up when requests share long prefixes, which is the normal case for agents, few-shot prompting and repeated queries over one document. The paper reports up to 6.4 times higher throughput than prior systems on those workloads.

Reach for it when your traffic has heavy prefix overlap or you need constrained output formats.

llama.cpp

It is written in C/C++ and runs anywhere, including CPU-only machines, Apple silicon and phones, and it consumes GGUF. It is optimised for one or a few users and not for a busy batch.

Reach for it for local, edge and single-user work. Page 30 covers this path in full.

Not the tool for a multi-tenant service.

TensorRT-LLM

This is NVIDIA's engine. It compiles the model into a hardware-specific plan ahead of time, which buys performance that runtime-dispatched engines cannot match on the same card. The cost is a build step, tighter version coupling, and less flexibility when you swap models.

Reach for it when you are locked to NVIDIA hardware, the model is stable, and the last increment of performance is worth real engineering time.

The honest default is vLLM. Move to SGLang if you can show your traffic shares prefixes, and to TensorRT-LLM only after you have measured that you are leaving enough on the table to justify the build pipeline.

09What breaks

Failures in serving are rarely a crash. They are a latency number quietly going wrong for reasons that are not in your logs.

  • The KV cache runs out before the weights do. A server that loads fine at startup can refuse requests an hour later, because concurrency times context length exceeded the block pool. Capacity depends on traffic shape as well as model size.
  • One user's long prompt degrades everyone. This is section 05's head-of-line blocking, and it looks like random latency spikes until you correlate against input length.
  • Preemption thrash. When the cache fills, schedulers evict and later recompute a request's prefill. Under sustained pressure a server can spend a serious fraction of its time redoing work it already did.
  • The average hides the failure. Mean TTFT looks fine while the 99th percentile is many times worse. Serving is a tail-latency discipline; report percentiles or you are not reporting anything.
  • Benchmarks measured at the wrong point. A throughput number from batch 256 tells you nothing about how the product feels, and a latency number from batch 1 tells you nothing about what it costs.
  • Speculative decoding is not free. It reduces latency when the draft model is accepted often, and wastes compute when it is not. Its win depends on the workload, and page 23 covers the mechanism.

10Where you meet this in the wild

The same four numbers, weighted differently, produce very different systems.

A chat product

TTFT matters most here, since users forgive slow streaming far more readily than a blank screen. Prefix caching on the system prompt and chunked prefill to protect against long pastes both pay for themselves immediately.

A batch pipeline

Summarising a million documents overnight. Nobody is waiting, so latency is irrelevant and throughput is the only metric. Run concurrency far past the knee, deep into the region a chat product could never tolerate.

An agent loop

Every step replays the accumulated history, so requests share enormous prefixes. Prefix caching was built for this workload, and the difference between having it and not having it is large.

Code completion in an editor

ITL barely matters because completions are short, while TTFT matters a great deal, because a suggestion that arrives after the developer has typed the next character is worthless. Use small models with aggressive caching, and skip the big models served well.

11Interview questions

BeginnerWhat are prefill and decode, and why does the distinction matter?

Prefill processes the entire input prompt in one parallel forward pass and produces the first output token. Decode then generates each subsequent token in its own pass. The distinction matters because they have opposite hardware profiles: prefill has many tokens per weight load and saturates compute, while decode has one token per request per pass and is limited by memory bandwidth. Almost every serving optimisation targets one phase or the other, so naming the phase is the first step in any performance discussion.

BeginnerWhat is TTFT, and what is inter-token latency?

Time to first token is how long a user waits before any output appears, and it is dominated by prefill, so it grows with prompt length. Inter-token latency is the gap between successive streamed tokens, dominated by decode, and roughly constant per token. They matter to different products: TTFT decides whether a chat interface feels responsive, while inter-token latency decides whether the stream reads comfortably once it starts.

IntermediateWhy is batching so much more effective during decode than during prefill?

Because decode is memory-bandwidth-bound and prefill is not. A decode step reads every weight to compute one token per request, so the weights are already in flight and adding another request costs almost no extra memory traffic. Its arithmetic intensity works out to roughly the batch size, independent of model size, so more requests move you along the roofline toward the compute limit. Prefill already has thousands of tokens per weight load and is compute-bound, so extra requests mostly queue for the same saturated arithmetic units.

IntermediateWhat is continuous batching and what problem does it solve?

It schedules at the granularity of a single decode step rather than a whole request, so finished requests leave the batch and waiting ones join it after every step. The problem it solves is that generation lengths vary a great deal. With static batching, a batch cannot retire until its longest member finishes, so short requests occupy slots while producing nothing. Iteration-level scheduling, introduced in the Orca paper, keeps the batch full instead and is the largest single throughput win in modern serving.

IntermediateUsers report their output freezing mid-stream at random. What would you check?

Correlate the freezes against other requests' input lengths. The likely cause is head-of-line blocking: a long prefill occupies the GPU for one long scheduler step, and every request already streaming stalls for its duration. The affected users did nothing unusual, which is why it looks random from their side. The fix is chunked prefill, which splits a long prefill into pieces interleaved with ongoing decodes, so no single step is long enough to starve anyone. Preemption and recompute under cache pressure can produce a similar signature and is worth ruling out.

DeepWhy is goodput a better target than throughput?

Because throughput can be raised without limit by accepting more concurrency, and past a certain point every additional request is served too slowly to be worth anything. A server can post an excellent tokens-per-second figure while missing its latency target on every request. Goodput counts only the requests that met their target, so it cannot be improved by degrading service. It also makes capacity planning honest, since it answers how many users you can serve acceptably rather than how much work the hardware can be made to do.

DeepYour system prompt is 4,000 tokens and TTFT is too high. Walk through the options.

Start with prefix caching, since a fixed system prompt is identical across every request and its KV entries can be computed once and reused, removing that prefill work entirely on a hit. Check the prompt's ordering next: cache reuse stops at the first differing token, so anything variable near the top, a timestamp or a user ID, invalidates everything below it and must move to the end. If caching is already in place, chunked prefill will not reduce this request's own TTFT but stops it inflating everyone else's latency. Only then consider shortening the prompt or serving a smaller model, and measure the tail rather than the mean throughout, because TTFT problems usually live in the 99th percentile.

DeepWhen would you choose SGLang or TensorRT-LLM over vLLM?

SGLang when the traffic has heavy prefix overlap, which is normal for agent loops, few-shot prompting and repeated questions about one document, because RadixAttention shares cached prefixes through a radix tree rather than requiring exact matches. It is also the stronger choice when constrained or structured output matters. TensorRT-LLM when the hardware is fixed to NVIDIA, the model is stable, and an ahead-of-time compiled plan is worth the build step and version coupling it introduces. vLLM remains the sensible default, and the case for moving should come from a measurement on your own traffic rather than from published benchmarks run on someone else's.

12Go deeper

🔧
Interactive
LLM Visualization
Brendan Bycroft — walk a token through every matrix of a real GPT, one operation at a time. The best way to make the forward pass concrete.
📝
Article
Making Deep Learning Go Brrrr From First Principles
Horace He — compute, bandwidth and overhead as three separate regimes. The clearest explanation of why decode is bandwidth-bound.
📝
Article
All About Transformer Inference — How To Scale Your Model
The arithmetic of prefill, decode and sharding, worked out properly rather than asserted.
📄
Paper
Efficient Memory Management for LLM Serving with PagedAttention
Kwon et al. — the vLLM paper — arXiv:2309.06180
📄
Paper
Orca: A Distributed Serving System for Transformer-Based Generative Models
Yu et al., OSDI 2022 — where iteration-level (continuous) batching comes from.
📄
Paper
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Agrawal et al. — the chunked-prefill paper — arXiv:2403.02310
📄
Paper
SGLang: Efficient Execution of Structured Language Model Programs
Zheng et al. — RadixAttention and prefix reuse — arXiv:2312.07104
📄
Paper
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving
Zhong et al. — where the goodput framing comes from — arXiv:2401.09670
📄
Paper
Fast Inference from Transformers via Speculative Decoding
Leviathan et al. — exact same output distribution, no retraining — arXiv:2211.17192
🔧
Tool
vLLM
The default GPU serving engine, and the reference implementation of PagedAttention.
🔧
Tool
SGLang
RadixAttention and structured generation, for workloads with heavy prefix sharing.
🔧
Tool
TensorRT-LLM
NVIDIA's ahead-of-time compiled engine, for when the hardware is fixed and the model is stable.
🔧
Docs
Prompt caching
What prefix caching looks like as a product API, including why ordering your prompt correctly is what makes it hit.

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

Other techniques for this problem

A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.

My Notes — 30 Running Models Locally

Free notes

Highlights on this page