LLM Inference at Scale
Page 31 served one model on one GPU. This is what happens when the unit stops being a device and becomes a fleet: the two phases move to separate machines, the weights drop to four bits, and one base model carries a thousand jobs.
Page 31 had one dial: batch size on one device. At fleet scale there are three, and they are about placement, precision and sharing.
Placement. Prefill and decode want opposite things, so run them on separate GPU pools and ship the KV cache between them. Precision. Weights go to FP8 or four-bit block formats, and the KV cache goes with them. Sharing. One base model serves hundreds of LoRA adapters, and one MoE layer spreads its experts across many GPUs.
One quantity keeps reappearing: KV bytes per token divided by prefill FLOPs per token. It sets a break-even bandwidth, and that number decides whether disaggregation, offload and cache reuse are worth doing. Section 02 derives it and the rest of the page spends it.
One sentence for an interview: at fleet scale you stop tuning a server and start deciding which bytes move, how wide they are, and who they belong to.
01What changes when the unit is a fleet
Everything on page 31 still applies. It just stops being sufficient.
Continuous batching and chunked prefill both work inside one engine on one set of GPUs. They interleave prefill and decode so that neither starves the other, which is the best you can do while the two phases share a device.
But they still share a device. A prefill chunk and a decode step compete for the same SMs, the same HBM bandwidth and the same KV block pool. Chunking makes the interference smaller and more even. It does not remove it.
The move at fleet scale is to stop sharing. Give prefill its own GPUs, give decode its own GPUs, and pay a transfer cost to connect them.
Three more consequences follow from a fleet, and each gets a section below:
- Bytes are the budget. Weights in FP8 or four bits, KV cache in FP8 or lower, and both of those change what fits and what moves.
- One model has to serve many jobs. Hundreds of LoRA adapters on one base, or hundreds of MoE experts across many GPUs.
- Throughput stops being the target. A fleet is provisioned against a latency contract, so the metric is goodput.
02Prefill/decode disaggregation, and where the crossover is
The argument for splitting is easy. The arithmetic for when to stop is the useful part.
The two phases contend for the same GPU and also want different machines. Prefill is compute-bound and latency-sensitive, so it wants wide tensor parallelism and small batches. Decode is bandwidth-bound, so it wants the largest batch the KV pool will hold.
A single pool has to pick one configuration for both phases, while two pools can each be sized, sharded and batched for its own phase, and the prefill:decode GPU ratio can follow the traffic.
The systems literature is consistent about the size of the win. DistServe reports serving "7.4x more requests or 12.6x tighter SLO" than the systems it compared against, while staying inside latency constraints for over 90% of requests. Splitwise reports "1.4x higher throughput at 20% lower cost", or "2.35x more throughput with the same cost and power budgets".
At production scale, Moonshot's Mooncake describes a KV-cache-centric disaggregated architecture and reports that it "enables Kimi to handle 75% more requests" under real workloads. NVIDIA's Dynamo packages the same idea as an orchestration layer above vLLM, SGLang and TensorRT-LLM.
vLLM's documentation states plainly that "Disaggregated prefill DOES NOT improve throughput", and marks the feature experimental. Splitting buys you control over TTFT and inter-token latency separately. It does not manufacture FLOPs.
If your problem is cost per million tokens rather than a latency contract, disaggregation is the wrong tool and a bigger batch is the right one.
The cost, derived
Prefill produces a KV cache and decode consumes it, so splitting the phases means that cache has to cross a wire, and the size of the cache is fixed by the architecture:
The leading 2 is for K and V. $b$ is bytes per element, so 2 in FP16 and 1 in FP8.
Prefill for a prompt of $L$ tokens costs roughly $2PL$ floating-point operations, where $P$ is the parameter count active per token. If the pool achieves $F$ FLOP/s and the link runs at $W$ bytes/s, then the transfer as a fraction of the prefill is:
The prompt length cancels, so the ratio does not depend on how long the prompt was.
This cancellation is why disaggregation scales. Both sides grow linearly in $L$, so a 100,000-token prompt pays the same proportional transfer cost as a 1,000-token one. Attention makes prefill superlinear in $L$, so at very long contexts the ratio gets better.
Setting the ratio to 1 gives the link speed at which shipping the cache costs as much as making it:
Call this the break-even bandwidth. It depends on the model and the GPU and does not depend on the workload.
Try it Compute the break-even bandwidth, then price four real interconnects against it
import numpy as np
# KV bytes per token = 2 (K and V) x layers x kv_heads x head_dim x dtype
def kv_per_token(layers, kv_heads, head_dim, nbytes):
return 2 * layers * kv_heads * head_dim * nbytes
F = 4.0e14 # assumed achieved prefill rate, FLOP/s. Yours differs.
# name, layers, kv_heads, head_dim, bytes/elem, params active per token
cfgs = [("Llama-3.1-8B ", 32, 8, 128, 2, 8.0e9),
("Llama-3.1-70B", 80, 8, 128, 2, 70.0e9),
("70B, FP8 KV ", 80, 8, 128, 1, 70.0e9)]
print("model KV per token break-even link")
for name, L, h, d, nb, P in cfgs:
c = kv_per_token(L, h, d, nb)
print(f"{name} {c/1024:6.0f} KiB {c*F/(2*P)/1e9:6.2f} GB/s")
c70 = kv_per_token(80, 8, 128, 2)
print("\n70B, an 8k-token prompt: KV =",
round(c70 * 8192 / 2**30, 2), "GiB to move.")
print("shipping it costs this share of the prefill that made it:")
for link, bw in [("NVLink 4, one direction", 450e9),
("PCIe Gen5 x16 ", 64e9),
("InfiniBand NDR 400G ", 50e9),
("25 GbE ", 3.125e9)]:
print(f" {link} {c70*F/(2*70e9*bw)*100:7.1f} %")
0.94 GB/s, and every datacentre interconnect sits one to three orders of magnitude above it. Hence 0.2 % on NVLink and 1.9 % on InfiniBand NDR. At 25 GbE it reads 30.0 %, and disaggregation has stopped being free. Note the 8B's higher break-even of 3.28 GB/s: a small model makes prefill cheap, which makes the same transfer look expensive.The link bandwidths above come from vendor specifications: NVLink 4 at 900 GB/s total on an H100 SXM, halved here for one direction; PCIe Gen5 x16 at 128 GB/s bidirectional, also halved; and InfiniBand NDR at 400 Gb/s per port.
Four cases, and only the first is about bandwidth. A slow link, meaning anything within about 10x of $W^{*}$. Short prompts, where per-transfer overhead is not amortised and the prefill was cheap anyway.
A small fleet, because two pools of one GPU each cannot rebalance, and a traffic shift strands one of them. Uniform traffic, where no request is long enough to block anyone and chunked prefill already solved the problem for free.
The last point deserves emphasis. Disaggregation is the expensive answer to interference. Chunked prefill is the cheap one, and page 31 covers it. Measure that the cheap one has run out before buying the expensive one.
You have the placement decision and the number that settles it. $W^{*} = cF/2P$ is the bandwidth at which moving a token's KV costs as much as computing it.
Keep this number in mind. The next section is about the same cache, and the same quantity decides a completely different question there: whether to fetch a cached prefix back from host memory or just recompute it.
03The KV cache: make it narrower, or move it off the GPU
Two independent attacks on the same problem. One changes the bytes per token, the other changes where those bytes live.
NVIDIA's inference optimization guide states the size directly: "Size of KV cache per token in bytes = 2 * (num_layers) * (num_heads * dim_head) * precision_in_bytes". Under grouped-query attention the head count there is the number of KV heads, which is what section 02 called $n_{\text{kv heads}}$. Three of the four terms are fixed by the architecture. Only the last is yours.
Quantizing the cache
vLLM's quantized KV cache supports fp8_e4m3 and fp8_e5m2, which halves $c$ against FP16. This doubles how many tokens fit in the block pool, and it halves the break-even bandwidth from section 02 at the same time.
Below eight bits the cache stops behaving like a generic tensor. KIVI found that the two halves need opposite treatment: "the key cache should be quantized per-channel" while "the value cache should be quantized per-token". With that asymmetry it reports 2-bit KV at "2.6x less peak memory" and "2.35x ~ 3.47x throughput".
KVQuant attacks the same problem from the accuracy side and reports "< 0.1 perplexity degradation with 3-bit quantization", which it converts into context length: "serving LLaMA-7B with a context length of up to 1 million on a single A100-80GB GPU".
vLLM recommends calibrating FP8 KV scales and not leaving them at 1.0, and notes that "some attention layer types (e.g. sliding-window) are more sensitive to KV-cache quantization". It publishes no blanket accuracy figure, and neither should you.
KV quantization is the one compression on this page whose damage is invisible in a smoke test. It degrades long-context recall first, and a 200-token chat exchange will not show it. Evaluate it at the context length you serve.
Offloading the cache
The second attack leaves the precision alone and moves the bytes to host memory or NVMe. Whether that helps is decided by arithmetic, and there are two completely different cases.
Case one: KV that decode is actively reading. Every decode step reads the whole cache for every request in the batch. Take 32 requests at 8,192 tokens on the 70B, at 320 KiB per token. That is 85.9 GB read per token generated.
Over H100 HBM3 at 3.35 TB/s, that read takes 25.6 ms. Over PCIe Gen5 x16 at 64 GB/s it takes 1.34 seconds. A 52x slowdown on the inner loop is not a tradeoff, it is a different product.
Case two: KV that nobody is reading yet. A cached prefix from a previous turn, sitting in host RAM. The question is not bandwidth against HBM, it is fetch against recompute, and that is exactly the comparison section 02 already solved:
The same break-even bandwidth, for a question that looks nothing like the first one.
For the 70B, $W^{*}$ was 0.94 GB/s and PCIe Gen5 x16 runs at 64 GB/s, about 68 times higher. Fetching a cached prefix over PCIe is far cheaper than recomputing it, which is the entire case for KV offload.
LMCache productised this. Its paper describes storing KV "out of the GPU memory" and sharing it "across engines and queries", supporting "both cache offloading (prefix reuse across queries) and prefill-decode (PD) disaggregation", and reports "up to 15x improvement in throughput" on workloads like multi-round question answering.
The extreme version of case two is FlexGen, which aggregates GPU, CPU and disk to run OPT-175B on a single 16 GB GPU at "1 token/s ... with an effective batch size of 144". It is a throughput system for offline work, and its latency is what the arithmetic above predicts.
Offload KV that will be read once, later. Never offload KV that will be read on every step. The first is a fetch racing a recompute, and it wins comfortably. The second is a memory bus racing a peripheral bus, and it loses by a factor of 50.
04Multi-LoRA: hundreds of models, one set of weights
A fleet rarely serves one model. It serves one base and a long tail of customer-specific adaptations, and the only affordable way to do that is to not duplicate the base.
Page 29 covers what a LoRA adapter is. The serving question is different: given a batch of requests where each one wants a different adapter, what does that cost?
Start with memory: a LoRA on a weight of shape $(d_{\text{out}}, d_{\text{in}})$ adds two matrices, $A \in \mathbb{R}^{r \times d_{\text{in}}}$ and $B \in \mathbb{R}^{d_{\text{out}} \times r}$, so $r(d_{\text{in}} + d_{\text{out}})$ parameters per adapted matrix.
Try it What a thousand adapters cost, and what a batch of distinct ones costs per step
import numpy as np
# LoRA on a weight of shape (d_out, d_in) adds A (r, d_in) and
# B (d_out, r): r * (d_in + d_out) new parameters.
def adapter_params(r, shapes):
return sum(r * (d_in + d_out) for d_out, d_in in shapes)
H, KV, LAYERS = 4096, 1024, 32 # Llama-3-8B, 32 q heads, 8 kv
attn = [(H, H), (KV, H), (KV, H), (H, H)] # q, k, v, o
for r in (8, 16, 64):
p = adapter_params(r, attn) * LAYERS
print(f"rank {r:3d}: {p/1e6:6.2f}M params, {p*2/1e6:6.1f} MB in BF16"
f" 1000 of them = {p*2*1000/1e9:6.1f} GB")
print(f"\nbase model in BF16 = {8e9*2/1e9:.0f} GB;"
f" 1000 full fine-tunes = {8e9*2*1000/1e12:.0f} TB")
print("\ndecode step reads 16.0 GB of base weights, plus, for n"
"\nDISTINCT adapters in the batch:")
for r in (16, 64):
mb = adapter_params(r, attn) * LAYERS * 2 / 1e6
for n in (1, 32, 128):
print(f" rank {r:3d}, {n:3d} distinct: {n*mb/1000:6.2f} GB"
f" -> step is {1 + n*mb/16000:5.3f}x longer")
27.3 MB, so a thousand of them is 27.3 GB against 16 TB for a thousand full fine-tunes. The second table is the one people miss: adapter bytes scale with the number of distinct adapters in the batch, not with the batch size, because a shared adapter is read once. At rank 64 with 128 distinct, the step is 1.872x longer.That last result is the whole design constraint. The base path is one big GEMM shared by the entire batch. The LoRA path is not shareable: each sequence needs its own $A$ and $B$ gathered and applied.
Done naively that becomes a Python loop over requests, which destroys the batch. Punica solved it with a kernel: "Segmented Gather Matrix-Vector Multiplication (SGMV)", which "allows batching GPU operations for the concurrent execution of multiple, different LoRA models". It reports "12x higher throughput ... while only adding 2ms latency per token".
S-LoRA added the memory-management half, and its "Unified Paging" puts adapter weights of different ranks and KV tensors of different lengths in one pool, which is PagedAttention generalised to a second kind of variable-size object. It reports up to 4x throughput over prior LoRA serving.
vLLM's documentation is explicit about the cost: set --max-lora-rank "to the maximum rank among all LoRA adapters you plan to use", and if your ranks are 16, 32 and 64, "use --max-lora-rank 64 rather than 256".
The buffers are sized for the maximum, so one customer who trained at rank 256 makes every other adapter in the fleet pay for it. Cap the rank at the training stage, because by serving time it is no longer negotiable.
05FP8 and MXFP4: where the scale factors live
Four-bit floating-point formats differ from plain INT4 in how many numbers share a scaling factor.
FP8 came first and is the easy case. The format paper defines two encodings, "E4M3 (4-bit exponent and 3-bit mantissa) and E5M2 (5-bit exponent and 2-bit mantissa)", and reports "effectively matching the result quality achieved by 16-bit training sessions".
The hardware reason to care is on the datasheet. An H100 SXM lists 1,979 TFLOPS for BF16 tensor cores and 3,958 for FP8, both quoted with sparsity and both halved without it. That is exactly 2x, and it comes from the format and not from the silicon generation.
What gets quantized
"The model is in FP8" is three separate decisions, and they have different risk profiles:
- Weights. These are static and quantized once offline. vLLM's llm-compressor FP8 scheme is "static, per-channel quantization on the weights", meaning one scale per output channel, stored alongside them.
- Activations. Their range depends on the input, so the scale is computed at runtime. vLLM uses "dynamic, per-token quantization on the activations".
- KV cache. See section 03. It has a separate flag and a separate risk, and it is the one that breaks long context first.
TensorRT-LLM's precision reference names the same three granularities: "per-tensor" with a single scaling factor for all elements, "per-token", and "per-channel". Finer granularity costs storage and buys tolerance to outliers, and that tradeoff explains the four-bit formats.
Microscaling: the scale moves inside the tensor
The OCP Microscaling (MX) specification v1.0 fixes the granularity in the format itself. MXFP4 is an E2M1 element, four bits, in blocks of 32 that share one E8M0 scale. E8M0 is a bare Float32 exponent with no sign and no mantissa, so the scale is a raw power of two.
Count the bits: four per element plus eight per block of 32 is $4 + 8/32 = 4.25$ bits per weight. NVIDIA's NVFP4 halves the block "from 32 values to 16" and upgrades the scale to E4M3, giving 4.5 bits, plus a per-tensor FP32 scale on top.
Try it One outlier weight, and what block scaling does about it that a tensor scale cannot
import numpy as np
# MXFP4 (OCP Microscaling v1.0): 32 elements share one E8M0 scale
# (a raw power of two); each element is E2M1, four bits.
E2M1 = np.array([0., .5, 1., 1.5, 2., 3., 4., 6.]) # magnitudes
def snap(q, grid):
i = np.abs(np.subtract.outer(np.abs(q), grid)).argmin(axis=-1)
return np.sign(q) * grid[i]
def mxfp4(x, blk=32):
y = x.reshape(-1, blk)
s = 2.0 ** np.floor(np.log2(np.abs(y).max(1, keepdims=True) / 6))
return (snap(y / s, E2M1) * s).ravel()
def int4_tensor(x): # one scale for the whole tensor
s = np.abs(x).max() / 7.0
return np.clip(np.round(x / s), -7, 7) * s
rng = np.random.default_rng(0)
w = rng.normal(0, 1, 4096)
w[137] = 50.0 # one outlier, as real weights have
for name, q in [("MXFP4, 32-elem blocks ", mxfp4(w)),
("INT4, one tensor scale", int4_tensor(w))]:
e = q - w
print(f"{name} rms err {np.sqrt((e**2).mean()):6.3f}"
f" weights flushed to zero {(q == 0).mean():6.1%}")
print("\nbits per weight, scale included:")
for nm, elem, blk, sc in [("MXFP4", 4, 32, 8), ("NVFP4", 4, 16, 8)]:
b = elem + sc / blk
print(f" {nm} {b:.2f} bits {16/b:.2f}x smaller than BF16")
99.9% of the tensor rounds to zero and the RMS error is 0.997, the standard deviation of the data. The model is ruined. MXFP4 confines that outlier to its own block of 32 and lands at 0.220. Both formats store four bits per element; one has 128 scales and the other has one.NVIDIA reports NVFP4 against FP8 on DeepSeek-R1-0528 at "1% or less accuracy degradation", with MMLU-Pro at 85 against 84 and GPQA Diamond at 81 against 80.
The clearest production example is OpenAI's gpt-oss. Its model card states the models "were post-trained with MXFP4 quantization of the MoE weights, making gpt-oss-120b run on a single 80GB GPU ... and the gpt-oss-20b model run within 16GB of memory".
The arithmetic behind that sentence is the 4.25 bits above. 120 billion weights at 4.25 bits each is 63.75 GB, and only the MoE weights are at that width, so the real footprint is somewhat higher. In BF16 the same weights would be 240 GB and would need four cards.
In a sparse MoE the expert FFNs hold the overwhelming majority of the parameters, and each one is read by only the tokens routed to it. They are the biggest, coldest, most redundant weights in the model, which makes them the cheapest place to spend accuracy.
Attention projections, embeddings and the router are left at higher precision. The router in particular decides which experts run at all, and an error there costs a whole expert, where other layers lose a few least significant bits.
Placement, precision and sharing are all covered now. What is left is the machinery that makes those three run: the kernels that decode uses, the communication that expert parallelism needs, and the scheduler that decides who is served.
Those three sections are also where most of the remaining performance is, and where the failure modes in section 10 come from.
06Decode needs its own attention kernel
FlashAttention is a prefill kernel. Running it at decode time leaves most of the GPU idle, and the reason is a single dimension collapsing to 1.
FlashAttention parallelises over batch, heads and blocks of queries. At decode there is one query per request, so the query dimension is gone and only batch times heads is left to fill the machine.
The Flash-Decoding authors put the consequence plainly: "if the batch size is smaller than the number of streaming multiprocessors (SMs) on the GPU (108 for an A100), the operation will only use a small part of the GPU! ... With a batch size of 1, FlashAttention will use less than 1% of the GPU!"
Count it yourself: a 32-head model at batch 1 offers 32 independent work units against an A100's 108 SMs, under a third of the machine, and no kernel tuning recovers the other two thirds because the work is not there.
If there is not enough parallelism across the output, create some across the reduction. Flash-Decoding splits the keys and values into chunks, attends to each chunk in parallel, and writes "1 extra scalar per row and per split: the log-sum-exp of the attention values".
A second, cheap pass rescales the partial outputs by those log-sum-exps and sums them. Splitting the KV into 8 chunks turns 32 work units into 256, and the reported result is "up to 8x faster generation for very long sequences".
The split only pays when the KV sequence is long relative to the batch, because that is precisely when parallelism over batch and heads is scarce and parallelism over sequence is abundant. At batch 256 and 128 tokens of context there is nothing to fix.
FlashInfer generalised this into a kernel library built for serving and not for training, and reports "29-69% inter-token-latency reduction compared to compiler backends for LLM serving benchmark, 28-30% latency reduction for long-context inference". It won a best paper award at MLSys 2025.
Two further decode-specific facts shape these kernels. Grouped-query attention gives each KV head several query heads, so the kernel can load the KV once and issue a small matrix multiply instead of several separate vector ones. And PagedAttention scatters the KV across non-contiguous blocks, so the kernel has to walk a block table rather than stride through an array.
07Expert-parallel MoE serving
A sparse mixture of experts is a model that does not fit on one GPU and does not want to. Serving it well is a communication problem with a load-balancing problem inside it.
DeepSeek-V3 is the reference point: "671B total parameters with 37B activated for each token". Each MoE layer has "1 shared expert and 256 routed experts", of which "8 experts will be activated for each token".
Tensor parallelism would shard every expert across every GPU, which means every GPU does a slice of all 256 experts. Expert parallelism does the opposite: each GPU owns a few whole experts, and tokens travel to them.
This changes what moves. Instead of all-reducing activations after every matmul, you dispatch each token's hidden state to the GPUs holding its 8 experts, compute, then combine the results back, which takes two all-to-all collectives per MoE layer, per step.
The all-to-all, priced
DeepEP benchmarks this at DeepSeek-V3's configuration: "8K tokens per batch, 7168 hidden dimensions, top 8 experts, FP8 dispatching, and BF16 combining". Those numbers are enough to size the traffic.
Dispatch moves $8192 \times 8 \times 7168$ bytes at one byte each, which is 470 MB. Combine moves the same count at two bytes, so 940 MB. DeepEP reports 726 GB/s dispatch and 740 GB/s combine over NVLink, against 90 and 81 GB/s over RDMA.
Divide through and the intra-node case costs about 1.9 ms per MoE layer, while the cross-node case costs about 16.8 ms. DeepEP notes these are logical bandwidths that include local traffic, so treat the ratio as the finding and neither absolute as exact.
That roughly 9x gap is why DeepSeek-V3's router constrains each token to "at most 4 nodes". Node-locality is not a deployment detail bolted on afterwards, it is a term in the routing function.
Imbalance is the cost
Routing is learned, and learned routing is not uniform. The LMSYS large-scale expert-parallelism report states the consequence: "This imbalance forces the system to wait for the slowest GPU computation or communication ... As the number of GPUs (EP size) increases, the imbalance issue gets more severe."
A step finishes when the busiest expert finishes. Mean load is irrelevant; the maximum sets the clock.
Try it Watch a skewed router create a straggler, then fix it with one bias vector
import numpy as np
E, K, T = 256, 8, 8192 # experts, experts per token, tokens
rng = np.random.default_rng(0)
def load(scores): # tokens per expert after top-K routing
top = np.argpartition(-scores, K, axis=1)[:, :K]
return np.bincount(top.ravel(), minlength=E)
def report(tag, c):
print(f"{tag:22s} mean {c.mean():6.1f} busiest {c.max():5d}"
f" imbalance {c.max()/c.mean():5.2f}x")
aff = rng.normal(0, 1, (T, E)) # router affinities
report("uniform router", load(aff))
pop = rng.normal(0, 0.6, E) # some experts are popular
report("popularity-skewed", load(aff + pop))
# DeepSeek-V3's auxiliary-loss-free fix: a per-expert bias, added to
# the score for ROUTING ONLY, nudged down when an expert is overloaded.
b, gamma = np.zeros(E), 0.05
for step in range(60):
c = load(aff + pop + b)
b -= gamma * np.sign(c - c.mean())
report("after bias correction", load(aff + pop + b))
8.75x imbalance. The busiest expert takes 2,241 tokens against a mean of 256, and every other GPU in the group waits for it. Sixty rounds of a per-expert bias bring it to 1.14x, below what a uniform router manages by chance. The bias touches routing only, so the gating values are untouched, and DeepSeek-V3 is explicit about that.Two production mitigations sit on top of that. DeepSeek deploys "redundant experts, which duplicates high-load experts", and SGLang's report finds that "SGLang with default expert imbalance is 20% slower than DeepSeek's profile, while the simulated perfect EPLB case narrows the gap to 6%".
For scale, that same deployment ran on 12 nodes of 8 H100s and reported "52.3k input tokens per second and 22.3k output tokens per second per node for 2000-token input sequences", using prefill/decode disaggregation from section 02 alongside expert parallelism.
08SLA-aware scheduling, and why throughput is the wrong target
Throughput rises monotonically with concurrency, but goodput does not, so running at the top of the first curve puts you well past the top of the second.
DistServe defines the target precisely: "per-GPU goodput, defined as the maximum request rate that can be served adhering to the SLO attainment goal (say, 90%) for each GPU provisioned". Tokens served too slowly to be useful do not count.
The gap between the two metrics is not small, and it is easy to see with the arithmetic already on this page.
Try it Sweep concurrency and watch goodput peak while throughput is still climbing
import numpy as np
P, W, BW, F = 8e9, 16e9, 3.35e12, 4.0e14 # 8B in BF16, one H100
L_out = 256 # output tokens per request
TTFT_SLO, ITL_SLO = 2.0, 0.025 # the contract with the user
rng = np.random.default_rng(0)
L_in = rng.lognormal(np.log(900), 1.0, 4000) # prompt lengths
wait = (np.arange(4000) + .5) / 4000 # place in the cycle
print(" N thr tok/s ITL ms met SLO goodput tok/s")
for N in (4, 8, 16, 24, 32, 48, 64, 128, 256):
# one cycle = N prefills plus L_out decode steps at batch N
cycle = L_out * max(W / BW, 2 * P * N / F) \
+ N * 2 * P * L_in.mean() / F
thr = N * L_out / cycle # tokens/s
itl = cycle / L_out
ttft = wait * cycle + 2 * P * L_in / F
met = ((ttft <= TTFT_SLO) & (itl <= ITL_SLO)).mean()
print(f"{N:4d} {thr:10.0f} {itl*1e3:8.1f} {met:9.0%} {thr*met:14.0f}")
1741 tokens/s. Throughput is still climbing there and reaches 3736 at 128, where goodput is zero. Running at the throughput maximum serves 2.1x more tokens and zero useful ones. This is a model and not a measurement: the decode step is the larger of weights-over-bandwidth and compute, prefills spread evenly through the cycle, prompt lengths lognormal. Sweep yours to find it.Once you accept goodput as the target, the scheduler gets choices it did not have before. Four systems illustrate the range:
- FastServe attacks head-of-line blocking with "a novel skip-join Multi-Level Feedback Queue scheduler", reporting up to 31.4x throughput at the same average latency requirement.
- Llumnix reschedules requests across model instances at runtime and reports improving "tail latencies by an order of magnitude" and accelerating "high-priority requests by up to 1.5x".
- QLM manages the queue and not the batch, reporting "SLO attainment by 40-90% and throughput by 20-400%".
- Andes optimises perceived quality of experience directly, reporting up to 4.7x average QoE at fixed GPU count, or "61% GPU resources" saved at fixed QoE.
These policies are built in: vLLM ships a policy setting where "fcfs" means arrival order and "priority" means "requests are handled based on given priority (lower value means earlier handling)".
vLLM's max_num_queued_tokens documentation gives the formula outright: set it to target_TTFT * prefill_throughput and "you reject requests when the prefill backlog would exceed the latency target". Over the limit, new requests get an HTTP 503.
Refusing a request you cannot serve on time is strictly better than accepting it and missing. The rejected client can retry elsewhere immediately, and the accepted ones keep their contract. A queue with no bound is a queue that converts one slow minute into an hour of broken promises.
Preemption is the other half. When the KV pool fills, the scheduler evicts a request and later recomputes its prefill, which is throughput spent twice. vLLM exposes a watermark that keeps a fraction of blocks free specifically to "avoid frequent KV cache eviction and the resulting repeated preemption of requests".
09Build this
Section 02's break-even bandwidth is three measurements away from being a number about your own hardware and not just a number on this page.
The formula $W^{*} = cF/2P$ has exactly three inputs. Two of them you can read off a config file, and the third is the one everybody guesses wrong. Measure all three, predict the transfer cost, then check the prediction against a real split.
- Compute $c$ from the model's
config.json: layers, KV heads, head dimension, dtype. Cross-check it against the number of tokens your engine says the KV pool holds, because a factor-of-two error here is easy and silent. - Measure $F$ by sending one request with a long prompt and no output, timing the first token, and dividing $2PL$ by that time. This gives your achieved prefill rate. Expect it to be well under the datasheet, and use yours everywhere below.
- Measure $W$ on the link that would connect the pools. Not the vendor figure, the one you get: a point-to-point send and receive of a few gigabytes, timed.
- Compute $W^{*}$ and the predicted transfer share $cF/2PW$. Write the number down before you run anything else.
- Now run vLLM's disaggregated prefill across the two GPUs and measure TTFT against a single-instance baseline at the same load.
fp8_e4m3 KV: $c$ halves, so the predicted share should halve too. If it does not, the transfer is not what is costing you.10What breaks
At fleet scale the failures are not crashes. They are a number that got worse somewhere you were not looking, for a reason that is nowhere in the logs.
- The KV transfer is serialised, not overlapped. Section 02's percentage assumes the send of layer $i$ overlaps the prefill of layer $i+1$. If it does not, you pay the whole transfer as added TTFT, and disaggregation looks like a regression.
- The pool ratio is tuned for yesterday's traffic. Prompts get longer, the prefill pool saturates, and the decode pool idles while queueing grows. The symptom is rising TTFT with flat inter-token latency, which points straight at the split.
- KV quantization passes every eval you ran. Short-context benchmarks will not show it. The damage lands on long-context retrieval, so it ships, and then a customer with a 100,000-token document finds it for you.
- One hot expert is setting the step time. Throughput falls by the imbalance factor and no component reports an error, because every GPU is healthy and most of them are waiting. Log per-expert token counts or you cannot see it.
- A new router checkpoint invalidates the expert placement. Yesterday's redundant-expert plan was fitted to yesterday's routing distribution. Rebalancing has to be part of deployment, not a one-off.
- One customer's rank-256 adapter taxes the whole fleet. LoRA buffers are sized for the maximum rank, so everyone pays for the largest adapter admitted. Cap it at training time.
- The decode kernel silently fell back. An unusual head dimension, dtype or block size drops you onto a generic path without the split-K, and inter-token latency worsens by a factor you will attribute to anything else.
- Throughput improved and goodput did not. The most common outcome of a batching change. If you are not measuring SLO attainment alongside tokens per second, you cannot tell an improvement from a degradation.
- KV was offloaded to the wrong side of the line. Offloading cached prefixes is a large win. Offloading the cache that decode reads every step costs a factor of about 50, and both are enabled by flags with similar names.
11Where you meet this in the wild
The same techniques, in wildly different combinations, because the binding constraint is different in each.
Huge model, enormous traffic, published latency targets. Everything on this page is in play at once: disaggregated pools, FP8 or four-bit weights, expert parallelism, cross-request prefix caching, and goodput as the provisioning metric.
Every paper cited here was written for this workload.
Thousands of customers, each with a small adapter, most of them idle most of the time. The base model is shared, adapters are paged in and out, and the scarce resource is adapter bandwidth and not model memory.
Multi-LoRA is the product, so cap the rank fleet-wide.
Disaggregation buys little here, because prompts are usually short.
A 600B-class sparse model on a handful of nodes. Expert parallelism is not optional, the all-to-all is the bottleneck, and load balancing decides whether you get the throughput the model card implies.
Budget engineering time for expert placement as well as for the deployment.
A few hundred requests an hour behind a company firewall. Continuous batching and prefix caching are already enough, and the KV pool will never fill.
Almost none of this page applies. Page 31 is the whole answer, and adding a second pool here buys complexity and nothing else.
12Interview questions
BeginnerWhat is prefill/decode disaggregation, and what does it cost?
It runs the two phases of a request on separate GPU pools instead of sharing one. The motivation is that the phases want different machines: prefill is compute-bound and latency-sensitive, so it prefers wide tensor parallelism and small batches, while decode is memory-bandwidth-bound and prefers the largest batch the KV pool allows. A single pool must pick one configuration for both, and two pools do not. The cost is that prefill produces the KV cache and decode consumes it, so that cache has to cross an interconnect. Whether the cost is acceptable is arithmetic rather than opinion: the transfer time as a fraction of the prefill time is the KV bytes per token times the achieved FLOP rate, divided by twice the parameter count times the link bandwidth. On NVLink that fraction is well under one percent; on commodity Ethernet it is tens of percent.
BeginnerWhy is FP8 weight quantization roughly twice as fast rather than just smaller?
Because the tensor cores have a separate, faster path for it. An H100 datasheet lists 1,979 TFLOPS for BF16 and 3,958 for FP8, a factor of exactly two, and that ratio comes from the numeric format rather than from any other hardware difference. The memory saving is real too and matters more during decode, which is bandwidth-bound, but the compute doubling is what helps prefill. The two encodings are E4M3, with four exponent bits and three mantissa bits, and E5M2, which trades a mantissa bit for range. E4M3 is the usual choice for weights and activations because precision matters more than range once a scaling factor is applied, and E5M2 appears more often for gradients and occasionally for KV cache.
IntermediateWhy does MXFP4 work when plain INT4 with one scale does not?
Because of where the scaling factor lives. A single per-tensor scale has to cover the largest magnitude in the tensor, and weight distributions have outliers, so one large value forces a step size that rounds almost everything else to zero. MXFP4 as specified by OCP puts an E8M0 scale on every block of 32 elements, so an outlier only distorts its own block and the other blocks keep sensible scales. The cost is the scale storage, which works out to four bits per element plus eight bits per thirty-two, or 4.25 bits per weight. NVFP4 halves the block to 16 elements and uses an E4M3 scale, reaching 4.5 bits with finer adaptation, and NVIDIA reports around one percent or less accuracy degradation against FP8 on reasoning benchmarks. The general principle is that low-bit quantization is a question about granularity, not about bit width alone.
IntermediateYou want to serve 500 customer LoRA adapters on one base model. What dominates the cost?
Not storage. A rank-16 adapter on an 8B model with adapted attention projections is about 13.6M parameters, or 27 MB in BF16, so 500 of them are under 14 GB and sit comfortably beside a 16 GB base model. What dominates is bandwidth during decode, and specifically the number of distinct adapters in a batch rather than the batch size, because requests sharing an adapter read it once. At rank 16 with 32 distinct adapters the extra read is about five percent of the base weights; at rank 64 with 128 distinct it approaches ninety percent, and the adapters cost nearly as much as the model. The second constraint is that the LoRA path cannot be one GEMM, since each sequence needs its own matrices, which is what Punica's SGMV kernel and S-LoRA's unified paging exist to solve. Practically, cap the maximum rank fleet-wide, because serving buffers are sized for the largest adapter admitted.
IntermediateWhy does decode need a different attention kernel from prefill?
Because FlashAttention parallelises over batch, heads and blocks of queries, and at decode time there is exactly one query per request, so that third dimension disappears. Only batch times heads is left to fill the GPU, and the Flash-Decoding authors note that if the batch is smaller than the SM count, 108 on an A100, the kernel uses a small part of the device, with batch one using under one percent of it. The fix is split-K applied to the sequence: divide the keys and values into chunks, attend to each in parallel, store one extra log-sum-exp scalar per row per split, then rescale and combine in a cheap second pass. The win is largest exactly when the context is long and the batch is small, which is when parallelism over batch is scarcest, and the reported speedup reaches eight times for very long sequences.
DeepYou have 32 requests at 8k context on a 70B model and 64 GB of host RAM free. Should you offload the KV cache?
It depends entirely on whether that KV is being read every step, and the two answers are opposite. Active decoding KV is read in full on every token: 32 requests at 8,192 tokens and 320 KiB per token is 85.9 GB per step, which takes about 26 ms over H100 HBM at 3.35 TB/s and about 1.34 seconds over PCIe Gen5 at 64 GB/s. That is a factor of fifty on the inner loop and it destroys the product. Cached prefixes that nobody is currently reading are the opposite case, because the real comparison is fetching them against recomputing them. Fetching costs the KV bytes divided by the link bandwidth, recomputing costs twice the parameter count times the tokens divided by the achieved FLOP rate, and fetching wins whenever the link exceeds the same break-even bandwidth as before. For this model that break-even is under 1 GB/s, so PCIe is roughly seventy times faster than it needs to be. The rule is to offload KV that will be read once, later, and never KV that will be read on every step.
DeepYour expert-parallel MoE deployment is at a third of its expected throughput. How do you diagnose it?
Start by logging tokens per expert per step, because expert load imbalance is the default explanation and it is invisible in ordinary metrics. A step finishes when the busiest expert finishes, so the maximum load sets the clock while the mean looks fine, and every idle GPU reports as healthy. A modest popularity skew in the router produces imbalance factors approaching an order of magnitude, which is precisely the gap you are describing. Second, check whether tokens are crossing nodes: intra-node all-to-all over NVLink and cross-node over RDMA differ by roughly a factor of nine in DeepEP's own benchmarks, which is why DeepSeek-V3 constrains each token to at most four nodes. Third, compare against a profile with simulated perfect balance; the LMSYS report found default imbalance cost twenty percent against DeepSeek's profile while perfect balance narrowed the gap to six. Mitigations are per-expert routing bias, which DeepSeek applies to routing scores only so the gating values are undistorted, and duplicating high-load experts as redundant copies.
DeepYour server's throughput went up 40% after a change and users are complaining. Explain and fix.
Almost certainly the change raised concurrency, and throughput and goodput diverge past a point. Throughput rises monotonically with concurrency until the hardware saturates, but goodput, meaning the tokens delivered inside the latency contract, peaks much earlier and then falls to zero. In a simple model of an 8B on one H100 with a two-second TTFT target, goodput peaks around twenty-four concurrent requests while throughput keeps climbing to roughly twice that level at a concurrency where SLO attainment is zero. So the 40% is real and useless. The fix is to make goodput the reported metric, sweep concurrency to find its peak on your own traffic, and then defend that operating point with admission control rather than a longer queue. vLLM exposes this directly: cap queued prefill tokens at the target TTFT times the prefill throughput and reject beyond it, so clients get an immediate rejection they can retry rather than a slow response they cannot use. Keep a KV watermark as well, since preemption and recompute under cache pressure is the other way a throughput gain turns into a latency loss.
13Go deeper
--max-lora-rank is a fleet-wide decision.●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.
Other techniques for this problem
A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.