Running Models Locally
Will this model run on my machine? About a minute of arithmetic answers that, and the same arithmetic tells you how fast it will be once it does.
Three things occupy your VRAM: model weights, the KV cache, and framework overhead. Weights dominate at short context and the KV cache overtakes them at long context, so compute both. Guessing is how a model that loaded fine at 2k context runs out of memory at 32k.
Quantization is the main lever, and your runtime picks the format for you, so quality does not decide it: GGUF for llama.cpp and Ollama, AWQ or GPTQ for vLLM, bitsandbytes NF4 for QLoRA fine-tuning (that one isn't a serving format at all).
Names mislead: Q4_K_M isn't 4 bits per weight and measures 4.89. Once the model fits, memory bandwidth sets your speed and FLOPs barely matter, so a smaller quantization generates tokens faster on identical hardware. One sentence for an interview: local inference is a memory-capacity problem first, a memory-bandwidth problem second, and a compute problem almost never.
01What everybody asks first
"Will this run on my machine?" is answerable in advance. Most people answer it by downloading 40 GB and watching it fail.
Page 23 covers why quantization works and how PagedAttention manages a KV cache. This page is the practical counterpart: what fits, what to install, and what to expect once it's running.
Your GPU has one fixed pool of memory, and three things compete for it:
- Model weights. Fixed once you choose a model and a quantization level. Usually the largest single item.
- KV cache. Grows linearly with context length and with the number of concurrent requests. Small at 2k context, dominant at 128k.
- Overhead. Activations, CUDA context, the framework itself. Small, at a few hundred megabytes to about a gigabyte, and people forget it.
If the sum exceeds your VRAM you get one of three outcomes: the load fails outright, the runtime quietly offloads layers to CPU and slows down by an order of magnitude, or you hit an out-of-memory error partway through a long conversation. That last one is the KV cache growing past your headroom, and it's the one that catches people.
02Do the arithmetic once
Learn to compute this and you'll never have to trust a compatibility table again. Two real terms and one fudge factor.
Worked example Does Llama 3 8B fit in 8 GB of VRAM at 8k context?
Every number below comes from the model's published config.json and from llama.cpp's own measurements. Your framework overhead will differ, so treat that one term as illustrative. The weight and cache terms are exact.
-
$$\text{weights} = N_{\text{params}} \times \frac{\text{bits per weight}}{8} \;\text{bytes}$$The naive version, and why it's wrong. People assume 4-bit means half a byte per parameter. Llama 3 8B has 8,030,261,248 parameters, so that predicts 4.0 GB. The real Q4_K_M file is 4.58 GiB. A name isn't a footprint.
-
$$8.03 \times 10^9 \times \frac{4.8944}{8} \;=\; 4.91 \times 10^9 \;\text{bytes} \;=\; 4.58\ \text{GiB}$$Why 4.89 and not 4.00: k-quants store a scale and a minimum per block of weights, and the
_Mvariant deliberately keeps some tensors at higher precision. Both cost bits. llama.cpp publishes a measured 4.8944 bits per weight for this model, and that matches the observed file size exactly. -
$$\begin{aligned}\text{KV bytes/token} = 2 \;\times\;& n_{\text{layers}} \times n_{\text{kv heads}} \\ \times\;& d_{\text{head}} \times \text{bytes}\end{aligned}$$Why the leading 2: you cache a key and a value for every token, at every layer. The term everybody gets wrong is $n_{\text{kv heads}}$. With grouped-query attention it's far smaller than the number of attention heads, and picking the wrong one inflates your estimate several-fold.
-
$$\begin{aligned}2 \times 32 \times 8 \times 128 \times 2 \;&=\; 131{,}072 \;\text{bytes} \\ &=\; 128\ \text{KiB per token}\end{aligned}$$Where these come from: Llama 3 8B's config gives 32 layers, 8 key-value heads, and a head dimension of 128 (hidden size 4096 over 32 attention heads). The final 2 is bytes per element for an FP16 cache. Note the 32 attention heads against only 8 KV heads.
-
$$131{,}072 \times 8192 \;=\; 1{,}073{,}741{,}824 \;\text{bytes} \;=\; 1.00\ \text{GiB}$$A clean coincidence worth remembering: this model at 8k context needs exactly one gibibyte of KV cache. Had it used full multi-head attention, with 32 KV heads instead of 8, the same calculation gives 4 GiB. Grouped-query attention is quietly doing a lot of work for local users.
-
$$4.58 + 1.00 + \approx\!0.7 \;\approx\; 6.3\ \text{GiB total}$$The fudge factor. Activations, the CUDA context and framework allocations are real but modest. Budget several hundred megabytes to roughly a gigabyte. That total leaves usable headroom inside 8 GB, so yes, it fits. Doubling the context to 16k costs exactly one more gibibyte, and that is the term to watch.
Practitioners quote roughly 0.6 to 0.7 GB per billion parameters at Q4_K_M, and that agrees with the calculation above (4.58 GiB is 4.92 GB, over 8.03B parameters, so 0.61), so it's a fine sanity check. It stays a rule of thumb because it silently assumes a 4-bit k-quant, ignores context length, and says nothing about how many KV heads the architecture has.
03Formats you'll meet
Quantization format looks like a quality decision. In practice it's a compatibility decision, because each runtime eats its own format. Pick the runtime first and the format usually follows.
- GGUF is the native format for llama.cpp, and therefore for Ollama and LM Studio, which build on it. It's the only one of these designed for CPU and for hybrid CPU-plus-GPU execution.
- GPTQ is a post-training weight quantization algorithm aimed at GPU inference, and it is widely supported, vLLM included.
- AWQ is activation-aware weight quantization. It protects the weights that matter most (the ones whose observed activations are largest), and it's a common choice for GPU serving.
- bitsandbytes NF4 quantizes on load and not ahead of time. It is the standard path for QLoRA fine-tuning and is not a serving format. See 29.
Q4_K_M reads like "4 bits per weight." It measures 4.8944. Q8_0 reads like 8 bits and measures 8.50.
Q4 is the base block type. K means k-quants, a block-wise scheme carrying a scale and a minimum per block. M is the medium mix, and it keeps selected tensors at higher precision than the name suggests.
Budget from measured bits per weight, never from the digit in the filename. A 22% underestimate is exactly enough to make a model you'd budgeted for refuse to load.
llama.cpp publishes measured figures for Llama 3.1 8B. Here's the real size ladder:
| Quant | Q2_K | Q3_K_M | Q4_K_M | Q5_K_M | Q6_K | Q8_0 | F16 |
|---|---|---|---|---|---|---|---|
| Bits per weight | 3.16 | 4.00 | 4.89 | 5.70 | 6.56 | 8.50 | 16.00 |
| Size (GiB) | 2.95 | 3.74 | 4.58 | 5.33 | 6.14 | 7.95 | 14.96 |
On quality, degradation is usually gradual, and how gradual seems to depend on both the model and the task. Q4_K_M is the widely used default because it sits at a good size-to-quality point.
Users start reporting noticeable losses below roughly 4 bits per weight. The very low quants exist for when fitting at all matters more than fidelity. Measure on your own task and do not trust a general claim, this one included.
04Which tool to pick
Tools split cleanly by what you're doing. Ask whether you're serving one user or many.
- One user, your own machine. Ollama for the shortest path from nothing to a running model. LM Studio if you want a GUI. llama.cpp directly when you need control over build flags, quantization and runtime settings. The first two are wrappers around the third.
- Many concurrent users. vLLM, TGI or SGLang. These implement continuous batching and a paged KV cache, and that's what turns idle GPU time into throughput. Page 23 explains the mechanisms.
- Apple Silicon. MLX is built for it, though llama.cpp also runs well on Metal.
A common mistake is reaching for a serving framework to run a single local chat. vLLM's advantages are throughput advantages. With one user and a batch size of one, you're paying setup complexity for benefits you can't reach.
05Hardware tiers
Everything below follows from the arithmetic in section 02, applied at Q4_K_M with a modest context. Treat the model sizes as what fits comfortably and not as hard ceilings, since context length moves these more than anything else.
- 8 GB VRAM. 7B and 8B models fit at 4-bit with room for a normal context (4k to 8k, going by the worked example above). That's roughly the entry point where local inference stops being painful.
- 12 to 16 GB. This range runs 7B to 14B comfortably, with longer contexts and no need to watch the cache, and the extra headroom mostly buys context and not parameters.
- 24 GB. This fits 30B-class models at 4-bit, or a smaller model with a large context, which is a meaningful step up.
- 48 GB and above. This is 70B-class territory. llama.cpp measures a 70B at 43.1 GB in Q4_K_M, so 48 GB works with care and multi-GPU is more comfortable.
- Apple Silicon. Unified memory changes the calculation, because CPU and GPU share one pool. A machine with plenty of it can hold models that'd otherwise need a very expensive discrete GPU. Bandwidth varies by chip tier, and bandwidth is what governs speed.
- CPU only. This works, thanks to GGUF, but generation is slow enough to feel like waiting and not conversing, because you're limited by system RAM bandwidth.
For scale at the top end, llama.cpp's own table puts Llama 3.1 405B at 249.1 GB even in Q4_K_M, against 1,625.1 GB unquantized. Some models aren't a local option at any consumer tier.
06What makes it slow
Once a model fits, the question becomes speed, and the answer surprises people: your bottleneck is memory bandwidth (how fast bytes move between memory and the compute units) and arithmetic throughput rarely limits you.
Generating one token means reading every weight in the model. So a 4.58 GiB model moves 4.58 GiB per token, and that cost repeats for every token of every request. At batch size one each of those bytes does a tiny amount of arithmetic before the next one's needed, so the GPU spends most of its time waiting. Page 23 formalises this as arithmetic intensity.
In llama.cpp's published benchmarks for Llama 3.1 8B, text generation gets faster as the quantization gets smaller: roughly 72 tokens/sec at Q4_K_M against 51 at Q8_0 and 29 at F16.
Prompt processing does the opposite. It stays broadly flat across quantizations, and F16 is the fastest of them.
That divergence is the thing to take away. Prefill processes many tokens at once against each loaded weight, so it's compute-bound and quantization doesn't help it. Generation produces one token at a time, so it's bandwidth-bound and shrinking the weights shrinks the work directly. Same model, same hardware, two opposite conclusions.
The source table doesn't say which hardware produced these figures, so read the pattern and not the absolute numbers.
Two things follow.
- Offloading layers to CPU is expensive. When a model doesn't fit, runtimes put some layers in system RAM. Every token then crosses the PCIe bus (far slower than VRAM), so a partial offload can cost more speed than just dropping to a smaller quantization would have.
- Batch size one is the worst case. It's also the normal case locally. Batching amortises weight loading across requests, so serving frameworks help enormously at volume and barely at all in a single chat session.
07Where you meet this in the wild
Local inference usually isn't about saving money. It's about constraints an API can't satisfy.
- Data can't leave. Regulatory or contractual, and no pricing argument overrides it.
- You need to work below the API surface: logits, activations, custom sampling, your own fine-tuned weights.
- Volume is high and per-call value is low, so per-token pricing dominates your costs.
- You need offline operation, or latency a network round trip can't meet.
- You want frontier quality, and the largest models aren't a local option at consumer scale.
- Traffic is spiky or low, and idle hardware still costs money while idle API calls don't.
- You'd be paying an engineer to babysit inference infrastructure that isn't your product.
- You need elastic concurrency, and one local GPU serving many users behaves like a queue.
08Interview questions
BeginnerHow do you estimate whether a model fits on a given GPU?
Add three terms. Weights are parameter count times bits per weight divided by eight, using the measured bits per weight, not the number in the quantization name. KV cache is 2 × layers × KV heads × head dimension × bytes per element, per token, multiplied by your context length and concurrency. Then add several hundred megabytes to about a gigabyte of framework overhead. As a sanity check, Q4_K_M lands near 0.6 to 0.7 GB per billion parameters before cache.
BeginnerWhy is Q4_K_M not 4 bits per weight?
Because k-quants store a scale and a minimum for each block of weights, and the medium mix deliberately keeps some tensors at higher precision. llama.cpp measures Q4_K_M at 4.8944 bits per weight for Llama 3.1 8B, about 22% above the naive assumption. Budget from the filename instead of the measured figure and you'll discover the model doesn't fit only after downloading it.
IntermediateWhy does the KV cache formula use KV heads and not attention heads?
Because grouped-query attention shares each key-value head across several query heads, so only the KV heads are actually stored. Llama 3 8B has 32 attention heads but 8 KV heads, so its cache is four times smaller than the multi-head equivalent: 128 KiB per token instead of 512 KiB. Use attention heads in the formula and you overestimate cache memory by exactly that ratio, which is enough to reject a configuration that would've worked.
IntermediateA model loads fine but runs out of memory after a long conversation. What happened?
Your KV cache grew past the remaining headroom. Weights are a fixed cost paid at load; the cache grows linearly with every token in the context. A model that leaves a gigabyte free after loading will exhaust it once the conversation's cache passes that gigabyte. Fixes: cap context length, drop to a smaller quantization to free capacity, or quantize the cache itself if your runtime supports it.
IntermediateWhy does a smaller quantization generate tokens faster on the same hardware?
Because generation is memory-bandwidth-bound. Producing each token requires reading every weight, so halving the bytes per weight roughly halves the bytes moved per token. llama.cpp's benchmarks show this directly: about 72 tokens/sec at Q4_K_M against 29 at F16 for Llama 3.1 8B. Prompt processing doesn't follow the same pattern and stays roughly flat across quantizations, because prefill processes many tokens per weight load and is compute-bound instead.
DeepYour model needs 26 GB and you have a 24 GB card. Walk through the options.
Take the cheapest quality loss first. Drop one quantization step, since Q5_K_M to Q4_K_M is roughly a 14% size reduction and usually closes a 2 GB gap on its own. Next, reduce maximum context length, which shrinks the cache term without touching weights at all, and check whether your runtime can quantize the KV cache to 8-bit. Offloading layers to CPU comes near-last, because every offloaded layer crosses PCIe on every token and usually costs more throughput than a smaller quantization would have. A second GPU works, but it adds communication between devices, and that's only worth it when quality genuinely can't be compromised.
DeepWhen is running locally genuinely cheaper than an API?
When utilisation is high and quality requirements are modest. Hardware is a fixed cost you pay whether or not it's busy, while API pricing is proportional to use, so local wins on sustained high-volume workloads and loses on spiky or low ones. Include engineering time too, since that's usually the largest hidden cost. Most local deployments are justified by privacy, offline operation or model access, not by cost, and treating cost as the main argument usually means nobody's done the arithmetic.
09Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.
Other techniques for this problem
A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.