←Home KnowML
Hands-onChapter 30

Running Models Locally

Will this model run on my machine? About a minute of arithmetic answers that, and the same arithmetic tells you how fast it will be once it does.

20 min read Assumes: KV cache (08), quantization theory (23)
Start reading
TL;DR

Three things occupy your VRAM: model weights, the KV cache, and framework overhead. Weights dominate at short context and the KV cache overtakes them at long context, so compute both. Guessing is how a model that loaded fine at 2k context runs out of memory at 32k.

Quantization is the main lever, and your runtime picks the format for you, so quality does not decide it: GGUF for llama.cpp and Ollama, AWQ or GPTQ for vLLM, bitsandbytes NF4 for QLoRA fine-tuning (that one isn't a serving format at all).

Names mislead: Q4_K_M isn't 4 bits per weight and measures 4.89. Once the model fits, memory bandwidth sets your speed and FLOPs barely matter, so a smaller quantization generates tokens faster on identical hardware. One sentence for an interview: local inference is a memory-capacity problem first, a memory-bandwidth problem second, and a compute problem almost never.

01What everybody asks first

"Will this run on my machine?" is answerable in advance. Most people answer it by downloading 40 GB and watching it fail.

Page 23 covers why quantization works and how PagedAttention manages a KV cache. This page is the practical counterpart: what fits, what to install, and what to expect once it's running.

Your GPU has one fixed pool of memory, and three things compete for it:

  • Model weights. Fixed once you choose a model and a quantization level. Usually the largest single item.
  • KV cache. Grows linearly with context length and with the number of concurrent requests. Small at 2k context, dominant at 128k.
  • Overhead. Activations, CUDA context, the framework itself. Small, at a few hundred megabytes to about a gigabyte, and people forget it.

If the sum exceeds your VRAM you get one of three outcomes: the load fails outright, the runtime quietly offloads layers to CPU and slows down by an order of magnitude, or you hit an out-of-memory error partway through a long conversation. That last one is the KV cache growing past your headroom, and it's the one that catches people.

The same 24 GB card, four ways — drawn to scale, 20 px per GB
weights KV cache overhead free 24 GB 70B · FP16 4k context 140 GB of weights — bar runs ~6× off the chart 70B · Q4_K_M 4k context ~40 GB — quantized hard and still 1.7× too big 8B · Q4_K_M 4k context 4.9 GB 17.6 GB free 8B · Q4_K_M 128k context 4.9 GB 16 GB of KV cache Rows 3 and 4 are identical weights on an identical card. Only the context length moved, and it took the card from 27% full to nearly full.
Everything here is drawn at one scale, 20 px per GB, so the bars are comparable and the top two do run off the chart instead of being drawn short. KV figures assume an 8B GQA model (32 layers, 8 KV heads, head dim 128, FP16), so 128 KB per token: 0.5 GB at 4k, 16 GB at 128k. That's why "will model X fit in Y GB" is unanswerable without a context length attached, and why an out-of-memory error can arrive an hour into a conversation instead of at load time.

02Do the arithmetic once

Learn to compute this and you'll never have to trust a compatibility table again. Two real terms and one fudge factor.

Worked example Does Llama 3 8B fit in 8 GB of VRAM at 8k context?

Every number below comes from the model's published config.json and from llama.cpp's own measurements. Your framework overhead will differ, so treat that one term as illustrative. The weight and cache terms are exact.

  1. $$\text{weights} = N_{\text{params}} \times \frac{\text{bits per weight}}{8} \;\text{bytes}$$
    The naive version, and why it's wrong. People assume 4-bit means half a byte per parameter. Llama 3 8B has 8,030,261,248 parameters, so that predicts 4.0 GB. The real Q4_K_M file is 4.58 GiB. A name isn't a footprint.
  2. $$8.03 \times 10^9 \times \frac{4.8944}{8} \;=\; 4.91 \times 10^9 \;\text{bytes} \;=\; 4.58\ \text{GiB}$$
    Why 4.89 and not 4.00: k-quants store a scale and a minimum per block of weights, and the _M variant deliberately keeps some tensors at higher precision. Both cost bits. llama.cpp publishes a measured 4.8944 bits per weight for this model, and that matches the observed file size exactly.
  3. $$\begin{aligned}\text{KV bytes/token} = 2 \;\times\;& n_{\text{layers}} \times n_{\text{kv heads}} \\ \times\;& d_{\text{head}} \times \text{bytes}\end{aligned}$$
    Why the leading 2: you cache a key and a value for every token, at every layer. The term everybody gets wrong is $n_{\text{kv heads}}$. With grouped-query attention it's far smaller than the number of attention heads, and picking the wrong one inflates your estimate several-fold.
  4. $$\begin{aligned}2 \times 32 \times 8 \times 128 \times 2 \;&=\; 131{,}072 \;\text{bytes} \\ &=\; 128\ \text{KiB per token}\end{aligned}$$
    Where these come from: Llama 3 8B's config gives 32 layers, 8 key-value heads, and a head dimension of 128 (hidden size 4096 over 32 attention heads). The final 2 is bytes per element for an FP16 cache. Note the 32 attention heads against only 8 KV heads.
  5. $$131{,}072 \times 8192 \;=\; 1{,}073{,}741{,}824 \;\text{bytes} \;=\; 1.00\ \text{GiB}$$
    A clean coincidence worth remembering: this model at 8k context needs exactly one gibibyte of KV cache. Had it used full multi-head attention, with 32 KV heads instead of 8, the same calculation gives 4 GiB. Grouped-query attention is quietly doing a lot of work for local users.
  6. $$4.58 + 1.00 + \approx\!0.7 \;\approx\; 6.3\ \text{GiB total}$$
    The fudge factor. Activations, the CUDA context and framework allocations are real but modest. Budget several hundred megabytes to roughly a gigabyte. That total leaves usable headroom inside 8 GB, so yes, it fits. Doubling the context to 16k costs exactly one more gibibyte, and that is the term to watch.
The habit worth forming. Compute the weight term and the cache term separately, because they scale differently. Weights are fixed the moment you pick a file. Cache grows with every token of context and every concurrent request. A model that loads comfortably at 2k can run out of memory at 32k on the same card, and that's probably the most common local-inference surprise there is.

Practitioners quote roughly 0.6 to 0.7 GB per billion parameters at Q4_K_M, and that agrees with the calculation above (4.58 GiB is 4.92 GB, over 8.03B parameters, so 0.61), so it's a fine sanity check. It stays a rule of thumb because it silently assumes a 4-bit k-quant, ignores context length, and says nothing about how many KV heads the architecture has.

03Formats you'll meet

Quantization format looks like a quality decision. In practice it's a compatibility decision, because each runtime eats its own format. Pick the runtime first and the format usually follows.

  • GGUF is the native format for llama.cpp, and therefore for Ollama and LM Studio, which build on it. It's the only one of these designed for CPU and for hybrid CPU-plus-GPU execution.
  • GPTQ is a post-training weight quantization algorithm aimed at GPU inference, and it is widely supported, vLLM included.
  • AWQ is activation-aware weight quantization. It protects the weights that matter most (the ones whose observed activations are largest), and it's a common choice for GPU serving.
  • bitsandbytes NF4 quantizes on load and not ahead of time. It is the standard path for QLoRA fine-tuning and is not a serving format. See 29.
The naming trap that costs people a GPU

Q4_K_M reads like "4 bits per weight." It measures 4.8944. Q8_0 reads like 8 bits and measures 8.50.

Q4 is the base block type. K means k-quants, a block-wise scheme carrying a scale and a minimum per block. M is the medium mix, and it keeps selected tensors at higher precision than the name suggests.

Budget from measured bits per weight, never from the digit in the filename. A 22% underestimate is exactly enough to make a model you'd budgeted for refuse to load.

llama.cpp publishes measured figures for Llama 3.1 8B. Here's the real size ladder:

QuantQ2_KQ3_K_MQ4_K_MQ5_K_MQ6_KQ8_0F16
Bits per weight3.164.004.895.706.568.5016.00
Size (GiB)2.953.744.585.336.147.9514.96

On quality, degradation is usually gradual, and how gradual seems to depend on both the model and the task. Q4_K_M is the widely used default because it sits at a good size-to-quality point.

Users start reporting noticeable losses below roughly 4 bits per weight. The very low quants exist for when fitting at all matters more than fidelity. Measure on your own task and do not trust a general claim, this one included.

04Which tool to pick

Tools split cleanly by what you're doing. Ask whether you're serving one user or many.

  • One user, your own machine. Ollama for the shortest path from nothing to a running model. LM Studio if you want a GUI. llama.cpp directly when you need control over build flags, quantization and runtime settings. The first two are wrappers around the third.
  • Many concurrent users. vLLM, TGI or SGLang. These implement continuous batching and a paged KV cache, and that's what turns idle GPU time into throughput. Page 23 explains the mechanisms.
  • Apple Silicon. MLX is built for it, though llama.cpp also runs well on Metal.

A common mistake is reaching for a serving framework to run a single local chat. vLLM's advantages are throughput advantages. With one user and a batch size of one, you're paying setup complexity for benefits you can't reach.

05Hardware tiers

Everything below follows from the arithmetic in section 02, applied at Q4_K_M with a modest context. Treat the model sizes as what fits comfortably and not as hard ceilings, since context length moves these more than anything else.

  • 8 GB VRAM. 7B and 8B models fit at 4-bit with room for a normal context (4k to 8k, going by the worked example above). That's roughly the entry point where local inference stops being painful.
  • 12 to 16 GB. This range runs 7B to 14B comfortably, with longer contexts and no need to watch the cache, and the extra headroom mostly buys context and not parameters.
  • 24 GB. This fits 30B-class models at 4-bit, or a smaller model with a large context, which is a meaningful step up.
  • 48 GB and above. This is 70B-class territory. llama.cpp measures a 70B at 43.1 GB in Q4_K_M, so 48 GB works with care and multi-GPU is more comfortable.
  • Apple Silicon. Unified memory changes the calculation, because CPU and GPU share one pool. A machine with plenty of it can hold models that'd otherwise need a very expensive discrete GPU. Bandwidth varies by chip tier, and bandwidth is what governs speed.
  • CPU only. This works, thanks to GGUF, but generation is slow enough to feel like waiting and not conversing, because you're limited by system RAM bandwidth.

For scale at the top end, llama.cpp's own table puts Llama 3.1 405B at 249.1 GB even in Q4_K_M, against 1,625.1 GB unquantized. Some models aren't a local option at any consumer tier.

06What makes it slow

Once a model fits, the question becomes speed, and the answer surprises people: your bottleneck is memory bandwidth (how fast bytes move between memory and the compute units) and arithmetic throughput rarely limits you.

Generating one token means reading every weight in the model. So a 4.58 GiB model moves 4.58 GiB per token, and that cost repeats for every token of every request. At batch size one each of those bytes does a tiny amount of arithmetic before the next one's needed, so the GPU spends most of its time waiting. Page 23 formalises this as arithmetic intensity.

What llama.cpp's benchmarks show

In llama.cpp's published benchmarks for Llama 3.1 8B, text generation gets faster as the quantization gets smaller: roughly 72 tokens/sec at Q4_K_M against 51 at Q8_0 and 29 at F16.

Prompt processing does the opposite. It stays broadly flat across quantizations, and F16 is the fastest of them.

That divergence is the thing to take away. Prefill processes many tokens at once against each loaded weight, so it's compute-bound and quantization doesn't help it. Generation produces one token at a time, so it's bandwidth-bound and shrinking the weights shrinks the work directly. Same model, same hardware, two opposite conclusions.

The source table doesn't say which hardware produced these figures, so read the pattern and not the absolute numbers.

Two things follow.

  • Offloading layers to CPU is expensive. When a model doesn't fit, runtimes put some layers in system RAM. Every token then crosses the PCIe bus (far slower than VRAM), so a partial offload can cost more speed than just dropping to a smaller quantization would have.
  • Batch size one is the worst case. It's also the normal case locally. Batching amortises weight loading across requests, so serving frameworks help enormously at volume and barely at all in a single chat session.

07Where you meet this in the wild

Local inference usually isn't about saving money. It's about constraints an API can't satisfy.

Coding assistants
Completion that never leaves the laptop
Proprietary code that can't be sent to a third party, plus latency low enough for inline completion. Small models compete well here, because the task is narrow and the context is already on the machine.
Regulated data
When the data legally can't leave
Health records, legal discovery, defence work. The requirement is that no bytes cross an organisational boundary, and that rules out a hosted API whatever its terms say.
Edge and offline
No network to depend on
Vehicles, field equipment, remote sites, air-gapped networks. The model has to fit a fixed hardware budget nobody's going to upgrade, so quantization stops being an optimisation and becomes a requirement.
Bulk processing
Millions of small, tolerant jobs
Classification, tagging and extraction over a large corpus. Per-token API pricing scales linearly with volume and owned hardware doesn't, so high-volume low-stakes work is where local economics win.
Experimentation
Poking at the actual weights
Interpretability, custom sampling, activation inspection, evaluating your own fine-tune from 29. An API gives you tokens. Local gives you the model.
Run it locally when
  • Data can't leave. Regulatory or contractual, and no pricing argument overrides it.
  • You need to work below the API surface: logits, activations, custom sampling, your own fine-tuned weights.
  • Volume is high and per-call value is low, so per-token pricing dominates your costs.
  • You need offline operation, or latency a network round trip can't meet.
Use an API when
  • You want frontier quality, and the largest models aren't a local option at consumer scale.
  • Traffic is spiky or low, and idle hardware still costs money while idle API calls don't.
  • You'd be paying an engineer to babysit inference infrastructure that isn't your product.
  • You need elastic concurrency, and one local GPU serving many users behaves like a queue.

08Interview questions

BeginnerHow do you estimate whether a model fits on a given GPU?

Add three terms. Weights are parameter count times bits per weight divided by eight, using the measured bits per weight, not the number in the quantization name. KV cache is 2 × layers × KV heads × head dimension × bytes per element, per token, multiplied by your context length and concurrency. Then add several hundred megabytes to about a gigabyte of framework overhead. As a sanity check, Q4_K_M lands near 0.6 to 0.7 GB per billion parameters before cache.

BeginnerWhy is Q4_K_M not 4 bits per weight?

Because k-quants store a scale and a minimum for each block of weights, and the medium mix deliberately keeps some tensors at higher precision. llama.cpp measures Q4_K_M at 4.8944 bits per weight for Llama 3.1 8B, about 22% above the naive assumption. Budget from the filename instead of the measured figure and you'll discover the model doesn't fit only after downloading it.

IntermediateWhy does the KV cache formula use KV heads and not attention heads?

Because grouped-query attention shares each key-value head across several query heads, so only the KV heads are actually stored. Llama 3 8B has 32 attention heads but 8 KV heads, so its cache is four times smaller than the multi-head equivalent: 128 KiB per token instead of 512 KiB. Use attention heads in the formula and you overestimate cache memory by exactly that ratio, which is enough to reject a configuration that would've worked.

IntermediateA model loads fine but runs out of memory after a long conversation. What happened?

Your KV cache grew past the remaining headroom. Weights are a fixed cost paid at load; the cache grows linearly with every token in the context. A model that leaves a gigabyte free after loading will exhaust it once the conversation's cache passes that gigabyte. Fixes: cap context length, drop to a smaller quantization to free capacity, or quantize the cache itself if your runtime supports it.

IntermediateWhy does a smaller quantization generate tokens faster on the same hardware?

Because generation is memory-bandwidth-bound. Producing each token requires reading every weight, so halving the bytes per weight roughly halves the bytes moved per token. llama.cpp's benchmarks show this directly: about 72 tokens/sec at Q4_K_M against 29 at F16 for Llama 3.1 8B. Prompt processing doesn't follow the same pattern and stays roughly flat across quantizations, because prefill processes many tokens per weight load and is compute-bound instead.

DeepYour model needs 26 GB and you have a 24 GB card. Walk through the options.

Take the cheapest quality loss first. Drop one quantization step, since Q5_K_M to Q4_K_M is roughly a 14% size reduction and usually closes a 2 GB gap on its own. Next, reduce maximum context length, which shrinks the cache term without touching weights at all, and check whether your runtime can quantize the KV cache to 8-bit. Offloading layers to CPU comes near-last, because every offloaded layer crosses PCIe on every token and usually costs more throughput than a smaller quantization would have. A second GPU works, but it adds communication between devices, and that's only worth it when quality genuinely can't be compromised.

DeepWhen is running locally genuinely cheaper than an API?

When utilisation is high and quality requirements are modest. Hardware is a fixed cost you pay whether or not it's busy, while API pricing is proportional to use, so local wins on sustained high-volume workloads and loses on spiky or low ones. Include engineering time too, since that's usually the largest hidden cost. Most local deployments are justified by privacy, offline operation or model access, not by cost, and treating cost as the main argument usually means nobody's done the arithmetic.

09Go deeper

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

Other techniques for this problem

A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.

My Notes — 30 Running Models Locally

Free notes

Highlights on this page