←Home KnowML
Model AtlasChapter 33

Model Atlas: Language Models

Which model for which job. The frontier tier, the open-weight tier and the tier that fits on your laptop, plus how to read the leaderboards that rank them.

Reference 26 min read Snapshot: 6 September 2026 Assumes: none
Start reading
TL;DR

The right model is decided by constraints more than by a leaderboard rank. Ask in this order: can the weights leave your network, what latency can you tolerate, how much context do you need, what is the cheapest model that passes your own eval. The answer usually falls out before you ever look at a benchmark score.

The open-weight field is no longer a consolation prize. The strongest open models now sit within a few points of the closed frontier on published aggregates, and the gap that remains is narrower than the gap between a good and a bad prompt.

Everything numeric on this page is a dated snapshot. Treat the tables as a starting shortlist and the leaderboard links as the source of truth.

Read this before you trust a number below

This page was assembled on 6 September 2026. Model releases land weekly and leaderboard positions move within days of publication.

Two habits keep you honest. Check the date on any model comparison before you quote it. Re-derive the shortlist from a live leaderboard and not from a blog post, including this one.

The parts of this page that stay true are the structural ones: what each benchmark measures, how licences differ, and how cost is computed. Those are the parts worth memorising.

01The order to ask the questions in

Most model choices are over-thought at the top and under-thought at the bottom. Filtering in this order collapses a field of two hundred models to about three.

  1. Can the weights leave your network? If the answer is no, every closed API drops out immediately and you are choosing among open-weight models only. This single question eliminates most of the field, so ask it first.
  2. What licence can you accept? Open weights are not one thing. Some are Apache 2.0, some carry revenue thresholds, some forbid commercial use outright. Section 05 breaks this down.
  3. What is your latency budget? Time to first token and tokens per second are set by the serving stack as much as the model. Page 31 is the mechanism; here it is a filter.
  4. How much context do you need? Advertised context and useful context are different numbers. Long context also costs prefill time and cache memory on every request.
  5. What is the smallest model that passes your eval? Look for the smallest one that clears your bar, measured on your own data, and skip the search for the best model.
  6. What does it cost at your volume? Last, because the answer is meaningless until the four questions above have narrowed the field.
Why "which model is best" is the wrong opening question

Ranked lists answer a question almost nobody has. They rank general capability, averaged over tasks you do not run, on inputs that look nothing like yours.

Your actual question is narrower: which model clears my quality bar, inside my latency budget, at a price I can pay, under a licence my lawyer will sign. Ranked highest and good enough are frequently different models, and the second one is usually five times cheaper.

02Reading the benchmark board

You cannot use these tables without knowing what the columns mean. Each benchmark below measures something specific, and each one fails in a specific way.

What each one tests

  • MMLU-Pro. Ten-option multiple choice across academic subjects, a harder rebuild of the original MMLU. It measures recall plus light reasoning. Strong models cluster near the top, which makes it poor at separating them.
  • GPQA Diamond. A small set of graduate-level science questions written so that a non-expert with a search engine still fails. It is a genuine reasoning test, but the set is small enough that a handful of items swings the score.
  • AIME. Competition maths, fifteen problems per contest. One problem is worth about 6.7 points, so run-to-run variance is large and single-run numbers deserve suspicion.
  • SWE-bench Verified. Real GitHub issues from real repositories, human-validated as solvable. The model has to produce a patch that passes the project's tests.
  • Terminal-Bench. End-to-end tasks in a real shell. It measures the whole loop: planning, running commands, reading the output, recovering from failure.
  • LiveCodeBench. Programming problems published after each model's training cutoff, which is the design that makes it resistant to contamination.
  • Humanity's Last Exam (HLE). Deliberately built to stay unsaturated. Useful precisely because scores are low, so there is headroom to distinguish models.
  • LMArena Elo. Humans vote blind between two responses. It captures things static benchmarks miss, and it rewards formatting and tone alongside correctness.

The four traps

The reasoning-effort confound. A modern model is not one point on the chart. The same weights at low, high and maximum reasoning effort produce different scores, different latencies and different prices. Artificial Analysis lists these as separate rows for exactly this reason. Comparing one vendor's maximum setting against another's default is not a comparison.

The scaffold does half the work. On SWE-bench and Terminal-Bench, the harness around the model matters enormously. Retry policy, tool definitions, context management and file-editing format can move a score by double digits with the weights unchanged. A number without its scaffold described is not reproducible.

Contamination. Public benchmarks leak into training corpora. A high score on a five-year-old benchmark may mean the model memorised it, and benchmarks with rolling post-cutoff problems exist to get around this.

Self-reporting. Vendor-published tables are run by the vendor, on their scaffold, with their prompt. Independent re-runs routinely come in lower. Prefer a third-party harness when the decision matters.

The number that should decide it

None of the above. Build a set of fifty prompts drawn from your own traffic, with a grader you trust, and run every candidate against it. Section 07 is that project.

A public benchmark tells you which models are worth putting in the bake-off. Your fifty prompts tell you which one to ship. Teams that skip the second step spend the next quarter debugging a model that ranked well and does not fit.

03The frontier tier

Closed models behind an API. You get the strongest general capability available and no ability to inspect, host or pin the weights.

The table is a snapshot of the Artificial Analysis Intelligence Index on 6 September 2026. The index is an aggregate over several benchmarks, so it smooths out single-benchmark noise. Prices are blended per million tokens at the effort setting shown.

ModelVendorIndexBlended $/1MOutput tok/s
Claude Fable 5.1 (max)Anthropic57$6.1269
GPT-6 Astra (max)OpenAI55$2.5770
Claude Opus 5 (max)Anthropic54$4.2157
Muse Spark 1.3 (max)Meta53$0.96190
GPT-5.6 Sol (max)OpenAI51$1.2587
Grok 4.6 (high)xAI51$1.2563

Two things in that table are worth more attention than the ranking. The spread in price across the top six is more than six-fold, while the spread in index is four points. The spread in output speed is nearly three-fold, and for an interactive product that is felt more directly than a benchmark point.

Meta's Muse Spark 1.3, released on 2 September 2026, is the clearest illustration. It sits three points below the leader and runs at almost three times the tokens per second for a sixth of the price. For most products that is the better trade.

04The open-weight tier

These are weights you can download, host, quantise, fine-tune and pin. Most teams with a data-residency requirement end up here, and the ecosystem's centre of gravity now sits here too.

ModelParams (total / active)ContextLicenceModality
Kimi K32.8T / 104B MoE1MKimi K3 (revenue-tiered)text + vision
DeepSeek-V4-Pro~1.6T MoE—open weightstext
Qwen3.8 flagship2.4T / 95B MoE—Apache 2.0text + vision
GLM-5.3753B MoE1MGLM-5.3 (custom)text
GLM-5.3-Flash320B / 18B MoE300KMITtext + vision
DeepSeek-V4-Flash~291B MoE—open weightstext (vision variant)
Qwen3.8-Flash-Next180B MoE—Apache 2.0text + vision
Qwen3.8-27B27B dense262K native, 1M with YaRNApache 2.0text + vision + video
Gemma 4 31B31B dense256KApache 2.0text + vision

The column that changes your hardware bill is the pair of total and active parameters, because the first number alone does not tell you the cost.

Total parameters buy the rack, active parameters buy the speed

In a mixture-of-experts model, every expert has to be resident in memory, so total parameters set your VRAM requirement. Only a few experts run per token, so active parameters set your decode arithmetic.

Kimi K3 is the extreme case: 2.8T total against 104B active. It decodes at roughly the cost of a 104B dense model and needs the memory of a 2.8T one. That is a good deal if you own a rack and out of reach if you own a workstation.

Mixture-of-experts — one token's route through Kimi K3's 896 experts
token router picks 16 896 experts, all resident in GPU memory on every pass 112 cells drawn · each cell = 8 experts resident, idle for this token firing — 16 of 896 MEMORY BILL 2.8T params every expert loaded, whether it fires or not ARITHMETIC BILL 104B params the 16 chosen experts, plus attention, which always runs one token, one route
The grid is drawn to scale: 112 cells standing for 896 experts at 8 each, so the two lit cells are 16 experts firing. Everything in the grid occupies memory on every pass; only the lit part multiplies anything. Note that the expert ratio and the parameter ratio are not the same number — 16 of 896 experts is 56× sparsity, while 104B of 2.8T params is about 27×, because attention and the shared layers run for every token regardless of routing. The total tells you what hardware you must own; the active count tells you how fast it runs once you do.

Note the licences, because they do not all mean the same thing. GLM-5.3-Flash and the Qwen3.8 line ship under MIT and Apache 2.0, which impose almost nothing. Kimi K3 ships under a bespoke licence with revenue thresholds. Section 05 covers what that means in practice.

One structural change is worth naming. Open-weight releases are now natively multimodal by default and no longer a later variant. Kimi K3 carries its own vision encoder, and the Qwen3.8 and GLM-5.3-Flash lines accept images as a matter of course.

05The tier that runs on your own machine

Small models are not scaled-down frontier models. They are a separate design target, tuned for a fixed memory budget and a fixed power envelope.

The sizing arithmetic from page 30 is what matters here. Weights in memory are roughly parameters times bytes per parameter, so a 27B model at 4-bit needs about 14 GB before the KV cache. That number decides what fits, and a benchmark cannot.

ModelSizeLicenceFits onNotes
Qwen3.8-27B27B denseApache 2.024 GB GPU at 4-bitVision and video input, 262K context
Gemma 4 31B31B denseApache 2.024 GB GPU at 4-bitText and images, 256K context
Gemma 4 26B-A4B26B / 3.8B MoEApache 2.024 GB GPU at 4-bitMoE, so decodes far faster than its size
Gemma 4 12B12B denseApache 2.012 GB GPU at 4-bitAdds audio and video input
Gemma 4 E4B4.5B effectiveApache 2.0phone / 8 GB laptop128K context, audio and video
Gemma 4 E2B2.3B effectiveApache 2.0phoneSmallest of the family
Granite 4.2 3B / 8B3B / 8BApache 2.0laptopEnterprise-tuned, tool calling
Qwen3-0.6B0.6BApache 2.0anythingDraft model, classifier, router

A note on Gemma 4's effective parameters. The E-series reports an effective count lower than the raw one because part of the network is not resident per token. Read it as the memory figure that matters.

The most common mistake at this tier is reaching for the largest model that fits. Often the better move is a smaller model plus retrieval, or a smaller model plus one fine-tune on your task. Page 29 covers when the fine-tune is worth it.

A sub-1B model is not a chat model, and judging it as one misses the point. Its jobs are narrow: draft model for speculative decoding, intent classifier, router in front of a larger model, structured extraction with a fixed schema. Each of those is a real production role.

06Licences: the section everyone skips

This is the most common way a model choice goes wrong six months later, and twenty minutes of reading avoids it.

Open-weight licences fall into four buckets, and the differences are not cosmetic.

  • Permissive and standard. Apache 2.0 or MIT. These allow commercial use, modification and redistribution with no thresholds. The Qwen3.8 line, Gemma 4, Granite 4.2 and GLM-5.3-Flash sit here. Gemma moved to Apache 2.0 at version 4, replacing the custom terms earlier Gemma releases carried.
  • Revenue-tiered. Free until your business crosses a threshold, then a separate agreement or an attribution requirement kicks in. The Kimi K3 licence requires an agreement to run it as a paid model-as-a-service above roughly $20M annual revenue from that use.
  • Research and non-commercial. Free to evaluate but not to ship. Common in speech and audio, where several strong models are gated this way.
  • Closed. API access under terms of service. No weights, no pinning, no offline use.
"Open weights" is not "open source"

The two phrases are used interchangeably and they should not be. Open weights means the parameter file is downloadable. Open source, in the sense the term has carried for twenty-five years, implies a licence with no field-of-use restrictions and, under the OSI's AI definition, meaningful information about the training data.

Most models described as open source in press coverage are open weights under a custom licence. Read the licence file and ignore the headline. The question is narrow: does this licence permit the specific thing my company is about to do?

Two practical checks before you commit. Confirm the licence covers distillation if you plan to train on the model's outputs, because several licences restrict this. Confirm what happens to your fine-tuned derivative, because a derivative usually inherits the base licence.

07Embeddings and rerankers

Retrieval quality is usually the ceiling on a RAG system, and the embedding model sets that ceiling. It gets a fraction of the attention the generator gets.

An embedding model maps text to a vector so that similar meanings land close together. A reranker is a second, slower model that scores query and document together and reorders the top candidates. The standard pipeline retrieves fifty with the embedder and reranks to five.

ModelSizeDimsContextLicence
Qwen3-Embedding-8B8Bup to 409632KApache 2.0
Qwen3-Embedding-4B4Bup to 256032KApache 2.0
Qwen3-Embedding-0.6B0.6Bup to 102432KApache 2.0
BGE-M30.6B10248KMIT
multilingual-e5-small118M384512MIT
bge-reranker-v2-m30.6Breranker8KApache 2.0

The Qwen3-Embedding family reported the leading multilingual MTEB score at release, and it ships in three sizes with the same training recipe. The three sizes matter more than the headline score, since they let you profile the 0.6B, and upgrade only if measurement says the larger one earns its cost.

Three practical points that MTEB rank will not tell you.

  • Dimension is a storage decision. 4096 floats per chunk against 384 is more than a ten-fold difference in index size and query cost. Several models support truncating the vector, so you can trade a little accuracy for a lot of memory.
  • MTEB is a public benchmark like any other. Models are tuned against it. Rankings compress at the top and rarely survive contact with a domain corpus.
  • Rerankers buy more than a bigger embedder. Adding a small cross-encoder over the top fifty results usually beats upgrading the embedding model, at a fraction of the reindexing cost.

Page 11 covers where this sits in a retrieval system, including chunking, which is the other half of the ceiling.

08The cost model

Two arithmetic facts settle most build-versus-buy arguments, and both fit in a paragraph.

API pricing is quoted separately for input and output tokens, with output typically several times the price of input. A blended price assumes a fixed ratio, commonly three input tokens to one output token. Your ratio is probably not that, so recompute:

$$\text{cost per request} = T_{\text{in}} \cdot p_{\text{in}} + T_{\text{out}} \cdot p_{\text{out}}$$

With reasoning models, $T_{\text{out}}$ includes the tokens spent thinking, which you pay for and never see.

That last point is the one that surprises finance. A reasoning model at maximum effort can emit ten times the visible output in hidden reasoning tokens. The advertised per-token price is unchanged, and the bill is not.

For self-hosting, the comparison is a throughput calculation:

$$\text{cost per 1M tokens} = \frac{C_{\text{gpu-hour}}}{\text{tokens/sec} \times 3600} \times 10^{6}$$

Tokens per second is the aggregate across all concurrent requests and is different from the single-stream figure.

This distinction is where self-hosting estimates usually go wrong. Single-stream decode is memory-bandwidth-bound and slow, while aggregate throughput at high concurrency can be twenty times higher on the same card, so quoting the wrong one puts you out by an order of magnitude in either direction. Page 31 derives why.

Two more levers that beat model choice. Prompt caching removes prefill cost on a repeated prefix, which is most of the bill for agent loops. Routing sends easy requests to a small model and only escalates the hard ones, which routinely cuts spend by more than half without a quality change anyone notices.

09Build this

The single highest-value artefact in this whole area is an eval you trust. It takes a day and it outlives every model on this page.

Project The fifty-prompt bake-off ~1 day · any three candidate models

Build the small, ugly, specific eval that decides your model choice, and instrument it so it reports quality, latency and cost together. Public leaderboards pick the candidates, and this eval picks the winner.

  1. Collect fifty real inputs from your own traffic or your own backlog. Avoid synthetic ones. Include the five that currently fail, because those are the ones a model change has to fix.
  2. Write the grading rule before you look at any output. Exact match where you can, a rubric where you cannot, a stronger model as judge only where a human would agree with it.
  3. Run three candidates from different tiers: one frontier API, one large open-weight model, one small model you could host. Record output, wall-clock latency and token counts for every request.
  4. Grade blind by shuffling and stripping the model names before a human looks at anything.
  5. Plot quality against cost per request, with latency as the point size, and leave rank off the plot.
You'll know it worked when the plot has a model in the lower-left that is close enough on quality to the top-right one that you hesitate. That hesitation is the entire value of the exercise: you have found the cheap option that is good enough, and you found it with evidence instead of a hunch.
The stretch that pays for itself. Run the same fifty prompts weekly against your production model and keep the results. When a vendor silently updates a model behind a stable name, your grader catches it in a week rather than a customer catching it in a month. This is the cheapest regression test in the stack, and almost nobody has one.

10What breaks

Model selection fails quietly, and these failures show up after the decision has been made.

  • The leaderboard model does not fit your task. Aggregate indices average over tasks you do not run. A model ranked fourth overall can be first on your workload, and the only way to find out is to measure.
  • The API model changed underneath you. Vendors update models behind stable names. Formatting, tone and refusal behaviour shift without a version bump. Pin dated model identifiers where the vendor offers them.
  • Advertised context is not usable context. A million-token window does not mean a million tokens of reliable reasoning. Test retrieval accuracy at the depth you plan to use.
  • Quantisation changed the model you benchmarked. The 4-bit build you deploy is not the FP16 model whose scores you read. Re-run your eval on the exact artefact you will ship.
  • The licence at scale. A revenue-tiered licence is free until the quarter it stops being free, and by then the model is load-bearing.
  • Hidden reasoning tokens broke the budget. Cost was estimated from visible output length. Reasoning models bill for thinking.
  • The scaffold was the bottleneck. Teams upgrade the model when the retry logic, the tool schemas or the chunking was the problem. Change one thing at a time.

11Pick by use case

Six common shapes, and the shortlist each one implies. Start here, then run section 09.

Customer-facing chat

Time to first token is the product. Output speed matters more than the last two index points, and prompt caching on the system prompt pays for itself immediately.

Shortlist: a fast frontier model, or GLM-5.3-Flash / Qwen3.8-Flash-Next self-hosted.

Coding agent

Long-horizon tool use, where Terminal-Bench and SWE-bench are the relevant signals. It is the workload where the frontier tier still earns its price, and the scaffold is half the score.

Shortlist: top-of-index frontier models; Kimi K3 or GLM-5.3 if weights must stay in-house.

High-volume extraction

Millions of documents to a fixed schema. Nobody is waiting, so latency is irrelevant and cost per million tokens is the only metric that survives.

Shortlist: the smallest model that passes your schema check, batched at high concurrency.

Not a reasoning model at maximum effort.

Regulated or air-gapped

The weights cannot leave, so the choice reduces to the open-weight table, filtered by what your hardware budget can hold and what your licence review will approve.

Shortlist: Apache 2.0 and MIT models first, because the review is short.

On-device

Fixed memory, fixed battery, no network. Quality expectations have to be set by what fits, and the task usually has to be narrowed to match.

Shortlist: Gemma 4 E2B/E4B, Granite 4.2 3B, Qwen3-0.6B for narrow jobs. Page 30 for the sizing.

RAG over your own corpus

The generator is rarely the bottleneck. Retrieval quality is, and that is an embedding, chunking and reranking problem before it is a model-choice problem.

Shortlist: spend the effort on section 07 and a mid-size generator.

12Interview questions

BeginnerWhat is the difference between open weights and open source?

Open weights means the trained parameter file can be downloaded and run locally. Open source, in the established sense, means a licence with no field-of-use restrictions, and under the OSI's AI definition it also implies meaningful disclosure about training data and code. Most models the press calls open source are open weights under a custom licence, which may carry revenue thresholds, attribution requirements or restrictions on training other models from their outputs. The practical question is never which label applies, but whether the specific licence permits the specific deployment you have in mind.

BeginnerIn a mixture-of-experts model, why do total and active parameters both matter?

Total parameters set the memory requirement, because every expert must be resident even though only a few run per token. Active parameters set the compute and memory traffic of a single decode step, so they determine speed and arithmetic cost. A model with 2.8 trillion total and 104 billion active parameters decodes at roughly the cost of a 104-billion dense model but needs the memory footprint of a 2.8-trillion one. That combination is excellent economics on a large served cluster and unusable on a single workstation, which is why quoting only one of the two numbers is misleading.

IntermediateA vendor reports 90% on a benchmark and an independent harness reports 78%. What explains the gap?

Several things, and usually more than one at once. The vendor may run at a higher reasoning-effort setting, which is a different price and latency point on the same weights. The scaffold differs: retry policy, tool definitions, context management and output parsing can move agentic scores by double digits. Prompt formatting and few-shot examples differ. The evaluation harness may parse answers more strictly, counting near-misses as failures. None of these imply dishonesty, but they do mean a vendor number is not directly comparable to a third-party number, and the decision should rest on a run you control.

IntermediateHow would you choose between a frontier API and a self-hosted open-weight model?

Start with the constraints that are not negotiable. If the data cannot leave the network, the decision is already made. Otherwise compare on three axes: quality measured on your own eval rather than a leaderboard, total cost at your projected volume including the engineering time to run a serving stack, and operational risk, which cuts both ways since an API can change under you and a self-hosted model is yours to keep running. Self-hosting tends to win at high, steady volume with a narrow task, because you can use a small model and saturate the hardware. APIs tend to win at low or spiky volume and at the frontier of capability.

IntermediateWhy is a single AIME score a weak signal?

Because the contest has only fifteen problems, so each one is worth about 6.7 percentage points. Sampling noise alone can move a single-run score by more than the gap between two models being compared. Reasoning models are also stochastic, so repeated runs on the same weights disagree. A defensible AIME number is an average over many runs with a reported variance, and comparing two single runs tells you very little. The same argument applies with less force to any small benchmark, including GPQA Diamond.

DeepYour inference bill tripled after switching to a reasoning model, but request volume is flat. Why?

Because reasoning models bill for tokens you never see. The model emits an extended internal chain before the visible answer, and those tokens are charged at the output rate, which is typically several times the input rate. A cost estimate built from visible response length will therefore understate the bill by a large multiple, and the multiple grows with the reasoning-effort setting. The fixes are to lower the effort setting where the task does not need it, to route only hard requests to the reasoning model, and to estimate cost from measured total output tokens rather than from what appears on screen.

DeepHow would you build an evaluation that actually decides a model choice?

Take fifty real inputs from production traffic, deliberately including current failures, and fix the grading rule before seeing any output. Use exact match or a programmatic check wherever the task permits it, a written rubric where it does not, and a model judge only where a human would agree with the judge. Run every candidate through the identical scaffold, record latency and token counts alongside quality, and grade blind with model names stripped. Then plot quality against cost rather than ranking, because the decision is a trade, not an ordering. Re-run it on a schedule so that a silent vendor-side model update is caught by your harness rather than by a customer.

DeepWhat does a one-million-token context window actually buy you?

Less than the number suggests. Retrieval of a single fact from deep in the window is now largely solved, but reasoning that requires combining many facts spread across the window degrades well before the advertised limit. Long contexts also cost real resources: prefill time grows with prompt length and directly sets time to first token, and the KV cache grows linearly with the context, so concurrency falls as context rises. In practice a retrieval step that puts five relevant pages in a short context usually beats putting five hundred pages in a long one, and it is cheaper on every axis. Test at the depth you intend to use before designing around the headline number.

13Go deeper

These are the live sources. When this page and a leaderboard disagree, the leaderboard is right.

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

My Notes — 33 Model Atlas: Language Models

Free notes

Highlights on this page