Model Atlas: Vision & Generative Media
Understanding images, finding objects, reading documents, and generating pictures and video. Four jobs, four model families, and one benchmark culture that rewards looking good over being right.
"Computer vision" now covers four separate jobs: describe an image in language, locate things in it, read a document, and embed it for search. Each has its own best model, and the most common design error is using a vision-language model for a job a detector does better and a hundred times cheaper.
Generation is the one area where open weights are at parity for most work. A 6B model under Apache 2.0 produces publishable images in eight steps on a 16 GB card.
Generative benchmarks measure preference, and preference includes style. A model can win an arena by being prettier while ignoring half your prompt.
Assembled on 6 September 2026. Image and video generation move faster than any other area on this site, and a six-month-old comparison here is obsolete.
What stays true is the shape of each problem, the metric definitions, and the licence structure. Check the boards in section 13 before quoting a rank.
01Four jobs people call "vision"
Naming your job correctly eliminates most of the field before you compare anything. These four want different models and have different costs by orders of magnitude.
- Describe. Answer a question about an image in natural language. A vision-language model, which is a language model with an image encoder attached. Flexible, expensive, and vague about exactly where things are.
- Locate. A detector or segmenter produces boxes or pixel masks for objects. It is fast, cheap and precise, and limited to what it was trained or prompted to find.
- Read. A document model converts a page into structured text with its layout preserved. The work goes beyond character recognition into layout and reading-order reconstruction.
- Embed. A contrastive encoder turns an image into a vector for search, deduplication or zero-shot classification. It takes milliseconds per image and generates nothing.
Asking a vision-language model to count boxes on a shelf, or to return coordinates, for every frame of a video feed. It will produce plausible-looking answers, at a hundred to a thousand times the cost of a detector, with worse accuracy and no calibration.
The rule of thumb: if the output has a fixed shape, a specialist model wants that job. Use the VLM for the open-ended part and only for that part.
Generation is a fifth category, and it splits again into image, video and 3D. Sections 06 and 07 cover the first two.
02Vision-language models
Multimodality stopped being a separate product line. The strongest open language models now accept images by default, so the VLM table and the LLM table have largely merged.
| Model | Params | Licence | Inputs | Notable numbers |
|---|---|---|---|---|
| Kimi K3 | 2.8T / 104B MoE | Kimi K3 (revenue-tiered) | image + text | MoonViT-V2 401M encoder; MathVision 94.3; MMVU 82.1; OSWorld-Verified 84.8 |
| GLM-5.3-Flash | 320B / 18B MoE | MIT | image + text | Natively multimodal; 300K context |
| DeepSeek-V4-Flash-Vision | 305B MoE | open weights | image + text | Vision variant of the V4-Flash line |
| Qwen3.8-27B | 27B dense | Apache 2.0 | image + video + text | MathVision 94.6; OmniDocBench 1.5 91.1; OSWorld-Verified 84.3 |
| Gemma 4 31B | 31B dense | Apache 2.0 | image + text | MMMU-Pro 76.9; 256K context |
| Gemma 4 12B | 12B dense | Apache 2.0 | image + audio + video + text | Audio to 30s, video to 60s; fits a 12 GB card at 4-bit |
| Gemma 4 E4B | 4.5B effective | Apache 2.0 | image + audio + video + text | Runs on a phone |
What the benchmarks measure
- MMMU and MMMU-Pro. These are college-level questions across disciplines, with diagrams and figures, and they serve as the general capability aggregate. MMMU-Pro is the harder rebuild with harder distractors.
- MathVista and MathVision. Mathematical reasoning where the necessary information is in the figure. These separate models that read a chart from models that guess from the caption.
- DocVQA, ChartQA and OCRBench. These measure document and chart reading, the practical proxy for whether a model can handle your PDFs.
- OSWorld-Verified. Real tasks in a real desktop environment. It is the benchmark for GUI agents, and it requires accurate grounding and not just description.
- MMVU. Video understanding, which is a different problem again because it adds temporal reasoning and a much larger token budget.
A model can score in the nineties on visual mathematics and still miscount six identical objects, or return a bounding box that is twenty pixels off. Describing what is in an image and saying precisely where it is draw on different capabilities, and benchmarks weight the first far more heavily than the second.
If your task needs coordinates, test coordinates directly. Grounding scores are the ones that matter, and OSWorld-style agentic numbers are a better proxy for them than MMMU.
03Detection and segmentation
The oldest branch of applied vision, and still the cheapest way to answer "where is it" at scale. The 2026 change is that you no longer have to choose between open vocabulary and real-time speed.
Three families, and the choice between them comes down to whether you know your class list in advance.
YOLO and its relatives are trained on a fixed class list and are tiny and fast enough for video on embedded hardware. Adding a new class means collecting data and retraining.
Reach for it when the class list is fixed and the frame rate is the constraint.
SAM 3 detects, segments and tracks every instance of a concept described by a text phrase or an example image, in images and in video. No retraining for a new class.
Reach for it when the classes are open-ended, or when you are exploring before you commit to a list.
The Segment Anything lineage, prompted with a point, a box or a mask, producing pixel-accurate outlines. The tool when the shape matters and not just the location.
Reach for it for measurement, editing and annotation work.
The pattern that has become standard combines two of them. Use SAM 3 to auto-label a large image set with high-quality masks, then train a small YOLO on that output for production. You get open-vocabulary flexibility at labelling time and real-time speed at inference time, which neither model gives you alone.
mAP is the metric, usually averaged over intersection-over-union thresholds from 0.5 to 0.95 on COCO. Two things it hides matter before you compare numbers. Small objects drag the average down heavily, so a model tuned for them looks worse overall. And the non-maximum-suppression threshold is a tuning knob that moves mAP without changing the network.
04Document AI and OCR
This became a solved-enough problem in 2026, and the reason is that the framing changed. The task is no longer reading characters. It is reconstructing a document's structure.
A modern document model takes a page image and emits Markdown or HTML with headings, reading order, tables, figures and formulas intact. Character accuracy on clean text was never the hard part. Reading order on a two-column page with a sidebar, and table structure with merged cells, are the hard parts.
| Model | Approach | Notable |
|---|---|---|
| PaddleOCR-VL 1.5 | Document VLM | Around 94.5 on OmniDocBench v1.5 and 96.33 on v1.6; the accuracy leader, and the safe pick when the language is unknown |
| GLM-OCR | Document VLM | Around 94.62 on OmniDocBench v1.5 at roughly 1.86 pages per second; the throughput pick |
| DeepSeek-OCR 2 | MoE decoder, ~570M active | Optical context compression, tuned for grounded Markdown at high throughput |
| dots.ocr | Layout-aware VLM | Strong on complex layouts |
| olmOCR | VLM pipeline | Robust on messy scans and handwriting |
| Tesseract, classic PaddleOCR | Traditional OCR | Orders of magnitude cheaper on clean, simple text. Still the right answer sometimes. |
OmniDocBench is the benchmark that made this field comparable. It scores end-to-end page parsing across document types, with separate breakdowns for text, formulas, tables and reading order. Look at the sub-scores, because a model can lead overall and be weak on exactly the element type your corpus is full of.
One economic note: document VLMs are language models, so they cost language-model money per page. If ninety percent of your corpus is clean single-column text, route that ninety percent to a classical engine and send only the hard pages to the VLM. The saving is large, and easy pages lose no quality.
05Visual embeddings and retrieval
This is the quiet workhorse: it needs no generation and no prompt, takes milliseconds per image, and solves more production problems than anything else on this page.
A contrastive encoder in the CLIP and SigLIP lineage maps images and text into one shared space. That single property gives you four things at once:
- Search by text over an image corpus, with no labels anywhere.
- Zero-shot classification, by embedding class names and taking the nearest one.
- Deduplication and near-duplicate detection, which is usually the first thing a large image pipeline needs.
- Moderation triage, by distance to reference concepts, as a cheap first pass before an expensive model.
The newer multimodal retrievers go further. Qwen3-VL-Embedding-8B reports a mean task score around 67.9 on the multilingual MTEB, with a matching reranker, and it handles image and text queries in one model.
For documents specifically, late-interaction retrievers changed the standard recipe. Instead of running OCR and then embedding the text, they embed the page image directly and match at the patch level. That removes the OCR stage from the retrieval path entirely, along with every error it would have introduced.
06Image generation
The area where open weights are closest to parity, and where the interesting axis stopped being quality some time ago. It is now steps, memory and licence.
| Model | Size | Licence | Steps | Notes |
|---|---|---|---|---|
| Z-Image-Turbo | 6B | Apache 2.0 | 8 | Single-stream DiT, fits in 16 GB, bilingual text rendering |
| FLUX.1-schnell | 12B | Apache 2.0 | 1–4 | The permissive distilled default |
| FLUX.1-dev | 12B | non-commercial | ~20–50 | The most fine-tuned base model in the ecosystem, and the largest LoRA library |
| Krea 2 Turbo | 12B DiT | Krea 2 Community | 8 | Distilled from Krea 2, tuned for photographic realism |
| LLaDA-Image | 7B | open weights | — | Diffusion-language approach rather than a standard latent diffusion stack |
| SDXL | 3B | OpenRAIL++ | ~30 | Older, smallest, still the widest tooling support |
Closed models still lead the human-preference boards. GPT Image 2 held a clear lead on the LMArena text-to-image board through mid-2026, with Gemini's image models, Recraft, Midjourney and Ideogram behind it. The gap is real and it is narrower than the headline Elo spread suggests, because Elo rewards a house style as much as it rewards correctness.
A diffusion model's cost is the number of denoising passes, and parameter count matters less. Z-Image-Turbo at 6B and 8 steps is cheaper per image than SDXL at 3B and 30 steps, despite being twice the size.
Distillation is what collapsed that number from fifty to under ten, and it is the single largest cost reduction in generative imaging. It is not free: distilled models produce noticeably less variety across seeds. If you need twenty different options, the undistilled model is still the right tool.
Two capabilities now matter more than raw fidelity for most products. Text rendering inside the image, which was broken until recently and is the difference between a usable poster and a re-shoot. And instruction-based editing, where you hand the model an image and a sentence describing the change, which is a far more common production need than generating from nothing.
07Video generation
Two years behind image generation and closing fast. The open tier crossed from demo to production in 2026, largely because of one architectural change: audio and video generated together.
| Model | Size | Licence | Notes |
|---|---|---|---|
| LTX-2.5 | 22B DiT | LTX-2.x Community (free under $10M revenue) | Native synchronised audio and video in one pass; up to 4K via spatial upscaling; FP8 for smaller cards |
| MiniMax-H3 | 33B | open weights | Camera moves exposed as prompt tokens, so shot planning is controllable rather than accidental |
| Wan 2.2-TI2V-5B | 5B | open weights | The small one. Text and image to video on modest hardware |
| HunyuanVideo 1.5 | 8.3B | open weights | Long clips on a single consumer card |
| FastVideo distillations | varies | varies | Four-step distillations of larger models. Same steps-versus-diversity trade as images |
LTX-2.5 is worth studying beyond its outputs. It uses a Gemma 4 12B model as its text encoder, which is a good illustration of how the tiers on these three pages compose: a language model from page 33 is a component inside a video model here.
OpenAI notified developers on 24 March 2026 that the Sora 2 model and the Videos API would be removed. The consumer app closed on 26 April 2026, and the API is removed on 24 September 2026. Anything built on it had six months to migrate.
This is the clearest argument for open weights, and it has nothing to do with cost or quality: a downloaded checkpoint cannot be deprecated. If a product depends on a specific look, keeping the weights that produce it is a business continuity decision.
The practical selection criteria for video have little to do with benchmark scores. Ask about seconds of output per GPU-minute, peak VRAM at your target resolution, whether audio comes free or needs a second model, and what control you get: first and last frame, camera path, reference images. Those four decide whether a model fits a production pipeline.
08Reading a generative benchmark
Generative evaluation is harder than discriminative evaluation because there is no correct answer. Everything here is a proxy, and knowing which proxy is which prevents most bad decisions.
- Arena Elo. Humans vote blind between two outputs for the same prompt. It is the best available signal, but it conflates prompt adherence with aesthetic style, so a model with a pleasing house look wins votes it did not earn on correctness.
- GenEval. Automated compositional checks: is there the right number of objects, the right colours, the right spatial relations. It measures obedience and ignores beauty, which is why it is worth reading alongside Elo.
- DPG-Bench. Adherence to long, dense prompts with many simultaneous constraints. The closest proxy for real production prompts.
- VBench. Video, broken into dimensions such as subject consistency, motion smoothness and temporal flicker. Read the individual dimensions and treat the total with caution.
Fréchet Inception Distance compares the statistics of generated images against a reference set. It answers "do these look like that distribution", which was the right question when models produced obvious artefacts.
It is the wrong question now. FID cannot see whether the model followed the prompt, it shifts with sample count and feature extractor, and a model can lower it by generating safe, average images. Once outputs became broadly realistic, the useful axis moved to instruction-following, and FID does not measure that at all.
One last correction applies to every published example. Demo images and videos are best-of-N, chosen by people whose job is to make the model look good. Your first generation is a sample from the distribution and is not from the highlight reel. Judge on twenty unedited outputs from your own prompts.
09Build this
Two hours with a fixed prompt set will tell you more than a week of reading comparisons, and it exposes a cost that no leaderboard reports.
Run a fixed prompt set across models, rate blind, and then measure the thing arenas never test: how much variety you lose when you drop from fifty steps to eight.
- Write twenty prompts covering the cases you need. Include at least three with counting ("exactly five"), three with spatial relations, and three with text to render inside the image.
- Generate four images per prompt per model, with fixed seeds, across a distilled model, an undistilled one and one closed API. Log step count, wall-clock time and peak VRAM.
- Rate blind on two separate scales: did it follow the prompt, and does it look good. Keep the scores apart, because they will disagree and the disagreement is informative.
- For the diversity test, generate sixteen images from one prompt with sixteen different seeds, for the distilled and undistilled model.
- Compare the two grids of sixteen side by side.
10What breaks
The failures here are mostly framing errors made early and discovered late.
- A VLM was used where a detector belonged. The answers are plausible, nothing is calibrated, and the bill scales with frames and not with objects.
- Grounding was never tested. The model describes the scene beautifully and its coordinates are unusable. Description scores do not predict grounding accuracy.
- Every page went through the expensive path. Clean single-column text does not need a document VLM. Routing by difficulty is usually a large, easy saving.
- The distilled model killed variety. Discovered during a shoot when every seed returns the same composition.
- The licence blocked the launch. FLUX.1-dev is non-commercial, several video models carry revenue thresholds, and the best-looking option is often the one you cannot ship.
- The API was deprecated. Sora 2 is the concrete case. Products whose visual identity depends on one hosted model inherit that model's lifecycle.
- Demo quality was assumed to be typical. Published samples are curated, so budget for a rejection rate and measure it before you promise a delivery schedule.
- Video VRAM was measured at the wrong resolution. Memory grows with frames times resolution. A model that fits at 480p and five seconds may not fit at 1080p and ten.
11Pick by use case
Six shapes, and the model family each one implies.
Structure matters more than prose: tables, totals, dates, reading order. A document VLM with a classical fallback for the easy pages.
PaddleOCR-VL for accuracy, GLM-OCR for throughput, Tesseract for the clean ninety percent.
Fixed classes, many frames per second, often on embedded hardware. A closed-set detector, auto-labelled with an open-vocabulary model during development.
SAM 3 to label, a small YOLO to serve.
Not a VLM per frame.
Text queries against unlabelled images, plus deduplication. A contrastive encoder and a vector index, with no generation anywhere in the path.
A SigLIP-class encoder, or a multimodal retriever if queries mix image and text.
Brand consistency, text inside the image, and many variations per brief. Fine-tuning and LoRA ecosystem matter more than base quality.
FLUX.1-dev if non-commercial terms permit the workflow, Z-Image-Turbo or FLUX.1-schnell if they do not.
Clips with synchronised audio, produced on a schedule. Seconds per GPU-minute is the metric that decides whether the pipeline is viable.
LTX-2.5 for joint audio and video, MiniMax-H3 when the camera move is part of the brief.
An agent that reads an interface and clicks it needs grounding and not just description, so the relevant benchmark is OSWorld and not MMMU.
A VLM with strong agentic scores. Test coordinate accuracy on your own screenshots first.
12Interview questions
BeginnerWhen should you use a detector rather than a vision-language model?
Whenever the output has a fixed shape and the class list is known. A detector returns boxes or masks with confidence scores, runs in milliseconds, and can be calibrated and thresholded, while a vision-language model returns free text that has to be parsed and cannot be calibrated the same way. The cost difference is one to three orders of magnitude per image, which matters enormously on video. The reasonable division of labour is to use the detector for locating and counting, and reserve the language model for open-ended questions where no fixed schema exists.
BeginnerWhat does an open-vocabulary detector give you that a YOLO does not?
The ability to find classes it was never trained on, specified at inference time by a text phrase or an example image. A closed-set detector like YOLO can only produce the labels in its training set, so a new class means collecting data and retraining. An open-vocabulary model such as SAM 3 accepts a concept prompt and finds every instance of it in an image or a video. The trade is speed and size: closed-set detectors remain far faster and smaller, which is why the common production pattern uses the open-vocabulary model to generate labels and then trains a small closed-set detector on them.
IntermediateWhy has FID fallen out of use for evaluating image generation?
Because it answers a question that is no longer the interesting one. FID measures the distance between the feature statistics of generated images and those of a reference set, so it detects unrealistic outputs, which was the dominant failure mode when it was introduced. It cannot see whether the model followed the prompt, and prompt adherence is where models now differ. It is also unstable with respect to sample count and the choice of feature extractor, and it can be improved by generating safe, average images, which is the opposite of what anyone wants. Evaluation moved to human preference arenas for overall quality and to compositional benchmarks like GenEval and DPG-Bench for adherence.
IntermediateWhat does distillation cost a diffusion model, beyond the obvious?
Diversity. A distilled model reaches a good image in four to eight denoising steps instead of thirty to fifty, which is a very large cost reduction because a diffusion model's expense is dominated by the number of passes rather than by parameter count. The hidden cost is that the output distribution collapses: generating sixteen images from sixteen seeds yields variations on one composition rather than several genuinely different ideas. That is fine when you know what you want and bad when you are exploring. It is also invisible on arena leaderboards, which score single outputs, so it has to be measured with a seed grid.
IntermediateWhy is modern OCR framed as document parsing rather than character recognition?
Because character accuracy on clean text stopped being the bottleneck years ago. What downstream systems actually need is structure: which text is a heading, what order the columns are read in, which cells belong to which row of a table, where a formula begins and ends. A stream of correct characters in the wrong order is unusable, and merged table cells destroy the meaning of every number in the row. So the current generation of models emits Markdown or HTML with layout preserved, and benchmarks like OmniDocBench score end-to-end page reconstruction with separate breakdowns for text, formulas, tables and reading order.
DeepA model tops the image arena but your users complain it ignores their prompts. Explain.
Arena Elo is a human preference score, and human preference conflates two things: whether the output satisfies the request, and whether it looks appealing. Voters see two images for a prompt and choose the one they like, so a model with a strong, attractive house style accumulates wins even when it quietly drops constraints, especially counting, spatial relations and text rendering, which are easy to overlook at a glance. The diagnosis is to score adherence separately with a compositional benchmark such as GenEval or DPG-Bench, or with your own rubric applied to your own prompts. It is common for adherence and aesthetic rankings to put different models first, and a product with specific requirements should be selecting on adherence.
DeepHow would you build a production pipeline that reads ten million PDFs?
Route by difficulty rather than sending everything through one model. Classify each page cheaply first: clean single-column text with an embedded text layer needs no vision model at all, and a classical engine handles simple scans at a tiny fraction of the cost. Send only complex layouts, tables, formulas and poor scans to a document VLM, since those are language models and cost language-model money per page. Validate structurally rather than by eye, checking that tables parse, that numbers reconcile and that reading order is monotonic, and sample a few hundred pages against a hand-built ground truth broken down by element type. Finally, consider whether retrieval even needs the text: a late-interaction retriever that embeds page images directly removes the parsing stage from the search path and every error it would have contributed.
DeepWhat are the real selection criteria for a video generation model?
Very little to do with published quality scores. Four operational questions decide it. Throughput, expressed as seconds of finished output per GPU-minute at your target resolution, because that sets both cost and schedule. Peak memory at that resolution and duration, since memory grows with frames times resolution and a model that fits at 480p may not fit at 1080p. Whether audio is generated jointly, because a separate audio model adds a synchronisation problem that joint models avoid. And what control surface exists: first and last frame conditioning, camera path, reference images, since a model that cannot be directed is a novelty rather than a tool. Licence sits alongside these, as several strong open video models carry revenue thresholds that convert to a paid agreement at scale.
13Go deeper
Live boards first, then the papers that define the metrics.
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.