Modern Vision Foundation Models
How attention took over vision: patches-as-tokens, contrastive image-text pretraining, promptable segmentation, and the encoder-projector-LLM recipe that turns all of it into a vision-language model.
Four ideas carry this page, and they build on each other:
- Vision Transformer (ViT). Chops an image into fixed-size patches, treats each patch as a token, and runs the exact same transformer encoder from page 08 over the resulting sequence. It uses no convolutions and has no built-in notion that nearby pixels matter more, and given enough data it still works.
- CLIP. Trains an image encoder and a text encoder together, so matching image-caption pairs land close in a shared embedding space and mismatched pairs land far apart. This makes zero-shot classification possible: compare an image's embedding to a handful of candidate text prompts, pick the closest, no fine-tuning required.
- Segment Anything (SAM). Applies the same "prompt instead of retrain" philosophy to segmentation. Click a point or draw a box, get a mask, for objects the model never saw labeled during training.
- The VLM recipe. Stack an image encoder (often CLIP's), a small projector layer, and an LLM. It is the dominant recipe behind almost every modern vision-language model.
01Intuition
ViT's whole trick: stop thinking of an image as a grid of pixels a convolution slides over, and start thinking of it as a sentence made of patches.
Page 08 covers a transformer that takes a sequence of word tokens and lets every token attend to every other token. ViT doesn't invent a new mechanism for vision. It invents a new way to turn an image into a sequence, then hands that sequence to the identical transformer encoder.
Cut a 224×224 image into a grid of 16×16 patches, flatten each patch into a vector, project it into an embedding, and you have a sequence of "visual words" a transformer can process exactly like a sentence. That is the conceptual leap, and everything else is bookkeeping.
CLIP's trick is different: during training it classifies nothing, and instead learns a shared embedding space between pictures and words.
Feed it an image of a dog and the caption "a photo of a dog," and it nudges the two embeddings closer together. Feed it that same image paired with an unrelated caption like "a photo of a mountain," and it pushes those embeddings apart. Do this over hundreds of millions of naturally occurring image-caption pairs scraped from the web.
What comes out is a space where "visually similar to X" and "described similarly to X" become the same notion. Classification then reduces to a nearest-neighbor lookup: embed the image, embed a few candidate class names as text, see which one it is closest to.
SAM's trick is the simplest to state and the hardest to have built. Do not train a model to output "the segmentation mask for a cat." Train it to output "the segmentation mask for whatever the user is pointing at." Then train it on so many diverse point-mask pairs that it generalizes to objects, textures, and domains it never explicitly saw.
ViT, CLIP and SAM look like three unrelated papers. Read them together and the same move shows up in each.
Every one of them replaces a fixed output vocabulary with an open one. A classifier's list of 1000 classes is fixed at training time. CLIP's list of candidate prompts is whatever you type at query time. A segmenter trained on a fixed category set is fixed. SAM's target is wherever you click.
ViT does the same thing to the input side. A convolution assumes a fixed spatial layout. A sequence of patch tokens assumes almost nothing, which is what lets the same encoder later swallow text tokens sitting next to image tokens.
That is why these three compose into a VLM so easily. They were all already speaking in things you can point at, and not in things someone enumerated in advance.
Vision transformers don't read pixels one by one. They cut the image into square patches and treat each patch as one token.
Try it Cut an image into patches the way a vision transformer does
import numpy as np
img = np.arange(64, dtype=float).reshape(8, 8) # a tiny 8x8 "image"
p = 4 # patch size
grid = img.reshape(8 // p, p, 8 // p, p).swapaxes(1, 2)
patches = grid.reshape(-1, p * p)
print("patches:", patches.shape, "(4 tokens of 16 numbers each)")
print("first patch, flattened:", patches[0].astype(int))
print("a 224x224 image at patch 16 gives", (224 // 16) ** 2, "tokens")
0 1 2 3, then 8 9 10 11, and so on. The same arithmetic on a 224x224 image with 16x16 patches gives 196 tokens, which is the sequence length the transformer sees.02Timeline
Classification, detection and segmentation each needed their own architecture and labeled dataset, and every new domain, from medical imaging to defect inspection, meant fresh labels and close to fresh training.
ViT (2020) showed attention alone works on images given enough data and compute, CLIP (2021) produced zero-shot embeddings from hundreds of millions of noisy web image-caption pairs, and SAM (2023) segmented anything from a point, box or rough mask.
"Encoder + projector + LLM" became the default way to add vision to a language model, and vision now follows the pretrain-once, adapt-cheaply playbook that reshaped NLP.
03Architecture: how the data flows
Start with ViT, since everything downstream in this page either uses it directly (CLIP's image encoder is usually a ViT) or borrows its "cut into patches, treat as tokens" trick.
CLIP wraps two encoders, typically a ViT (or a CNN, in the original paper's smaller variants) for images and a transformer for text. It trains them jointly so their outputs land in the same shared embedding space:
SAM's architecture follows the same "encode once, prompt cheaply" shape. Three pieces:
- A heavy image encoder runs once per image and produces a rich image embedding.
- A lightweight prompt encoder turns a click, box, or rough mask into a prompt embedding.
- A small, fast mask decoder combines the two and outputs a segmentation mask.
Only the last two run again when the user gives a new prompt. That is cheap enough for interactive speed, because the expensive image encoder never re-runs.
04The equations
ViT patch embedding
- x_p^i the $i$-th image patch, flattened from a $P \times P \times 3$ block of pixels into a single vector of length $P^2 \cdot 3$. For $P=16$, that is a 768-length vector per patch.
- E a single learned linear projection matrix, shared across every patch, mapping that 768-length vector into the model's working dimension $d_{model}$. This is the entire "convolution" ViT uses: one linear layer, applied identically to every non-overlapping patch.
- x_class a learnable embedding prepended to the sequence, unrelated to any patch. Its output after the full encoder is what the classification head reads.
- E_pos learned position embeddings, one per sequence position, added so the model can tell patch 1 from patch 197. Without them the sequence is just an unordered set of patches.
The result, $z_0$, is a (197, d_model) sequence: exactly the shape a transformer encoder expects. Everything from here is identical to page 08: multi-head self-attention, residual connections, layer norm, feed-forward blocks, stacked N times.
CLIP's contrastive loss
Stated at the same level as the softmax you've already seen in 01 and 08. This is InfoNCE, the same contrastive-loss family used elsewhere in self-supervised learning, just applied across two encoders instead of two augmented views of one input:
- sim(I,T) cosine similarity between an image embedding and a text embedding.
- τ a learned temperature that sharpens or softens the softmax distribution over similarities.
- N the batch size. Every image in the batch is contrasted against every caption in that same batch, not just its own match.
- symmetric loss this is computed once treating images as the anchor (image-to-text) and once treating captions as the anchor (text-to-image); CLIP trains on the average of both directions.
Bigger batches make this loss harder and more informative, because each image has more negative captions to be distinguished from, which is a large part of why CLIP-style training is done at very large batch sizes.
05Why this architecture
The obvious alternative to ViT was "just keep using a CNN, it already works." Page 05's closing point is exactly why that alternative does not win by default.
Convolution bakes in locality and translation equivariance as hard architectural constraints. A kernel that detects an edge in the top-left corner detects the same edge anywhere else in the image, for free, because it is the same weights sliding across the whole input.
ViT has none of that built in. A patch in the top-left and a patch in the bottom-right are just two tokens. The model has to learn from data that nearby patches tend to be related, because the architecture doesn't hand it over.
That is a real cost at small data scale. On ImageNet-1k alone, a ViT trained from scratch underperforms a comparably-sized ResNet. The convolutional inductive bias acts as a strong, free regularizer, and the ViT has not seen enough examples to learn an equivalent notion of locality on its own.
The ViT paper's finding is the opposite once you pretrain on hundreds of millions of images. At that scale, the inductive bias that helped a CNN generalize from little data becomes a constraint. It stops the CNN from finding patterns a less-constrained model can discover.
ViT catches up and then passes CNNs specifically as pretraining data grows. The crossover is a data/compute-scale story, and "ViT is just better" misses it.
CLIP's obvious alternative was "keep training on ImageNet-style labeled datasets, just make them bigger." That does not scale the way natural language supervision does.
A fixed label set caps out twice over: at however many classes someone was willing to define, and however many images someone was willing to hand-annotate. ImageNet's 1000 classes took enormous manual effort to label.
Captions are different: they are a byproduct of how the web already works. Alt text, photo captions, and image descriptions exist in effectively unlimited free-form quantity, already paired with images, with nobody doing extra annotation work for this specific purpose.
This is the bet behind CLIP: trade clean, expensive, fixed-vocabulary supervision for messy, free, open-vocabulary supervision, and let scale make up the difference.
06Tradeoffs and what breaks
| CNN | ViT | CLIP-style contrastive | |
|---|---|---|---|
| Inductive bias | Strong — locality, translation equivariance built in | None — must be learned from data | Inherits its encoder's bias (often a ViT) |
| Data efficiency (small data) | Better — bias regularizes with little data | Worse — needs large-scale pretraining to catch up | N/A — designed for web-scale data by default |
| Zero-shot capability | None — needs a task-specific head and labels | None on its own — still needs a labeled head | Native — compare embeddings to text prompts, no fine-tuning |
| Compute at scale | Efficient per-FLOP at moderate size | Scales cleanly with more data/compute | Expensive to pretrain — needs huge batches and web-scale pairs |
Failure modes
- CLIP's zero-shot accuracy is sensitive to prompt phrasing. "dog" versus "a photo of a dog" versus "a photo of a dog, a type of animal" can shift accuracy meaningfully on the same image set. The text encoder was trained on caption-like sentences, not bare nouns. Production zero-shot pipelines use "prompt ensembling", averaging embeddings over several prompt templates, specifically to reduce this sensitivity.
- SAM segments the wrong instance when a prompt is ambiguous. A single click on a person wearing a striped shirt could plausibly mean "the shirt," "the person," or "the stripe pattern." SAM is explicitly designed to return multiple candidate masks with confidence scores in exactly this situation, rather than silently guessing one.
- VLMs hallucinate visual details. Encoder + projector + LLM systems describe an object, color, or count that is not present in the image. The LLM component is a strong, fluent language generator. It can produce a plausible-sounding sentence even when the visual evidence feeding it is thin, ambiguous, or was compressed too aggressively by the projector.
"ViT beats CNNs, full stop, so CNNs are obsolete." Not what the evidence shows. ViT wins when pretraining data and compute are large; at small-to-moderate scale, well-tuned CNNs remain competitive and often more data-efficient, precisely because of the inductive bias ViT lacks.
Plenty of production systems in 2026 still use CNN or hybrid CNN-transformer backbones where labeled data is limited or latency is tight. "Attention replaced convolution" is a scale-dependent claim and not a universal one.
07What gets built on top: grounding, video, 3D, agents
Open-vocabulary and grounding detection extend CLIP-style alignment. The task moves from "classify the whole image" to "localize the region described by this free-text phrase."
Instead of training a detector on a fixed list of categories (person, car, dog), these models align region-level image features with text embeddings. So a query like "the red backpack on the left" can be matched against candidate regions even if "backpack" never appeared as a training label. The model generalizes through the same shared embedding space CLIP builds, applied at the region level instead of the whole image.
Video understanding extends ViT's patch idea into time. A tubelet is a spatiotemporal patch. Instead of slicing a single frame into 16×16 squares, you slice a short clip into 16×16×T chunks spanning a few frames at once.
Flatten each tubelet, then feed the resulting sequence through a transformer with attention operating across both space and time. A single attention operation can now directly relate "this patch, three frames ago" to "this patch, right now". That is the same one-hop, order-independent relevance-scoring idea from page 08, extended with a temporal axis.
3D vision, at a conceptual level, has two dominant representations right now.
- NeRF (Neural Radiance Field). A learned implicit scene representation. A small neural network takes a 3D point and a viewing direction, and outputs color and density at that point. It is trained so that rendering rays through the network reproduces the training photos. Once trained, you can render viewpoints that were never photographed, because the network learned a continuous function over the whole volume rather than a lookup table of the input images.
- Gaussian splatting. Represents the same kind of scene explicitly: a large collection of 3D Gaussian "blobs" with position, size, color, and opacity. Rendering means projecting and compositing those Gaussians directly, with no neural network queried per pixel per ray.
The consequence is a clean tradeoff. Gaussian splatting renders dramatically faster than a NeRF at similar visual quality. It pays for that with a more literal, less compact scene representation.
Vision agents close the loop: perception (an encoder or VLM turns pixels into a description or set of detections), reasoning (an LLM decides what that perception implies and what to do next), and action (a tool call: click here, move the robot arm, crop and re-inspect this region, query a database).
This perception-reasoning-action loop is the same agent architecture covered in depth on page 11, with a vision model supplying the perception step instead of, or in addition to, text.
08Vision-language models: encoder + projector + LLM
Suppose you already have a strong image encoder (CLIP's) and a strong language model (any modern LLM). The cheapest way to combine them is not training a new model from scratch. It is building an adapter between two models that already work.
That adapter is a projector: usually a small MLP, or a handful of cross-attention layers. Its entire job is to map the image encoder's output features into the same embedding space where the LLM's token embeddings already live.
Once the image features look like token embeddings to the LLM, dimensionally and distributionally, you can concatenate them with the text prompt's token embeddings and let the LLM process both together. The LLM does not need to "know" that some of its input tokens originated from pixels.
The training recipe that makes this cheap has three parts:
- Image encoder: freeze it, or fine-tune lightly. It is already good at producing useful visual features from CLIP-style pretraining.
- LLM: freeze it, or fine-tune lightly. It is already good at reasoning and generating fluent text.
- Projector: train mostly this, on paired image-text examples. It is small, fast, and comparatively cheap to learn.
This is why so many VLMs appear in quick succession once a good open image encoder and a good open LLM both exist: the expensive parts get reused and only the connective tissue needs training.
This adapter-based, bolt-on-a-projector approach differs from a native multimodal architecture. There, images and text are first-class inputs from the start of pretraining: tokenized (or patch-embedded) and fed through a single model trained jointly on both modalities from the ground up, rather than stitching two independently-pretrained models together after the fact.
- Native multimodal can in principle learn deeper cross-modal interactions, since the model never has to work around features that were optimized for a different objective.
- The adapter approach is cheaper and faster to iterate on. It also lets you swap in a better image encoder or a better LLM independently, as each one improves.
Both patterns are in active production use. Which one a given system chose is exactly the kind of design-tradeoff question worth being able to explain in an interview, beyond recognizing it by name.
09Build this
Zero-shot classification sounds like magic until you write the ten lines that do it. Then it looks like a dot product, and its failure modes stop being surprising.
Take a frozen CLIP checkpoint and classify images against a label set nobody trained it on. Pick something narrow and personal: the plants on your windowsill, six shapes of pasta, the pieces on a chessboard. You define the classes by typing them, and no gradient ever runs.
- Load a pretrained CLIP image encoder and text encoder. Take nothing else from the library. You are writing the classifier head yourself.
- Write that head from the similarity term in section 04: encode the image, encode each candidate prompt, L2-normalise both, take cosine similarity, divide by $\tau$, softmax.
- Collect fifty or so images across five classes. Label them yourself, run the classifier, and build a confusion matrix.
- Now break it on the text side. Rerun the identical images three times, changing only the phrasing: bare nouns (
"cat"), the standard template ("a photo of a cat"), then something wordier ("a blurry phone snapshot of a cat"). Keep every confusion matrix. - Add a distractor class that no image belongs to. Watch how confidently the softmax hands out probability anyway.
- Finally, average the text embeddings across several templates before comparing, which is prompt ensembling in one line.
Where this runs in production
Visual search and content moderation pipelines lean on CLIP embeddings directly. Embed a large catalog of images once, offline. Store the embeddings in a vector index. At query time, embed either a text query or a reference image, then run nearest-neighbor search over that index.
The payoff is that two apparently different features collapse into one. "Find images similar to this one" and "find images matching this text description" become the same operation: a similarity search over the same embedding space.
Image-captioning and visual-QA products typically use the encoder + projector + LLM pattern from the previous section. A CLIP-style image encoder plus a small projector feeds an LLM that generates the caption, answer, or reasoning trace. The product team can then swap in incrementally better open-source encoders or LLMs as they ship, without retraining the whole system from scratch.
10The nuances a vision specialist needs
Everything above is the standard story. What follows separates having read the papers from having shipped the models.
What patchification throws away
Cutting an image into 16×16 patches is not a neutral operation. Three things get given up at that exact moment:
- Sub-patch detail. Everything inside a 16×16 block collapses into one vector. A ViT never sees structure finer than its patch size, and no later layer can recover what the projection discarded.
- Grid geometry. Position embeddings are learned. The model is told "you are token 37" and is never told "you sit two rows below token 23", so adjacency is something it has to infer.
- Scale. A plain ViT runs every layer at one resolution. A CNN gets a feature pyramid for free: fine detail early, coarse semantics late.
That third loss is the one that bites hardest in practice. Detection and segmentation both want multi-scale features, and a single-resolution token sequence does not provide them.
Swin Transformer is the standard answer. It computes attention inside local windows rather than globally, then shifts the window boundaries on alternating layers so information still crosses between them. Two things fall out of that design:
- Cost becomes linear in image size instead of quadratic, because each token only attends within its window.
- The network can build a hierarchy, merging patches between stages. That produces exactly the pyramid an FPN or U-Net head expects.
This is why hierarchical windowed designs, not plain ViTs, became the common backbone for dense prediction.
The original ViT result needed hundreds of millions of pretraining images to beat a ResNet, and that number gets quoted as if it were a property of the architecture, which it is not.
DeiT trained a competitive ViT on ImageNet-1k alone, on a single machine in under three days. The fix was not more data. It was a stronger augmentation recipe plus a distillation token: an extra learnable token whose job is to match a CNN teacher's prediction.
So the honest version of the claim is narrower. Naive ViT training is data-hungry. The inductive bias a CNN gets from its architecture can instead be injected through the training recipe.
How a VLM is wired
Section 08 described the projector as a small MLP or a handful of cross-attention layers. That choice is not cosmetic. It decides how many tokens the LLM has to read, and how much visual detail survives the trip.
Two families dominate:
- Projection connectors. LLaVA uses a single linear layer; LLaVA-1.5 uses a two-layer GELU MLP. Every patch token is kept and handed to the LLM. Maximum fidelity, maximum token count.
- Resampler connectors. Flamingo introduced the Perceiver Resampler: a fixed set of learned queries cross-attends to the visual grid and compresses it to a fixed budget, 64 tokens in Flamingo's case. BLIP-2's Q-Former is the same idea.
The trade is information fidelity against token efficiency. A resampler gives you a constant cost regardless of input size, which is what makes long video tractable. A projection connector keeps detail a resampler would have thrown away, which is what makes reading small text in a document possible.
Notice the second-order effect: every visual token occupies LLM context, so a high-resolution image through a projection connector can consume more context than the question being asked about it.
Open vocabulary: replacing the class ID with a sentence
A classical detector ends in a fixed-width classification layer. Eighty COCO classes means eighty output slots. Adding an eighty-first means retraining.
Open-vocabulary detectors delete that layer. Region features get scored against text embeddings instead of learned class IDs. The consequence is that the label set becomes a runtime argument and no longer a training-time constant.
- OWL-ViT keeps it minimal: take a ViT, drop the final token pooling, attach lightweight box heads directly to the output tokens, and align them with CLIP-style text embeddings.
- Grounding DINO fuses harder. Swin image features and BERT text features meet in a feature enhancer, text guides the object queries, and cross-attention decodes the boxes. Fusion happens at every layer rather than only at the scoring step.
The practical payoff is querying for things nobody defined as a class: "cracked windshield", "the red backpack on the left". You are no longer limited to a vocabulary someone froze at annotation time.
The other backbone: supervision-free features
CLIP learns from image-text pairs. DINOv2 learns from images alone, through self-distillation across views, on a curated corpus of roughly 142 million images.
What makes it useful is where its features land. Freeze the backbone entirely, train only a linear probe or a small convolutional head, and it stays competitive on classification, semantic segmentation and depth estimation. No labels touched the backbone.
DINOv2 is a different tool from CLIP. CLIP gives you a shared space with language, which is what you want for zero-shot and retrieval. DINOv2 gives you stronger dense, patch-level features, which is what you want when the task is spatial and your labelled data is thin.
Where this still breaks
The most useful failure result in this area is Eyes Wide Shut? (CVPR 2024), and its method is worth understanding because it isolates the problem cleanly.
The authors search for CLIP-blind pairs: two images CLIP embeds as near-identical despite an obvious visual difference. They then build the MMVP benchmark by asking simple questions about exactly those differences.
The result is uncomfortable: every model tested except GPT-4V and Gemini scored below the 25% random-guess level on those questions. Those two did better, but the authors still report a large gap to human performance. The models also produce fluent, confident explanations for their wrong answers.
If your VLM's eyes are a frozen CLIP encoder, then anything CLIP cannot distinguish is invisible to the LLM downstream. No amount of language-side scaling fixes it. The information was destroyed before the LLM was ever called.
This is the strongest argument for the architectural questions above being worth caring about: which encoder, which connector, how many tokens, at what resolution.
The open problems cluster where you would predict from that analysis: fine-grained localisation, counting, spatial relations such as left-versus-right and in-front-versus-behind, and dense high-resolution documents where the answer depends on small text.
11Where you meet this in the wild
The shift these models caused is less about accuracy than about who gets to define the task. The label set moved from training time to inference time.
Once the class list is a runtime argument, whole categories of project stop needing a training run at all. That is the change worth internalising, and it is what most of the uses below have in common.
- The label set is unstable, or you cannot enumerate it up front, which is the single strongest signal.
- You have very few labels. A frozen backbone plus a small head beats training from scratch on a thousand images.
- You need language in the loop: query by description, explain a decision, follow an instruction about an image.
- You are still exploring. Zero-shot gets you a baseline this afternoon instead of after an annotation contract.
- The task is fixed and narrow. Five defect classes on one production line does not need open vocabulary.
- Latency or hardware is binding. A small CNN on an embedded device will beat a VLM API call, and it works offline.
- You need high accuracy on a specialised domain and you have the labels. A fine-tuned specialist usually beats a generalist on its own narrow distribution.
- The decision is fine-grained or spatial. Counting, precise localisation and small-detail discrimination are exactly the MMVP weaknesses above.
12Interview questions
BeginnerHow does ViT turn an image into a transformer input?
Cut the image into fixed-size, non-overlapping patches (16×16 is the common choice — for a 224×224 image that's 196 patches). Flatten each patch into a vector and pass it through one shared learned linear projection to get a token embedding. Prepend a learnable [CLS] token, add learned position embeddings so the model knows patch order, and feed the resulting sequence into a standard transformer encoder — identical to the one used for text.
BeginnerWhat does "promptable" mean in Segment Anything?
Instead of training a model to segment one fixed set of categories, SAM is trained to take a prompt at inference time — a point, a box, or a rough mask — and produce a segmentation mask for whatever that prompt indicates. It generalizes to objects and domains never explicitly labeled in training, because the task it learned is "segment what's pointed at," not "recognize this specific category."
IntermediateWhy does CLIP enable zero-shot classification?
CLIP's contrastive pretraining puts image embeddings and text embeddings in the same space, trained so matching image-caption pairs are close and mismatched pairs are far apart. Classification then reduces to comparing one image embedding against a handful of candidate text embeddings (one per class, phrased as a prompt like "a photo of a <class>") and picking the closest — no gradient update, no labeled examples for the new classes, because the model never learned a fixed set of output classes in the first place. It learned a general-purpose similarity function between images and language.
IntermediateWhat's the "encoder + projector + LLM" pattern, and why is it so popular?
A pretrained image encoder (often CLIP's) produces visual features; a small projector — usually an MLP or a few cross-attention layers — maps those features into the same embedding space the LLM's own token embeddings live in; the LLM then processes the projected image features alongside text tokens and does the actual reasoning and generation. It's popular because it reuses two already-expensive, already-good pretrained models and trains only the comparatively cheap connective piece, instead of pretraining a new multimodal model from scratch.
DeepWhy does ViT need more data than a CNN to reach the same accuracy, and what changes at large scale?
Convolution hard-codes locality and translation equivariance into the architecture — a kernel that detects a pattern in one location detects the same pattern anywhere, for free, because the weights are shared across every spatial position. ViT has no such constraint; a patch in one corner and a patch in another are just two tokens the model must learn to relate from data. At small data scale, that missing inductive bias hurts — the CNN's built-in regularization wins. At large pretraining scale, the same constraint that helped with little data starts limiting what patterns a CNN can discover, while ViT, unconstrained, keeps improving with more data — which is why the ViT paper's headline result only shows up once pretraining data reaches into the hundreds of millions of images, not on ImageNet-1k alone.
DeepYour CLIP-based zero-shot classifier's accuracy swings noticeably depending on the exact prompt template you use. Why does this happen, and what do you do about it?
CLIP's text encoder was trained on naturalistic, caption-like sentences scraped from the web, not on bare category names — so a prompt like "a photo of a dog" sits closer, in the learned embedding space, to how dog-related captions actually looked during training than the single word "dog" does. Sparse or unusual phrasing produces a noisier, less representative text embedding, which shows up directly as noisier similarity scores. The standard fix is prompt ensembling: embed the same class through several different prompt templates ("a photo of a {}", "a blurry photo of a {}", "a close-up photo of a {}", etc.) and average the resulting embeddings before comparing — this smooths out the sensitivity to any one specific phrasing.
13Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.