Multimodal AI
This page pulls the earlier ones together: it shows how vision, language, audio, and video models get wired into one system that perceives, reasons, and generates across modalities.
This page introduces no new modality-specific math. It is the synthesis page. How pixels become tokens is 06. How attention works is 08. How diffusion and autoregressive generation work is 12. How speech gets encoded is 13.
Multimodal AI is the layer on top: how you combine models that already work well within one modality into a single system that reasons across all of them. Three decisions do almost all of the work. Fusion, meaning where in the network the modalities meet.
Shared embedding spaces built by contrastive learning, which is how you make different modalities' representations directly comparable in the first place. It also covers adapter-based versus native multimodal, which is whether you bolt a pretrained encoder onto an existing LLM or train everything jointly from raw tokens. One sentence for an interview: multimodal AI is mostly a fusion and systems problem and brings no new kind of math.
01Intuition
Every modality speaks a different native language — so the first problem multimodal AI solves is translation, and fusion comes after.
An image, as data, is a grid of pixel intensities. A sentence is a sequence of discrete token ids. There is no shared coordinate system connecting them. You can't take a dot product between a pixel grid and a token sequence and expect it to mean anything.
So before you can combine them, you have to translate both into a common space where distance and similarity are meaningful across modalities. An image of a dog and the sentence "a photo of a dog" should land near each other. An image of a cat should land far from that sentence. It shouldn't matter that one started as pixels and the other as text.
This is the self-supervised contrastive recipe from 03, generalized across modalities, in four steps:
- Take a large batch of matched pairs: image and caption, audio clip and transcript, video and description.
- Run each half of every pair through its own encoder.
- Project both into the same dimensional space.
- Train with a loss that pulls the embeddings of true pairs together while pushing every other combination in the batch apart.
CLIP (Radford et al., 2021) did exactly this for image-text, and it is the reference implementation almost everyone means by "contrastive multimodal alignment". 06 covers the full derivation and training details.
Swap in an audio encoder in place of the image encoder and you get audio-text alignment. Swap in a video encoder and you get video-text alignment. The loss stays the same, and only the encoder that produces each half of the pair changes.
Once that shared space exists, it becomes the common language every downstream multimodal system relies on. Retrieval becomes a nearest-neighbor search across modalities: find the image whose embedding is closest to this text query. Zero-shot classification becomes comparing an image embedding against a handful of candidate label embeddings and picking the closest one.
It also becomes the raw material that fusion operates on, which is the subject of the next section. You can't usefully combine two modalities' representations until you have first made those representations comparable.
Multimodal models put images and text into the same vector space, so that a matching pair lands close together. Here the vectors are hand-picked, two numbers each, to show the matching step.
Try it Match images to captions with cosine similarity
import numpy as np
# Hand-picked [furry, wheels] embeddings: three images, three captions
images = np.array([[0.9, 0.1], [0.1, 0.9], [0.5, 0.5]])
texts = np.array([[1.0, 0.0], [0.0, 1.0], [0.6, 0.4]])
names = ["a photo of a dog", "a photo of a car", "a photo of a robot dog"]
def norm(a): return a / np.linalg.norm(a, axis=1, keepdims=True)
sim = norm(images) @ norm(texts).T # cosine similarity
print(np.round(sim, 2))
print("best caption per image:", [names[i] for i in sim.argmax(axis=1)])
0.99, 0.99 and 0.98, and the off-diagonal entries are lower, such as 0.11 for the dog against the car. The third image, half furry and half wheeled, picks "robot dog" at 0.98, ahead of "dog" and "car" at 0.71 each. Real models learn these vectors from data, and here they were hand-picked.02Timeline
Through the 2010s vision, language and audio were separate fields, and multimodal systems glued independent unimodal models together with hand-written rules for combining their scores, which was brittle and blind to interactions below the final prediction.
CLIP-style contrastive alignment (2021) trained an image encoder and a text encoder into a shared space with one loss, and the encoder + projector + LLM recipe (2022–2023) bolted a pretrained vision encoder onto a pretrained LLM through a small trainable projector at a fraction of the cost.
Today adapter-based VLMs inherit improvements from better encoders and LLMs, natively multimodal models train jointly across modalities at much higher cost, and cross-modal generation and multimodal agents push toward systems that perceive, reason and act in one loop.
03Architecture: where modalities meet
Once you have per-modality representations, you have to decide where in the network they meet. This is the fusion taxonomy, and it is the most consequential architectural decision in a multimodal system. It trades off how much cross-modal interaction the model can learn against how expensive and inflexible the resulting system becomes.
| Early fusion | Late fusion | Intermediate / cross-attention fusion | |
|---|---|---|---|
| Where modalities meet | At the input, before any real per-modality processing | At the very end — only final predictions or scores are combined | Partway through — separate encoders, then cross-attention layers mid-network |
| Typical example | Concatenate raw or near-raw pixel and audio features into one input sequence | Average an image classifier's and a text classifier's output scores | A VLM where text tokens cross-attend into image patch tokens inside the transformer |
The tradeoff moves in a straight line with how early you merge.
Early fusion can, in principle, learn the finest-grained cross-modal correlations. A single joint model sees raw signal from both modalities from layer one, so nothing about their interaction is ever forced through a bottleneck.
But it is expensive, because you are training the whole thing jointly and can't just reuse a pretrained unimodal encoder. And it is inflexible: add a third modality, or swap in a better vision encoder, and you are often retraining from scratch.
Late fusion is the opposite. It is cheap and modular: you plug in any pretrained image model and any pretrained text model, then train a small combination layer on top. But it misses fine-grained interaction entirely, because by the time the two modalities' outputs are combined, all the low-level detail that made them useful together is already gone.
Suppose answering a question depends on relating one specific word to one specific pixel region.
Late fusion structurally cannot do it, because the word and the pixel never appear in the same computation.
Intermediate fusion, specifically cross-attention fusion, is the compromise that won for most systems built since 2022. Here is what it looks like as data flows through one block:
The recipe, generalized beyond vision
06 covers the dominant VLM recipe in depth. A pretrained vision encoder produces a sequence of patch embeddings. A small trainable projector maps those into the same dimensional space the LLM's token embeddings live in. The LLM, usually pretrained and often frozen or only lightly fine-tuned, then treats the projected image tokens exactly like any other tokens in its context, attending over them with the self-attention machinery it already has.
The same recipe generalizes directly to every other modality. An audio encoder plus a projector feeding an LLM gives you a speech-language model, using the identical architecture and identical training logic. Only the pretrained encoder at the bottom changes.
Training usually happens in two stages either way: an alignment stage that teaches the projector to map the new modality into the LLM's space, then instruction tuning on multimodal task data. This is a training-recipe detail and adds no new architecture.
Native multimodal tokenization: the alternative to bolt-on adapters
The adapter recipe above always has a seam. A pretrained encoder is trained on its own objective, then adapted afterward to work inside an LLM it never saw during its own pretraining.
Native multimodal tokenization removes that seam. Every modality is converted into tokens from the start, where the encoder-then-project design converts later. Text stays as subword tokens, images are chopped into patch tokens or discrete codes from a learned image tokenizer, and audio is chopped into frame tokens. All of it is interleaved into one shared sequence that a single transformer is trained on jointly, with no separate modality-specific pretraining stage bolted on afterward.
The real tradeoff cuts both ways:
- The upside. A natively multimodal model can, in principle, learn richer cross-modal reasoning, because nothing about how it represents an image was ever optimized for a different, unimodal objective first. Every parameter was trained with the other modalities in the loop from day one.
- The cost. You give up reusing a great pretrained unimodal encoder and improving it independently. You are training the whole stack jointly, which is dramatically more expensive and harder to get right, and a bad tokenization choice for one modality can hurt every other modality sharing the same model.
Cross-modal generation and multimodal reasoning
Two more categories are worth naming explicitly, since interviewers use these exact terms.
Cross-modal generation is any system that generates in one modality conditioned on another: text-to-image, text-to-audio, text-to-video, image-to-video, speech-to-speech. None of these needs new generative machinery, since they reuse the diffusion and autoregressive techniques from 12 (see The Illustrated Stable Diffusion for the mechanics of the text-to-image case specifically), just conditioned on an embedding from a different modality's encoder in place of, or in addition to, a text prompt.
Multimodal chain-of-thought, or visual reasoning, is a different thing from generation. It is about how a model reasons over an image or video it is already looking at. The naive approach captions the image once, in text, then reasons purely over that caption, so every downstream reasoning step is blind to the actual image and can only be as good as that first, one-shot caption.
Real multimodal reasoning re-attends to the image, or to specific regions of it, or specific moments in a video, at each reasoning step. This is the same cross-attention pattern as the diagram above, invoked repeatedly during reasoning and not once at the input.
04The equations
Cross-attention is not a new mechanism. It is the identical formula from 08, with one detail changed: where each matrix comes from.
- Q_text a linear projection of the querying modality's tokens, here text. Same role as Q in 08's self-attention formula.
- K_image, V_image linear projections of the modality being attended over, here image patch tokens. Same role as K and V in 08, just sourced from a different sequence than Q.
- everything else the dot product, the $\sqrt{d_k}$ scaling, the row-wise softmax, the weighted sum over V. All identical to self-attention in every detail. The only thing that changed is that Q and K/V no longer come from the same sequence.
Cross-attention fusion is architecturally cheap to add because it is a new wiring pattern and needs no new operation. Any transformer block that already has self-attention can grow a second attention sub-layer where Q still comes from the primary modality's residual stream, but K and V are swapped in from a different modality's encoder output.
This mechanism is all that the diagram above shows, and it swaps in for any pair of modalities (text querying image, text querying audio, video querying text) without changing the math at all.
The contrastive alignment loss that builds the shared embedding space in the first place (03, 06) is the same InfoNCE-style objective regardless of which two modalities you are aligning. Maximize the similarity $\text{sim}(a_i,b_i)/\tau$ of a true pair relative to every mismatched pair in the batch, symmetrically in both directions.
Swap in an audio encoder and a text encoder in place of an image encoder and a text encoder, and the loss itself doesn't change. See 06 for the full derivation, which would read the same here with different variable names.
05Why this architecture
Why did intermediate fusion, specifically cross-attention, become the default over pure early or pure late fusion for most systems? Cost and reuse.
Cross-attention fusion lets you start from two or more already-excellent pretrained unimodal models: a vision encoder trained on huge amounts of image data, an LLM trained on huge amounts of text. Then you add a comparatively small number of new parameters, the cross-attention layers or the projector, to connect them.
You get most of the benefit of early fusion's fine-grained interaction, since a text token can attend directly to a specific image patch mid-network, before any prediction is made. And you pay closer to late fusion's training cost, since you are mostly fine-tuning a small bridge and not training two giant encoders jointly from scratch.
This is the accuracy-versus-cost sweet spot. Early fusion's ceiling is higher in principle, but almost nobody can afford to pay for it from scratch for every new modality combination. Late fusion is cheap but structurally incapable of fine-grained cross-modal reasoning. Cross-attention fusion is the point on that curve where most production systems land.
Native multimodal tokenization is the closest thing to deliberately paying for early fusion's ceiling. It is rarer for that reason, and mostly pursued by teams with the budget to train large joint models from scratch and not adapt existing ones.
06Tradeoffs, and what breaks
| Early fusion | Late fusion | Intermediate / cross-attention fusion | |
|---|---|---|---|
| Interaction quality | Highest possible — learned jointly from raw signal | None below the final prediction | Good — direct token-to-token attention mid-network |
| Compute cost | Highest — joint training from scratch | Lowest — reuse frozen pretrained encoders | Moderate — mostly training a small bridge |
| Modularity / reuse | Low — hard to swap either modality's model | High — swap either side freely | High — swap encoders, retrain only the bridge |
Failure modes
- Grounding failures. A model describes something plausible-sounding but not present in the input. Language priors learned from huge amounts of text pretraining are strong, and a fluent, wrong description often reads just as convincingly as a correct one. It is the most common real complaint about VLMs in production.
- Temporal inconsistency. Video generation and video understanding both have to maintain consistency across time that image models never had to deal with. An object's identity, appearance, and position have to stay coherent across dozens or hundreds of frames. Both generative and discriminative video-language models still struggle with this over long clips.
- Evaluation is hard. Unlike a text benchmark where you can check an answer string, checking a multimodal answer means checking several harder-to-automate things at once: perception (did it correctly perceive what's in the input), grounding (does a specific claim correspond to a specific part of the image or video, not just sound plausible), reasoning (given correct perception, is the inference valid), generation quality (for cross-modal generation tasks), and temporal/spatial consistency (for video specifically). A single benchmark number rarely captures all five, which is why real multimodal evaluation usually needs several benchmarks plus a human or rubric-based check for grounding specifically.
"The model looked at the image, described it, then reasoned about the description — so it did visual reasoning." Captioning once and then reasoning over that caption is not the same thing as multimodal reasoning.
Every error in that first, one-shot caption propagates uncorrected through every later step, because the model never looks back at the actual image. Real multimodal chain-of-thought re-attends to the image, or a specific region or moment in it, at each reasoning step. Not just once, up front, before reasoning goes blind.
Both lanes from the timeline are active and shipping today. Adapter-based VLMs remain more common because they're cheaper to build and let you reuse best-in-class unimodal encoders as those keep improving independently.
Natively multimodal models trained jointly from the start are the harder, more expensive bet, made by teams aiming for tighter cross-modal reasoning than an adapter can provide. Neither has definitively displaced the other. Treat any claim that one has fully replaced the other as an overclaim.
07Build this
Cross-attention is the same formula as section 04, wired differently. What the formula cannot show you is a single word picking out the patch it belongs to.
Train a tiny image-text matching model on shapes you draw yourself, then plot what each caption word attends to. The task is deliberately trivial: two coloured shapes per image, a template caption naming both. Because you drew the shapes, you can look at one attention map and say immediately whether the alignment is real or wishful.
- Draw the dataset with NumPy. Each 64×64 image gets two solid shapes of different colours in two random quadrants. The caption is a template string:
"a red square and a blue circle". A few thousand images is plenty. - Cut each image into a 4×4 grid of 16 patches and push each patch through one linear layer. On the text side, an embedding table plus one small transformer block. Keep both towers deliberately small.
- Write the cross-attention layer by hand from the equation in section 04. Q from the text tokens, K and V from the patch tokens,
softmax(QKᵀ/√d_k)V. Avoidnn.MultiheadAttention. - Pool the fused text tokens and score the pair with a linear head. Train it to separate true image-caption pairs from pairs made by shuffling the captions within a batch. A few minutes on CPU.
- Keep the attention weights from that layer for one pair. For each caption word, reshape its 16 weights into a 4×4 grid and draw it over the image.
- Now break it deliberately by re-initialising the text tower randomly, freezing it, retraining everything else on the same data, and plot the same words again.
red brightens over the quadrant holding the red shape, and circle brightens over the circle. No patch was ever labelled. The only supervision was whether a whole caption matched a whole image, and the grounding fell out of that.Where this runs in production
Two applications are worth calling out first, since they show up constantly in real systems and in interviews.
Multimodal retrieval takes the shared embedding space from the intuition section and uses it directly. A query in any modality (text, an image, a voice clip) gets embedded and compared against a database of embeddings from any other modality.
So "find images similar to this description" and "find documents containing a chart like this one" are the same nearest-neighbor operation with different inputs, and it is the retrieval half of 11's RAG pipeline with the embedding space extended past text.
Multimodal agents extend 11's tool-use loop with a perception step. An agent that can look at a screenshot, decide in text what action to take, call a tool (click a button, run OCR on a region, zoom into part of an image), then look again at the result before deciding the next step.
The loop is identical to a text-only agent's plan-act-observe cycle. The only difference is that "observe" can now mean look as well as read a tool's text output.
A real document (an invoice, a scanned contract, a medical form) mixes three kinds of information that no single pretrained model handles well alone:
- Text content, which OCR can pull out.
- Spatial layout: which text belongs to which table cell, which field a value answers.
- Purely visual content that text extraction throws away entirely: a logo, a handwritten signature, a checkbox, a chart with no underlying text at all.
A production document AI pipeline typically runs OCR to extract raw text, a layout model to recover structure (reading order, table boundaries, key-value regions), and a vision-language model that reasons over the raw document image directly.
Combining all three is what makes it work, so it can answer "what's the total on this invoice" correctly even when the total sits inside a visually formatted table that OCR text alone would jumble into the wrong reading order. Or answer "is this contract signed", a question pure text extraction can't answer at all, because a signature isn't text.
08Interview questions
BeginnerWhat's the difference between early, late, and intermediate fusion?
Early fusion merges raw or near-raw modality inputs before any real per-modality processing — expensive and inflexible but can learn the finest-grained interactions. Late fusion processes each modality fully separately and combines only the final predictions or scores — cheap and modular but misses fine-grained interaction entirely. Intermediate fusion combines partway through, typically with cross-attention layers letting one modality's tokens attend to another's mid-network — most of the interaction benefit of early fusion without forcing everything through one early merge.
BeginnerWhat is a shared embedding space, and why do you need contrastive learning to build one?
An image (a pixel grid) and a sentence (a token sequence) live in completely different, incomparable representations by default — there's no meaningful way to compute a similarity between them directly. A shared embedding space is a learned space where representations from different modalities are directly comparable by something as simple as a dot product. Contrastive learning builds it by encoding matched pairs (image and caption, audio and transcript) separately, projecting both into the same space, and training with a loss that pulls true pairs together and pushes mismatched pairs apart — that's CLIP's recipe, generalized to any two modalities.
IntermediateEarly vs. late vs. intermediate fusion — when would you actually choose each?
Late fusion for a quick system combining two off-the-shelf models where you don't need fine-grained cross-modal interaction — e.g. separately scoring image quality and text relevance and combining the scores. Intermediate/cross-attention fusion for most production VLM and audio-language work, since it reuses strong pretrained unimodal encoders cheaply while still allowing real token-level interaction. Early fusion, or its modern form — native multimodal tokenization — only when you have the budget to train jointly from scratch and specifically need the highest possible ceiling on cross-modal reasoning, and you're willing to give up easy reuse and independent upgrades of each modality's encoder.
IntermediateWalk me through the encoder + projector + LLM recipe, and how does it generalize past vision?
A pretrained modality encoder (vision, audio, whatever) produces a sequence of embeddings. A small trainable projector maps those embeddings into the same dimensional space the LLM's token embeddings already live in. The LLM then treats the projected tokens exactly like any other tokens in its context, attending over them with its existing self-attention machinery — no architectural change to the LLM itself. This generalizes directly: swap the vision encoder for an audio encoder and a projector, and you have a speech-language model using the same recipe. Only the encoder at the bottom changes.
DeepHow would you evaluate whether a VLM's answer is actually grounded in the image, rather than just plausible-sounding?
You can't trust a fluent answer at face value, because language priors alone can produce a convincing-sounding wrong answer. Concretely: construct probes where swapping the actual image content should change the correct answer, and check whether the model's answer actually changes accordingly — if it doesn't, the model is answering from a language prior, not from the image. Use benchmarks specifically built so that guessing from text priors alone gives close to random accuracy, forcing real visual evidence to matter. And separate the evaluation into distinct axes — perception (did it see the right thing), grounding (does the specific claim map to a specific part of the image), and reasoning (is the inference from correct perception valid) — since a model can fail at any one of these while looking fine on an aggregate score.
DeepWhy does context length become a bigger systems problem for multimodal models than for pure text models?
One image contributes dozens to hundreds of patch tokens; one video contributes many frames, each contributing its own patch tokens, so a few seconds of video can dwarf the token count of a long text document. Because attention cost is quadratic in sequence length (08), that token count exploding compounds the cost far faster than the equivalent scene described in a paragraph of text would. This is why the token count an image or video gets compressed into is a real, load-bearing systems decision in multimodal models — how many tokens a vision encoder emits per image, or how aggressively a video encoder subsamples frames — in a way it simply isn't for text-only models, and it's directly downstream of everything 23 covers on serving cost.
09Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.