←Home KnowML
Sequence, Attention & LLMsChapter 09

NLP Evolution

Bag-of-words to TF-IDF to Word2Vec to ELMo to BERT — the same question asked four times, with a slightly better answer each time: how do you turn a word into a number that means something?

18 min read Assumes: none, but pairs with sequence modeling (07) and attention (08)
Start reading
TL;DR

Every idea on this page is the same experiment run four times, with a better instrument each time: represent a word by the company it keeps. Bag-of-words and TF-IDF count words with zero notion of meaning. "Good" and "great" are as unrelated as "good" and "aardvark." Word2Vec and GloVe fix that.

They learn a dense vector per word from co-occurrence statistics, so meaning becomes geometry. But every word still gets exactly one vector, no matter which of its senses is in play. ELMo breaks that: a deep bidirectional LSTM produces a vector that depends on the whole sentence, so "bank" the riverbank and "bank" the financial institution finally get different vectors.

BERT and GPT (07/08) go one step further: they pretrain the entire network on that objective, where earlier methods pretrained only an embedding table, and they use a transformer in place of an LSTM. T5 then reframes every NLP task as reading one string and writing another.

One sentence for an interview: representation learning in NLP is a 60-year march from hand-counted, context-free, symbolic word representations toward fully learned, context-dependent, continuous ones. And subword tokenization is the one piece from this entire lineage still running, unchanged in spirit, inside every LLM you'll use today.

01Intuition

"You shall know a word by the company it keeps." — J. R. Firth, 1957

That one sentence, written decades before any of these models existed, is the idea behind every technique on this page.

The disagreement between generations was never about whether a word's meaning comes from its context. Everyone from TF-IDF to GPT agrees on that. The disagreement was about two narrower questions: how much of that context to use, and how to turn "context" into a number.

Bag-of-words and TF-IDF use almost none of it. They look at which words co-occur in a whole document, and even then only as raw or re-weighted counts, so "context" means little more than which bucket a word fell into. The vector for "good" cannot sit close to the vector for "great," because there are no vectors at all, only counts.

Why the byproduct is the thing you keep

Word2Vec's insight was to turn "predict the neighbors of a word" into a training objective.

Train a shallow network to predict the words that tend to surround "good." Separately, train it to predict the words that surround "great." Those two prediction tasks end up wanting very similar internal representations, because the two words do appear in similar contexts.

The prediction itself is not the goal. The goal is the vector that falls out of solving the prediction task.

GloVe reaches a similar place by a different road. Instead of predicting anything, it directly factorizes the global count matrix of how often every word co-occurs with every other word.

Both of those still hand you one fixed vector per word, decided once at training time and frozen forever after.

That's the crack ELMo exploits. "Bank" the riverbank and "bank" the financial institution are spelled identically but mean nothing alike. A single static vector has to awkwardly average both senses together.

ELMo's fix is to change the question. Stop asking "what's the vector for this word?" Start asking "what's the vector for this word in this sentence?" Run the whole sentence through a language model, then read the word's representation off the model's internal hidden state at that position.

BERT and GPT push that same idea to its natural conclusion. If the representation should depend on the whole sentence, don't bother with a separate embedding-lookup stage at all. Pretrain one big network, end to end, and let every layer contribute to what "bank" means right here.

An early way to turn text into numbers was to count the words. Try it on two sentences that mean opposite things.

Try it Count words in two sentences that mean different things
from collections import Counter

a = "dog bites man".split()
b = "man bites dog".split()

print("bag of words a:", dict(Counter(a)))
print("bag of words b:", dict(Counter(b)))
print("same counts:", Counter(a) == Counter(b))
print("same sentence:", a == b)
The two sentences produce identical counts. dog bites man and man bites dog each contain one dog, one bites and one man, so a bag of words can't tell them apart. Most of what came next in NLP is a way to keep word order and meaning.

02Timeline

Before

Through the 2000s, text became numbers by counting: bag-of-words, TF-IDF and n-gram models, which had no notion that "good" and "great" are related.

→
Innovation

Word2Vec (Mikolov et al., 2013) and GloVe (2014) learned a dense vector per word from co-occurrence, so similar words landed near each other, and ELMo (2018) made the vector depend on the whole sentence.

→
After

BERT and GPT (2018; see 08 and 10) moved to transformers pretrained at scale, T5 (2019) cast every task as text in, text out, and instruction tuning and RLHF turned next-token predictors into models that follow requests.

03Architecture: how each generation works

Bag-of-words and TF-IDF: counting, dressed up

Bag-of-words builds a vocabulary of every distinct word across a corpus. Each document then becomes a vector of length |vocabulary|, where each entry is how many times that word appears in the document.

Word order is thrown away entirely. "Dog bites man" and "man bites dog" produce the identical vector.

TF-IDF keeps the same vector shape but replaces raw counts with a weighted score. Multiply how often a term appears in this document by how rare that term is across all documents. Common words like "the" get downweighted toward zero, while a document-specific term gets to dominate the vector.

Word2Vec: predicting your way to meaning

Word2Vec comes in two mirror-image forms:

  • Skip-gram takes one center word and tries to predict each of its surrounding context words.
  • CBOW (continuous bag-of-words) does the reverse. Average the context word vectors, then try to predict the center word.

Both are, mechanically, a two-layer network. A one-hot input selects a row from an embedding matrix, and that row is the word vector being learned. A linear layer projects it back up to vocabulary size. Softmax turns that into a probability distribution over every word in the vocabulary.

The two forms trade off against each other. Skip-gram tends to do better on rare words, since each occurrence generates multiple training pairs, one per context word. CBOW trains faster, because it smooths several context words into one input.

Computing a full softmax over a 100,000+ word vocabulary for every single training step is expensive. So in practice Word2Vec uses negative sampling. Instead of scoring every word in the vocabulary, training becomes a binary classification task: is this a real (center, context) pair, or one of k randomly sampled "negative" pairs? Dramatically cheaper, and it works about as well.

GloVe: skip the prediction, factorize the counts directly

GloVe doesn't predict anything. It first builds a global word-word co-occurrence matrix: for every pair of words, how often do they appear near each other across the entire corpus?

Then it learns vectors whose dot product approximates the log of that co-occurrence count, via a weighted least-squares objective that downweights extremely frequent (and extremely rare) pairs.

The contrast with Word2Vec is clean. Word2Vec is local and predictive, working one context window at a time. GloVe is global and count-based, factorizing the whole corpus's statistics at once.

In practice the two produce vector spaces with very similar properties. That suggests the co-occurrence statistics carry the semantic signal, and the particular training method matters less.

Subword tokenization: the fix for words you've never seen

Every model above has to first decide what counts as "a word." Whole-word vocabularies have a hard failure mode here: any word not seen during training, whether a typo, a new slang term, or a name, has no vector at all.

Subword tokenization sidesteps this: you build a vocabulary of word pieces in place of whole words, and any string can be represented as some sequence of pieces, even if the whole word is unseen.

  • BPE (Byte-Pair Encoding) starts from individual characters as the vocabulary. It then repeatedly finds the single most frequent pair of adjacent symbols in the training corpus and merges it into one new symbol, repeating until it reaches a target vocabulary size (commonly 30,000–50,000 tokens). Common words end up as one token; rare words fall back to a handful of subword pieces.
  • WordPiece (used by BERT) runs the same merge-based procedure, but it merges whichever pair most increases the training corpus's likelihood under the current tokenization, where BPE merges the most frequent pair. The objective differs a little and the overall shape is the same.
  • SentencePiece treats input as a raw stream of Unicode characters, including whitespace itself as an ordinary symbol. So it needs no language-specific pre-tokenization step before BPE/WordPiece can run, which is essential for languages like Japanese or Chinese that don't delimit words with spaces at all. It's the tokenizer behind T5, ALBERT, and many modern LLMs.

ELMo: the same word, a different vector every time

ELMo trains a two-layer bidirectional LSTM language model. One LSTM reads the sentence left-to-right, predicting the next word. A separate one reads right-to-left, predicting the previous word. The two are trained jointly.

A word's ELMo vector is a learned, task-specific weighted combination of three things, all evaluated on this specific sentence:

  • This is the raw input embedding.
  • The first LSTM layer's hidden state at that position.
  • The second layer's hidden state.

Lower layers tend to carry more syntactic information, higher layers more semantic and contextual information. Letting each downstream task learn its own weighting over the three lets it pick whichever blend helps most.

One more detail matters: ELMo's input representation is built from a character-level CNN and has no word lookup table. So unlike Word2Vec or GloVe, it can produce a reasonable representation for a word it never saw during training, just from its spelling.

Same word, two very different pipelines
Static — Word2Vec / GloVe "...money at the bank" "...sat by the river bank" bank [0.2, -0.4, 0.8…] one fixed vector, always — both senses blurred together Contextual — ELMo / BERT "...money at the bank" "...sat by the river bank" biLSTM / Transformer bank₁ [0.9, 0.1, -0.2…] bank₂ [-0.3, 0.7, 0.4…] two different vectors — one per sense, resolved from context
Illustrative vector values, not real model output. Word2Vec/GloVe commit to one vector for "bank" at training time and never revisit it. ELMo and BERT re-derive the vector for every occurrence, using the surrounding sentence — the exact same input word, two different outputs.

BERT, GPT, and T5: pretrain the whole network, then reuse it

BERT and GPT are covered in full architectural depth on 08 and 10. The short version for this lineage:

  • Both replace ELMo's biLSTM with a transformer.
  • Instead of just producing a reusable embedding table, they pretrain the entire stack, dozens of layers, on a self-supervised objective. Masked language modeling for BERT, next-token prediction for GPT.
  • Then they fine-tune, or later instruction-tune, the whole thing per task.

T5 (Raffel et al., 2019) makes one further structural change. Instead of a different output head per task, it reframes every NLP task as reading one text string and writing another, using a short task prefix to tell the model what to do: "translate English to German: ...", "summarize: ...", "cola sentence: ..." for grammatical-acceptability classification.

T5 uses one encoder-decoder architecture and one training objective, predicting corrupted spans of text, so every downstream task becomes the same kind of example.

04The equations

TF-IDF

$$\text{tfidf}(t, d) = \text{tf}(t, d) \times \log\!\left(\frac{N}{\text{df}(t)}\right)$$
  • tf(t, d) how many times term $t$ appears in document $d$.
  • N total number of documents in the corpus.
  • df(t) document frequency: in how many documents does $t$ appear at all.
  • log(N / df(t)) is the inverse-document-frequency term. A word appearing in every document (df ≈ N) drives this toward $\log(1) = 0$, killing its weight. A word appearing in only a few documents keeps a large weight.

Skip-gram objective

$$\frac{1}{T}\sum_{t=1}^{T}\ \sum_{-c \le j \le c,\ j \ne 0} \log\, p(w_{t+j} \mid w_t)$$
  • T total number of words processed; c the context window size (how many neighbors on each side count as "context").
  • p(w_{t+j} | w_t) the probability the model assigns to the actual context word $w_{t+j}$, given center word $w_t$. The exact objective computes this with a softmax over the dot product of the two words' vectors. In practice it's approximated far more cheaply via negative sampling.
  • Maximizing this sum means: whatever vectors make the real (center, context) pairs in your corpus look likely are the vectors you keep.

The BPE merge rule

BPE has no loss function to minimize. It is a greedy, deterministic algorithm and not a training objective:

$$\begin{aligned}(a^*, b^*) &= \underset{(a,b)}{\text{arg max}}\ \ \text{count}(a, b) \\[3pt] &\Rightarrow\ \text{merge } (a^*, b^*) \text{ into one new symbol}\end{aligned}$$

At every iteration, count every pair of adjacent symbols across the whole corpus, still character-level at first, and merge whichever pair occurs most often into a single new symbol. Then repeat.

After enough iterations, common whole words have been merged all the way into single tokens, while rare words remain split into a handful of frequent subword pieces. This property lets a fixed-size vocabulary represent effectively unbounded text.

05Why this design

The obvious alternative at every step was "just use a bigger lookup table." It fails twice over. A bigger table of hand-built synonym rules doesn't scale past a few thousand words and can't be learned from data. A bigger one-hot vocabulary still gives you zero notion of similarity between "good" and "great," no matter how large you make it.

The real axis of progress across this whole page is a single question: how much does the representation have to be told, versus how much can it learn on its own from raw co-occurrence?

One-hot / BOWWord2Vec / GloVeELMoBERT / GPT
Captures semantic similarityNoYesYesYes
Vector depends on contextNoNo — one per word typeYesYes
Handles unseen wordsNoNo (fixed vocabulary)Partially (char-CNN input)Yes (subword tokenization)
What's pretrainedNothingAn embedding table onlyA 2-layer biLSTMThe entire network, dozens of layers
The pattern hiding in that last row

Read the bottom row left to right. Nothing pretrained, then an embedding table, then a 2-layer biLSTM, then the entire network.

At every step, more of the model gets pretrained, beyond the input layer, and that trend runs through the whole page.

It's the same story as the residual-stream idea in 04 and the attention idea in 08, wearing a different hat. Push more of the computation into something learned end-to-end from data, and the hand-designed stages keep disappearing: a fixed synonym list, then a fixed embedding table someone else trained.

06Complexity, failure modes, and what breaks

  • Polysemy is the fundamental limit of static embeddings. Word2Vec and GloVe are mathematically incapable of giving "bank" two different vectors. The training objective produces exactly one vector per word type, so it ends up as an uneasy average over every sense the word is ever used in. This isn't a bug that more training data fixes. It's a structural property of "one vector per word."
  • Sparse count vectors don't compose. A TF-IDF vector for a document about "cars" and one about "automobiles" can score as nearly unrelated by cosine similarity, purely because the exact words don't overlap. There's no mechanism for "car" and "automobile" to be recognized as related at all, unlike in a learned embedding space.
  • Embeddings inherit the biases of their training corpus. Bolukbasi et al. (2016) showed that Word2Vec vectors trained on Google News exhibit measurable gender stereotypes in their analogy geometry, with the well-known example: "man is to computer programmer as woman is to homemaker." The vectors are a faithful compression of co-occurrence statistics in real text, and real text carries real social biases forward into the geometry.
  • Subword tokenization has its own sharp edges. How a string gets split into subword pieces is sensitive to whitespace, casing, and even whether a number appears at the start or middle of a word. This is the mechanical reason LLMs are bad at character-level tasks like counting letters in a word: the model never sees individual characters as separate tokens to begin with.
Common misconception

"Word embeddings capture meaning." More precisely, they capture the distributional statistics of how words are used together in a specific training corpus. That is usually a very good proxy for meaning, but it is still a proxy, and it silently encodes whatever the training text encodes, biases included. The same distinction shows up again on 25 (Evaluation, Reliability & Safety) at model scale.

2026 status

Bag-of-words, TF-IDF, Word2Vec, GloVe, and ELMo are all functionally retired as the primary text representation in new systems. Nobody is training a fresh Word2Vec model to power a 2026 product.

But two pieces of this lineage never left. TF-IDF/BM25-style sparse retrieval is still one half of almost every production hybrid search system (see 11). And subword tokenization (BPE/WordPiece/SentencePiece, or close variants) is still the very first thing that happens to your prompt before it reaches any modern LLM.

The representations died; two of the mechanisms that produced them didn't.

07Build this

"Meaning becomes geometry" is the claim this page rests on. It stops being a slogan the first time you type a word into a space you trained yourself and read back its neighbours.

Project Train word vectors, then interrogate the space they live in ~4 hours · PyTorch

Train skip-gram with negative sampling on a small corpus, roughly 100MB of plain text. A slice of a Wikipedia dump is the classic choice. Small is deliberate: you want a vocabulary you personally recognise, so you can tell a good neighbour list from a bad one at a glance.

  1. Build the vocabulary and the training pairs by hand. Tokenise, drop words below a minimum count, and emit every (center, context) pair inside a window of 5.
  2. Implement skip-gram with negative sampling from the objective in section 04. Two embedding matrices, a dot product, a sigmoid, and $k$ negatives drawn from the unigram distribution raised to the 3/4 power.
  3. Train, then write a nearest-neighbour lookup by cosine similarity. Query words you know well. Then query a polysemous one, bank or bat, and read what comes back.
  4. Do the analogy arithmetic yourself: king − man + woman, then find the nearest vector to that result, excluding the three inputs. Try capitals, tenses and plurals too.
  5. Project a few hundred frequent words down to 2D with t-SNE and look at what clusters.
  6. Now break it by retraining with a context window of 1, and again with a window of 10, then re-run the same neighbour queries across all three models side by side.
You'll know it worked when the neighbour lists read like something a person would have written. Query a country and get other countries. Query a verb and get its other tenses. Nothing in the training signal named those categories. They fell out of counting co-occurrences.
What the breakage teaches. Window size is not a tuning knob. It is a definition of what "similar" means. At window 1 the neighbours drift toward words that are grammatically interchangeable, because only the adjacent slot counts as evidence. At window 10 they drift toward words from the same topic. The bank query is the other lesson: one list, both senses, blurred into a single vector. That is the polysemy limit from section 06, staring back at you.

Where this runs in production

Two pieces of this page are still doing real work in 2026 production systems.

Subword tokenization is the literal first step of the inference pipeline for every LLM API call you make. Before any transformer layer runs, your raw text string is split into BPE or SentencePiece tokens. The vocabulary size of that tokenizer is a real design decision every model provider makes: a bigger vocabulary means shorter sequences for the same text, at the cost of a bigger embedding table.

TF-IDF and its close relative BM25 remain a standard "sparse" retrieval signal in production RAG systems. They earn their place precisely because they need no training and match exact terms that dense embeddings can blur together: a product SKU, an error code, a rare proper noun. See 11 for how sparse and dense retrieval get combined in practice.

08Interview questions

BeginnerWhat's the difference between bag-of-words and Word2Vec?

Bag-of-words represents a document as raw word counts, with no learned structure and no notion of similarity between words — "good" and "great" are unrelated columns in a sparse vector. Word2Vec learns a dense, low-dimensional vector per word by training a model to predict context from a center word (or vice versa); the resulting vectors place similar-usage words near each other in the space, which bag-of-words can never do.

BeginnerSkip-gram vs. CBOW — what's actually different?

Skip-gram predicts each surrounding context word from one center word; CBOW averages the context words and predicts the center word from that average. Skip-gram tends to work better on rare words because each occurrence produces multiple training signals (one per context word); CBOW trains faster because it compresses several context words into a single averaged input.

IntermediateWhy can't Word2Vec or GloVe handle polysemy (a word with multiple meanings)?

Both learn exactly one vector per word type, fixed at training time. That single vector has to serve every sense the word is ever used in, so it ends up as an average over all of them — there's no mechanism in the architecture for the vector to change based on which sentence the word appears in. Fixing this requires a fundamentally different approach: making the representation a function of the surrounding context, which is exactly what ELMo does.

IntermediateWhy does modern NLP use subword tokenization instead of whole-word vocabularies?

A whole-word vocabulary has no way to represent a word it never saw during training — typos, new terms, rare names all become an out-of-vocabulary token with no useful signal. Subword tokenization (BPE, WordPiece, SentencePiece) builds a vocabulary of frequent word pieces instead, so any string can be represented as some sequence of known pieces, even if the whole word is unseen — common words collapse to one token, rare or unseen ones fall back to a handful of subword fragments.

IntermediateHow does ELMo actually produce a "contextual" vector?

ELMo trains a deep bidirectional LSTM language model over the whole sentence — one direction predicting the next word, the other predicting the previous word. A word's ELMo representation is a learned, task-specific weighted combination of the raw input embedding and both LSTM layers' hidden states at that position in that sentence. Because it's derived from running the whole sentence through the model, the same word produces a different vector depending on what sentence it's actually in.

DeepT5 frames every task as text-to-text. What's the actual benefit of doing that over BERT's approach of a separate task-specific head per problem?

BERT needs a different output layer architecture per task type — a classification head for sentiment, a span-extraction head for QA, a token-tagging head for NER — each learned somewhat independently on top of the shared encoder. T5 needs exactly one architecture and one training objective (predict corrupted spans of text) for every task, because the task itself is specified in the input string via a prefix rather than baked into the model's output layer. That uniformity is what let the T5 paper run one of the largest systematic comparisons of pretraining objectives, architectures, and dataset choices in NLP up to that point — every experiment used literally the same evaluation harness regardless of task type.

DeepWord2Vec and GloVe often produce vector spaces with similar analogy-solving properties despite very different training procedures. What does that suggest?

It suggests the semantic signal lives primarily in the corpus's co-occurrence statistics themselves, not in the particular algorithm used to extract them — a predictive local-window method (Word2Vec) and a global count-factorization method (GloVe) converging on similarly-shaped spaces is evidence that both are approximating the same underlying statistical structure of the language, just via different optimization paths. It's a useful sanity check to reach for whenever two very differently-motivated methods keep producing suspiciously similar results.

09Go deeper

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

My Notes — 09 NLP Evolution

Free notes

Highlights on this page