Learning Roadmaps
Thirty-seven numbered pages is a library, not a course. These are four routes through it — in order, with a time cost, a checkpoint you have to pass, and a list of what to leave unread.
A suggested order, not a prerequisite chain. Nothing is locked — go straight to the page that serves you and backfill later.
Every stage ends in a checkpoint: something you should be able to do. The skip lists matter most — people stall by trying to read everything.
01Which one is yours
Pick by where you are now and what you want to be able to do, not by what sounds most advanced.
Ends when you can frame a problem as a learning task, train something on it, and say honestly how well it works. Four stages, mostly classical, one week on neural nets.
Total beginner → can train a model · ~8 weeks of evenings
The most common gap on this site. Ends when you can open an arXiv transformer paper and follow the architecture section without stopping.
Classical ML → modern LLMs · ~6 weeks of evenings
The AI-engineering track: what fits in your VRAM, how serving actually works, fine-tuning, multi-GPU, and keeping it alive once real traffic arrives.
Understands LLMs → can ship one · ~9 weeks, plus GPU access
Shorter and denser. Only the material that actually gets asked, ordered by how often it comes up, wired into the drill pages.
Interview prep & consolidation · ~4 weeks
They assume an hour or two on weekday evenings, and that you run the code rather than read it. Reading alone is roughly three times faster and does not produce the same result.
No stage here promises mastery. "You can now do X" is the claim, and X is written out in the checkpoint so you can check it yourself.
02Roadmap A · Total beginner → can train a model
You know Python and no ML. Four stages. At the end you can take a messy CSV, build something that predicts, and defend the number you report.
One page, and you are not reading all of it. Sections 02 and 03 give you matrix shapes, gradients, and what a distribution is. That is the whole prerequisite for stage 2.
Do not try to "finish maths first." Nobody does. You come back to this page three more times, each time for one specific thing you now need.
Given a 32×784 matrix and a 784×128 matrix, state the output shape of the product without computing anything, and say which dimension is the batch.
Then: a test is 95% sensitive and 95% specific for a disease with 1% prevalence. You test positive. Write Bayes' rule and compute the probability you have it. If your answer is not roughly 16%, you have found the thing this page exists for.
Section 04's term-by-term equations, the information-theory material past "cross-entropy is the loss classifiers use", and every matrix-calculus identity. Section 06's failure modes will mean nothing to you yet.
Why: none of it is load-bearing until page 04 makes you want it, and at that point it takes twenty minutes instead of a week.
Page 02 first, in full. Gradient-boosted trees still beat deep learning on tabular data, which means your first genuinely useful model is on this page and not a neural network.
Page 03 second, and only the clustering and dimensionality-reduction half. You want to know what learning without labels even means before page 10 tells you that is how every LLM is trained.
On any tabular dataset you like: train logistic regression as a baseline, then beat it with gradient boosting. Report both numbers.
Then write one sentence explaining why your train/validation split is honest — what could leak across it, and why it does not. If you cannot write that sentence, the improvement you just measured is not real.
On page 02: the SVM kernel derivation in section 04. Know that kernels map data somewhere it becomes separable, and move on. On page 03: VAEs and contrastive learning — stop after PCA.
Why: both are genuinely important and neither helps you train your first model. VAEs make far more sense after page 04, and contrastive learning after page 09.
This is the page everything after it leans on. Read sections 02 through 04 slowly, then do the Build this project, which is a network written without a framework.
The forward pass is easy and the backward pass is the point. If you only half-follow backprop here, pages 08 and 10 turn into memorisation.
Derive dL/dW2 for a two-layer MLP with MSE loss by hand. Then check it numerically: perturb one weight by 1e-5, recompute the loss, and compare the finite difference to your analytic gradient. They should agree to about six decimal places.
If they do not, the bug is almost always a transpose or a missing activation derivative — find it. Finding it is the checkpoint.
The normalization zoo in section 06 — batch vs layer vs group vs RMS. Know LayerNorm exists and that it stabilises training. Skip the optimizer taxonomy too: use Adam, understand why later.
Why: these are tuning knobs. They matter enormously when you are debugging a training run that will not converge, and not at all before you have one.
Read sections 02 to 04 of page 25 and stop. Calibration, base rates, and why accuracy is the wrong metric almost every time you reach for it.
Then go to practice and enter a Kaggle Playground competition. A leaderboard is an honest verifier, and it is the fastest way to find out that your validation scheme was optimistic.
Take your stage-2 model and plot a calibration curve. At predicted probability 0.8, how often is it actually right? Say whether it is over- or under-confident.
Then name the single metric you would report to someone non-technical, and explain in one sentence why accuracy is not it on your dataset.
Everything LLM-specific on page 25: hallucination, red-teaming, LLM-as-judge, jailbreaks. Roughly sections 03 onward that mention generation.
Why: it is excellent material that assumes you know what an LLM does. It is stage 5 of roadmap B, not stage 4 of this one.
03Roadmap B · Knows classical ML → understands modern LLMs
You are comfortable with regression, trees and cross-validation, and transformers are a fog. Five stages. Ends when you can read a transformer paper's architecture section without stopping.
Skipping straight to attention is the classic mistake. Attention is an answer, and this page is the question — the fixed-size vector that an entire sentence had to squeeze through.
Read sections 02 and 03. You need the shape of the problem, not the recurrence equations.
In two sentences: name the one tensor that every token of the source sentence has to pass through in an encoder-decoder RNN, and say what happens to translation quality as the sentence grows from 10 tokens to 40.
The LSTM gate equations, one by one. And GRUs entirely.
Why: "gates decide what is kept and what is forgotten" is the whole idea, and it is all that pages 08 onward ever use. The six equations are historically interesting and cost you an evening.
The single highest-value page on this site, and the one worth rereading. Sections 02 to 05, then the Build this project.
Then spend an hour in the Attention Lab, which computes every matrix live. Attention stops being abstract the moment you watch the mask land on the scores.
From memory, write the shape at every step of one attention head for batch 2, sequence 5, d_model 64, 8 heads: Q, K, V, the score matrix, the post-softmax weights, the output. Check yourself in the lab.
Then answer: why is the causal mask added to the scores before the softmax rather than applied to the weights after it? Give the numerical reason, not the hand-wave.
FlashAttention, and the KV-cache variants — GQA, MLA, PagedAttention. Also positional-encoding archaeology: know that RoPE is what is used now and move on.
Why: every one of those is a serving optimisation. They make sense once you care about latency and VRAM, which is roadmap C. Reading them now costs a week and changes nothing about your understanding of the mechanism.
A short stage that pays for itself. Every acronym in every paper you are about to read is on this page, in the order it was invented and with the problem it solved.
Read it once, quickly. You are building an index, not learning a technique.
Put these in chronological order and give each one sentence on what it added: word2vec, ELMo, BERT, GPT-2, InstructGPT.
Then say which of them are encoder-only, and what that rules out. If you cannot explain why BERT cannot generate text, reread section 03.
The n-gram and count-based prehistory, beyond knowing what TF-IDF computes.
Why: it is real history and it is dead weight for reading a 2026 paper. The one case it matters — sparse retrieval in RAG — is covered again on page 11 when you need it.
Page 10 in full. This is where attention becomes a language model: the block, the scaling laws, the data pipeline, and the three-stage training run.
Then page 32, sections 01 to 06 only. Stop after the DPO derivation — watching the reward model cancel out of the objective is the single most satisfying piece of algebra on this site.
Given a 7B model trained on 300B tokens, say whether it is Chinchilla-undertrained and by roughly what factor. Then say why a lab might do that deliberately anyway.
Then draw the three stages — pretrain, SFT, preference optimization — and for each one name the data it needs and who has to produce it.
On page 32: PPO's four-model loop (section 05), GRPO and RLVR. On page 10: MoE routing internals and the distributed-training sections.
Why: DPO is the one you will actually meet, and it is the cleanest entry point. PPO makes far more sense read backwards from DPO, and reads as unmotivated machinery if you take it first.
Page 11 explains the two things bolted onto every deployed LLM: retrieval, because weights go stale and cannot cite; and the act–observe loop, because a single forward pass cannot check its own work.
Page 33 is a reference, not a read. Skim it once so you know which models exist, then return to it whenever a paper names one you do not recognise.
Open any recent architecture paper on arXiv. Read the architecture and training sections end to end without looking anything up. You should be able to name what is standard and what is the paper's actual contribution.
Then explain why RAG reduces hallucination without eliminating it, and name the two points where retrieval quality caps the final answer.
The agent-framework tour on page 11, and page 33's benchmark tables beyond understanding what one or two of them measure.
Why: plan → act → observe → repeat is the durable idea; the frameworks implementing it turn over every few months. Benchmark numbers go stale faster than anything else on this site.
04Roadmap C · Understands LLMs → can ship one
You can read the papers. You have never put a model behind an endpoint and watched p95 latency. Five stages, and this one needs a GPU you can actually use.
Start here because every later decision is downstream of one number. Section 02 does the VRAM arithmetic once, properly — weights, KV cache, activations, overhead.
Do the arithmetic yourself for a model you want to run. Most "why is this so slow" questions dissolve the moment you see that the weights do not fit.
Compute the VRAM for a 14B model at 4-bit with an 8k context and a batch of 4. Give three separate numbers: weights, KV cache, everything else.
Then compare that to the card you have, and state what you can actually run and at what context length. Write the number down; stage 2 uses it.
The full quantization-format comparison. Know that GGUF, AWQ and GPTQ exist, know which runtime wants which, and pick one.
Why: format choice is a twenty-minute decision when you need it and an hour of reading when you do not. Section 03 is a reference table, not a chapter.
Page 31 in full — prefill versus decode is the split that every other serving fact hangs off. Do its Build this project; the latency wall is more convincing when you hit it yourself.
Then page 23, sections 07 to 11 only. That is the inference half: arithmetic intensity, why memory bandwidth rather than FLOPs is the ceiling, and the compression options.
For one model on one GPU, say whether prefill or decode dominates at batch size 1, and again at batch size 64 — and explain why the answer flips.
Justify it with the arithmetic-intensity number from page 23 section 08, not with intuition. Then say what happens to TTFT and to inter-token latency as you raise the batch size, and why they move in opposite directions.
Page 23 sections 04 to 06 — tensor parallelism, ZeRO/FSDP, mixed precision. Also MoE serving and edge inference in section 12.
Why: those are training-time concerns and they come back as stage 4. Reading them now mixes two different problems and makes both harder to hold.
Page 29 opens with the decision before the decision: whether to fine-tune at all. Take it seriously. Prompting and retrieval solve more cases than fine-tuning does, and cost nothing to try.
When you do run it, run it end to end once — data, chat template, LoRA, eval. Then read page 32 properly, including PPO and GRPO, now that you have seen SFT fall short.
Given 800 training examples and a 24 GB card: choose LoRA, QLoRA or full fine-tuning, justify the rank you picked, and name the eval that will tell you it worked.
The eval has to be decided before training starts. If you cannot state it up front, you are going to declare victory on vibes.
Preference optimization as a first move, unless you already hold real preference data. Also the full variant tour in page 32 section 10.
Why: most teams who say they need RLHF need better SFT data. Collecting pairwise preferences is expensive, and doing it before your supervised baseline is clean means you will be optimising against noise.
Page 28 explains the hardware that page 23 kept referring to: the memory hierarchy, warps, tensor cores, and the interconnect that decides whether your multi-GPU run scales.
Then go back for page 23's training half. Tensor parallelism, ZeRO and FSDP make sense once you know what an all-reduce costs on the wire.
Explain ring all-reduce's per-step communication cost and why it does not grow with the number of GPUs. Page 28 section 11 has the derivation; do it without looking.
Then: at what point would you reach for tensor parallelism instead of ZeRO-3? Name the specific constraint that forces the switch.
Writing raw CUDA. Read section 08 on Triton and skip the kernel-level walkthrough in section 07 unless you are targeting a kernel-engineering role.
Why: almost nobody hand-writes CUDA any more, and the ones who do were hired to. You need to read a profiler and reason about memory movement — both of which section 13 gives you without a single kernel.
Page 24 is the loop your model lives inside: shadow deploys, canaries, drift monitoring, rollback. Page 25 is how you decide whether the new version is better, which is harder than it sounds.
Read page 25 in full this time, including the LLM-specific half — hallucination measurement, red-teaming, and the sharp limits of LLM-as-judge.
Write the rollback criterion for your deploy as one sentence an on-call engineer could act on at 3am without asking you anything.
Then name one offline metric and one online metric that can disagree on the same model, say which you would trust, and say what you would ship while you worked out why they disagreed.
Feature-store vendor comparisons, and the orchestration-tool landscape.
Why: the concept — one definition of a feature, used identically in training and serving — is what prevents training-serving skew. Which product implements it is a procurement question, and it will have changed before you need to answer it.
05Roadmap D · Interview prep & consolidation
Shorter and denser, because you mostly know this. Four weeks, ordered by how often the material actually comes up, and wired into the drill pages throughout.
Page 27's first three sections give you the six question types and the shape of a system-design and a debugging answer. Read them, then run one pass of the review deck without preparing.
The point is a cold measurement. Whatever you rate "again" on tonight is your actual syllabus, and it is rarely the topic you were worried about.
Rate yourself honestly on each of the six question types in page 27 section 01. Write down the two lowest.
Those two decide which of the stages below you spend double time on. Everyone's list is different, and skipping this step is why generic prep plans waste weeks.
Any attempt to "finish" the review deck tonight. Do 30 cards and stop.
Why: you are sampling to locate gaps, not studying. A long first session gives you a worse signal because fatigue looks exactly like a knowledge gap.
Read only section 08 — Interview questions — on each page first. Anything you cannot answer, go back into that page for the section it came from. Anything you can, move on.
This inverts the usual order deliberately. Reading a whole page to discover you already knew it is the most common way prep time disappears.
Use the 60-second explanation drill from page 27 section 04 on four things: backprop, self-attention, the bias-variance tradeoff, and why accuracy misleads on imbalanced data.
Record yourself. Over 90 seconds means you do not have it compressed yet, and compression is what is being tested.
Any derivation you cannot reproduce under time pressure. Learn the shape of the argument and where the key term comes from instead.
Why: interviewers ask "why does this term appear" far more than "reproduce this proof". A half-remembered derivation on a whiteboard reads worse than a clear verbal argument.
Page 16 is here because the retrieval–rank–rerank funnel is the most reused design answer there is, and it transfers directly to search and to RAG. Page 24 supplies the deploy-and-monitor half that candidates forget.
Practise speaking, not reading. A design answer you have only thought about collapses at minute three.
Out loud, in 20 minutes, using page 27's structure: design a recommender for 10M items under a 30 ms budget. Cover retrieval, ranking, features, metric, and how you would A/B test it.
Then do the same for an LLM serving stack with a p95 TTFT target. Different domain, same five moves — that is the thing being checked.
Page 16's ANN index internals. "HNSW is a navigable graph, IVF-PQ trades memory for recall" is enough for any interview that is not specifically a retrieval role.
Why: the funnel and the feedback loop are what get probed. Nobody has ever been rejected for not knowing HNSW's efSearch parameter.
This block has grown fastest in real interviews. Scaling laws, the RLHF-to-DPO story, why RAG helps, and the KV cache — those four come up constantly, across research and engineering loops alike.
Page 23's inference sections are the ones that separate candidates. Most people can define a KV cache; far fewer can say how fast it grows.
On a whiteboard, in under 10 minutes: start from the RLHF objective and arrive at the DPO loss, saying out loud where the reward model cancels.
Then, cold: give the KV cache size in GB for a 70B model at 32k context and batch 8, and name one technique that reduces it and what it costs you.
GRPO and RLVR unless you are interviewing at a lab that trains reasoning models. Also the MoE internals on page 10.
Why: they are asked in a narrow band of roles. If that is your band you already know it, and if it is not, the week is better spent on stage 2.
The review page schedules every interview question on the site by spaced repetition. Answer before revealing, and rate honestly — rating generously is the only way to break it.
If your bottleneck turns out to be writing code rather than knowing ML, practice has the fix, and it is a cheaper problem than it feels like.
Two consecutive passes of the review deck, a week apart, with 90% rated "good" or better, and no card that keeps coming back.
Plus: two full system-design answers delivered out loud to another person, in under 25 minutes each, without notes.
New topics, in the last week. Pick nothing up that you have not already seen.
Why: a half-learned topic you volunteer in an interview is worse than one you never mention. The last week is for compression and retrieval speed, not coverage.
06The pages no roadmap sends you to
Roughly half the site is not on any route above. That is deliberate, and worth being explicit about.
The four roadmaps cover the load-bearing core. Everything else is a destination page: you read it when your job, your project or your curiosity takes you there, and reading it earlier buys nothing.
| Pages | What they are | When to read them |
|---|---|---|
| 05, 06, 12, 35 | Vision and generative media, from convolutions to diffusion. | The moment you work with images. Page 12 also stands alone as the best explanation of sampling on the site. |
| 13, 14, 34 | Speech, audio and multimodal fusion. | When you build a voice product. Page 34's metrics section is worth it on its own if anyone ever quotes you a WER. |
| 15 | MDPs, Q-learning, policy gradients, PPO. | Read sections 01–04 after roadmap B stage 4. Page 32's PPO section assumes it, and pretending otherwise is how RLHF stays mysterious. |
| 17, 18, 19 | Data shapes with their own rules — temporal, relational, physical. | When your data has that shape. Page 17's "you cannot shuffle" is worth an afternoon even if you never forecast anything. |
| 20, 21, 22 | Embodied AI: perception, control, and learned simulators. | Career-choice reading. Take them as a block, in that order, if the area appeals. |
| 26, 36, 37 | Where the field is going, and what it has actually delivered. | Any time. Page 37 is the antidote to a week of reading benchmark numbers. |
The Technique Map organises every method on the site by the problem it solves rather than by page order. It is the better entry point when you have a specific problem and no interest in a curriculum.
Search (⌘K) indexes every page, section and term. For a single definition, that beats any roadmap.