Home · Browse
Everything on the site, in one place
All 40 sections grouped by theme, the timeline they sit on, and the order they depend on each other in. ⌘K if you already know what you want.
The learning graph
All 40 sections
Grouped by theme. The numbers are a reading order, not a ranking.
Foundations
Neural Nets & Vision
Sequence, Attention & LLMs
07
Sequence Modeling Pre-Transformer
RNNs, LSTMs, seq2seq — the bottleneck attention removed.
Full
08
Attention & Transformers
Query, key, value. The mechanism behind almost everything.
Full
09
NLP Evolution
Bag-of-words to BERT to instruction tuning.
Full
10
LLM Architecture & Training
Decoder-only stacks, scaling laws, RLHF and DPO.
Full
Generative & Multimodal
Decision & Retrieval Systems
15
Reinforcement Learning
MDPs, Q-learning, PPO — and what RLHF borrows.
Full
16
Recommenders, Ranking & Search
Retrieval, ranking, rerank — and the feedback loop that poisons it.
Full
17
Time Series & Forecasting
Why you cannot shuffle. Stationarity and honest backtests.
Full
18
Graph ML
Message passing as the one primitive.
Full
Embodied & Frontier
Systems, Safety & Interview
23
Efficient AI & Systems
FlashAttention, quantization, parallelism.
Full
24
ML Engineering & MLOps
Feature stores, drift, canary deploys.
Full
25
Evaluation, Reliability & Safety
Calibration, hallucination, red-teaming, LLM-as-judge.
Full
26
2026 Frontier Map
Reasoning models, agents, world models.
Full
27
Interview Mastery
Question types, drills, how to structure an answer.
Full
28
GPU Architecture, CUDA & Distributed Training
SIMT, warps, memory hierarchy, ring all-reduce.
Full
Hands-on
Inference & Serving
31
LLM Inference & Serving
Prefill vs decode, TTFT, continuous batching.
Full
38
LLM Inference at Scale
Disaggregation, KV quantization, multi-LoRA, FP4.
Full
39
Real-Time Voice AI
The turn-latency budget, endpointing, barge-in.
Full
40
Diffusion & Video Inference
Cutting the sampling loop, caching, video cost.
Full
The Labs about the labs →
Backprop Lab
Forward, backward, then gradients checked numerically.
Interactive
Tokenizer Lab
Build a BPE vocabulary merge by merge.
Interactive
Attention Lab
Seven attention mechanisms on one worked example.
Interactive
Sampling Lab
Temperature, top-k and top-p on the same row.
Interactive
Quantization Lab
INT8 error cell by cell, and what one outlier costs.
Interactive
Model Atlas
00 · The timeline
Seven eras, 1950s → 2026
Scroll sideways.
1950s – 1970s
Symbolic roots
- Perceptron
- Symbolic AI & expert systems
- Bayesian foundations
- Classical control theory
1980s – 1990s
The statistical turn
- Backpropagation popularized
- Decision trees
- SVMs
- HMMs & graphical models
2000s
Kernels & ensembles
- Kernel methods
- Boosting & random forests
- CRFs
- Matrix factorization
2012 – 2017
Deep learning breaks out
- AlexNet
- Word2Vec / GloVe
- Seq2seq + attention
- GANs, ResNet
2017 – 2020
The Transformer era begins
- Transformer, BERT, GPT
- ViT
- Contrastive learning
- Diffusion beginnings
2020 – 2023
Scale and alignment
- Scaling laws
- Instruction tuning, RLHF
- Multimodal foundation models
- Diffusion explosion
2023 – 2026
Reasoning & agents
- Reasoning models, tool use
- Long context, efficient attention
- Multimodal-native models
- Video/world models, VLA robotics
Suggested path
Core learning order
If you are starting cold, this is the dependency order.
Math→
Classical ML→
Unsupervised→
NN Fundamentals→
CNN/Vision→
RNN/Seq2seq→
Attention→
Transformers→
NLP/LLMs→
Generative→
Multimodal→
Speech→
RL→
RAG/Agents→
Rec/TS/Graph→
3D/Spatial→
Robotics→
World Models→
Efficient AI→
MLOps/Eval/Safety→
2026 Frontier
Vocabulary check
AI vs ML vs deep learning vs foundation models vs agents
Nested, not synonymous.
Scope, nested from broadest to narrowest
AI — any system that performs tasks we'd call intelligent, including hand-coded expert systems, no learning required. ML — a subset that learns its behavior from data. Deep learning — a subset of ML using multi-layer neural networks specifically. Foundation models — a subset of deep learning: large models pretrained on broad data, adaptable to many downstream tasks. Agents aren't a subset at all — they're a system built around a foundation model, wiring it to tools, memory and a planning loop.
Beyond this site
Watch, practice, discuss
Creators worth following.
Creators worth following
Priyam Mazumdar — Exploratory Data Adventures
PyTorch and deep learning built up from the ground up, aimed squarely at demystifying the field.
Umar Jamil
Transformers, LLaMA, and modern architectures implemented from scratch, line by line.
Yannic Kilcher
Reads the actual papers, start to finish, and unpacks the parts that don't hold up.
Plus 3Blue1Brown, StatQuest, and Andrej Karpathy — linked directly from the "Go deeper" section of the pages they're most relevant to.
Practice
Deep-ML
LeetCode, but for ML — implement linear regression, attention, backprop, and dozens more from scratch, implemented from scratch.
Start solving →