←Home KnowML
Home / Quantization Lab

Quantization Lab

One small weight matrix, squeezed into 8-bit and 4-bit integers four different ways. Every scale, every integer code and every last decimal of error, including the single outlier that quietly costs the other weights most of their precision, and the one-line change that gives it back.

Interactive Step through at your own pace · prerequisites: Efficient AI & Systems (23)
How to use this page ▾

Start with INT8 — the other three tabs reuse its vocabulary (α, the scale s, the integer codes q, the error matrix E) and its numbers as the baseline they are measured against. Every figure you read is arithmetic done in the page from one seeded 4×8 weight block, not a hand-typed illustration: the INT8 tab checks its own dequantized values against code × scale computed by hand, proves quantizing twice changes nothing, and confirms that no weight was moved by more than half a grid step. The Outlier tab is the one that explains why quantization is a research area rather than a utility function.

→What this page simplifies, on purpose

  • Only weights are quantized, never activations. Weight-only quantization is the easy half: weights are fixed, known offline, and can be measured exactly. Activations change with every input, so their scale has to be estimated from calibration data or computed on the fly — and they are where the really vicious outliers live. Everything on this page applies to activations too, but the scale would be a guess rather than a measurement.
  • Symmetric only — no zero-point. The scheme here maps 0 to code 0 and uses a single multiplier. Asymmetric (affine) quantization adds an integer offset z so the grid can be shifted, which fits skewed distributions better and costs one more stored number per scale. The error analysis is structurally identical; only W′ = (q − z)·s replaces W′ = q·s.
  • Round-to-nearest, with no error compensation. Every rounding decision here is made independently. GPTQ, AdaRound and friends deliberately round some weights the "wrong" way so that the error they introduce cancels against errors elsewhere in the same layer. That changes which codes get chosen; it does not change what a code means.
  • The block is tiny. Four output channels and eight input dimensions, so every grid fits on a screen and every cell can be checked by hand. A real layer is 4096×4096. Nothing in the arithmetic changes — but note that the per-channel scale overhead, which looks expensive at 8 weights per row, is 0.1% at 4096 weights per row.
  • No kernel, no hardware. This page shows what the numbers become, not how an INT8 matmul is actually executed, how 4-bit weights are packed two-to-a-byte, or which instructions accumulate in INT32. That side of the story is on page 28.