Evaluation, Reliability & Safety
A test-set accuracy number answers "did it do well on questions like these." It says nothing about the question nobody thought to ask. This page is the discipline of finding that question before a user, an attacker, or a regulator does.
Evaluation is a portfolio of different questions and not a single number, because a model can fail in categorically different ways: confidently wrong (calibration), fragile to inputs slightly unlike its training data (robustness), factually wrong while sounding fluent (hallucination), exploitable by adversarial input (security), or opaque about why it did what it did (interpretability). LLM-as-judge and pairwise preference comparison are used because human evaluation doesn't scale to the volume modern systems need.
They inherit real, well-documented biases (position, verbosity, self-preference) that you can mitigate but not eliminate. Red-teaming and monitoring run continuously and are not a one-off pre-launch gate, because new failure modes and attacks are usually discovered after launch. One sentence for an interview: every method here is a proxy for "will this behave well on cases I haven't seen," and the discipline is knowing which proxy's blind spots matter for your system.
01Intuition
You cannot fully test a system against failure modes you haven't imagined yet. Every technique on this page is a different strategy for imagining more of them, faster than an attacker or an unlucky user does.
Here is the uncomfortable core fact. A model's test-set performance tells you how it behaves on inputs statistically similar to that test set. It tells you nothing directly about three other kinds of input:
- Subtly different. Distribution shift, where the input population has drifted.
- Deliberately adversarial. Someone crafting an input specifically to break it.
- Rare. Cases too infrequent to show up in a few thousand held-out examples.
The test set is a sample of "questions like these." The whole point of evaluation as a discipline is that "questions like these" is never the full space of questions a deployed system receives.
Read as a list, these eighteen concepts look unrelated, but each one is a different strategy for the same job: finding failure modes before a user or an attacker does.
- Robustness testing and adversarial examples search near the edge of what the model handles well.
- Red-teaming searches for what a malicious user would craft.
- Calibration asks the model to know when it is in unfamiliar territory, even when it can't always get the answer right.
- Interpretability asks what the model computes internally, so you are not relying purely on input-output behavior to guess at untriggered failures.
None of them prove the system is safe. They give you a wider net for finding the ways it isn't, before those ways find your users.
Accuracy can look excellent for a model that does nothing useful. Suppose 1 case in 100 is the one you care about, and the model always answers no.
Try it Score a model that never finds the rare case
import numpy as np
y = np.array([0] * 990 + [1] * 10) # 1% of cases are the rare one
pred = np.zeros_like(y) # a "model" that always says 0
acc = (pred == y).mean()
recall = (pred[y == 1] == 1).mean()
print(f"accuracy: {acc:.1%}")
print(f"recall on the rare class: {recall:.1%}")
99.0% accuracy and finds none of the rare cases: 0.0% recall. Ninety-nine percent of its answers are right only because 99% of the cases are negative. On imbalanced problems, a metric that counts the rare class, such as recall, is the one to watch.02Timeline
Classical evaluation was static: hold out a test set, compute accuracy, F1 or MSE once, and trust that the test set resembled production.
Open-ended generation has no single correct string, which pushed the field to LLM-as-judge and pairwise preference comparisons; web-scale pretraining leaked benchmarks into training data and inflated scores; and agentic systems made the whole trajectory, not just the final answer, the thing to evaluate.
Modern practice treats reliability as ongoing and multi-dimensional, with benchmarks, calibration, robustness testing, continuous red-teaming and production monitoring running together, and the mindset has shifted from "did it pass" to "how do we keep finding the ways it fails."
03The landscape: five questions, five toolkits
Eighteen distinct concepts sounds unmanageable until you notice they cluster into five questions being asked about a model. Keep the grouping in your head and not the list.
1. How do you measure quality when there's no single right answer?
Benchmark design and contamination. A benchmark is only informative if the model hasn't already seen its answers during training. Contamination is benchmark questions and answers leaking into a pretraining corpus scraped from the same public internet the benchmark was published on. It is a real, documented problem for web-scale LLMs, and it inflates scores silently: nothing errors, and the number means nothing.
Three detection strategies, none conclusive alone:
- N-gram overlap. Check the training corpus against the benchmark text directly.
- Canary strings. Insert markers into a benchmark, then see whether a model has memorized them.
- Held-back variants. Compare the public version against a never-published variant of similar difficulty.
Calibration and uncertainty. A model is calibrated if its stated confidence matches its actual accuracy. When it says "80% confident" across many predictions, it should be right about 80% of the time. Modern neural networks are frequently overconfident by default.
Temperature scaling (dividing logits by a learned scalar before softmax) is the standard practical fix, covered in depth on 01. What is specific to this page is that calibration is also an evaluation target you measure, and not only a training-time fix. The Expected Calibration Error metric below is how.
LLM-as-judge. Use a strong LLM to score or compare another model's open-ended output against a rubric. Exact-match scoring doesn't work for free-form text, and human evaluation doesn't scale to continuous, high-volume testing. It works well enough to be useful. It also has three specific, well-documented failure modes:
- Position bias. Favoring whichever answer appears first or second in a pairwise comparison, regardless of quality.
- Verbosity bias. Rewarding longer answers independent of correctness.
- Self-enhancement bias. A judge scoring outputs from its own model family more favorably.
The mitigations are mechanical. Randomize answer order and average across orderings to cancel position bias. Use a judge from a different model family than the one being evaluated. Use multiple judges and not just one.
Pairwise preference evaluation. Instead of asking a judge for an absolute score ("rate this 1–10"), ask which of two outputs is better. Relative judgments are consistently more reliable than absolute ones, a documented pattern in human preference research as well as LLM-judge research. It is also exactly why RLHF preference data (see 15) is collected as pairwise comparisons and not as absolute scores.
Agent trajectory evaluation. An agent doesn't produce one output. It produces a sequence of tool calls, observations, and intermediate decisions (the ReAct-style loop from 11). Grading only the final answer misses failures mid-trajectory: reaching the right answer via an unsafe action, an unnecessarily expensive sequence of tool calls, or getting lucky after an early wrong turn.
Trajectory evaluation checks the path as well as the destination, and that is harder. There are usually many valid paths to a correct outcome, so "did it match a reference trajectory" is often the wrong question, and "was every step reasonable given what the agent knew at that point" is closer to the right one.
2. Does it still work on inputs unlike its training data?
Robustness to distribution shift means measuring performance on data that differs statistically from the training data: a new time period, a new demographic, a phrasing style users pick up after launch. It is the most common way production models degrade, and it happens without a crash, an error or a code change. The inputs drift away from what the model was validated on, and nothing tells you.
Adversarial examples. Small, often imperceptible, deliberately crafted perturbations that flip a model's prediction. The classic case is an image, altered by an amount invisible to a human, that a classifier suddenly misclassifies with high confidence. For LLMs the analogous attack is an adversarial suffix or crafted phrasing engineered to manipulate output. What separates these from ordinary distribution shift is intent: someone specifically searched for the input that breaks you.
OOD detection means recognizing when an input is unlike anything the model was trained on, so the system can abstain, flag for human review, or express low confidence instead of confidently emitting a wrong answer. It is the practical, deployable complement to calibration: calibration asks the model to know its confidence is low, and OOD detection asks the system to catch the case entirely, ideally before the model has to guess.
Fairness and subgroup evaluation. An aggregate accuracy number can look fine while hiding a large performance gap for a specific subgroup. Computing metrics sliced by group (by demographic, by region, by device type) is the only way to catch that a "97% accurate" model is failing 40% of the time for one slice of users.
A hard result: several intuitive fairness definitions can be mathematically incompatible with each other except in special cases:
- Equal false-positive rates across groups
- Equal calibration across groups
- Equal overall accuracy
You generally cannot satisfy all of them simultaneously. So "make the model fair" always requires picking which fairness definition matters for the decision at hand.
3. Can you trust what it says?
Hallucination and factuality. LLMs are trained to produce fluent, plausible continuations and not to fact-check themselves. Fluency and truth are correlated in training data, but they are different objectives.
That gap is why a model can generate a confident, well-formatted, entirely fabricated citation. Hallucination isn't a bug that a bigger model necessarily removes. It is a direct consequence of the training objective, so mitigation leans on retrieval grounding and verification as well as scale.
Groundedness and citation correctness. Specific to retrieval-augmented systems (11). Check that a claim in the generated answer is supported by the source it is attributed to, and that it is more than plausible-sounding text near a citation marker.
A groundedness checker (often another LLM call, or a dedicated smaller classifier) compares each claim against the retrieved passages it is supposed to rest on. This catches the common RAG failure where retrieval succeeded but generation still drifted away from what was retrieved.
4. Can someone break it on purpose?
Red teaming. Deliberately trying to elicit harmful, unsafe, or undesired behavior from a system before real adversaries do, using human testers, automated adversarial generation, or both. It has to run continuously and cannot be a one-time pre-launch gate, because new jailbreak techniques are discovered constantly after launch. A system red-teamed thoroughly in January is not necessarily still safe against attacks discovered in June.
Prompt injection and tool security. A prompt injection hides instructions inside content the model processes: a webpage, a document, an email. Those instructions hijack the model's behavior when it treats embedded text as commands and not as data to reason about.
Agentic and tool-using systems are especially exposed (see the ReAct loop on 11). An agent that reads a webpage as part of a task can have its goal silently overwritten by text on that page, and because the "attack" is ordinary-looking text, it is far harder to filter than a malicious API payload.
Data poisoning and model extraction. Poisoning is an attacker manipulating training data, inserting examples designed to implant a backdoor (a hidden trigger causing specific misbehavior) or bias the model in a chosen direction. Extraction runs the other way. An attacker repeatedly queries a deployed model to reconstruct a functional copy, or to recover memorized training data, without ever accessing the weights.
Privacy leakage. Models can memorize specific training examples and later regurgitate them, verbatim or near-verbatim, when queried in particular ways. When training data includes personal or sensitive information this is a genuine risk, because the model didn't "learn a pattern" from that example so much as store it outright.
5. What is it computing, and can you steer it?
Saliency, probing, activation analysis. Saliency methods attribute an output back to input features, usually via gradients or perturbation: which pixels, which tokens most influenced this specific prediction. Probing trains a small auxiliary classifier on internal activations to test whether specific information (part of speech, sentiment, a fact) is linearly recoverable from a layer. That tells you what is represented, even when you don't know how it is used.
Both are useful, and both share a limitation worth stating plainly: a saliency map can look plausible while not reflecting the computation that caused the output. It is a correlational tool and is no proof of causation.
Mechanistic interpretability is a more ambitious research program: reverse-engineering the algorithm a network has learned internally, the way you would reverse-engineer a compiled program back into readable source. That means finding specific circuits, subsets of attention heads and neurons across layers responsible for a particular capability, and going beyond correlating activations with outputs.
It is a young, active research area and not a routine production tool, but it is the direction the field is heading if "we don't really know why it did that" stops being an acceptable answer for high-stakes systems.
Safety alignment and controllability. Alignment means making a model's behavior reliably match intended values and constraints; the RLHF/DPO machinery from 10 and 15 is the primary current tool.
Controllability is separate: ensuring humans retain a meaningful ability to correct, constrain, or halt undesired behavior after deployment. Alignment happens at training time, controllability at run time. They are related but distinct goals, and a system can have one without fully having the other.
The reference-based metrics you will be asked to name
Before model-graded evaluation existed, generation quality was scored by comparing output against human-written references. These metrics are still the standard vocabulary in translation and summarisation, and they still appear in interviews.
- BLEU. Built for machine translation. Counts how many n-grams of the candidate appear in the reference, with a brevity penalty so a system cannot score well by emitting three safe words. Precision-oriented: it asks how much of what you produced was warranted.
- ROUGE. Built for summarisation, and recall-oriented by contrast: it asks how much of the reference you managed to cover. ROUGE-N counts n-gram overlap, ROUGE-L uses the longest common subsequence, so word order matters without demanding a contiguous match.
- Perplexity compares nothing against a reference. It is the exponentiated average negative log-likelihood the model assigns to held-out text, so it measures how surprised the model is by real data. Lower is better, and it is only comparable between models sharing a tokenizer, since it is computed per token.
The shared weakness is usually the actual interview question. All three reward surface overlap and not meaning. A correct paraphrase that shares no n-grams with the reference scores badly, and a fluent, confidently wrong answer that reuses the reference's vocabulary scores well.
They are cheap, deterministic and reproducible, which model-graded evaluation is not. An LLM judge captures meaning far better but costs money per call, drifts when the judge model is updated, and carries its own biases, including a documented preference for longer answers.
The practical answer is rarely one or the other: use n-gram metrics as a fast regression check that catches gross breakage, and reserve judged or human evaluation for the decisions that matter.
04The math that's load-bearing here
- Expected Calibration Error bucket predictions into $B$ bins by their predicted confidence (e.g. 0–10%, 10–20%, …), then for each bin $S_b$ compare the bin's actual accuracy against its average stated confidence, weighted by how many predictions fall in that bin.
- acc(S_b) the fraction of predictions in that confidence bin that were correct.
- conf(S_b) the average predicted confidence of predictions in that bin.
- Reading it ECE near zero means confidence tracks reality. The model's "80% sure" predictions are right about 80% of the time. A large ECE means the model is systematically over- or under-confident, and temperature scaling (01) is the standard fix once you've measured it.
Now consider agreement between two raters: two human annotators, or a human checked against an LLM judge. Plain percent agreement overstates reliability, because two raters agree by pure chance some of the time. More often still if one label is common. Cohen's kappa corrects for that:
Here $p_o$ is observed agreement and $p_e$ is the agreement expected purely by chance, given each rater's label distribution. A kappa near 0 means the raters are barely beating chance. Near 1 means they agree almost perfectly beyond what chance alone would produce.
Use it for one specific question: does my LLM judge agree with my human evaluators, or does it just look that way because it always picks the more common label?
Finally, any binary safety classifier (a jailbreak detector, a content filter, a red-team success/fail label) inherits the precision/recall tradeoff from 01 directly:
- Tuned for recall. Catches nearly everything unsafe (high recall), and flags legitimate content as unsafe more often (lower precision).
- Tuned for precision. Rarely bothers legitimate users (high precision), and lets more unsafe content through (lower recall).
No threshold maximizes both simultaneously. Deciding which error is more expensive for your system is a product decision.
05Why this design
Why trust an LLM to judge another LLM, given the known biases? Because the alternative doesn't scale: exhaustive human evaluation of every model change, at the rate modern systems iterate, isn't available. And the biases can be mitigated, so they are not fatal:
- Position bias cancels out if you randomize order and average.
- Self-enhancement bias shrinks when the judge comes from a different model family than the one under test.
- Verbosity bias is caught by naming it as an explicit rubric criterion, so "quality" is not left undefined.
Used carefully, with those mitigations applied and human spot-checks retained, LLM-as-judge is a useful scalable proxy. It isn't a perfect one, and saying so is the more honest claim.
Why pairwise comparison instead of absolute scoring? "Is A better than B" is a much easier and more consistent judgment than "rate this response 7 out of 10," for humans and LLM judges alike.
An absolute scale has no fixed anchor, and different raters silently use different internal scales for what a "7" means. Relative comparisons sidestep that: you don't need to agree on what a 7 means, only on which of two concrete things is better.
Why red-team continuously instead of once before launch? Because the set of known attacks is not fixed. A model considered safe against every jailbreak technique known in January can be broken by one discovered in June. The attack surface evolves after deployment, the same way a security system's threat model doesn't freeze the day it ships.
Treating red-teaming as a one-time gate treats an adversarial, ongoing process as a static test, and that category error shows up in production as a sudden surprise where a slow trend would have been easy to catch.
06Complexity, failure modes, and what breaks
A benchmark that has leaked into training data doesn't fail loudly. It produces an inflated, meaningless score that looks like genuine progress. The only way to catch it is actively checking for overlap between training data and benchmark content, or holding back a fresh, never-published variant to compare against; a suspiciously high score alone is not proof, but it's a reason to check.
If your judge model and your system-under-test are both updated over time, self-enhancement bias can silently shift. A judge from a newer model family might start favoring outputs that share that family's stylistic quirks, making "quality improved" and "outputs got more judge-flattering" indistinguishable without a periodic human-eval check to recalibrate against.
Because several reasonable fairness definitions are mathematically incompatible except in special cases, "make it fair" is underspecified until you pick which definition matters for the decision. Equal error rates across groups and equal calibration across groups can be in direct tension, and satisfying one can require violating the other.
A saliency map highlighting "the features that mattered" is a correlational attribution and proves nothing about the model's causal computation. It can look clean and interpretable while not accurately reflecting why the model produced that output, and that is the gap mechanistic interpretability tries to close by studying the computation itself as well as its attributed inputs.
Because the "malicious payload" is just ordinary-looking text on a webpage or in a document, standard input sanitization (built for detecting code injection, SQL injection, and similar structured attacks) doesn't generalize to catching it. The content is entirely well-formed natural language, which is precisely what makes it hard to filter without also filtering legitimate instructions.
07Build this
Overconfidence is the one failure on this page you can see with your own eyes in a single plot. Draw the reliability diagram once and the ECE equation stops being notation.
Train a small network on two deliberately overlapping Gaussian classes, so that no model, however good, is entitled to be certain anywhere near the boundary. That overlap is the reason for this dataset: it gives you a task where confident predictions are provably unjustified, which is exactly what a reliability diagram is built to expose.
- Generate the two overlapping classes. Split three ways: train, a held-out calibration split, and test.
- Train a small MLP well past the point where training accuracy stops improving. Overconfidence is something a network grows into late in training and does not start with.
- Build the reliability diagram by hand. Bin the test predictions by stated confidence, plot each bin's accuracy against its mean confidence, and draw the $y = x$ diagonal. Compute ECE from the equation in section 04 while you are already binning.
- Fit one scalar temperature $T$ on the calibration split, minimising negative log-likelihood over $\text{softmax}(z/T)$. It is one parameter fitted after training, and the weights never move.
- Replot the diagram and recompute ECE on the same untouched test split.
- Now break it: refit $T$ on the training split instead of the held-out one, and replot.
Where this runs in production
A team wants to change a chatbot's system prompt, and needs to know whether it is an improvement before rolling it out to all traffic. The practical pipeline:
- Generate. Run both prompt versions against a fixed set of representative queries.
- Compare. Use pairwise LLM-as-judge, order-randomized, judged by a model from a different family than the one being tested, to get a preference rate between old and new.
- Spot-check. Review a random sample of the judge's calls by hand. The purpose is to catch systematic judge errors before trusting the aggregate.
- Canary. If the new prompt wins clearly and the humans agree with the judge, roll out to a small percentage of real traffic. Watch engagement, escalation rate to a human agent, and explicit user feedback before widening.
The judge comparison tells you the new prompt is probably better in the way you tested.
The canary tells you it is better in the way that matters, on traffic you did not construct. Neither one substitutes for the other, which is why the pipeline has both.
08Interview questions
BeginnerWhat does it mean for a model to be "calibrated," and how would you measure it?
A calibrated model's stated confidence matches its actual accuracy — when it says 80% confident across many predictions, it's right about 80% of the time. You measure it with Expected Calibration Error: bin predictions by confidence, compare each bin's average confidence against its actual accuracy, and take the weighted average of that gap across bins. Neural networks tend to be overconfident by default; temperature scaling — dividing logits by a learned scalar before softmax — is the standard fix, applied after training without changing which class gets predicted, only how confidently.
BeginnerHow would you detect that a benchmark has been contaminated by training data?
A few converging signals, none conclusive alone: check for direct n-gram overlap between the benchmark's text and the training corpus; insert canary strings into a benchmark and check whether the model can reproduce them verbatim, which it shouldn't be able to do if it never saw them; and compare performance on the published benchmark against a fresh, never-released variant of similar difficulty — a large, unexplained gap between the two is a strong signal the published version leaked in.
IntermediateWhat are the known failure modes of LLM-as-judge, and how do you mitigate each?
Position bias — favoring whichever answer appears first or second regardless of quality — mitigated by randomizing order and averaging across orderings. Verbosity bias — rewarding length over correctness — mitigated by making conciseness or specificity an explicit rubric criterion instead of leaving "quality" undefined. Self-enhancement bias — a judge favoring outputs from its own model family — mitigated by using a judge model from a different family than the system under test. None of these are fully eliminable, which is why periodic human spot-checks against the judge's calls remain part of a responsible pipeline rather than a one-time validation step.
IntermediateWhy evaluate an agent's full trajectory instead of just its final answer?
An agent can reach a correct final answer through a path that was unsafe, wasteful, or lucky — e.g. it might delete and recreate a file to "fix" something when a safer edit existed, or make ten redundant tool calls to arrive at an answer a well-designed agent would reach in two. Grading only the final output misses all of that. Trajectory evaluation instead checks whether each step was reasonable given what the agent knew at that point — harder than final-answer grading because there's often more than one valid path to a correct outcome, so you're judging reasonableness of the process, not matching against one canonical reference trajectory.
IntermediateExplain indirect prompt injection and why it's hard to defend against.
An attacker embeds instructions inside content the model will process as part of a task — a webpage, a document, an email — rather than sending the model instructions directly. When an agent reads that content (say, to summarize a webpage) and the model doesn't reliably distinguish "text I should reason about" from "text telling me what to do," the embedded instructions can hijack its behavior. It's hard to defend against because the payload is just ordinary, well-formed natural language — there's no malicious syntax pattern to pattern-match against the way there is for SQL injection — so mitigation leans on architectural separation (clearly marking retrieved content as untrusted data, restricting what actions an agent can take based on content it merely read) rather than input filtering alone.
DeepDesign an evaluation pipeline for a new agentic coding assistant feature before launch. Walk through it.
Start with offline trajectory evaluation on a fixed suite of representative coding tasks — did it reach a correct, working solution, and separately, was the sequence of tool calls and edits reasonable (no destructive actions, no wasted steps) along the way. Add adversarial/red-team probes specifically for this feature: prompts trying to get it to run destructive commands, exfiltrate data, or ignore explicit user constraints. Add LLM-as-judge pairwise comparison against the previous version of the assistant on a broad query set, with human spot-checks on a sample. Before full rollout, run a canary with a small percentage of real users, monitoring for escalations, explicit negative feedback, and any destructive-action reports. Post-launch, keep red-teaming continuously rather than treating the pre-launch pass as final — new jailbreak and injection techniques targeting coding agents specifically are an active, evolving area, not a solved problem you test once.
DeepYour LLM-judge pairwise eval shows the new model version winning 65% of comparisons (illustrative), but a few users are reporting the new version feels worse. What do you check?
First, check for judge bias explaining the gap: is the new model simply more verbose, and is verbosity being rewarded rather than controlled for in the rubric? Is the judge from the same family as the new model, creating self-enhancement bias? Second, check the query distribution: the judge's comparison set may not represent the specific query types the complaining users are sending — a model can win in aggregate while regressing on a specific, underrepresented but important slice, which is exactly the subgroup-evaluation problem applied to query types instead of demographics. Third, do a targeted human review specifically on the query types the complaints are about, rather than trusting the aggregate win rate to generalize to every slice equally — an aggregate number, like an aggregate accuracy, can hide a real regression in a subset that matters.
BeginnerWhat do BLEU and ROUGE measure, and how do they differ?
Both compare generated text against human references by n-gram overlap. BLEU was built for translation and is precision-oriented: it asks how much of what the system produced is warranted by the reference, with a brevity penalty so a system cannot win by emitting a few safe words. ROUGE was built for summarisation and is recall-oriented: it asks how much of the reference the system managed to cover, with ROUGE-L using longest common subsequence so word order matters without requiring a contiguous match.
DeepA model paraphrases the reference correctly and scores badly on BLEU. Is the metric wrong?
The metric is measuring what it was designed to measure, which is surface overlap rather than meaning, so this is a known limitation rather than a bug. A correct paraphrase sharing no n-grams scores badly, and a fluent but wrong answer reusing the reference's vocabulary scores well. The practical response is to stop treating it as a quality score and start treating it as a cheap regression check that catches gross breakage deterministically. Reserve judged or human evaluation for decisions that matter, while remembering that an LLM judge costs money per call, drifts when the judge model changes, and has its own biases including a preference for longer answers.
09Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.
Other techniques for this problem
A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.