←Home KnowML
Systems, Safety & InterviewChapter 24

ML Engineering & MLOps

A trained model isn't finished. It has to keep being right while the world under it changes, and MLOps is the machinery that notices when it stops and the discipline that ships new versions safely.

20 min read Assumes: none specific, but training-loop basics (04) help
Start reading
TL;DR

A model's accuracy on the day you shipped it says little about its accuracy six months later, because the world it predicts on keeps moving while the model stays frozen. That gap is why MLOps exists.

In production the work runs as a loop, not a pipeline. Data becomes features, computed through a feature store so training and serving use identical logic. The features train a model, tracked so the run can be reproduced. The model is checked against a golden set before anyone trusts it, and it rolls out gradually so a bad version reaches few users.

Monitoring then watches for drift and decay, and that starts the next lap. If an interviewer asks for one sentence, most of MLOps is answering this: how do we find out our model went stale before our users do, and how do we fix it without breaking anything on the way?

01Intuition

A model isn't a fact about the world that you compute once and keep. It's a bet on a pattern that held in your training window, and the bet gets a little staler every day.

Why ML failures are worse than ordinary bugs

Ordinary software fails loudly: a null pointer throws, a 500 shows up in the logs, someone gets paged. A machine learning system can fail the other way. It keeps returning confident, well-formed, plausible predictions while it is quietly and progressively wrong, and nothing in a standard error log says so.

Two things can slip out from under a working model:

  • The inputs move. The data the model sees in production drifts away from what it trained on, because users change, a competitor launches a promotion, or a pandemic changes shopping behavior overnight.
  • The answer moves. The relationship between inputs and the right answer changes: what counted as fraud last year doesn't cover this year's new pattern.

The model notices neither, and keeps applying last year's pattern with full confidence.

That is why MLOps is a discipline of its own and not just software engineering applied to ML. You need infrastructure whose job is to notice decay that throws no exceptions and causes no crashes. The only symptom is a slow decline in a business metric, and unless you watch for it specifically, it looks like nothing until it has become a serious and expensive problem.

A model in production sees data that slowly stops resembling its training data. One simple way to notice is to compare the two distributions bin by bin.

Try it Detect drift in a feature with the population stability index
import numpy as np

rng = np.random.default_rng(0)
train = rng.normal(50, 10, 10_000)         # values at training time
live_same = rng.normal(50, 10, 10_000)     # live, same distribution
live_drift = rng.normal(58, 10, 10_000)    # live, the mean shifted

def psi(a, b, bins=10):
    edges = np.quantile(a, np.linspace(0, 1, bins + 1))
    edges[0], edges[-1] = -np.inf, np.inf
    pa = np.histogram(a, edges)[0] / len(a) + 1e-6
    pb = np.histogram(b, edges)[0] / len(b) + 1e-6
    return ((pb - pa) * np.log(pb / pa)).sum()

print(f"no drift: PSI = {psi(train, live_same):.3f}")
print(f"drifted:  PSI = {psi(train, live_drift):.3f}")
Live data from the same distribution scores a PSI of 0.001. Live data whose mean has shifted from 50 to 58 scores 0.595. A shift of less than one standard deviation in a single feature moves the score by orders of magnitude. The number itself matters less than having one to alert on.

02Timeline

Before

Early production ML was research code with a cron job: a notebook trained a model, someone copied the weights to a server by hand, nothing recorded how it was made, and monitoring meant noticing user complaints, which broke once there were dozens of models.

→
Innovation

Google's 2015 "Hidden Technical Debt in Machine Learning Systems" paper named the problem, and with the "ML Test Score" rubric that followed it changed the goal to a production system with versioned data, tracked experiments, gated rollouts and continuous monitoring.

→
After

Today's stack is a fairly standard set of layers (feature stores, experiment trackers, model registries, drift monitors, canary infrastructure), and LLMs extended it to versioned prompts and "LLM-as-judge" evaluation (see 25).

03Architecture: the production ML lifecycle as a loop

The naive version is to train once, deploy and forget. Each box below fixes a specific way that goes wrong. The loop closes because monitoring feeds back into retraining, so a pipeline that ends at "deploy" isn't finished.

The MLOps loop — from raw data back to itself
Data & feature store versioned, point-in-time Train tracked experiment reproducible run Evaluate golden set + offline metrics Registry & rollout shadow → canary → full traffic Monitor drift, performance, data/concept decay retrain — triggered by drift, or on a schedule
Feature stores keep training and serving computing features the same way (the single biggest source of train-serving skew otherwise). Experiment tracking makes every run in the "Train" box reproducible. The registry gates what's even eligible to serve traffic, and shadow → canary → full is how a new version earns more traffic instead of getting all of it at once. The dashed line is the loop closing: monitoring is what decides a retrain is needed, not a calendar.

Two ideas carry most of the weight in that loop.

  • A feature store is the system of record for features. The logic that computes "the user's average order value over the last 30 days" for training also computes it for a live request, from the same pipeline, so the model never sees a subtly different version of a feature at serving time than it saw in training.
  • Experiment tracking logs each run's code version, data version, hyperparameters and resulting metrics, so "reproduce the model currently in production" is something you can answer six months later and not an archaeology project.

That feature mismatch has a name, train-serving skew, and it is one of the most common reasons a model scores well offline and then quietly underperforms in production. A feature store is built to rule it out, so you don't have to remember to get it right every time.

04Rolling out a new model without betting the business on it

The obvious way to deploy a new model is to switch all traffic to it and watch. But "watch" hides a lot of unstated work.

If a bad model serves all traffic until someone notices, it can make thousands of bad recommendations, approve thousands of bad transactions or show thousands of users a broken experience. A rollback can't undo a recommendation that was already shown or a loan that was already denied.

The techniques below all limit that damage by letting the new model affect less before you commit.

Batch versus online inference is the first choice.

  • Batch runs predictions on a schedule over a stored dataset: score every user overnight and write the results to a table. It's simple, cheap and tolerant of a slow pipeline, and the predictions are only as fresh as the last run.
  • Online computes a prediction for each request in real time. You need it when the input only exists at request time, such as a live search query or a fraud check on a transaction happening now, and the cost is meeting a latency budget under real traffic.

Then comes the rollout itself, in increasing order of exposure:

  • Shadow deployment runs the new model on live production traffic alongside the current one. It computes predictions but never serves or acts on them, which gives you real traffic and real behavior with no user-facing risk, purely to compare its outputs with the incumbent's.
  • Canary deployment is the next step up. The new model serves a small, real slice of traffic (often 1–5%) end to end, and its effect on business metrics is measured directly before the slice is gradually increased.

Shadow tells you the new model's predictions look reasonable. Only a canary tells you whether users respond to them the way you hoped.

Rollback strategy is the part teams most often forget to design in advance. Reverting a model version is easy when it's just a registry pointer. But if the new version also needed a new feature or a new database schema, then "roll back the model" and "roll back everything the model depended on" are different operations, and an incident is a bad time to find that out.

Human-in-the-loop systems route a model's low-confidence or high-stakes predictions to a person instead of acting on them automatically. You use them where a wrong automated decision is costly, such as content moderation edge cases, large financial transactions and medical triage, trading some automation for a human backstop where errors are most expensive.

They have a second benefit: human decisions on the cases a model is least sure about are disproportionately useful training data for the next version.

05Monitoring: telling drift apart from noise

"Something changed" isn't actionable on its own, because the fix for a shifted input distribution is different from the fix for a changed relationship between inputs and outcomes. So the first job of monitoring is to tell apart three things that all look like "the numbers moved."

  • Data drift is a change in the distribution of the model's input features. Your fraud model was trained when the average transaction was $40 and it's now $65, even though what counts as fraud hasn't changed.
  • Concept drift is a change in the actual relationship between inputs and the correct output. The fraud patterns themselves evolved, so a transaction profile that used to be safe often isn't now.
  • Performance drift is the downstream symptom: accuracy, precision or whatever business metric you track visibly degrading. Data or concept drift eventually causes it, but you can only measure it directly once you have ground-truth labels, which in many production systems arrive late or never at scale.
$$\text{PSI} = \sum_{i=1}^{k} \left( p_i - q_i \right) \cdot \ln\!\left(\frac{p_i}{q_i}\right)$$
  • p_i, q_i the proportion of examples falling into bin i of a feature's distribution, for the current production window (p) versus the original training/reference window (q) — you bucket a continuous feature into k bins to compare two distributions this way.
  • PSI the Population Stability Index — zero when the two distributions are identical, growing as they diverge. It's one of several standard drift statistics (alongside KL divergence and Wasserstein distance), popular because it's simple, symmetric enough in practice, and easy to compute per feature.
  • Common thresholds (industry convention, not a law of nature) PSI below ~0.1 is usually treated as no meaningful shift, 0.1–0.25 as a moderate shift worth investigating, and above ~0.25 as a significant shift worth acting on. Treat these as starting points to calibrate against your own feature's normal volatility, not universal cutoffs.

In practice you can't eyeball this for hundreds of features across every model in a company. Model observability tooling does it for you with dashboards, automated per-feature drift scoring and alerting.

Why a noisy monitor is worse than no monitor

Alert design matters as much as detection. A drift monitor that fires constantly on features with normal seasonal variance teaches everyone to ignore it, so you pay the full cost of the monitor and get none of its benefit, and the team has learned to dismiss the one alert that mattered. This is the ML version of the general alert-fatigue problem.

06Evaluation: golden sets, and offline vs online metrics

A golden set is a fixed, carefully curated, labeled dataset held out to evaluate every candidate model version against the same yardstick. It isn't a random validation split that might itself drift. It's a benchmark you maintain deliberately, so that it means the same thing this quarter as last quarter.

Golden sets are also where a team deliberately overrepresents rare but critical cases that a random sample would barely include, such as a fraud pattern that is 0.1% of traffic but 40% of the financial risk.

Two families of metric answer two different questions:

  • Offline metrics (accuracy, AUC, F1, calibration) are computed against the golden set or a held-out historical dataset before anything ships. They answer "does this model appear to be at least as good as what it's replacing."
  • Online metrics (click-through rate, conversion, revenue per user, actual fraud losses) are measured from real canary traffic. They answer the commercial question: does this model change user behavior and business outcomes in the intended direction?

The two can disagree. A model with a better offline AUC can still do worse online if it's more confidently wrong on the cases users meet most. That gap is why canary deployment exists, as a gate between "looked good offline" and "fully trusted."

07Why this design

Why not just retrain on a fixed schedule (say, every Monday) instead of building drift detection?

A fixed schedule is wasteful or too slow, and usually both at different times. Retraining weekly when nothing has changed burns compute and engineering time for no benefit, while a sudden sharp shift, such as a fraud ring changing tactics overnight, can do real damage for days before the next scheduled retrain.

Drift-triggered retraining spends the budget when it's needed, at the cost of real monitoring infrastructure to know when that is. A fixed schedule is simpler to build and is a reasonable choice when drift is slow and predictable, which is something to verify and not assume.

Why not skip staging and go straight from "passed offline eval" to full production traffic?

Because offline metrics and online outcomes can disagree, and because a golden set, however carefully built, is a fixed snapshot. It can't fully capture live traffic's edge cases, adversarial behavior or genuine novelty. Shadow and canary catch the gap between "looks right on the benchmark" and "behaves correctly on today's real traffic" while catching it is still cheap.

08Complexity, failure modes, and what breaks

Train-serving skew is the recurring problem on this page. Wherever feature computation at training time differs even slightly from serving time, the model looks great offline and is silently miscalibrated the moment it sees production traffic.

The cause is usually small: a different library version, a slightly different time-window definition, a difference in how nulls are handled. A feature store exists to make that hard to do by accident.

Hidden feedback loops appear when a model's own predictions influence the data it is later trained on. A recommender that under-shows a category collects less engagement data for it, which the next training run reads as low demand, reinforcing the original bias with no outside cause.

This is one of the risk factors named in the "Hidden Technical Debt" paper. It is hard to detect from inside the system, because every signal you'd use to check is produced by the loop itself.

Data leakage is information about the label leaking into the features, usually through a feature that is only available after the outcome is known, such as using "did the customer cancel this month" to help predict "will the customer cancel this month."

It inflates offline metrics dramatically and then disappears in production, because the leaked signal isn't available at real prediction time. It is one of the most common reasons a model that looked excellent in evaluation performs close to random once deployed.

Alert fatigue and threshold tuning are an ongoing operational cost and not a one-time setup step. A threshold that's too sensitive teaches engineers to ignore the monitor, and one that's too loose lets real regressions through silently.

Rollback strategy failures are almost always a dependency problem found too late. Rolling back model weights is trivial, but if the new version shipped with a new required feature, a new schema or a new downstream consumer expecting a new field, then "roll back the model" and "restore the previous working system" can be very different amounts of work. Decide that in advance, not during an incident.

Privacy, security and governance add to all of the above. Training data and logged features often contain PII, which limits what a feature store can retain, for how long and who can query it. It also decides whether a user's deletion request has to propagate through every training set that ever touched their data, and not only the live database.

Getting this wrong is a compliance and trust failure, not a performance one, so it has to be designed into the data layer from the start and not retrofitted after an incident.

09Build this

The whole argument for monitoring inputs is that inputs move before accuracy does, and you don't have to take that on faith. You can simulate months of drift in an afternoon and read the lead time off your own plot.

Project Make a drift monitor fire before the accuracy metric does ~2 hours · scikit-learn, CPU only

Train a small classifier, freeze it, then push its input distribution away from the training data one simulated week at a time. The dataset is two overlapping Gaussian blobs, deliberately, so it is simple enough that you can see exactly what you moved and by how much. The model isn't the interesting part. The two curves you plot against each other are.

  1. Generate two overlapping Gaussian classes in 2-D. Train logistic regression on a reference window, then freeze it, because production models don't keep learning.
  2. Simulate 30 weekly windows. In each one, sample fresh data with one feature's mean nudged a little further from where it started.
  3. Implement PSI by hand from the equation in section 05. Bin the reference window, freeze those bin edges, and score every later window against them.
  4. Record accuracy per window too, pretending its labels only arrive at the end of that window. That lag is the realistic part.
  5. Plot PSI and accuracy on one shared time axis. Mark the conventional 0.1 and 0.25 PSI lines from section 05.
  6. Now break the assumption. Hold the input distribution completely still and move the labels instead, rotating the true decision boundary a little each week. Rerun both curves.
You'll know it worked when the PSI curve lifts off its floor and crosses a threshold while the accuracy curve is still flat inside its own noise. Measure the gap between the two crossings in windows. That gap is the lead time, and it is the entire commercial argument for input monitoring, stated in a unit you just measured rather than asserted.
What the breakage teaches. With the inputs frozen and only the labels moving, PSI stays pinned to the floor while accuracy falls away underneath it. The monitor is not broken. It is blind by construction. PSI watches $p(x)$, and what changed was $p(y \mid x)$. This is why the page names three kinds of drift instead of one. Input monitoring buys you lead time on data drift and nothing whatsoever on concept drift, so it complements label-based monitoring rather than replacing it.

Where this runs in production

A team retrains a fraud-detection model with a new feature set. It runs in shadow for a week first: real transactions, real predictions computed, nothing acted on. They compare its flagged transactions with the current model's and with the chargebacks that arrive during that week.

Satisfied that it isn't obviously worse, they move to a 2% canary: real users and real blocking decisions, with the exposure capped. They track online metrics, such as the false-positive rate on real customers and the fraud losses prevented, against the incumbent's on matched traffic.

Only once the canary clears does traffic ramp to 25% and then 100%. When the model is fully live, per-feature drift monitoring on its inputs is what will eventually trigger the next retrain, which closes the loop back to where this example started.

10Interview questions

BeginnerWhat's the difference between data drift and concept drift?

Data drift is a change in the distribution of the model's input features — the inputs it sees look statistically different from training, even though the true relationship between inputs and outputs hasn't changed. Concept drift is a change in that underlying relationship itself — the same input now genuinely maps to a different correct output than it used to. They can happen independently or together, and they call for different fixes: data drift might just need feature recalibration, while concept drift usually means the model's learned pattern is actually outdated and needs retraining on newer labeled data.

BeginnerWhy does a feature store help prevent bugs that a normal database wouldn't?

A feature store's core guarantee is that the exact same transformation logic computes a feature for both training (often over historical, point-in-time-correct data) and online serving (a single live request), from one shared definition. Without it, teams often end up with two independently written implementations of "the same" feature — one in a training notebook, one in serving code — that drift apart in small ways over time, producing train-serving skew that's very hard to detect because both implementations look reasonable in isolation.

IntermediateWalk through the difference between shadow deployment and canary deployment, and when you'd use each.

Shadow deployment runs the new model on live traffic in parallel with the current one, computing predictions but never acting on them or showing them to users — zero business risk, used to sanity-check that the new model's outputs look reasonable on real traffic before trusting it with anything. Canary deployment goes further: the new model actually serves a small percentage of real traffic end to end, with real user impact measured on real business metrics. You'd use shadow first, as a cheap initial check, and canary next, because only canary can reveal whether users actually respond differently to the new model's live decisions — something shadow's non-acting predictions can't show you.

IntermediateA model's offline evaluation metrics look great, but online performance is worse than the model it replaced. What are the likely explanations?

Several candidates, in roughly the order worth checking: train-serving skew (a feature computed differently online than in the offline eval), the golden/validation set not representing current live traffic (drift between when it was built and now, or it never covered some real-world segment well), data leakage inflating the offline number artificially, or a genuine offline/online metric mismatch — a model can have a better AUC while being more confidently wrong on exactly the cases that matter most for the real business metric. Isolate by first confirming the offline features and serving features are computed identically, then checking whether the eval set is still representative of current traffic.

DeepDesign the monitoring and retraining strategy for a recommendation model at a fast-growing consumer product. What do you actually build, in order?

Start with per-feature drift monitoring (PSI or similar) on the model's key inputs, tuned against each feature's own normal seasonal variance so alerts are meaningful rather than constant. Add online metric tracking (CTR, conversion, session length) segmented by user cohort, since aggregate metrics can hide a regression that's concentrated in one segment. Watch specifically for feedback-loop symptoms — categories or items whose exposure is shrinking over time independent of any external signal, which suggests the model is reinforcing its own past decisions rather than reflecting real demand. Wire drift and performance alerts to trigger a retraining pipeline automatically for routine cases, but keep a human in the loop before any retrained model actually promotes past canary, since a bad automated retrain triggered by a false-positive drift alert is its own failure mode worth guarding against.

DeepYour canary is passing all offline and online metrics, but a rollback still causes an incident. What went wrong in the rollout design, and how do you fix it?

Most likely, the new model version wasn't actually independent of everything around it — it may have shipped alongside a new feature-store column, a new schema, or a downstream consumer that started depending on a new output field, so "roll back the model" silently left those dependent changes in place, breaking the previous version's assumptions. The fix is designing rollback as a first-class deliverable of every rollout, not an afterthought: version the model together with its exact feature schema and any downstream contract changes, so reverting the model reverts the whole unit, and test the rollback path itself before the rollout ships, the same way you'd test the rollout.

11Go deeper

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

Other techniques for this problem

A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.

My Notes — 24 ML Engineering & MLOps

Free notes

Highlights on this page