Free guide · Part of Grokking the Machine Learning Interview · Last updated July 2026
Production & MLOps

The MLOps questions that quietly sink candidates.

Anyone can train a model on a static dataset. The interview signal is whether you can ship one and keep it healthy — pipelines, registries, deployment patterns, monitoring, drift, and retraining.

FH By Fahim Ul Haq
· Last updated July 2026 · ~19 min read

Most candidates over-prepare for the part of the job they'll spend the least time on. You can train a beautiful model in a notebook — but in a real ML role, the model that ships is maybe 5% of the work. The other 95% is everything around it: getting clean data in reliably, versioning what you trained on, deploying without breaking the live system, and knowing the moment the model starts to rot. That last part is where production ML actually lives, and it's exactly what interviewers probe when they want to separate "can fit a model" from "can own a model in production."

This is why the MLOps round — sometimes its own loop, sometimes folded into ML system design — is where strong-on-paper candidates lose points. They nail the model architecture, then go blank on "how would you know this model is still working three months from now?" The honest answer involves data drift, concept drift, training-serving skew, and a monitoring stack that catches silent failures before users do. None of it is glamorous. All of it is testable.

This guide is the production-side map: the ML lifecycle, training pipelines and reproducibility, feature stores, the model registry, the deployment patterns interviewers expect you to name, and the centerpiece — monitoring and drift. Pair it with ML System Design (the serving half of the same coin) and Modeling Decisions (how to justify the trade-offs you make under production constraints).

01

Foundations

The ML lifecycle

Traditional software is deterministic — the same input gives the same output forever. ML systems decay even when the code never changes, because the world they model keeps moving.

The single most important framing for the whole MLOps round: an ML system is code plus data plus a model, and all three drift. In traditional software, you write logic, test it, ship it, and it behaves identically until you change it. An ML model is different — its behavior is learned from a snapshot of data, and the moment that snapshot stops matching reality, the model gets quietly worse without anyone touching a line of code. This is why "it passed all the tests at deploy" means much less for ML than it does for a backend service.

The lifecycle is also a loop, not a line. You don't ship a model and walk away; you ship it, watch it, and feed what you learn back into the next training run. Interviewers love to see candidates describe this loop explicitly — problem framing, data collection, training, evaluation, deployment, monitoring, and back to retraining — because it shows you understand that the work doesn't end at deployment. It begins there.

Traditional software — linear, deterministic Code Test Deploy Stable forever ML system — a loop that decays without intervention Frame & data Train Evaluate Deploy Monitor drift detected → retrain → redeploy

What changes vs traditional software

Behavior is learned, not written

You can't read the source and predict the output. Behavior depends on training data you may not fully control — so testing means validating data and predictions, not just code paths.

It decays on its own

A frozen model loses accuracy as the input distribution shifts. "No code changed" is not a guarantee of correctness. Performance is a moving target.

Three things to version

Code, data, and model artifacts must all be versioned together to reproduce a result. Versioning code alone — the default in software — is not enough.

INTERVIEW SIGNAL

When asked "how is deploying an ML model different from deploying a service?", the answer interviewers want is the data dependency and the decay: the model's correctness is tied to a data distribution that drifts, so deployment is the start of an ongoing monitoring-and-retraining loop, not the end of the project. Candidates who treat a model like a static binary are the ones who get marked down.

02

Build it once, rebuild it anytime

Training pipelines & reproducibility

If you can't reproduce a model — same data, same code, same result — you can't debug it, can't roll back to it, and can't trust it. Reproducibility is the foundation everything else sits on.

A training pipeline turns a one-off notebook into a repeatable, automated process: ingest data, validate it, transform features, train, evaluate, and register the resulting model. The reason this matters in interviews isn't tooling trivia — it's that a pipeline is what makes a model reproducible and auditable. Six months from now, when a stakeholder asks "why did the model make this decision?", you need to point to the exact data, code, and hyperparameters that produced it. A notebook on someone's laptop can't answer that question.

The two pillars are data versioning and experiment tracking. Data versioning (DVC, LakeFS, or a versioned table format like Delta/Iceberg) pins the exact dataset a model was trained on, because "the training data" silently changes as new rows arrive. Experiment tracking (MLflow, Weights & Biases) logs every run's hyperparameters, metrics, code commit, and artifacts, so you can compare runs and reproduce the winner. Together they let you answer the only question that matters when a model misbehaves: what exactly went into this?

train.py — a run you can reproduce
# Pin every input so the run is reproducible
data = dvc.get("features/train@v7")      # versioned dataset
set_seed(42)                              # seed numpy / torch / cuda

with mlflow.start_run():
    mlflow.log_params({"lr": 3e-4, "depth": 8})
    mlflow.log_param("data_version", "v7")
    mlflow.log_param("git_sha", current_commit())

    model = train(data)
    mlflow.log_metric("val_auc", evaluate(model))
    mlflow.sklearn.log_model(model, "model")   # artifact + lineage

The four things you must pin

Data version

A content hash or snapshot ID of the exact training set. "We trained on the prod table" is not reproducible — the table changed yesterday.

Code version

The git SHA of the training and feature code. Feature logic is code too — a changed transform is a changed model.

Hyperparameters & config

Learning rate, architecture, preprocessing flags. Logged automatically by the tracker, not remembered.

Environment & seed

Library versions (a pinned lockfile or container image) and the random seed. GPU nondeterminism is a real source of "I can't reproduce it."

INTERVIEW SIGNAL

"How would you make a training run reproducible?" is a fast credibility check. Strong answer: version the data, pin the code commit, log hyperparameters and the environment, and fix random seeds — then store all four with the model artifact so any past model can be rebuilt or rolled back. Mentioning that feature-engineering code must be versioned alongside model code is the detail that signals real production experience.

03

Offline / online parity

Feature stores & data pipelines

A feature store exists to solve one nasty problem: making sure the features you train on are computed exactly the same way as the features you serve. When they diverge, your model silently breaks.

Here's the failure mode a feature store prevents. During training, you compute features in a batch job over historical data — say, a user's average order value over the last 30 days. At serving time, you need that same feature computed live, on the fly, for a single user. If the two code paths differ even slightly — a different rounding rule, a different time window, a timezone bug — the model sees different inputs in production than it saw in training. The model didn't change; the inputs did. This is training-serving skew, and it's one of the most common silent killers in production ML.

A feature store gives you a single definition of each feature, materialized into two synchronized stores: an offline store (a data warehouse holding historical feature values for training, with point-in-time correctness) and an online store (a low-latency key-value store like Redis or DynamoDB for serving). You define the feature once; the store guarantees both paths use the same logic. It also enables feature reuse across teams and point-in-time joins that prevent label leakage — fetching feature values as they were at prediction time, not as they are now.

Feature definition avg_order_value_30d Offline store Warehouse · historical values point-in-time joins · training Online store Redis / DynamoDB · low latency single-row lookup · serving Training pipeline Real-time serving same logic → no skew

Why interviewers ask about this

It solves

Training-serving skew (one definition, two stores), feature reuse across teams, and point-in-time correctness so a training row only sees feature values that existed before the label. That last point prevents label leakage — a subtle bug that inflates offline metrics and collapses in production.

When it's overkill

A single model with batch-only scoring and a handful of features doesn't need a full feature store — the operational overhead isn't worth it. Reach for one when you have multiple models, real-time serving, or many teams sharing features. Naming this trade-off, rather than reflexively saying "use a feature store," signals judgment.

CONNECTS TO

The data and feature layer is the foundation of the ML System Design round — when you sketch "data → features → model → serve," the feature store is what makes the features box trustworthy. And the low-latency online store is a classic AI-infrastructure design problem in its own right.

04

The system of record for models

Model registry & versioning

The registry is the single source of truth for "which model is where." It's the bridge between a training run that produced an artifact and a serving system that needs to know what to load.

Once you're training models repeatedly, you accumulate artifacts fast — and "the latest model" stops being a meaningful phrase. A model registry is a versioned catalog that tracks every model, its version, its lineage (the data and code that produced it), its evaluation metrics, and its stage: staging, production, or archived. It's the inventory system that lets a deployment ask "give me the current production version of the fraud model" and get a precise, auditable answer.

The most valuable thing a registry buys you in interviews is the promotion workflow and instant rollback. A new model is registered, evaluated, promoted to staging for shadow testing, then promoted to production once it clears the bar. If it misbehaves, you flip the production pointer back to the previous version — no retraining, no scramble. That rollback story is exactly what interviewers want to hear when they ask "the new model is degrading in production at 2am, what do you do?"

Model lifecycle through the registry v9 · archived v10 · archived v11 · prod v12 · staging Registered Staging Production Serving rollback: flip prod pointer to v10 — instant, no retrain

INTERVIEW SIGNAL

A registry answers two questions interviewers love: "how do you roll back a bad model?" (flip the production-stage pointer to the prior version — seconds, not a retrain) and "how do you know what's running in production?" (the registry is the source of truth, with full lineage back to the data and code). If you can describe the staging → production promotion gate plus shadow evaluation before promotion, you sound like someone who has actually shipped models.

05

How the model reaches users

Deployment patterns

"How would you deploy this model?" is a near-guaranteed question. The answer starts with matching the serving mode to the latency and freshness the use case actually needs — then de-risking the rollout.

There are three serving modes, and choosing the right one is the first decision. Batch (offline) inference scores a large set of records on a schedule and writes predictions to a table — perfect when predictions don't need to be fresh-to-the-second (nightly churn scores, weekly recommendations). It's the simplest and cheapest to operate. Online (real-time) inference serves predictions synchronously behind an API, on the order of milliseconds, when the input only exists at request time (fraud check on a transaction, ranking a feed on page load). Streaming inference scores events continuously off a message bus (Kafka, Kinesis) when you need near-real-time reactions to a flow of events without a blocking request.

The second decision is how to roll out a new version safely. You almost never flip 100% of traffic to a fresh model. Interviewers expect you to name the de-risking patterns: shadow deployment (the new model receives real traffic but its predictions are logged, not served — pure observation, zero user risk) and canary release (send a small slice of traffic, say 5%, to the new model, watch the metrics, then ramp up gradually). Blue-green is the related infra pattern: two identical environments, switch traffic over, keep the old one warm for instant rollback.

Mode Latency When to use Example
Batch minutes–hours Predictions don't need to be fresh; score everything on a schedule Nightly churn scores, weekly recs
Online < 100 ms Input exists only at request time; user is waiting Fraud check, feed ranking, search
Streaming seconds React continuously to a flow of events, no blocking request Real-time anomaly detection on a stream

De-risking a rollout

Shadow deployment

New model gets real traffic; predictions are logged, not served. Compare against the live model offline. Zero user risk — the gold standard for validating a new model on production traffic before it touches a user.

Canary release

Route a small slice (1–5%) of live traffic to the new model, watch guardrail metrics, then ramp. Limits blast radius — if it's bad, only a few users saw it before you roll back.

Blue-green

Two identical environments. Cut traffic over to "green," keep "blue" warm. Instant rollback by switching back. The infra companion to canary and shadow.

INTERVIEW SIGNAL

The crisp distinction interviewers reward: shadow tests correctness without user risk (predictions logged, not served), while canary tests real impact on a small slice (predictions served to a few users, metrics watched). Most strong answers use both: shadow first to confirm the model behaves, then canary to confirm it actually moves the metric — before a full ramp. And always have a rollback plan ready.

06

The centerpiece of the MLOps round

Monitoring & drift

This is the section that decides the MLOps round. A model degrades quietly — no error, no exception, no page. The whole game is detecting that degradation before your users (or your revenue) do.

Traditional monitoring tells you whether the service is up: latency, error rate, throughput. ML monitoring has to answer a harder question — whether the model is still good — and the cruel part is that a model can be returning 200s with perfect latency while its predictions quietly turn to garbage. There's no stack trace for "the model is wrong now." You have to instrument it deliberately, on multiple layers, because the failure is statistical, not exceptional.

There are three layers worth monitoring, from easiest to detect to most valuable. Operational metrics (latency, errors, throughput) — same as any service. Data and prediction monitoring — the distribution of inputs and outputs, which you can watch in real time without labels. Model performance — accuracy, AUC, business KPIs, which require ground-truth labels that often arrive with a delay (sometimes days or weeks). Because labels lag, you lean on input/output drift as an early warning that something is wrong before the performance metrics confirm it.

The three kinds of drift

Data drift (covariate shift)

The input distribution P(X) changes — a new user segment, a feature whose range shifts, an upstream pipeline change. The relationship is intact, but the model now sees inputs it wasn't trained on. detect: PSI / KL-divergence on features

Concept drift

The relationship P(Y|X) changes — the same inputs now map to different outcomes. Fraudsters change tactics; user taste shifts. This is the dangerous one: the world's rules changed, so even perfectly-distributed inputs get wrong predictions. detect: performance decay vs labels

Training-serving skew

Not time-based at all — the serving inputs differ from training inputs from day one because of a pipeline mismatch (the feature-store problem). The model was wrong out of the gate, not degraded over time. detect: compare train vs serve feature stats

Performance decay & drift detection over time

model accuracy → time since deploy → alert threshold (SLA) healthy data drift slow input shift ① input drift PSI spikes — early warning concept drift event ② breach accuracy crosses SLA → page ③ retrain triggered on breach recovered ④ redeploy fresh model back above SLA input drift leads the performance breach — which is why you alert on distributions, not just accuracy

What to actually monitor — and how to alert

Input distributions

Per-feature stats (mean, variance, null rate, category mix). Quantify drift with PSI (population stability index) or KL-divergence against a training baseline. No labels needed — your earliest signal.

Prediction distributions

Watch the output distribution shift (e.g., the model suddenly predicts "fraud" 3× more often). A cheap, label-free proxy that often catches problems before ground truth arrives.

Performance vs ground truth

Accuracy, AUC, calibration, and the business KPI — computed once labels land. The ground truth, but delayed. Bridge the lag with the distribution signals above.

Alerting that isn't noise

Threshold + duration (don't page on a single blip), segment-level alerts (catch a regression in one slice), and route to a human with context. A monitor nobody acts on is decoration.

INTERVIEW SIGNAL

"How would you know your model is still working in production?" is the question that separates the field. The answer that lands: monitor on three layers — operational (latency/errors), data & prediction distributions (PSI/KL, label-free, early warning), and true performance once labels arrive — and name the drift type (data vs concept vs training-serving skew) because the fix differs. The killer detail: because labels lag, you alert on input drift as a leading indicator rather than waiting for accuracy to confirm the damage.

07

Closing the loop

Retraining strategies

Once you've detected drift, the obvious answer is "retrain." The interesting interview question is when and how — and how you avoid making things worse.

There are two triggers for retraining, and the right choice depends on how predictable your drift is. Scheduled retraining runs on a fixed cadence — nightly, weekly, monthly — and is simple and predictable. It fits domains where drift is gradual and steady. Triggered retraining fires when monitoring detects a problem: drift crosses a threshold, or performance drops below an SLA. It's more responsive and avoids retraining when nothing has changed, but it requires the monitoring infrastructure from the previous section to exist. The most mature setups combine both — a regular cadence as a baseline, plus a trigger for sudden shifts.

The trap interviewers probe is treating retraining as a no-op. It isn't. A retrained model is a new model and must go through the same gate as any deployment: train on fresh data, evaluate against a held-out set and against the current production model, shadow or canary test it, and only promote if it actually beats what's live. Blindly auto-deploying every retrain is how teams ship regressions. And retraining can't fix everything — if your labels are drifting or your feature pipeline is broken, a fresh model trained on bad data just relearns the problem.

Scheduled Triggered
Fires on A fixed clock (nightly/weekly) A drift or performance alert
Best for Gradual, predictable drift Sudden shifts, unpredictable drift
Pro Simple, predictable, easy to reason about Responsive; no wasted retrains
Con May retrain when nothing changed, or lag a sudden shift Needs solid monitoring to fire correctly

INTERVIEW SIGNAL

When asked "how often would you retrain?", resist a number. The strong answer ties cadence to how fast the data drifts and how costly staleness is, proposes a scheduled baseline plus a drift-triggered path, and — critically — insists that every retrained model is re-validated against the current production model before promotion. Mentioning that you'd guard against silently retraining on corrupted or drifting labels shows you've seen retraining go wrong.

08

Proving the model actually helps

Experimentation in production

Offline metrics improving doesn't mean the business metric improves. The only way to know a new model is actually better is to test it on real users — carefully.

A model with higher offline AUC can still lose in production — the offline metric may not align with what users do, or the gain may not survive contact with real traffic. So mature ML teams treat every model change as a hypothesis to be tested live. The default tool is the A/B test: randomly split users into control (current model) and treatment (new model), run long enough to reach statistical significance, and compare the business metric. The subtleties interviewers probe are the statistical ones — sample size and test duration, the multiple-comparisons problem, and novelty effects where a change looks great for a week purely because it's new.

Two refinements are worth naming. Interleaving is a more sensitive technique for ranking systems: instead of showing user A list 1 and user B list 2, you blend both rankers' results into one list for the same user and see which ranker's items get more clicks. Because it controls for user-level variance, it needs far less traffic to detect a difference. And guardrail metrics are the safety net: secondary metrics (latency, revenue, error rate, user complaints) you watch to make sure a win on the primary metric isn't quietly tanking something else. A recommender that boosts clicks but cuts revenue failed — the guardrail catches it.

Live traffic randomized split Control · 50% current production model Treatment · 50% new candidate model Compare metric significance test guardrails: latency · revenue · error rate must not regress

INTERVIEW SIGNAL

The point that separates senior answers: offline metrics are a filter, not a verdict — you ship behind an A/B test and judge on the business metric, with guardrails so a primary-metric win can't silently hurt latency or revenue. Bonus signal: knowing interleaving for ranking (controls user variance, needs far less traffic) and flagging novelty effects and sample-size/duration as the things that invalidate a naive test.

09

Shipping safely, serving affordably

CI/CD, cost & latency

CI/CD for ML is CI/CD for software plus two extra things to test — the data and the model — and a third pipeline most engineers forget: continuous training.

The MLOps community talks about CI/CD/CT: continuous integration, continuous delivery, and continuous training. CI for ML doesn't just run unit tests on code — it also validates data (schema checks, distribution checks, null-rate thresholds) and validates the model (does it clear a minimum metric bar? does it beat the current production model on a held-out set?). CD automates the promotion-and-deploy path with the safe-rollout patterns from Section 5. CT — continuous training — is the automated pipeline that retrains, re-validates, and re-registers on the schedule or trigger from Section 7. Naming all three, and noting that ML adds data and model tests on top of normal code tests, is the clean way to answer "what does CI/CD look like for ML?"

The other thing interviewers probe is cost and latency trade-offs, because a model that's accurate but too slow or too expensive doesn't ship. The levers worth knowing: model compression (quantization, pruning, distillation to a smaller student model) to cut latency and serving cost; batching requests to raise GPU throughput at the cost of a little latency; caching predictions for repeated inputs; and hardware choice (CPU vs GPU vs accelerator) matched to the latency budget. The framing that lands: serving cost and latency are first-class design constraints, not afterthoughts — sometimes the right answer is a slightly less accurate model that's 10× cheaper to serve.

CI/CD/CT at a glance

CI — integrate

Unit tests on code plus data validation (schema, distributions, nulls) plus model validation (metric bar, beats current prod). The two extra gates ML adds.

CD — deliver

Automated promotion through the registry stages and a safe rollout (shadow → canary → ramp) with one-flip rollback. Software CD plus ML-aware deployment.

CT — train

The pipeline that retrains, re-validates, and re-registers automatically on a schedule or drift trigger. The piece unique to ML — software has no equivalent.

INTERVIEW SIGNAL

Two phrases earn credit here. "CI/CD plus CT" — and the point that ML CI tests data and the model, not just code. And treating latency and serving cost as design constraints — reaching for quantization, distillation, batching, or caching, and being willing to trade a little accuracy for a large efficiency win when the budget demands it.

COMMON PITFALLS

The production mistakes that fail candidates.

Most MLOps interview misses aren't exotic — they're the same handful of blind spots. Recognizing them in your own answer is half the battle.

Pitfall 1

"Deploy and walk away" — no monitoring

Treating deployment as the finish line. Without monitoring, a degrading model fails silently — no error, no alert — until a business metric tanks weeks later. Always describe the post-deploy loop: monitor distributions and performance, alert, retrain.

Pitfall 2

Training-serving skew

Features computed one way in the batch training job and a different way in the live serving path. The model looks great offline and underperforms in production — not because it degraded, but because it never saw those inputs. A feature store with one definition is the fix.

Pitfall 3

Silent label drift & data leakage

Labels arrive late or quietly change meaning, so you retrain on a moving target and don't notice. Or a feature leaks future information (a non-point-in-time join), inflating offline metrics that collapse live. Both are invisible without explicit checks.

Pitfall 4

No reproducibility

A model nobody can rebuild — unversioned data, unlogged hyperparameters, a notebook on a laptop. You can't debug it, roll back to it, or audit its decisions. Version data, code, config, and environment together.

Pitfall 5

Auto-deploying every retrain

Shipping each retrained model straight to production with no gate. A bad retrain (corrupted data, a drifting label) becomes an instant regression. Always re-validate against the current production model before promotion.

Pitfall 6

Trusting offline metrics alone

Promoting a model because its offline AUC went up, without an A/B test or guardrails. Offline gains don't always survive contact with users — and a primary-metric win can quietly hurt revenue or latency. Test live, with guardrails.

See MLOps in real ML system designs.

The full course walks the production side end to end — pipelines, monitoring, drift, and retraining — woven into worked ML system designs for recommendation, ranking, and search, with the trade-offs senior interviewers actually probe.

Take Grokking the Machine Learning Interview

Free trial · 25+ lessons · no credit card required

Recent updates

See the latest promotion on Grokking the Machine Learning Interview