Free guide · Part of Grokking the Machine Learning Interview · Last updated July 2026
The senior-level differentiator

The ML system design round, decoded.

One repeatable framework — frame, metric, data, features, model, serve, monitor — plus worked designs for recommendation, ranking, search, and fraud. The round that separates a mid-level engineer from a senior one.

FH By Fahim Ul Haq & the Educative ML team
· Last updated July 2026 · ~24 min read
Recent updates
  • 2026-07-27Added online vs offline evaluation metrics distinction — When evaluating ML systems in interviews, distinguish between offline and online metrics. Offline metrics like AUC, precision at k, or NDCG measure model performance on historical data before deployment. Online metrics like click-through rate, engagement time, or conversion rate measure real user behavior in production. Interviewers expect you to discuss both: offline metrics guide model development and catch regressions, while online metrics via AB tests reveal actual business impact. A model with strong offline metrics can still underperform online due to distribution shift or user behavior changes.

The ML system design round is the one that decides your level. Classical ML and deep-learning rounds test whether you know things; the system design round tests whether you can turn an ambiguous business problem into a model that ships, performs, and survives contact with production. It's the closest thing in the loop to the actual job — and it's where the gap between a mid-level engineer and a senior one is most visible.

Here's the trap. Most candidates hear "design a system to recommend videos" and immediately start talking about models: "I'd use a two-tower network, maybe a transformer for sequence modeling…" That's exactly backwards. The interviewer hasn't heard you define what "good" means, what data you have, or how you'd know the thing works. Strong candidates spend the first third of the interview before the word "model" comes up — clarifying the problem, pinning down a metric, and reasoning about data and labels.

This guide gives you the framework that keeps you on the rails under pressure, then walks four canonical designs — recommendation, feed ranking, search, and fraud detection — that cover the overwhelming majority of what you'll be asked. Learn the framework cold and you'll never face a blank page again: every new prompt becomes "apply the loop to this domain."

WHY THIS ROUND EXISTS

It's the only round that tests judgment, not recall.

Every other round has a knowable answer. There's a correct definition of precision, a correct way to derive backprop, a correct output for a coding problem. The ML system design round has no single right answer — and that's the point. The interviewer is watching how you navigate ambiguity, what trade-offs you surface unprompted, and whether your instincts match what production actually demands.

That's why it's the strongest level signal in the loop. A new grad can memorize XGBoost's hyperparameters. Only someone who has shipped models knows to ask "what's the label latency?" before designing the training pipeline, or to flag that a recommendation model trained on logged clicks has a feedback loop baked into its data.

What separates the levels

Signal Mid-level answer Senior / staff answer
Opening move Jumps to a model architecture. Clarifies the problem and constraints first — who's the user, what's the scale, what's the latency budget, what counts as success.
Metrics Names one offline metric (e.g. accuracy). Distinguishes an offline proxy from the online business metric, and explains how they're validated against each other.
Data Assumes clean labeled data exists. Asks where labels come from, how delayed they are, and where leakage and bias could creep in.
Modeling Reaches for the most powerful model. Starts with a simple baseline, justifies each step up in complexity by the cost it buys down.
After launch Stops at "train the model." Designs monitoring, A/B measurement, and a retraining loop — knows the model degrades the day it ships.

THE ONE-LINE TELL

Interviewers can often place your level in the first three minutes. If you ask "what are we optimizing, and how will we know it worked?" before naming a model, you've already signaled seniority. If you open with "I'd use a transformer," you've signaled the opposite — regardless of how good the rest is.

★

The centerpiece

The ML system design framework

Nine stages, always in this order. Walk them out loud and you turn any prompt into a structured conversation instead of a guessing game.

The single most valuable thing you can bring into this round is a repeatable pipeline you walk every time. It does two jobs: it stops you freezing on an unfamiliar prompt, and it signals to the interviewer that you think about ML systems the way practitioners do — end to end, data-first, production-aware. Memorize the order. The mnemonic below is the spine of every answer in this guide.

The ML system design pipeline — walk it top to bottom, every time

1 · Frame problem & constraints 2 · Metrics offline AND online 3 · Data & labels sourcing · labeling leakage 4 · Features engineering · feature store 5 · Model baseline FIRST, then complexity 6 · Train splits · tuning · no leakage 7 · Evaluate offline against the baseline 8 · Serve batch vs online, latency budget 9 · Monitor drift · A/B · retrain iterate — production signal feeds the next model Stages 1–3: before the word "model" Spend the first third of the interview here. Most candidates skip it. Don't. the dark stage is where real systems quietly fail — and where interviewers probe hardest

The memorable walk — nine stages

Below is the ordered walk to narrate. Each stage has the one question a strong candidate asks before moving on.

01 · frame

Clarify the problem & constraints

Restate the goal as an ML problem. Is it classification, ranking, regression, retrieval? Who's the user, what's the QPS, what's the latency budget, what's the cost ceiling? Ask: "What does success look like to the business, and what can't we break?"

02 · metric

Define success metrics — offline and online

Pick an offline metric you can optimize during training (AUC, NDCG, recall@k, MAE) and the online metric the business judges you on (engagement, revenue, retention, fraud loss). Name the gap between them. Ask: "Is my offline metric a faithful proxy for the online one?"

03 · data

Data & labels — sourcing, labeling, leakage

Where do features and labels come from? Are labels explicit (a purchase) or implicit (a click)? How delayed are they? Watch for leakage (a feature that encodes the label) and sampling bias (you only log what the current system shows). Ask: "Would this feature be available at inference time?"

04 · features

Feature engineering

User, item, context, and interaction features. Embeddings for high-cardinality IDs. Decide what's computed in batch (slow-changing) vs real time (session signals). A feature store keeps training and serving consistent. Ask: "How do I avoid training/serving skew on this feature?"

05 · model

Model selection — baseline first

Always start with a simple baseline: a popularity heuristic, logistic regression, gradient-boosted trees. It sets the bar, ships fast, and tells you whether complexity is even worth it. Step up to deep models only when the baseline's ceiling is the bottleneck. Ask: "What's the simplest thing that could beat random?"

06 · train

Training

Split data by time, not randomly, to mimic prediction-on-the-future. Handle class imbalance, pick a loss aligned with the metric, tune honestly on a validation set. Plan for retraining cadence from day one. Ask: "Does my split leak the future into the past?"

07 · evaluate

Evaluation

Compare against the baseline on the offline metric, then slice by segment to catch failures the aggregate hides. A model that wins overall but tanks for new users may be a net loss. Ask: "Where does this model fail, and for whom?"

08 · serve

Serving — batch vs online / real-time

Can predictions be precomputed in batch (daily recommendations) or must they be online (per-request ranking under a 100ms budget)? Online serving forces choices: caching, model size, candidate pre-filtering. Ask: "What's the p99 latency budget, and what fits inside it?"

09 · monitor

Monitoring & iteration

The model starts decaying the day it ships. Monitor prediction distributions, feature drift, and the online metric. Wire an A/B test to measure real impact, and close the loop with retraining. Ask: "How will I know when this model has gone stale?" This stage is where interviewers separate the people who've shipped from the people who've only studied.

the-walk.txt — say this out loud
FRAME   → restate as an ML problem + pin down constraints (QPS, latency, cost)
METRIC  → offline proxy  ⇄  online business metric   (name the gap)
DATA    → sourcing · labels (explicit/implicit, delayed?) · leakage · bias
FEATURE → user / item / context / interaction · embeddings · batch vs real-time
MODEL   → baseline first (heuristic → LR → GBT) → deep only if justified
TRAIN   → time-based split · imbalance · loss aligned to metric · retrain cadence
EVAL    → beat the baseline offline · then slice by segment
SERVE   → batch (precompute) vs online (per-request, latency-bound)
MONITOR → drift · A/B test · close the loop with retraining
01

Canonical design 1

Recommendation system

"Design a system to recommend videos / products / people you may know." The most-asked ML system design prompt — and the cleanest illustration of the two-stage retrieve-then-rank pattern.

You can't score millions of candidate items per request under a latency budget. The universal answer is a funnel: a cheap candidate generation stage narrows millions of items down to a few hundred, then an expensive ranking stage scores just those few hundred precisely. This two-stage shape recurs across reco, feed, search, and ads — internalize it once and it transfers everywhere.

Candidate generation → ranking — the recommendation funnel

Item corpus ~10M+ items videos / products Candidate generation cheap · high recall two-tower embeddings (ANN) collaborative filtering recent / trending / co-view ~500 Ranking expensive · high precision GBT / deep ranker rich user × item features ~20 Re-rank & serve diversity · business rules millions → hundreds → tens — cost concentrates where the list is already short candidate gen optimizes recall; ranking optimizes precision — different jobs, different models

The two stages

Candidate generation (recall)

Reduce millions to hundreds, fast. The dominant technique is a two-tower model: a user tower and an item tower each produce an embedding, and you retrieve the nearest item embeddings to the user via approximate nearest-neighbor (ANN) search. Item embeddings are precomputed; only the user embedding is computed per request. Often blended with collaborative filtering and trending heuristics.

Ranking (precision)

Score the few hundred survivors precisely with a heavier model — gradient-boosted trees or a deep ranker — using rich user×item interaction features you couldn't afford across the whole corpus. Optimize a click/conversion objective, then re-rank for diversity, freshness, and business rules before serving the top ~20.

The cold-start problem

The question the interviewer is waiting for you to raise: what about brand-new users and brand-new items? A new user has no interaction history; a new item has no engagement signal. Collaborative filtering is useless for both. Bring it up unprompted — it's a senior tell.

New users

Fall back to popularity and trending, lightly personalized by whatever you do know — geo, device, referral, a quick onboarding survey. Exploit context until you have behavior.

New items

Lean on content features (text, metadata, image embeddings) so the item has an embedding before anyone has interacted with it. Inject some exploration so new items get exposure to gather signal.

Exploration vs exploitation

Pure exploitation starves new items and creates feedback loops. A small exploration budget (epsilon-greedy or bandits) keeps the catalog fresh and the training data unbiased.

FRAMEWORK CHECKPOINTS

Metric: offline recall@k for candidate gen, NDCG for ranking; online watch long-term engagement, not just CTR (which rewards clickbait). Data: implicit feedback (watches, clicks) with a feedback loop — you only see engagement on what you surfaced. Serving: candidate gen via precomputed ANN index; ranking online under budget. Pitfall: optimizing CTR alone breeds clickbait — pair it with a satisfaction or dwell-time signal.

Related → the embedding and two-tower mechanics live in Deep Learning Fundamentals; the metric trade-off (CTR vs satisfaction) is unpacked in Modeling Decisions.

02

Canonical design 2

Feed & ranking

"Design the news feed / home timeline ranking." Shares the funnel with reco, but the hard part is the objective: a feed has to balance several competing things at once.

The retrieval-then-rank structure carries straight over. What makes feed ranking its own problem is that "good" is not one number. A feed that maximizes clicks alone degrades into clickbait; one that maximizes time-spent can promote outrage; one that ignores recency feels stale. The senior move is to design a multi-objective ranking function and to be explicit about how you weight the pieces.

Designing the objective

The standard pattern: a ranker predicts several engagement probabilities per item — P(click), P(like), P(comment), P(share), P(hide) — and you combine them into a single score with tunable weights. The weights encode product strategy, and negative signals (hides, "see less") get negative weight.

feed_score.py — a multi-objective ranking score
# model heads predict per-action probabilities for each candidate item
score = (
    w_click   * p_click
  + w_like    * p_like
  + w_comment * p_comment
  + w_share   * p_share
  - w_hide    * p_hide        # negative signal → negative weight
) * freshness_decay(age)     # down-weight stale items

# weights are tuned via online A/B tests, not picked offline —
# they encode product strategy, and the right values shift over time

The three tensions

Engagement vs quality

Raw engagement rewards clickbait and outrage. Counterbalance with quality and satisfaction signals (surveys, dwell time, downstream retention) so the feed doesn't optimize itself into a worse product.

Relevance vs freshness

A purely relevance-ranked feed shows the same evergreen hits forever. A freshness decay term and recency features keep it timely — critical for news and social, less so for evergreen content.

Personal vs diverse

Over-personalization creates filter bubbles and fatigue. A diversity penalty (or MMR-style re-ranking) keeps the feed from becoming ten near-identical items in a row.

Two more points that read as senior. First, position bias: items at the top get more clicks regardless of quality, so your training labels are biased by where the old model placed each item — correct for it or you'll bake the bias in. Second, delayed feedback: a share or a follow can arrive hours after the impression, so your label window and training cadence have to account for it.

FRAMEWORK CHECKPOINTS

Metric: a weighted multi-objective offline score; online, the north-star (sessions, retention) plus guardrail metrics (reports, hides). Data: position bias and delayed feedback are the two traps. Serving: per-request ranking, tight budget — a multi-head model amortizes the cost of predicting many actions at once. Pitfall: tuning the weights offline; they only make sense when validated in an A/B test.

04

Canonical design 4

Fraud & abuse detection

"Design a system to detect fraudulent transactions / abusive accounts." A different beast from reco and ranking — defined by extreme class imbalance, a hard latency wall, and an adversary who adapts.

This is the design where naively reaching for accuracy gets you eliminated. If 0.1% of transactions are fraudulent, a model that predicts "never fraud" is 99.9% accurate and completely useless. The first thing out of your mouth should be that accuracy is the wrong metric here — and that this is a precision/recall and cost trade-off, not a guessing game.

The three things that make fraud hard

Extreme imbalance

Fraud is <1% of events. Accuracy is meaningless; optimize precision/recall and PR-AUC. Handle imbalance with class weighting or careful resampling — and frame it as a cost trade-off: a missed fraud (false negative) costs real money; a false positive blocks a legitimate customer.

Hard latency wall

A payment must be approved or declined in tens of milliseconds, inline with the transaction. That rules out heavy models for the synchronous decision. Common shape: a fast inline model for the block/allow call, plus a slower asynchronous model for review queues and investigations.

Adaptive adversary & feedback loops

Fraudsters change tactics the moment you block them — the data distribution shifts because of your model. And you only get labels on transactions you allowed; the ones you blocked never resolve, so the training data is censored. Both demand fast retraining and deliberate label collection.

The feedback-loop trap, made concrete

This is the subtlety that separates strong candidates. When your model blocks a transaction, you never learn whether it was actually fraud — there's no outcome. So you train on a biased sample: only the transactions the model let through. Over time the model becomes confident about a slice of reality it no longer sees. The fix is to deliberately let a small, controlled fraction of borderline cases through (or route them to manual review) so you keep collecting unbiased labels — accepting a little expected loss to keep the model honest.

Two more senior notes. Features in fraud are heavily behavioral and velocity-based — "how many transactions from this card in the last hour," "is this device new for this account," graph features over the entity network. These need a real-time feature store because the signal is the recent behavior. And a pure ML score is rarely the whole system: fraud teams pair the model with hard rules for known-bad patterns and a human review loop for high-stakes, low-confidence cases.

FRAMEWORK CHECKPOINTS

Metric: precision/recall & PR-AUC offline, expressed as a dollar cost (fraud loss vs false-decline cost); online, fraud-loss rate and false-positive rate as guardrails. Data: censored labels (blocked = no outcome) + adversarial drift. Serving: inline fast model (tens of ms) + async deep model for review. Pitfall: optimizing accuracy on an imbalanced problem — it's the classic disqualifier.

Related → imbalance, precision/recall, and PR-AUC are covered in Classical ML Fundamentals; framing the false-positive/false-negative cost as a business decision is in Modeling Decisions.

CROSS-CUTTING CONCERN

Online vs offline, feature stores & training/serving consistency.

Three of the most reliable senior signals don't belong to any one design — they cut across all of them. Bring them up wherever they fit and you'll consistently outscore candidates who only talk about models.

Batch vs online serving

  Batch (precomputed) Online (real-time)
When predictions run Ahead of time, on a schedule Per request, on the hot path
Latency Irrelevant — results are cached Hard budget (often <100ms, <50ms inline)
Freshness of inputs As stale as the last run Uses live session / context signals
Good for Daily reco emails, precomputed user→item lists Feed ranking, search, fraud, anything context-dependent
Cost shape Cheap per prediction, can over-compute unused ones Pay per request, but only compute what's asked for

Many real systems are hybrid: precompute candidate embeddings and slow features in batch, then do the final ranking online with fresh context. Naming that split is a strong answer.

Feature stores & training/serving skew

Training/serving skew is when a feature is computed one way in the training pipeline (batch SQL over a warehouse) and a subtly different way at serving time (streaming code in a service). The model learns one distribution and gets fed another in production — and the failure is silent: offline metrics look great, online performance quietly underperforms. It is one of the most common real-world causes of "the model worked in the notebook but not in prod."

A feature store is the standard fix. It serves the same feature definitions to both training and serving — an offline store materializes features for training, an online store serves them at low latency for inference — so a feature is computed once and reused everywhere. It also gives you point-in-time correctness (no leaking future values into past training rows) and feature reuse across teams.

feature_store.txt — one definition, two paths
              ┌──────────────────────────┐
              │  feature definition (1×)  │   # avg_spend_30d, txns_last_hour …
              └────────────┬─────────────┘
                  ┌────────┴────────┐
        offline store            online store
        (warehouse, bulk)        (low-latency KV)
              │                        │
        ┌─────▼─────┐            ┌─────▼─────┐
        │ TRAINING  │            │  SERVING  │   ← same values, no skew
        └───────────┘            └───────────┘

MEASURING IMPACT

A/B testing — how you prove the model actually helped.

A better offline metric is a hypothesis, not a result. The only way to know your model improved the product is to ship it to a fraction of real traffic and measure. Treating the online A/B test as the source of truth — and offline metrics as cheap proxies you validate against it — is a core senior reflex. Interviewers almost always ask "how would you know this is actually better?" Have the answer ready.

The mechanics

Randomly assign users (not requests) to control and treatment, hold everything else constant, and compare the online metric. Size the experiment for enough statistical power to detect the effect you care about, and run it long enough to clear novelty effects and weekly seasonality.

Guardrails & pitfalls

Track guardrail metrics (latency, error rate, revenue, complaints) so a win on the target metric doesn't hide a regression elsewhere. Watch for peeking (stopping the moment it looks significant inflates false positives) and network effects (treatment leaking into control) in social products.

Two refinements that read as experienced. Roll out gradually — 1% → 10% → 50% → 100% — so a bad model is caught while its blast radius is tiny. And before a full A/B, validate offline with backtesting or replay on logged data to avoid burning live traffic on a model that was never going to win. The offline funnel filters; the A/B test decides.

THE CLOSING-THE-LOOP STORY

A complete answer ties it together: offline metric improves → validate on logged data → ship behind a small A/B → confirm the online metric moves without tripping guardrails → ramp to 100% → keep monitoring for drift → retrain → repeat. Narrate that loop and you've demonstrated you understand ML as a continuous system, not a one-shot model.

PRODUCTION CONSTRAINTS

Scaling, latency budgets & cost.

The constraints are what make it system design. A model that's accurate but blows the latency budget or the infra bill isn't shippable. Senior candidates reason about the budget explicitly — and the funnel architecture you've already drawn is the main tool for living inside it.

The latency budget

State a number early ("let's say a 200ms p99 end to end") and decompose it: retrieval, feature fetch, model inference, re-ranking, network. The two-stage funnel exists precisely because of this budget — cheap retrieval shrinks the candidate set so the expensive ranker only runs on a few hundred items. When inference is still too slow, the standard moves are caching, smaller/quantized/distilled models, precomputing what you can, and batching requests on the GPU.

Scaling the read path

Replicate stateless model servers behind a load balancer; cache hot predictions and embeddings; shard the ANN and feature indices. Most serving systems are read-heavy and scale horizontally.

Cutting inference cost

Distillation (small model mimics big one), quantization, and caching repeated requests. Right-size the model to the budget — the biggest model is rarely the most cost-effective in production.

Cost vs accuracy

Every accuracy gain has a cost: compute, latency, complexity, maintenance. Frame the choice as "is this lift worth the bill?" — a senior framing that mid-level candidates skip entirely.

Related → the pure-infrastructure side of this — load balancing, sharding, caching, replication — is the home turf of Grokking the System Design Interview. In an ML loop you reason about it at a high level; in a generic system design loop it's the main event. Knowing which round you're in tells you how deep to go.

WHAT SINKS CANDIDATES

Common pitfalls.

Most ML system design interviews are lost the same handful of ways. Each maps to a stage of the framework you skipped or rushed. Run the loop in order and you sidestep all of them.

Jumping to a model before defining the metric

The number-one failure. Naming an architecture before you've said what "good" means, or what data exists, signals you think of ML as model-picking rather than problem-solving. Fix: spend the first third on frame → metric → data.

Ignoring where data and labels come from

Assuming clean labeled data appears by magic. Real systems have delayed labels, implicit/biased feedback, leakage risks, and cold-start gaps. Fix: always ask where labels come from and whether each feature exists at inference time.

No baseline — straight to the fanciest model

Reaching for a transformer when logistic regression or GBT would set the bar and ship in a week. Without a baseline you can't tell whether the complexity bought anything. Fix: propose the simplest model first, then justify each step up.

Stopping at "train the model" — no serving or monitoring

Treating the model as the finish line. Production is where models meet latency budgets, drift, training/serving skew, and feedback loops. Fix: always reach stages 8–9 — how it serves, how you measure impact, how it stays healthy.

Optimizing the wrong metric

Accuracy on an imbalanced fraud problem; raw CTR on a feed (clickbait); an offline metric you never tie to the business goal. Fix: distinguish the offline proxy from the online objective and name the gap between them.

See the framework on worked ML designs.

The full course drills the ML system-design round end to end — recommendation, ranking, and search systems designed start to finish, with the trade-offs senior interviewers actually probe.

Take Grokking the Machine Learning Interview

Free trial · worked ML system designs in Python · no credit card required

See the latest promotion on Grokking the Machine Learning Interview