WHY THIS ROUND EXISTS
It's the only round that tests judgment, not recall.
Every other round has a knowable answer. There's a correct definition of precision, a correct way to derive backprop, a correct output for a coding problem. The ML system design round has no single right answer — and that's the point. The interviewer is watching how you navigate ambiguity, what trade-offs you surface unprompted, and whether your instincts match what production actually demands.
That's why it's the strongest level signal in the loop. A new grad can memorize XGBoost's hyperparameters. Only someone who has shipped models knows to ask "what's the label latency?" before designing the training pipeline, or to flag that a recommendation model trained on logged clicks has a feedback loop baked into its data.
What separates the levels
| Signal | Mid-level answer | Senior / staff answer |
|---|---|---|
| Opening move | Jumps to a model architecture. | Clarifies the problem and constraints first — who's the user, what's the scale, what's the latency budget, what counts as success. |
| Metrics | Names one offline metric (e.g. accuracy). | Distinguishes an offline proxy from the online business metric, and explains how they're validated against each other. |
| Data | Assumes clean labeled data exists. | Asks where labels come from, how delayed they are, and where leakage and bias could creep in. |
| Modeling | Reaches for the most powerful model. | Starts with a simple baseline, justifies each step up in complexity by the cost it buys down. |
| After launch | Stops at "train the model." | Designs monitoring, A/B measurement, and a retraining loop — knows the model degrades the day it ships. |
THE ONE-LINE TELL
Interviewers can often place your level in the first three minutes. If you ask "what are we optimizing, and how will we know it worked?" before naming a model, you've already signaled seniority. If you open with "I'd use a transformer," you've signaled the opposite — regardless of how good the rest is.