Free guide · Part of Grokking the Machine Learning Interview · Last updated September 2026
The interview loop · 2026

The modern ML interview loop, round by round.

Coding, breadth & theory, ML system design, the project deep-dive, and behavioral — what each round actually tests, how loops differ by role and company, and a prep plan that fits the loop instead of fighting it.

FH By Fahim Ul Haq
· Last updated September 2026 · ~20 min read
Recent updates
  • 2026-09-19Added behavioral round coverage for ML interviews — Many ML interview loops now include a behavioral round focused on collaboration, handling ambiguity, and stakeholder communication. This round tests whether you can work cross-functionally with product managers, explain model decisions to non-technical partners, and navigate tradeoffs when business goals conflict with technical constraints. Interviewers look for concrete examples using the STAR format and often probe how you handled failed experiments or scope changes mid-project.
  • 2026-07-22Added note on take-home ML assignments between interview stages — Some companies insert a take-home ML assignment between the phone screen and onsite loop. These typically run three to seven days and ask you to build a small model, document your approach, and justify design choices. Interviewers assess code quality, clarity of explanation, and whether you considered trade-offs like runtime, interpretability, or fairness. If your loop includes one, treat it as a fifth round: your write-up often becomes the basis for follow-up questions in the project deep-dive.

Most candidates prepare for an ML interview as if it were one undifferentiated wall of "machine learning." They study a bit of everything, walk in, and discover the loop is actually five or six distinct rounds — each with its own rubric, its own failure modes, and its own idea of what a strong answer looks like. The coding round does not care how well you can derive backprop. The system-design round does not care whether your Python is idiomatic. The breadth round rewards exactly the rapid recall that the system-design round punishes if you lead with it.

The unlock is simple to state and hard to internalize: know which round tests what, and calibrate to it in real time. The same sentence — "I'd use a transformer here" — is a great answer in the breadth round and a red flag in the system-design round, where jumping to a model before framing the problem signals inexperience. Strong candidates read the room and switch modes; weak ones run the same playbook through every round and wonder why their strongest skill keeps landing flat.

This guide walks the modern (2026) loop end to end. For each round you'll get what it tests, what a good answer looks like, and the mistakes that quietly sink otherwise-qualified people. Then we cover how the loop changes by role — ML engineer vs. applied scientist vs. data scientist vs. research — and by company, from big-tech rubrics to the frontier-lab style at places like OpenAI and Anthropic. We close with a concrete prep plan keyed to the loop, not to a generic "study ML" instinct.

00

Orientation

The ML loop, end to end

A typical loop is a recruiter screen, a technical phone screen, and then a four-to-six-round onsite — each round a different lens on whether you can build, reason about, and ship ML systems.

The shape is remarkably consistent across companies, even as the content shifts. After a recruiter conversation, you'll usually do one technical phone screen — most often a coding problem, sometimes a quick ML-concepts conversation — and if that goes well, an onsite (now usually virtual) of four to six back-to-back rounds. The onsite is where the loop fans out into its distinct components: coding, breadth/theory, ML system design, a deep-dive on your past work, and behavioral. Senior loops add rounds and raise the bar on system design; junior loops compress and lean harder on coding and fundamentals.

The single most useful mental model: each round is a different question about the same candidate. Can you write correct code under time pressure (coding)? Do you actually understand the models you name-drop (breadth)? Can you take an ambiguous product goal and design an ML system around it (system design)? Have you done real work, and can you reason about the decisions you made (deep-dive)? And will people want to work with you (behavioral)? A no in any one round can sink the loop — which is why "I'm great at modeling" is not a strategy.

A typical ML interview loop Recruiter screen fit · level · logistics Phone screen coding / ML concepts ONSITE · 4–6 rounds 1 · Coding DSA in Python 2 · Breadth ML + DL theory 3 · ML system design the differentiator 4 · Deep-dive your past work 5 · Behavioral collaboration Debrief & hire/level decision one weak round can sink the loop

What each round is really asking

RoundThe literal taskThe real question
CodingSolve a DSA or data-manipulation problem in PythonCan you write correct, clean code under pressure?
Breadth & theoryAnswer rapid-fire ML / DL concept questionsDo you actually understand what you use?
ML system designDesign an end-to-end ML system for a product goalCan you turn ambiguity into a shippable system?
Project deep-diveWalk through something you builtDid you do real work — and own the decisions?
BehavioralTell stories about past collaboration & conflictWill people want you on the team?

THE UNLOCK

Treat the loop as five separate exams, not one. Before each round, name to yourself which question it's really asking and switch modes. The candidates who fail despite strong fundamentals almost always run their favorite mode — usually deep modeling talk — through every round, including the ones that punish it.

01

Round 1

The ML coding round

Yes, ML roles still have a coding round. It's DSA in Python — sometimes with an ML twist: implement a metric, vectorize a computation, or manipulate a dataset without reaching for a library.

The most common surprise for ML candidates is that the coding bar is real and it's in Python. For most ML and applied-science loops it's a standard data-structures-and-algorithms problem — arrays, strings, hashing, trees, occasionally dynamic programming — graded the same way a software engineer would be graded: correctness, clarity, complexity, and how you handle edge cases. The pattern library you'd use here is exactly the one from the coding-interview world, which is why most candidates pair this round's prep with Grokking the Coding Interview for the algorithm patterns and just translate them to Python.

The ML-flavored variant shows up more at applied-science and DS-leaning loops. Instead of (or alongside) a pure algorithm problem, you might be asked to implement a metric from scratch (precision/recall, AUC, log-loss), code a simple algorithm (k-means assignment step, a single gradient-descent update, k-NN classification), or do data manipulation — group-bys, joins, rolling windows — sometimes in raw Python, sometimes in pandas/NumPy. The skill being tested is whether you understand the math well enough to turn it into correct code, and whether you can write vectorized, readable Python rather than nested loops.

// A representative "implement this metric" prompt

# Implement precision@k from scratch — no sklearn. def precision_at_k(ranked_ids, relevant_ids, k): """Fraction of the top-k predictions that are relevant.""" if k <= 0: raise ValueError("k must be positive") top_k = ranked_ids[:k] relevant = set(relevant_ids) hits = sum(1 for i in top_k if i in relevant) return hits / min(k, len(top_k)) # guard k > len(predictions)

What interviewers actually test

Correctness first

A working brute force beats an elegant solution that doesn't run. State your approach, get something correct, then optimize. Edge cases (empty input, k larger than n) are part of the grade.

Pythonic, vectorized code

For ML variants, reaching for NumPy vectorization instead of triple-nested loops signals you've written real numerical code. Know broadcasting, argsort, and boolean masking cold.

Math → code fluency

"Implement AUC" or "code the E-step" tests whether the formula you can recite is one you actually understand. Define the quantity in one sentence before you type.

Common mistakes

Do this

Treat it like a real SWE coding round. Communicate as you go, name the pattern, state complexity, test on a small example. If it's an ML metric, write down the definition first — half the failures are coding the wrong formula confidently.

Avoid this

Assuming "I'm an ML person, I don't need to grind coding." Many strong modelers fail the loop here. Also: silent coding, ignoring edge cases, and writing un-vectorized NumPy that betrays you haven't shipped numerical code.

CROSS-LINK

The algorithm half of this round is the same material covered in Grokking the Coding Interview — the reusable patterns (sliding window, two pointers, BFS/DFS, top-K, DP) that solve the DSA portion. Prep the patterns there; practice them in Python here.

02

Round 2

ML breadth & theory

Rapid-fire concept questions across classical ML and deep learning. Less "derive this" than "explain it like you've used it" — depth is probed by follow-ups, not by the first question.

The breadth round is a conversation that moves fast and goes one level deeper than you expect. The opening question is usually broad — "explain the bias-variance trade-off," "how does a random forest reduce variance," "what is attention" — and then the interviewer follows up until you hit the edge of your understanding. The signal isn't whether you know the textbook answer; it's whether you can keep going. "Regularization prevents overfitting" is a memorized phrase. Being able to explain why L1 produces sparsity while L2 doesn't, and when you'd pick each, is understanding.

In 2026 the topic mix has shifted decisively. Classical ML — logistic regression, trees, SVMs, k-means, PCA, the metrics zoo, cross-validation — is now table stakes: you're expected to be fluent, and being shaky there is disqualifying, but being strong there is no longer differentiating. The deep-learning and LLM half is where loops increasingly separate candidates: backprop, optimizers, normalization, the transformer architecture, attention, embeddings, and the GenAI-era questions (fine-tuning vs. RAG vs. prompting, what RLHF optimizes, KV caches, quantization, how to evaluate a generative model). Treat classical ML as the floor and invest your marginal prep hour in the modern stack.

What the breadth round draws from TABLE STAKES · classical ML bias-variance regularization trees / ensembles metrics / ROC cross-val PCA / k-means DIFFERENTIATOR · deep learning & LLMs backprop optimizers normalization transformers attention embeddings RAG vs fine-tune RLHF / alignment KV cache / serving quantization eval gen The follow-up ladder: "what is dropout?" → "why does it work?" → "train vs. inference behavior?" → "interaction with batch norm?" they keep climbing until you stop. depth, not the first answer, is the signal.

What interviewers actually test

Depth on follow-up

Anyone can recite a definition. The grade comes from the second and third question. Strong answers volunteer the "why" and the trade-off before being asked.

Intuition over derivation

Most rounds want the mental model, not a chalkboard proof — though research roles do ask you to derive. "Explain it to a smart colleague" is the right register.

Modern-stack fluency

Comfort with transformers, attention, and the LLM/GenAI vocabulary now reads as a proxy for "has shipped recently." Gaps here stand out more than a fuzzy SVM answer.

Common mistakes

Do this

Answer in layers: one-sentence intuition, then the mechanism, then a trade-off or failure mode. Say "I'm not certain, but here's how I'd reason about it" when you hit your edge — honest reasoning beats confident wrong.

Avoid this

Buzzword answers with no mechanism ("dropout regularizes the network" — how?). Bluffing — interviewers probe exactly to find the bluff. And neglecting the deep-learning/LLM half because classical ML feels safer.

GO DEEPER

The two halves of this round each have a dedicated guide: Classical ML Fundamentals for the table-stakes layer (bias-variance, regularization, trees, metrics) and Deep Learning Fundamentals for the differentiator layer (backprop, CNNs, RNNs, transformers, attention).

03

Round 3 · the differentiator

The ML system design round

"Design the feed-ranking system for a social app." This is the round that separates levels — and the one most candidates under-prepare. It rewards a framework and disciplined sequencing, not a model name.

At mid and senior levels, this is the round that decides your offer and your level. You're given a deliberately under-specified product goal — "design a recommendation system for a marketplace," "build a system to detect fraudulent transactions," "rank a news feed" — and asked to design the whole ML system around it, out loud, in 45 minutes. The interviewer is watching how you handle ambiguity, whether you reason about trade-offs, and whether you understand that an ML system is far more than a model: it's data, labels, features, training, serving, evaluation, and a feedback loop that keeps it alive in production.

The number-one failure mode is jumping straight to "I'd use a transformer / gradient-boosted trees" before framing anything. Strong candidates run a disciplined sequence: clarify the goal and constraints → define success metrics (an offline metric you can optimize and an online metric the business cares about) → reason about the data and labels you'd actually have → design features → pick a model that fits the latency and scale budget → specify training and serving → and crucially, plan evaluation, monitoring, and the feedback loop. Interviewers consistently probe what happens after the model ships, because that's where real systems decay — through drift and training/serving skew.

The ML system design sequence — walk it in order 1 · Frame goal · constraints 2 · Metrics offline + online 3 · Data labels · sources 4 · Features eng · leakage 5 · Model fit the budget 6 · Serve latency · scale 7 · Evaluate · monitor · feedback A/B tests · drift · training-serving skew the loop closes — production signal flows back into data and reframing candidates who skip step 7 lose the most points

What interviewers actually test

Framing before modeling

The first five minutes — clarifying the goal, the constraints, and the metric — predict the grade more than the model choice. Pin down "what does success even mean" before anything else.

Trade-off reasoning

Every choice has a cost: a bigger model vs. latency budget, precision vs. recall, fresh features vs. serving complexity. Naming the trade-off and choosing deliberately is the senior signal.

The full lifecycle

Data, labels, training/serving consistency, monitoring, drift, and the feedback loop. Treating the model as one box in a larger system is what distinguishes levels.

Common mistakes

Do this

Lead with clarifying questions and a stated metric. Drive the session — propose the sequence, check in, manage your time so you reach serving and monitoring. Quantify: "100ms p99 budget means I can't run a 7B model per request, so I'd precompute embeddings."

Avoid this

Jumping to a model name in minute one. Ignoring data and labels (where do they come from?). Never mentioning evaluation or monitoring. Designing a research-grade model that can't meet the latency or cost budget the problem implies.

THE DEDICATED GUIDE

This round has its own deep-dive: ML System Design walks the framework across real problems — recommendation, ranking, feeds, fraud — with the trade-offs senior interviewers probe and worked examples of the full lifecycle.

04

Round 4

The project / paper / résumé deep-dive

"Walk me through a project you're proud of." The most underestimated round: a relaxed conversation that is actually a forensic audit of whether you did the work and owned the decisions.

This round feels like the easy one — you're talking about your own work, on home turf — which is exactly why candidates under-prepare and lose it. The interviewer picks one project (or a paper, for research roles) and drills relentlessly into the decisions: Why that model and not a simpler baseline? How did you choose the metric? What was the hardest bug, and how did you find it? What would you do differently? The goal is to distinguish people who genuinely drove the work from people who were nearby when it happened or are inflating a team effort into a solo one.

What separates strong candidates is ownership and honest reflection. They can articulate the alternatives they rejected and why, name the things that went wrong, and quantify impact ("cut p99 latency 40%," "lifted conversion 3% in the A/B test"). They use "I" and "we" precisely — claiming their part without erasing the team. For applied scientists and researchers, expect the interviewer to push on methodology, baselines, ablations, and the limitations of the work — sometimes adversarially, to see whether you can defend choices without getting defensive.

How to structure the project walkthrough Problem why it mattered Your role "I" vs "we" Decisions + alternatives Impact quantified Reflection what you'd change the interviewer will live in the "Decisions" box "why this model?" · "why this metric?" · "what was the baseline?" · "what broke?" prepare 2–3 projects this deeply — you don't choose which they pick apart for research roles, swap "project" → "paper": baselines, ablations, limitations

What interviewers actually test

Real ownership

Did you drive the decisions, or narrate someone else's work? Precise "I did X, the team did Y" framing — and the ability to go arbitrarily deep on your part — is the tell.

Decision quality

Not "did it work" but "did you reason well." Naming the baseline you compared against and the alternatives you rejected shows judgment, not luck.

Honest reflection

What went wrong, what you'd change, the limitations. Candidates who can critique their own work read as senior; those who claim everything was perfect read as junior or evasive.

Common mistakes

Do this

Prepare two or three projects to extreme depth — you don't pick which gets dissected. Lead with the problem and your role, then go deep on decisions and impact with numbers. Rehearse the "what would you do differently" answer; it always comes.

Avoid this

Vague "we built a model that improved things." Taking sole credit for team work (a fast trust-killer). Picking a project too old or too thin to withstand drilling. Getting defensive when an interviewer challenges a choice — they're testing composure, not just correctness.

05

Round 5

The behavioral round

Not filler. A real round that gates offers — and at some companies, levels. It tests collaboration, ownership, how you handle conflict and ambiguity, and whether the team wants you in the room.

Technical candidates routinely treat the behavioral round as a formality and walk in cold. It isn't, and they shouldn't. At many companies a weak behavioral signal can sink an otherwise strong loop, and at places with explicit leadership rubrics it directly informs your level. The questions are predictable — a time you disagreed with a teammate, a project that failed, how you handled an ambiguous goal, a moment you influenced without authority — but the answers separate people who reflect from people who improvise.

The expectation is structured, specific stories. The STAR shape — Situation, Task, Action, Result — keeps you concrete and stops you from drifting into generalities. For ML roles specifically, the strongest behavioral stories sit at the seam between the technical and the human: a time you had to communicate a modeling trade-off to a non-technical stakeholder, push back on an unrealistic metric, or decide whether to ship a model that was good-but-not-perfect. Those answers double as evidence for both the behavioral rubric and your judgment as an ML practitioner.

Questions to have ready

PromptWhat it's probing
Tell me about a conflict with a teammateMaturity, empathy, how you disagree and recover
Describe a project that failedOwnership, honesty, what you learned
A time you handled an ambiguous goalAutonomy and structure under uncertainty
When you influenced without authorityCommunication and cross-functional credibility
Explaining a model trade-off to a non-expertThe ML-specific seam: technical + human

Do this

Pre-write five or six STAR stories that flex to cover most prompts. Be specific and own the result. Pick at least two stories that showcase ML judgment and stakeholder communication, not just generic teamwork.

Avoid this

Winging it. Vague, hypothetical, or hero-narrative answers with no conflict. Blaming teammates for failures. Treating the round as small talk — interviewers are scoring it against a rubric.

THE ML ANGLE

Your strongest behavioral stories often come from modeling decisions — choosing recall over accuracy because the classes were imbalanced, or pushing back on a vanity metric. See Modeling Decisions & Communication for how to frame those trade-offs in a way that lands with both technical and non-technical interviewers.

HOW LOOPS DIFFER · BY ROLE

Same rounds, different weights.

"ML role" hides four fairly different jobs, and the loop reweights its rounds for each. The titles aren't standardized across companies — one firm's "applied scientist" is another's "ML engineer" — but the underlying axes are consistent: how much engineering vs. research, how much product vs. methodology. Read the job description and the recruiter's framing to calibrate which rounds will carry the most weight, then prep accordingly.

RoleCodingBreadth / theorySystem designEmphasis
ML engineerHeavy — SWE barApplied, practicalHeavy — productionBuilding & shipping ML systems; MLOps, serving, scale
Applied scientistMediumDeep — incl. derivationsModel-centricModeling depth + research applied to product
Data scientistSQL + Python; lighter DSAStats & experimentation heavyLighter / analyticsAnalysis, A/B testing, causal & product sense
Research scientistLighterDeepest — math & proofsRareNovel methods, publications, paper deep-dive

ML engineer

The most engineering-forward loop. Expect a real coding bar and a production-grade system-design round. Strong MLOps and serving instincts matter — pair this with MLOps & Production.

Applied scientist

Modeling depth is the differentiator. Breadth goes deeper — be ready to derive, not just describe — and the deep-dive probes methodology hard. Coding is present but secondary.

Data scientist

Statistics, experimentation, and product sense dominate. Expect SQL and A/B-test design over heavy DSA. System design, if present, leans analytics and measurement.

Research scientist

The most theory-heavy loop. Deep math, derivations, and a rigorous paper deep-dive — baselines, ablations, limitations. System design is rare; novelty and publications carry the day.

HOW LOOPS DIFFER · BY COMPANY

Big tech, frontier labs, and startups want different things.

The same five rounds get assembled differently depending on where you interview. Big-tech loops are standardized and rubric-driven; frontier labs lean harder into depth and research taste; startups compress everything and weight practical, ship-it ability. Knowing the house style lets you emphasize the right things instead of running a generic playbook.

Standardized & rubric-driven

Big tech (FAANG-scale)

  • →Structured loops, explicit rubrics, calibrated leveling. Predictable round set.
  • →Real coding bar plus a production-grade ML system-design round at senior levels.
  • →Behavioral is formalized and can affect level (e.g., leadership-principle rubrics).
  • →Prep for: breadth, scale, and the full lifecycle in system design.

Depth & research taste

Frontier labs (OpenAI / Anthropic-style)

  • →Deeper on modern deep learning and LLMs; genuine fluency with the stack expected.
  • →Research taste and reasoning from first principles weigh heavily; strong paper deep-dive for research roles.
  • →Often a strong coding/engineering bar too — research and infra blur together.
  • →Prep for: transformers, training/scaling, evaluation, and clear reasoning under open-ended questions.

Compressed & practical

Startups

  • →Fewer, broader rounds — sometimes one person covering several areas at once.
  • →Weighted toward practical, end-to-end ability: can you ship a useful model fast with limited data and infra?
  • →Less rubric, more "would I want to build alongside this person." Breadth and pragmatism over specialization.
  • →Prep for: pragmatic trade-offs, doing more with less, and owning the whole pipeline.

PUTTING IT TOGETHER

A prep plan keyed to the loop, not to "study ML."

"Study machine learning" is not a plan — it's how prep stays infinite. Prep per round instead, and front-load your weakest one. The schedule below assumes roughly eight weeks; compress or stretch it, but keep the order: foundations first, then the differentiating rounds, then full mock loops to integrate everything. The biggest mistake is spending all your time on the round you already enjoy.

An ~8-week plan (adjust to your timeline) wk1wk3wk5wk7wk8 Phase 1 · Foundations Phase 2 · Differentiating rounds Phase 3 · Mock loops Phase 1 — weeks 1–2 · Coding: drill DSA patterns in Python (pair with Grokking Coding Interview) · Breadth floor: lock classical ML — bias-variance, regularization, trees, metrics Phase 2 — weeks 3–6 · Deep learning + LLMs: transformers, attention, RAG vs fine-tune, evaluation · ML system design: run the framework on 8–10 problems, end to end incl. monitoring · Deep-dive: write up 2–3 projects to drilling depth; behavioral: pre-write STAR stories Phase 3 — wks 7–8 · full mock loops, timed, then fix gaps

PHASE 1 · FOUNDATIONS

Weeks 1–2

Re-establish the coding bar in Python and lock classical-ML fluency. These are the floors — get them solid so they're not where the loop breaks, then move on. Don't over-invest here if it's already your strength.

PHASE 2 · DIFFERENTIATORS

Weeks 3–6

The bulk of your time: deep learning + LLMs for breadth, and the system-design framework run end-to-end across many problems. In parallel, write up your projects and pre-write behavioral stories so nothing is cold.

PHASE 3 · INTEGRATION

Weeks 7–8

Full, timed mock loops — ideally with someone who'll push back. Integration is its own skill: switching modes between rounds in a single day is what you're rehearsing. Triage your weakest round and fix it.

Two cross-cutting habits pay off across every round. First, practice switching modes — the loop's hidden difficulty is doing five different exams in one day. Second, always tie modeling choices back to product and production reality: a metric exists to serve a business goal, and a model has to survive deployment. Those instincts are sharpened in Modeling Decisions & Communication and MLOps & Production, and they're exactly what separates a hire from a strong hire.

Prep the whole loop, round by round.

The full course walks every round in depth — coding, breadth & theory, ML system design, and behavioral — with interactive lessons, worked problems in Python, and ML system-design walk-throughs, modernized for 2026.

Take Grokking the Machine Learning Interview

Free trial · 25+ lessons · no credit card required

Prep every round of the ML loop with Grokking the Machine Learning Interview