Free guide · Part of Grokking the Machine Learning Interview · Last updated August 2026
The breadth round

The classical ML an interviewer expects you to explain cold.

Bias–variance, regularization, regression, trees and boosting, SVMs, clustering, PCA, and the metrics that decide every model — with the framing that makes you sound fluent, not memorized.

FH By Fahim Ul Haq
· Last updated August 2026 · ~20 min read
Recent updates
  • 2026-08-04Added missing data imputation strategies — Missing data is common in real-world datasets and interviewers expect you to handle it thoughtfully. Simple imputation fills missing values with mean, median, or mode, while model-based imputation uses algorithms like k-NN or regression to predict missing values. Some tree-based methods like XGBoost handle missing values natively. Dropping rows with missing data is only viable when the missing proportion is small and missingness is random. Always check whether data is missing completely at random, missing at random, or missing not at random—the pattern affects which strategy is valid.
  • 2026-07-16Added feature scaling guidance for distance-based algorithms — Feature scaling matters most for distance-based algorithms like k-NN, SVM with RBF kernels, and k-means clustering. These methods measure similarity using Euclidean or other distance metrics, so features with larger numeric ranges dominate the calculation. Standardization (zero mean, unit variance) or min-max scaling ensures all features contribute proportionally. Tree-based models like random forests and gradient boosting are scale-invariant because they split on thresholds, not distances. Interviewers often ask when scaling is required and which algorithms need it. Always scale after splitting train and test sets to avoid data leakage from test statistics.

Most ML loops open with a breadth round — a fast, conversational sweep across the classical toolkit before anyone touches deep learning or system design. The interviewer isn't trying to catch you on an obscure derivation. They're checking whether the fundamentals are load-bearing in your head: can you explain the bias–variance tradeoff without a whiteboard, say when L1 beats L2, and justify a model choice in one clean sentence?

The trap is sounding memorized. A candidate who recites "L1 does feature selection" but can't say why reads as someone who studied flashcards. A candidate who says "L1's constant gradient pushes small weights exactly to zero, so it both regularizes and selects features — I'd reach for it when I suspect most inputs are noise" reads as someone who has actually trained models. Fluency, not recall. That's the whole game in this round.

This guide walks the topics in roughly the order they come up, each with the crisp explanation an interviewer wants plus the common mistakes that quietly cost points. Where a topic is really a judgment call — which metric, which model — we cross-link Modeling Decisions & Communication, since the breadth round and the modeling-judgment round test the same knowledge from different angles.

01

Framing

Supervised, unsupervised & reinforcement learning

The first fork in any ML conversation: what kind of signal is the model learning from? Get the framing right and every model below slots into place.

The cleanest way to organize the whole field is by what supervises the learning. In supervised learning, every training example carries a label — the right answer — and the model learns a mapping from inputs to that target. In unsupervised learning, there are no labels; the model finds structure in the data itself. In reinforcement learning, there's no fixed answer key at all — an agent takes actions in an environment and learns from a reward signal that may arrive much later.

Interviewers like this question because your answer reveals whether you think in terms of problem shape. The right move is to anchor each paradigm to the task it solves and the kind of data it needs, then give one concrete example. The most common stumble is blurring the line between classification and clustering — both produce groups, but classification learns from labels and clustering discovers groups with none.

Supervised

Labeled data → learn input→output. Classification (discrete labels: spam / not-spam) and regression (continuous targets: house price). The bulk of the breadth round lives here.

Unsupervised

No labels → find structure. Clustering (k-means), dimensionality reduction (PCA), density estimation. Used for segmentation, compression, and exploratory analysis.

Reinforcement

Reward signal → learn a policy by trial and error. Sequential decisions, delayed reward, exploration vs exploitation. Robotics, game-playing, and RLHF for LLMs.

Two refinements worth dropping in to sound current. Self-supervised learning — the engine behind modern LLMs — is technically unsupervised data turned into a supervised task: the label is part of the input itself (predict the next token, fill the masked word). And semi-supervised learning mixes a small labeled set with a large unlabeled one, which is the realistic regime in most companies, where labels are expensive. If the conversation drifts toward how foundation models are trained, that's your bridge to the Deep Learning guide.

WHAT THEY'RE CHECKING

Can you place a novel problem into the right bucket on the fly? "We want to group customers but have no labels" → unsupervised, clustering. "We want to predict churn and have historical churn flags" → supervised classification. Sorting the problem correctly is the prerequisite for choosing a model — and it's the first thing the rest of the round builds on.

02

The most-tested idea

The bias–variance tradeoff

If you internalize one idea before the breadth round, make it this one. Nearly every other classical-ML question is a special case of managing bias against variance.

A model's expected error on unseen data decomposes into three parts: bias, variance, and irreducible noise. Bias is error from wrong assumptions — the model is too simple to capture the real pattern (a straight line trying to fit a curve). Variance is error from sensitivity to the particular training set — the model is so flexible it fits the noise, and a different sample would produce a wildly different fit. Irreducible noise is the floor you can't beat no matter what.

The tradeoff: making a model more flexible lowers bias but raises variance, and simplifying it does the reverse. Underfitting = high bias (poor on train and test alike). Overfitting = high variance (great on train, poor on test). The whole craft of modeling is finding the sweet spot where total error is minimized — and naming where a given symptom sits is exactly what interviewers probe.

Error Model complexity / flexibility → simple complex Bias² Variance Total error optimal underfit (high bias) overfit (high variance)

How to diagnose & fix each

High bias (underfitting)

Symptom: high error on both training and validation sets, and the two are close. Fixes: a more expressive model, more / better features, less regularization, train longer. More data won't help — the model can't represent the pattern in the first place.

High variance (overfitting)

Symptom: low training error but a large gap to validation error. Fixes: more training data, stronger regularization, a simpler model, feature selection, early stopping, or ensembling (bagging) to average out the variance.

The advanced beat — drop it if the interviewer is probing — is that the classic U-shaped curve isn't the whole story for very large models. Double descent shows that as you push past the interpolation threshold (enough capacity to memorize the training set), test error can fall again. It's why massively over-parameterized deep nets generalize despite "overfitting" by the classical definition. You don't need the math, but knowing the phenomenon exists signals you've kept up with how the field has moved.

COMMON MISTAKE

Saying "I'd add more data to fix underfitting." More data fixes variance, not bias — a model that's too simple stays too simple. The candidates who score connect a symptom to the right lever: train/val both bad → bias → more capacity; train good, val bad → variance → more data or more regularization.

03

Controlling variance

Regularization — L1 vs L2

Regularization is the most direct lever you have on the variance side of the tradeoff: penalize complexity so the model can't memorize noise.

The idea: add a penalty term to the loss that grows with the size of the model's weights, so optimization has to balance fitting the data against keeping weights small. A strength parameter (often λ or alpha) dials how hard you push. Bigger penalty → simpler model → less variance, more bias. The two workhorses differ in which norm they penalize, and that single choice changes their behavior in a way interviewers love to test.

L2 (Ridge) penalizes the sum of squared weights. Its gradient shrinks weights proportionally to their size, so big weights get pulled hard but nothing is forced to exactly zero — you get a model with many small, diffuse weights. L1 (Lasso) penalizes the sum of absolute weights. Its gradient is constant regardless of weight size, so it keeps pushing small weights all the way to exactly zero — producing a sparse model that does automatic feature selection.

L1 (Lasso) — diamond hits corner → w₁=0 L2 (Ridge) — circle tangent → both small, neither zero The diamond's corners sit on the axes — that geometry is why L1 zeros out weights and L2 doesn't.
AspectL1 — LassoL2 — Ridge
PenaltyΣ|wᵢ| (absolute)Σwᵢ² (squared)
Effect on weightsDrives many to exactly zero → sparseShrinks all toward zero, none reach zero
Feature selectionYes — built inNo
Correlated featuresTends to pick one, drop the restSpreads weight across the group
Reach for it whenMany features, you suspect most are irrelevant; you want an interpretable, sparse modelMany useful, correlated features; you want stability and smaller weights, not selection

The third option worth naming is Elastic Net, which blends both penalties. It's the pragmatic default when you have groups of correlated features and want some sparsity: L1 alone behaves erratically with correlated predictors (arbitrarily picking one), and the L2 component stabilizes that. In deep learning the same instinct shows up as weight decay (essentially L2) and dropout — different mechanisms, same goal of fighting variance.

COMMON MISTAKE

Reciting "L1 = feature selection, L2 = ridge" without the why. Always have a one-liner on the mechanism: L1's gradient is constant, so it keeps nudging weights to zero; L2's gradient vanishes as weights shrink, so they level off near (but never at) zero. One more: regularization assumes scaled features — forget to standardize and the penalty hits large-scale features unfairly.

04

The baselines

Linear & logistic regression

The two models everyone underrates — and exactly where interviewers go to see if you understand assumptions, loss functions, and interpretation, not just APIs.

Linear regression fits a weighted sum of features to a continuous target and is trained by minimizing mean squared error. With the right assumptions it has a closed-form solution (the normal equation); in practice you fit it with gradient descent at scale. The reason it's an interview staple isn't the math — it's the assumptions, because knowing them is what separates "I can call .fit()" from "I understand when this model is valid."

Linearity & additivity

The target is a linear function of the features. Curved relationships need feature transforms (polynomials, logs) or a different model.

Independent, homoscedastic errors

Residuals are independent with constant variance. Funnel-shaped residuals (heteroscedasticity) break inference; correlated errors break time-series fits.

No severe multicollinearity

Highly correlated features make coefficients unstable and uninterpretable. Check VIF; fix with L2 or by dropping redundant features.

Normal residuals (for inference)

Needed for valid confidence intervals and p-values, not for point prediction. Useful to know which assumptions matter for which goal.

Logistic regression is the classification counterpart, and the name trips people up — it predicts a probability, not a continuous value. It passes the same linear combination through the sigmoid to squash it into [0, 1], then thresholds for a class. Crucially, it is not trained with MSE; it's trained with log loss (binary cross-entropy), which is convex for this model and penalizes confident-and-wrong predictions sharply. Being able to say "MSE here is non-convex and gives weak gradients, so we use cross-entropy" is a reliable signal that you understand the model.

// logistic regression — the parts interviewers ask you to write
# linear score, then squash to a probability
z = w · x + b
p = sigmoid(z) = 1 / (1 + e^(−z))

# loss: binary cross-entropy (NOT mean squared error)
L = −[ y·log(p) + (1−y)·log(1−p) ]

# interpretation: a unit increase in xᵢ multiplies the
# ODDS by e^(wᵢ)  →  coefficients are log-odds, not probabilities

The interpretation point is where strong candidates pull ahead. In linear regression a coefficient is "the change in the target per unit change in the feature, holding others fixed." In logistic regression a coefficient is a log-odds — exponentiate it and you get an odds ratio. Saying "this weight means the odds of the positive class multiply by e^w" shows you actually read the model's output rather than just trusting its predictions. Both models stay popular precisely because they're interpretable and make rock-solid baselines — always worth proposing one before you reach for something heavier.

COMMON MISTAKE

Calling logistic regression a "regression model" that outputs continuous values — it's a classifier that outputs probabilities. And don't claim its coefficients are probabilities; they're log-odds. Naming the loss correctly (cross-entropy, not MSE) and the interpretation correctly (odds ratios) is a fast, reliable credibility check.

05

The tabular workhorses

Trees → random forests → gradient boosting

The single most useful model family for structured data — and a near-guaranteed question, because the progression from one tree to boosting tells a clean story about bias and variance.

A decision tree splits the data with a sequence of feature thresholds, choosing each split to maximize purity (lower Gini impurity or entropy for classification, lower variance for regression). It's wonderfully interpretable and needs no feature scaling — but a single deep tree overfits hard: it's a low-bias, high-variance model that memorizes the training set. Everything that follows is a strategy for taming that variance.

Bagging — Random Forest independent trees, in parallel → vote tree 1 tree 2 tree 3 average / majority vote cuts VARIANCE Boosting — XGBoost / LightGBM sequential — each tree fixes the last's errors tree 1 tree 2 tree 3 residuals→ residuals→ weighted sum cuts BIAS (and variance)

Random forests attack the variance with bagging: train many trees, each on a bootstrap sample of the rows and a random subset of features at each split, then average their predictions. The trees are deliberately decorrelated (that's what the random feature subset buys you), so averaging cancels much of the individual-tree variance without raising bias. Trees train independently — embarrassingly parallel — and you get feature importances and out-of-bag error nearly for free.

Gradient boosting takes the opposite tack: build trees sequentially, where each new (shallow) tree is fit to the residual errors of the ensemble so far. Bagging reduces variance by averaging; boosting reduces bias by relentlessly correcting mistakes, and with shrinkage (a small learning rate) it controls variance too. XGBoost and LightGBM are the optimized implementations — regularized objectives, clever histogram-based split finding, native missing-value handling — and they remain the default winners on tabular data.

Random Forest (bagging)Gradient Boosting (XGBoost / LightGBM)
How trees combineParallel & independent, then averagedSequential — each corrects the prior
Primarily reducesVarianceBias (variance too, via shrinkage)
Tree depthDeep, fully grownShallow "weak learners"
Overfitting riskLow — hard to overfit by adding treesHigher — too many trees / high LR overfits
Tuning effortForgiving, strong out of the boxMore knobs (LR, depth, rounds) but higher ceiling

Why boosting dominates tabular data

This is the question behind the question, and the freshness note up top previews the answer. Boosted trees thrive on the messy reality of structured data: mixed feature types (numeric and categorical side by side), missing values handled natively, non-smooth, non-linear decision boundaries, and near-zero preprocessing — no scaling, no one-hot gymnastics required. Deep nets, by contrast, need a lot of data and care to match that, and their real edge is on unstructured signals (images, audio, text) where learned representations matter. The crisp interview answer: "For tabular, I start with gradient-boosted trees; I reach for deep learning when the input is raw and high-dimensional." See Deep Learning for that other half.

WHAT THEY'RE CHECKING

Do you understand why bagging and boosting both help, but for different reasons? "Bagging averages independent high-variance trees to cut variance; boosting fits trees to residuals to cut bias." If you can also say when a single interpretable tree beats either (when explainability outranks accuracy), you've shown the judgment the round is really after.

06

The rest of the toolkit

SVMs, k-NN, k-means, PCA & Naive Bayes

The classics you should be able to summarize in two or three correct sentences each — including the one-line "gotcha" that proves you actually understand them.

Support Vector Machines & the kernel trick

An SVM finds the hyperplane that separates classes with the maximum margin — the widest gap to the nearest points (the support vectors). The kernel trick lets it draw non-linear boundaries by implicitly mapping data into a higher-dimensional space (via a kernel like the RBF) without ever computing those coordinates. Gotcha: strong on small/medium high-dimensional data; scales poorly to very large datasets, and you must standardize features.

k-Nearest Neighbors (k-NN)

A lazy, non-parametric method: no training step — at prediction time it finds the k closest training points and takes a vote (classification) or average (regression). Simple and surprisingly strong as a baseline. Gotcha: prediction is slow on big data, it's sensitive to feature scale and the curse of dimensionality, and choosing k trades bias (large k) against variance (small k).

k-Means & clustering

Unsupervised: partition points into k clusters by alternating "assign each point to its nearest centroid" and "recompute centroids," minimizing within-cluster variance. Gotcha: you must pick k up front (use the elbow method or silhouette score), it assumes roughly spherical, equal-size clusters, and it's sensitive to initialization — use k-means++. For arbitrary shapes or noise, reach for DBSCAN.

PCA — dimensionality reduction

Finds the orthogonal directions (principal components) of maximum variance and projects data onto the top few, compressing many correlated features into a handful while preserving most of the signal. Useful for visualization, denoising, and speeding up downstream models. Gotcha: it's linear and unsupervised (ignores labels), components are hard to interpret, and you must standardize first or high-variance-by-scale features dominate.

Naive Bayes

A probabilistic classifier applying Bayes' theorem with one bold simplifying assumption: features are conditionally independent given the class — the "naive" part. That assumption is almost never true, yet it works remarkably well for high-dimensional text (spam filtering, document classification), trains in a single fast pass, and needs little data. Gotcha: it's a great fast baseline but its probability estimates are poorly calibrated, and you need Laplace smoothing so an unseen feature value doesn't zero out the whole prediction.

COMMON MISTAKE

Forgetting that SVM, k-NN, k-means, and PCA all assume scaled features — distance- and variance-based methods are wrecked by unscaled inputs. Mentioning standardization unprompted is a quiet but reliable signal. The other classic slip: describing k-means as if it guarantees a global optimum. It's a local-optimum method — that's why initialization and multiple restarts matter.

07

Measuring success

Evaluation metrics — and when to use which

"Why not just use accuracy?" is one of the most common breadth-round questions — and one of the fastest ways to separate candidates who've shipped models from those who haven't.

Everything starts at the confusion matrix: true positives, false positives, true negatives, false negatives. Every classification metric is some ratio of those four cells, and choosing the right one is a statement about which errors hurt. Accuracy — fraction correct — is the default, and the default trap: on a 99%-negative dataset, a model that predicts "negative" every time is 99% accurate and completely useless. The instant a class is rare or the cost of errors is asymmetric, accuracy stops being meaningful.

Precision asks: of the items I flagged positive, how many really were? (It punishes false alarms.) Recall asks: of all the actual positives, how many did I catch? (It punishes misses.) They trade off — pushing the threshold to catch more positives raises recall but lowers precision. F1 is their harmonic mean, a single number for when you care about both. The whole skill is matching the metric to the business cost of each error type.

Confusion matrix — the source of every metric Predicted + Predicted − Actual + Actual − TPhit FNmiss — hurts recall FPfalse alarm — hurts precision TNcorrect reject Precision = TP/(TP+FP) Recall = TP/(TP+FN) F1 = harmonic mean of precision & recall = 2·P·R / (P+R)
MetricWhat it measuresReach for it when
AccuracyFraction of all predictions correctClasses are roughly balanced and all errors cost the same. Rarely the right call beyond that.
PrecisionOf predicted positives, fraction correctFalse positives are costly — spam filters, fraud flags a human must review, anything where a false alarm is expensive.
RecallOf actual positives, fraction caughtFalse negatives are costly — disease screening, security threats, anything where a miss is the dangerous outcome.
F1Harmonic mean of precision & recallYou need a single number balancing both, and the positive class is the rare one you care about.
ROC-AUCRanking quality across all thresholds (TPR vs FPR)Roughly balanced classes; you want a threshold-independent summary of separability.
PR-AUCPrecision vs recall across thresholdsHeavy class imbalance. Far more informative than ROC-AUC when positives are rare.

The threshold-free metrics are where senior signal shows up. ROC-AUC measures how well a model ranks positives above negatives across every threshold — the probability that a random positive scores higher than a random negative. It's intuitive and threshold-independent, but it has a blind spot: on heavily imbalanced data it's misleadingly optimistic, because the huge number of true negatives keeps the false-positive rate low even for a mediocre model. That's exactly when PR-AUC earns its keep — it focuses on the rare positive class and exposes weaknesses ROC-AUC hides. The crisp rule: balanced → ROC-AUC; rare positives → PR-AUC.

For regression, the trio is MAE (robust to outliers, same units as the target), RMSE (penalizes large errors more — use it when big misses are disproportionately bad), and R² (fraction of variance explained, a unitless "how much better than predicting the mean"). Choosing among these — and articulating the trade-off out loud — is squarely the judgment tested in Modeling Decisions & Communication.

COMMON MISTAKE

Reporting accuracy on an imbalanced problem, or quoting ROC-AUC when positives are 1% of the data. The model answer ties the metric to the cost of each error: "Positives are rare and a miss is expensive, so I'd optimize recall and report PR-AUC, then tune the threshold to the precision the business can tolerate." That sentence alone signals real modeling experience.

08

Trusting your numbers

Cross-validation & data leakage

How you estimate performance matters as much as the model — and data leakage is the bug that makes a broken model look brilliant right up until production.

A single train/test split gives one noisy estimate of how a model generalizes, and it wastes data. k-fold cross-validation fixes both: split the data into k folds, train on k−1 and validate on the held-out fold, rotate so every fold is the validation set once, then average. You get a more stable performance estimate with a variance band, and every example is used for both training and validation. Five or ten folds is the usual default.

Two refinements interviewers expect you to know. Stratified k-fold preserves the class distribution in each fold — essential for imbalanced or multiclass problems so a rare class doesn't vanish from a fold. And for time-series data, ordinary k-fold is wrong: shuffling lets the model train on the future to predict the past. Use forward-chaining (expanding-window) splits where the validation fold always comes after the training data in time.

Data leakage — the silent killer

Data leakage is when information from outside the training set sneaks into the model — usually information you wouldn't actually have at prediction time. The tell is a result that's too good to be true: stellar validation scores that collapse in production. It's one of the most common real-world ML bugs, which is why it's a favorite interview probe.

Preprocessing on full data

Scaling, imputing, or selecting features using statistics computed over the whole dataset leaks test info into training. Fit transforms on the training fold only — use a Pipeline so CV does this correctly.

Target leakage

A feature that's a proxy for, or computed after, the label — e.g., "was_refunded" when predicting fraud. It won't exist at prediction time. Always ask: would I have this value before the event?

Temporal / group leakage

Random splits on time-series let the future leak backward; rows from the same user/group landing in both train and test inflate scores. Split by time or by group, not at random.

COMMON MISTAKE

Fitting the scaler or imputer on the entire dataset before splitting — the single most common leakage bug. The fix is to wrap preprocessing in a pipeline so each fold fits its transforms on training data alone. If you ever see a validation score that looks suspiciously perfect, your first instinct should be "where's the leak?" — and saying that out loud reads as production-hardened.

09

When positives are rare

Handling class imbalance

Fraud, churn, disease, anomalies — the most valuable problems are usually the most imbalanced. Interviewers ask how you'd handle a 99:1 dataset because it forces you to connect metrics, sampling, and thresholds.

The first move isn't a technique — it's a metric. As covered above, accuracy is meaningless here, so the very first thing to say is "I'd stop using accuracy and switch to precision/recall, F1, or PR-AUC." Only then do you reach for the levers that actually rebalance the learning. There are three families, and strong candidates know all three and pick by context.

Resampling

Oversample the minority (duplicate, or synthesize with SMOTE), or undersample the majority. SMOTE interpolates new minority points rather than copying. Caveat: resample inside the CV fold only, never before splitting, or you leak.

Class weights / cost-sensitive

Tell the loss function to penalize minority-class mistakes more (class_weight='balanced'). No data duplication, no leakage risk — often the cleanest first thing to try.

Threshold tuning

The 0.5 cutoff is arbitrary. Most models output probabilities — move the decision threshold along the precision-recall curve to hit the operating point the business needs. Cheap, powerful, and often overlooked.

The senior framing pulls these together: rebalancing the data and reweighting the loss both shift where the model draws its boundary, while threshold tuning shifts where you cut the probabilities afterward — and you can combine them. For many real problems, class weights plus threshold tuning beats aggressive oversampling, because synthetic minority points can introduce noise. And the closing thought interviewers reward: if the minority class is extremely rare, reframe it as anomaly detection rather than classification.

COMMON MISTAKE

Applying SMOTE or oversampling to the whole dataset before the train/test split — synthetic copies of test points then leak into training and the score balloons. Resample only within the training portion of each fold. The second slip is jumping to fancy resampling before trying the simplest fix: class weights and a tuned threshold.

10

Where most gains come from

Feature engineering basics

On classical models, better features beat a fancier model more often than not. Interviewers ask because it reveals whether you've actually worked with real, messy data.

Feature engineering is turning raw data into inputs a model can learn from well. With deep learning the model learns representations for you, but for the classical models in this round — regression, SVMs, even boosted trees — the features you hand it largely cap how well it can do. The practical toolkit is small and worth being able to rattle off cleanly.

Scaling & normalization

Standardize (zero mean, unit variance) or min-max scale. Mandatory for distance- and gradient-based models (SVM, k-NN, k-means, linear models); trees don't care. Fit on train only.

Encoding categoricals

One-hot for low-cardinality, ordinal when order is meaningful, target/frequency encoding for high-cardinality — but guard the latter against leakage with out-of-fold encoding.

Missing values & outliers

Impute (mean/median/model-based) or, for trees, leave them. A "was-missing" indicator often helps. Clip or transform extreme outliers; consider whether missingness itself is signal.

Transforms & interactions

Log skewed features, bin continuous ones, build interaction and domain features (ratios, time-since-event, day-of-week). This is where domain knowledge pays off most.

WHAT THEY'RE CHECKING

Two things: that you treat features as the highest-leverage lever on a classical model, and that you fold the leakage discipline in automatically — fit every transform on training data only, inside the CV loop. Mentioning a "was-missing" indicator or out-of-fold target encoding unprompted is the kind of detail that marks someone who's actually built features for a living.

PUTTING IT TOGETHER

What the breadth round is really testing.

Step back and the through-line is clear: almost everything above is a different face of one idea — managing bias against variance, and one habit — connecting a choice to its reason out loud. The interviewer isn't tallying definitions. They're listening for whether the fundamentals are wired together in your head.

  1. 1.Reason, don't recite. Every fact should come with a "because." "L1 zeros weights because its gradient is constant." That one word is the difference between memorized and understood.
  2. 2.Map symptoms to levers. Train and val both bad → bias → more capacity. Train good, val bad → variance → more data or regularization. This diagnostic loop is the spine of the round.
  3. 3.Pick metrics by the cost of errors. Never default to accuracy. Tie precision/recall/PR-AUC to which mistake actually hurts the business.
  4. 4.Justify the model in one sentence. "Tabular and messy → gradient-boosted trees; raw and high-dimensional → deep learning." Judgment over hype.
  5. 5.Guard against leakage reflexively. Fit transforms inside the CV fold; split time-series by time. A result too good to be true usually is.
  6. 6.Start simple. Propose a logistic-regression or single-tree baseline before anything heavy. Knowing when not to over-engineer is senior signal.

Get these reflexes down and the breadth round stops feeling like a quiz and starts feeling like a conversation between two people who build models. From here, the natural next steps are Deep Learning Fundamentals for the representation-learning half of modern ML, and ML System Design for the round where these models get wired into real, serving systems.

Go from fluent to unshakeable.

The full course turns every topic above into interactive lessons and worked problems in Python — bias-variance, boosting, metrics, and the rest — with the follow-up questions senior ML interviewers actually ask.

Take Grokking the Machine Learning Interview

Free trial · interactive lessons in Python · no credit card required

See the latest promotion on Grokking the Machine Learning Interview