Free guide · Part of Grokking the Machine Learning Interview · Last updated September 2026
The deep-learning round

The deep-learning foundations every ML interview probes.

Backpropagation, activations, optimizers, regularization, CNNs, RNNs, attention & Transformers, embeddings, and the modern LLM stack — what interviewers actually test, and the mistakes that sink candidates.

· Last updated September 2026 · ~22 min read

The deep-learning portion of an ML interview is where breadth meets depth. Interviewers aren't looking for someone who can recite the Transformer paper — they're checking whether you understand why each piece exists. Why does ReLU train faster than sigmoid? Why does batch norm let you use a higher learning rate? What actually goes wrong when gradients vanish, and which architectural choices fix it? The candidates who pass can reason from first principles, not just name-drop.

This guide walks the foundations in the order they build on each other: the neuron and the forward pass, then backprop and gradient descent, then the design choices that make deep nets trainable — activations, losses, optimizers, regularization. From there it covers the three architecture families interviewers test (CNNs, RNNs, Transformers), then embeddings and transfer learning, and finally a 2026-aware view of the LLM stack. Each section flags what interviewers actually test and the common mistakes that cost offers.

If you're still shoring up the non-neural side — bias-variance, regularization theory, trees, and metrics — read Classical ML Fundamentals first. Many of the same ideas (overfitting, regularization, the bias-variance trade-off) reappear here in their deep-learning form.

01

Foundations

The neuron, MLPs & the forward pass

Every deep network is a stack of the same primitive: a weighted sum, a bias, and a nonlinearity. Get this right and the rest of deep learning is composition.

A single artificial neuron computes a weighted sum of its inputs, adds a bias, and passes the result through a nonlinear activation: a = f(w·x + b). Without the nonlinearity f, stacking neurons gains you nothing — a composition of linear maps is still just a linear map. The activation is what lets a deep network represent curved decision boundaries and, with enough width, approximate any continuous function (the universal approximation theorem).

A multilayer perceptron (MLP) arranges these neurons into layers: an input layer, one or more hidden layers, and an output layer. The forward pass is just matrix multiplication followed by an activation, repeated layer by layer — h = f(Wx + b) at each step. The whole network is a differentiable function from inputs to outputs, which is exactly what makes gradient-based training possible.

x₁ x₂ h₁ h₂ h₃ h₄ h₅ ŷ input hidden 1 hidden 2 output forward: h = f(Wx + b) → repeat per layer → ŷ

What interviewers actually test

They want to hear

Why the nonlinearity is non-negotiable (stacked linear layers collapse to one), what each weight matrix's shape is, and that the forward pass is just batched matmuls. Bonus: parameter counting — (in × out) + out per dense layer.

Common mistakes

Saying "more layers always means more power" (without nonlinearities, no), confusing the bias term with regularization, and forgetting that weight initialization matters — all-zeros weights make every neuron in a layer identical and the network can never break symmetry.

02

The core algorithm

Backpropagation & gradient descent

Backprop is the chain rule applied efficiently. Gradient descent is what you do with the gradients it produces. Together they are how every neural network learns.

Training a network means minimizing a loss L by nudging the weights in the direction that reduces it. Gradient descent does this by computing the gradient of the loss with respect to every weight, then stepping each weight a small amount in the opposite direction: w ← w − η · ∂L/∂w, where η is the learning rate.

The hard part is computing ∂L/∂w for millions of weights efficiently. Backpropagation is the answer: a single backward pass that applies the chain rule layer by layer, reusing intermediate results. After the forward pass caches each layer's activations, backprop starts from the loss and propagates the gradient backward, multiplying local derivatives along the way. The whole thing is one forward pass plus one backward pass — O(parameters), not exponential.

In practice you don't compute gradients on the full dataset (too slow) or one example at a time (too noisy). You use mini-batch SGD: average the gradient over a batch of examples, take a step, repeat. One full sweep through the data is an epoch; the batch size trades gradient noise against hardware efficiency.

x Layer 1 Layer 2 Loss L forward pass → ← backward pass (chain rule): ∂L/∂w at every layer then update: w ← w − η · ∂L/∂w

What interviewers actually test

They want to hear

That backprop = chain rule + caching, why the forward pass must store activations (you need them in the backward pass), and the difference between batch, mini-batch, and stochastic gradient descent. A strong answer can derive the gradient for a one-hidden-layer net on a whiteboard.

Common mistakes

Confusing backprop (computing gradients) with gradient descent (using them) — they're separate steps. Forgetting to zero gradients between steps in PyTorch. Claiming a larger batch is "always better" — it's more stable but generalizes worse and costs memory.

03

Design choices

Activation functions

The nonlinearity is the whole point of a hidden layer. Which one you pick decides whether gradients flow — and that decides whether the network trains at all.

Sigmoid squashes inputs to (0, 1); tanh to (−1, 1). Both were standard in the 1990s and both saturate: for large-magnitude inputs the gradient is nearly zero, so backprop can't push useful signal through deep stacks. This is the original cause of the vanishing-gradient problem. Sigmoid has the extra flaw of being non-zero-centered, which slows convergence.

ReLU — max(0, x) — won the 2010s because it doesn't saturate for positive inputs: the gradient is exactly 1 there, so signal flows freely through deep networks, and it's trivially cheap to compute. Its weakness is "dead ReLUs" (a neuron stuck outputting 0 has zero gradient forever), which Leaky ReLU and variants patch. GELU — a smooth, probabilistic gate that weights inputs by how far they are above zero — has become the default inside Transformers and most modern LLMs because the smoothness helps optimization and it slightly outperforms ReLU empirically.

sigmoid / tanh ReLU GELU x f(x)

They want to hear

Why ReLU beat sigmoid/tanh (no saturation for x>0 → gradients flow → deep nets become trainable), what a "dead ReLU" is, and that GELU/SiLU are the modern defaults in Transformers. Keep sigmoid/softmax for the output layer where you need probabilities.

Common mistakes

Using sigmoid in hidden layers of a deep net (vanishing gradients). Putting ReLU on the output of a regression model when targets can be negative. Confusing softmax (an output normalizer over classes) with a hidden-layer activation.

04

Design choices

Loss functions — and when to use which

The loss defines what "good" means. Pick the wrong one and you optimize for the wrong thing — even with a perfect architecture.

Mean squared error (MSE) is the default for regression: it penalizes the squared distance between prediction and target, which makes it smooth and easy to differentiate but sensitive to outliers (a single large error dominates). When outliers are a concern, MAE (absolute error) or Huber loss (quadratic near zero, linear in the tails) are robust alternatives.

Cross-entropy is the default for classification. It measures the distance between the predicted probability distribution and the true label, and pairs naturally with a softmax (multi-class) or sigmoid (binary) output. The reason it beats MSE for classification: combined with softmax, its gradient is simply (ŷ − y), which stays large when the model is confidently wrong — exactly when you want a strong learning signal. MSE on a sigmoid output, by contrast, produces tiny gradients in that regime and learning stalls.

Regression → MSE / MAE / Huber

Continuous targets. MSE for smoothness, MAE/Huber when outliers would otherwise dominate the gradient.

Classification → Cross-entropy

Softmax + categorical CE for multi-class; sigmoid + binary CE for binary or multi-label. The default for almost every classifier.

Imbalance / ranking → specialized

Focal loss down-weights easy examples for heavy class imbalance; contrastive / triplet losses for embeddings and retrieval.

THE CLASSIC FOLLOW-UP

"Why not MSE for classification?" — Because squared error on a saturating output gives vanishing gradients when the model is confidently wrong, and it doesn't correspond to maximum-likelihood under a categorical distribution. Cross-entropy does both: a clean (ŷ − y) gradient and a proper probabilistic interpretation. Being able to state this crisply is a reliable signal of depth.

05

Design choices

Optimizers & learning-rate schedules

Plain SGD works but crawls through ravines and plateaus. Momentum and adaptive methods are how modern nets converge in reasonable time.

SGD takes a step proportional to the current gradient. Simple, well-understood, and still the best generalizer for some vision tasks — but it oscillates in narrow valleys and stalls on flat regions. Momentum fixes this by accumulating an exponentially-weighted average of past gradients, so the optimizer builds velocity in consistent directions and damps oscillation — like a ball rolling downhill instead of restarting each step.

Adam is the workhorse default. It combines momentum (a running mean of gradients) with an adaptive per-parameter learning rate (a running mean of squared gradients), so each weight gets a step size scaled to its own gradient history. It converges fast and is forgiving of the initial learning rate, which is why most people reach for it first. AdamW, which decouples weight decay from the gradient update, is the standard for training Transformers and LLMs.

The learning-rate schedule matters as much as the optimizer. A common recipe: a short warmup (ramp the LR up from near-zero to avoid early instability), then cosine decay down toward zero over training. Warmup-then-decay is near-universal in Transformer training.

warmup cosine decay → LR training step ramp up to avoid early instability · decay to settle near a minimum

They want to hear

What momentum buys you (velocity + damping), how Adam adapts the LR per parameter, why AdamW exists (decoupled weight decay), and that warmup-then-cosine is the standard Transformer schedule. The learning rate is the single most important hyperparameter to tune.

Common mistakes

Treating Adam as strictly better than SGD (well-tuned SGD+momentum often generalizes better on vision). Forgetting warmup when training Transformers (early instability). Confusing learning rate with batch size — they interact, but they're different knobs.

06

Generalization

Regularization in deep learning

Big networks overfit. The whole toolbox of dropout, normalization, weight decay, and early stopping exists to make a high-capacity model generalize.

Dropout randomly zeros a fraction of activations during training, forcing the network not to rely on any single neuron — effectively training an ensemble of subnetworks that share weights. At inference, dropout is turned off and activations are scaled appropriately. Weight decay (L2 regularization) penalizes large weights, nudging the model toward simpler functions; early stopping halts training when validation loss stops improving, capturing the model before it memorizes the training set.

Batch normalization normalizes each layer's pre-activations across the batch, which smooths the loss landscape, lets you use higher learning rates, and acts as a mild regularizer. Its catch: it depends on batch statistics, so it behaves differently at train vs. inference and struggles with tiny batches. Layer normalization normalizes across features within each example instead of across the batch — which is why Transformers use it: it's batch-size-independent and works cleanly for variable-length sequences.

Dropout

Randomly drop activations at train time. Implicit ensemble. Off at inference.

Weight decay (L2)

Penalize large weights → simpler functions, less overfitting.

Batch / layer norm

Normalize activations → stable, faster training. BN across batch; LN across features.

Early stopping

Stop when validation loss plateaus. Cheapest regularizer there is.

THE CLASSIC FOLLOW-UP

"Batch norm vs. layer norm — why do Transformers use layer norm?" Because batch norm couples examples through batch statistics (bad for variable-length sequences and small batches), while layer norm normalizes within a single example and is independent of batch size. Knowing why the choice differs — not just that it does — is the depth interviewers look for.

07

What goes wrong

Vanishing & exploding gradients

The reason deep networks were hard to train for decades — and the cluster of fixes that finally made depth practical.

Backprop multiplies gradients layer by layer. If those per-layer factors are consistently less than 1, the product shrinks exponentially with depth and the early layers get a near-zero gradient — they stop learning. That's the vanishing-gradient problem, and saturating activations (sigmoid, tanh) make it worse. If the factors are consistently greater than 1, the product blows up and weights diverge into NaNs — that's exploding gradients, especially common in RNNs over long sequences.

The fixes are a layered defense, and naming several signals real understanding: ReLU/GELU activations (gradient ≈ 1 for active units), careful weight initialization (Xavier/He, sized to keep activation variance stable across layers), batch/layer normalization (keeps signals well-scaled), residual (skip) connections (give gradients a direct path that bypasses the multiplicative chain — the key idea behind ResNets and Transformers), and, for exploding gradients specifically, gradient clipping (cap the gradient norm).

The fixes (name several)

ReLU/GELU · He/Xavier init · batch/layer norm · residual connections · gradient clipping (for explosion). Residual connections are the highest-leverage idea — they're why 100+ layer networks train at all.

Common mistakes

Citing only "use ReLU" (it helps but residuals + normalization matter more at depth). Confusing the two failure modes. Not connecting this to why LSTMs/GRUs were invented (to fight vanishing gradients in sequences) or why ResNets work.

08

Architecture family

Convolutional networks (CNNs)

The architecture built for grid-structured data. Weight sharing and locality make it parameter-efficient and translation-aware — which is why it dominated vision for a decade.

A convolution slides a small learnable filter (kernel) across the input, computing a dot product at each position to produce a feature map. Two properties make this powerful: parameter sharing (the same filter is reused everywhere, so a 3×3 kernel has 9 weights regardless of image size) and locality (each output depends only on a small input patch, the receptive field). Stacking convolutions grows the receptive field, so early layers learn edges and textures while deeper layers compose them into objects.

Pooling (typically max-pooling) downsamples feature maps, shrinking spatial dimensions while keeping the strongest activations — adding a degree of translation invariance and cutting computation. A classic CNN alternates conv → activation → pool blocks, then flattens into a few dense layers for the final prediction. Landmark architectures to know by name: LeNet (the original), AlexNet (kicked off the deep-learning era in 2012), VGG (deep, uniform 3×3 stacks), and ResNet (residual connections enabling 100+ layers).

image conv pool conv pool dense ŷ spatial dims shrink · channel depth grows · edges → textures → objects

They want to hear

Why convolutions beat dense layers for images (parameter sharing + locality + translation equivariance), what a receptive field is, the role of stride/padding, and how to compute output dimensions. Bonus: that ResNet's skip connections solved the depth problem.

Common mistakes

Saying convolutions are translation invariant (they're equivariant; pooling adds invariance). Botching the output-size formula. Forgetting that channels stack — a conv layer's filters span the full input depth.

09

Architecture family

RNNs, LSTMs & GRUs

The pre-Transformer answer to sequences. Worth knowing both for what they do well and for the exact limitations that motivated attention.

A recurrent neural network (RNN) processes a sequence one element at a time, maintaining a hidden state that carries information forward. The same weights are applied at every timestep, so an RNN can handle variable-length inputs. The catch: training requires backpropagation through time, which unrolls the network across the sequence — and that long multiplicative chain is exactly where vanishing/exploding gradients bite. In practice, vanilla RNNs struggle to learn dependencies more than a few dozen steps apart.

LSTMs fix this with a cell state and a system of gates (input, forget, output) that learn what to remember, what to discard, and what to expose. The cell state acts as a gradient highway, letting information persist over hundreds of steps. GRUs are a streamlined variant with two gates instead of three — fewer parameters, often comparable performance. Both were the state of the art for translation and speech until attention took over.

The defining limitation: recurrence is inherently sequential, so you can't parallelize across the time dimension, and even LSTMs lose signal over very long ranges. Those two constraints — no parallelism, limited long-range memory — are precisely what the Transformer was designed to eliminate.

They want to hear

How the hidden state carries context, why vanilla RNNs fail on long sequences (vanishing gradients through time), and how LSTM gates + the cell state create a gradient highway. Then the punchline: why this motivated attention (no parallelism + limited range).

Common mistakes

Not being able to name the LSTM gates or say what each does. Claiming LSTMs "solve" long-range memory (they extend it, don't solve it). Missing the connection between RNN limitations and the rise of Transformers.

10

The architecture behind modern AI

Attention & Transformers

The single most important architecture to understand in a 2026 ML interview. Every modern LLM is a Transformer — and interviewers will probe whether you actually grasp self-attention.

Self-attention lets every token in a sequence look at every other token and decide how much to weight each one. Each token is projected into three vectors — a query, a key, and a value. For a given query, you score it against every key (a dot product), normalize the scores with softmax into attention weights, and take the weighted sum of the values. The result: a representation of each token that is informed by the whole sequence, with the weights learned per context. Multi-head attention runs several of these in parallel so the model can attend to different relationships at once (syntax, coreference, position).

Because attention has no inherent notion of order, Transformers add positional encodings to inject sequence position. A Transformer block is then: multi-head attention → add & layer-norm → a feed-forward MLP → add & layer-norm, with residual connections around each sublayer. Stack dozens of these blocks and you have a Transformer. The decisive advantage over RNNs: every token is processed in parallel (no sequential bottleneck), and any token can reach any other in one step (no long-range decay). The cost is that attention is O(n²) in sequence length — the central scaling challenge that long-context research keeps attacking.

Attention(Q, K, V) = softmax(Q · Kᵀ / √dₖ) · V token Query Key Value Q·Kᵀ → softmax attention weights context vector weights from (Q,K) reweight the Values → one context-aware vector per token
scaled_dot_product_attention.py · torch
import torch
import torch.nn.functional as F

def attention(Q, K, V, mask=None):
    # Q,K,V: (batch, heads, seq_len, d_k)
    d_k = Q.size(-1)
    scores = Q @ K.transpose(-2, -1) / d_k ** 0.5
    if mask is not None:
        scores = scores.masked_fill(mask == 0, float("-inf"))
    weights = F.softmax(scores, dim=-1)   # attention distribution
    return weights @ V                  # context vectors

They want to hear

The Q/K/V mechanism in your own words, why we scale by √dₖ (keeps softmax gradients sane), why positional encodings are needed, the role of residuals + layer-norm, and the O(n²) cost. Knowing the encoder/decoder split and causal masking is a plus.

Common mistakes

Reciting "attention is all you need" without explaining attention. Forgetting positional information (attention is permutation-invariant without it). Not knowing why it scales O(n²). Confusing self-attention with cross-attention.

11

Representations & reuse

Embeddings, transfer learning & fine-tuning

How models turn discrete things into geometry — and how you reuse a pretrained model instead of training from scratch. Both come up constantly in applied-ML interviews.

An embedding is a learned dense vector that represents a discrete item — a word, a user, a product — in a continuous space where geometric distance encodes similarity. Items used in similar contexts end up nearby, which is what makes embeddings the backbone of search, recommendation, and retrieval. The classic intuition (word2vec) is that "king − man + woman ≈ queen" — relationships show up as consistent directions in the vector space. Modern systems use contextual embeddings from Transformers, where the same word gets different vectors depending on its sentence.

Transfer learning reuses a model trained on a large dataset for a new, usually smaller task. Instead of training from scratch, you take a pretrained backbone (a vision model on ImageNet, or an LLM on web text) and adapt it. Two common modes: feature extraction (freeze the backbone, train only a new head) when your dataset is small, and fine-tuning (continue training some or all weights at a low learning rate) when you have more data. For large models, parameter-efficient fine-tuning like LoRA updates only a small set of added weights — far cheaper, and the standard way teams adapt LLMs in 2026.

Embeddings

Discrete → dense vectors where distance = similarity. Power search, reco, and retrieval. Contextual ones come from Transformers.

Feature extraction

Freeze the pretrained backbone, train a new head. Best when your labeled data is small. Fast and hard to overfit.

Fine-tuning / LoRA

Update weights at a low LR for more data. LoRA tunes a tiny added set — the cheap, standard way to adapt LLMs.

THE MODERN STACK · 2026

LLMs & the modern deep-learning stack.

A 2026 ML interview will almost certainly touch large language models. You don't need to have trained one, but you should be able to sketch the lifecycle at a high level and reason about the trade-offs.

  1. PretrainingSelf-supervised next-token prediction on web-scale text. This is where the model learns language and world knowledge — enormously expensive, done once, and the source of the model's raw capability.
  2. Fine-tuningInstruction tuning + alignment (supervised fine-tuning, then preference optimization such as RLHF/DPO) turns a raw next-token predictor into a model that follows instructions and behaves the way you want.
  3. RAGRetrieval-augmented generation grounds answers in an external knowledge base: embed the query, retrieve relevant chunks via vector search, and feed them into the prompt. The standard way to give a model fresh or proprietary knowledge without retraining — and to reduce hallucination.
  4. ServingInference at scale — quantization, KV-caching, batching, and latency/cost trade-offs. This is where deep learning meets systems, and it's a round of its own.

Two of these branch into rounds of their own. For how LLM systems are designed and served end to end — retrieval pipelines, latency budgets, evaluation — see ML System Design. For how models are deployed, monitored, and kept healthy in production — drift, training/serving skew, A/B testing — see MLOps & Production. The deep-learning round tests that you understand the model; those two test that you can ship it.

Go deep on the whole ML loop.

The full course turns these foundations into interactive lessons and worked problems in Python — classical ML, deep learning, ML system design, and MLOps, continuously maintained for 2026.

Take Grokking the Machine Learning Interview

Free trial · interactive lessons in Python · no credit card required

Recent updates
  • 2026-09-10Added layer normalization as alternative to batch normalization — Layer normalization normalizes activations across features for each training example independently, unlike batch normalization which normalizes across the batch dimension. This makes it effective for variable-length sequences and small batch sizes where batch statistics are unstable. Layer norm is the standard choice in Transformer architectures and recurrent networks. During interviews, you may be asked to contrast it with batch normalization and explain when each approach is preferred. The key distinction is the normalization axis: layer norm computes mean and variance per example, batch norm per feature across examples.
  • 2026-07-20Added batch normalization technique for training stability — Batch normalization is a widely used technique that normalizes layer inputs during training, helping networks converge faster and reducing sensitivity to initialization. It computes mean and standard deviation across each mini-batch, then scales and shifts the normalized values with learnable parameters. During inference, it uses running statistics collected during training instead of batch statistics. Interviewers often ask when to use batch norm versus other normalization techniques, or how it interacts with dropout. The key tradeoff is that batch norm introduces dependence on batch size, which can be problematic for small batches or recurrent architectures.

See the latest promotion on Grokking the Machine Learning Interview