Foundations
The neuron, MLPs & the forward pass
Every deep network is a stack of the same primitive: a weighted sum, a bias, and a nonlinearity. Get this right and the rest of deep learning is composition.
A single artificial neuron computes a weighted sum of its inputs, adds a bias, and passes the result through a nonlinear activation: a = f(w·x + b). Without the nonlinearity f, stacking neurons gains you nothing — a composition of linear maps is still just a linear map. The activation is what lets a deep network represent curved decision boundaries and, with enough width, approximate any continuous function (the universal approximation theorem).
A multilayer perceptron (MLP) arranges these neurons into layers: an input layer, one or more hidden layers, and an output layer. The forward pass is just matrix multiplication followed by an activation, repeated layer by layer — h = f(Wx + b) at each step. The whole network is a differentiable function from inputs to outputs, which is exactly what makes gradient-based training possible.
What interviewers actually test
They want to hear
Why the nonlinearity is non-negotiable (stacked linear layers collapse to one), what each weight matrix's shape is, and that the forward pass is just batched matmuls. Bonus: parameter counting — (in × out) + out per dense layer.
Common mistakes
Saying "more layers always means more power" (without nonlinearities, no), confusing the bias term with regularization, and forgetting that weight initialization matters — all-zeros weights make every neuron in a layer identical and the network can never break symmetry.