Section 1
Choosing the right metric
Every modeling decision flows from one question the interviewer is silently grading: does the metric you optimize actually move the outcome the business cares about?
The single most common mistake candidates make is reaching for accuracy by reflex. On a fraud dataset that's 99.8% legitimate, a model that predicts "never fraud" scores 99.8% accuracy and catches zero fraud — useless. The interviewer wants to see you reason from the business goal backward to the metric, not forward from "what metric do I know."
The mapping has three steps. First, name the real-world outcome: catch more fraud, retain more subscribers, surface relevant search results. Second, identify the cost asymmetry: which mistake hurts more, a false positive or a false negative? Third, pick the metric that rewards the right behavior under that asymmetry. The definitions themselves — precision, recall, F1, ROC-AUC, PR-AUC, NDCG — are covered in Classical ML Fundamentals; here the skill is choosing among them out loud.
Business goal → ML metric
| Problem | Business goal | Cost asymmetry | Metric to optimize |
|---|---|---|---|
| Fraud detection | Block fraud without blocking real customers | Severe class imbalance; a missed fraud and a blocked good customer both hurt | PR-AUC and recall at a fixed precision; never raw accuracy |
| Churn prediction | Find at-risk users while retention budget lasts | You can only call/discount the top N users — ranking quality matters more than a threshold | Recall@K or lift in the top decile; AUC for overall ranking |
| Medical screening | Catch the disease early | A false negative (missed case) is far costlier than a false positive (extra test) | Recall / sensitivity with a precision floor |
| Spam filter | Block spam without losing real mail | A false positive (real email in spam) is far costlier than letting one spam through | Precision (high), recall as secondary |
| Search / ranking | Put the most relevant results at the top | Position matters — being right at rank 1 beats being right at rank 9 | NDCG, MRR, MAP — rank-aware, not flat accuracy |
| Demand forecasting | Predict units sold to plan inventory | Over- and under-forecasting have different dollar costs; outliers distort squared error | MAE / MAPE (or quantile loss when the asymmetry is explicit) |
Notice the pattern: the metric is downstream of the cost structure. When you can't get the costs from the interviewer, say so and assume — "I'll assume a missed fraud costs roughly 10× a false alarm, so I'll optimize recall at a precision I can defend, and I'd validate that ratio with the risk team." That single sentence shows more maturity than a perfect recitation of the F1 formula.
Two further moves separate strong candidates. They distinguish the offline metric they can compute on a held-out set (AUC, NDCG) from the online metric the business actually reports (fraud-dollars prevented, subscriber retention, revenue per search) — and they note that the two can diverge. And they remember to pick an operating threshold, not just a curve: AUC summarizes all thresholds, but production needs one, chosen from the precision/recall trade-off at the business's cost point.
SAY THIS IN THE ROOM
"Before I pick a model, let me pin the metric. The goal is X; the costly mistake is Y; so I'll optimize Z offline and watch the business KPI online. I'll choose the threshold from the precision/recall curve at our cost ratio." Leading with the metric reframes the whole answer around outcomes — exactly the altitude interviewers reward.