Bias-Variance Tradeoff
Problem Why does a model that performs perfectly on training data often fail in production?
Be ready to discuss
- Overfitting/high variance: the model memorized noise and idiosyncrasies of the training set rather than learning generalizable patterns, so training error keeps dropping while held-out error rises.
- Underfitting/high bias: the opposite failure — the model is too simple to capture the underlying signal at all, and does poorly on both train and test.
- The bias-variance tradeoff itself: how expected generalization error decomposes into bias, variance, and irreducible noise, and why lowering one typically raises the other.
- Diagnosis: train vs. validation learning curves, cross-validation, and reading the gap between the two to tell overfitting from underfitting.
- Remedies: regularization (L1/L2, dropout, early stopping), more or better-quality training data, feature pruning, and matching model complexity to the data available.
- Production-specific causes beyond variance: train/serve skew, data leakage inflating offline scores, and distribution shift over time.
asked …