Bagging vs Boosting
Problem Explain bagging and boosting.
Be ready to discuss
- Bagging (bootstrap aggregating): train many models independently and in parallel on bootstrap resamples, then average or vote. It attacks variance — averaging decorrelated high-variance learners cancels their individual errors. Random Forest is the canonical example, adding random feature subsets per split to decorrelate the trees further.
- Boosting: train models sequentially, each focused on what the ensemble has so far gotten wrong. AdaBoost reweights misclassified examples; gradient boosting fits each new learner to the residual/gradient of the loss. It attacks bias.
- Base-learner choice follows from that: bagging wants deep, low-bias/high-variance trees; boosting wants shallow, high-bias/low-variance stumps.
- Overfitting behaviour: bagging is hard to overfit by adding more trees; boosting will happily overfit given too many rounds, so it needs a learning rate, early stopping, and depth limits.
- Practical trade-offs: bagging parallelizes trivially and tolerates default settings; boosting usually reaches higher accuracy but is sequential and tuning-sensitive.
- Where stacking fits as the third ensembling family — a meta-learner trained over heterogeneous base models.
asked …