ZZomato·Tech KnowledgeL3DSA Round

Handling an Imbalanced Dataset

Problem How would you handle an imbalanced dataset?

Be ready to discuss

  • Resampling: oversampling the minority class (random duplication, or synthetic points via SMOTE/ADASYN) vs. undersampling the majority class — and the cost of each, overfitting duplicated points vs. discarding real signal.
  • Class-weighted loss: penalize minority-class errors more heavily so the optimizer stops ignoring the rare class; often the cleanest first move since it touches no data.
  • Threshold tuning: the default 0.5 cutoff is arbitrary — choose the operating point off the PR curve against the real cost of a false positive vs. a false negative.
  • Metric choice: accuracy is useless here; use F1, precision/recall, or PR-AUC, and evaluate on an untouched, naturally-distributed holdout.
  • Where resampling must happen: inside the cross-validation folds, never before the split, or synthetic neighbours leak across train and test.
  • Ensembles and alternatives: balanced bagging/EasyEnsemble, and treating extreme skew (e.g. 1:10000) as anomaly detection rather than classification.
asked …
LeaderboardSalaryAccount