Pre-training vs Fine-tuning in LLMs
Problem What is the difference between pre-training and fine-tuning a language model?
Be ready to discuss
- Pre-training teaches general language capability — "how to speak" — through self-supervised learning (next-token prediction) over massive, broad, unlabelled corpora. No human labels, trillions of tokens, and the bulk of the compute bill.
- Fine-tuning specializes that pre-trained model for a narrower task or domain — "how to be a medical assistant" — using far smaller curated or labelled datasets, and hours-to-days of compute rather than months.
- The alignment stages layered on top: instruction tuning to follow directions, then RLHF/DPO to match human preferences on helpfulness and safety.
- Why the split works at all: pre-training learns transferable representations, so a downstream task only has to re-shape existing knowledge instead of learning language from scratch.
- Parameter-efficient alternatives: LoRA/adapters that update a small fraction of weights, and the case for prompting or RAG instead when you need facts rather than behaviour.
- Failure modes of fine-tuning: catastrophic forgetting of general ability, overfitting a small dataset, and data quality mattering far more than volume.
asked …