BERT vs GPT vs T5: Architecture Choice
Problem Why would you choose BERT (encoder-only), GPT (decoder-only), or T5 (encoder-decoder) for a given project?
Be ready to discuss
- BERT: encoder-only, trained with masked-language-modeling so every token sees both left and right context. Suited to understanding tasks — classification, sentiment, NER, extractive QA, and producing embeddings. It cannot generate text autoregressively.
- GPT: decoder-only, trained autoregressively with causal masking so each token sees only the past. Suited to open-ended generation, chat, and in-context/few-shot learning.
- T5: encoder-decoder, framing every task as text-to-text with a task prefix. Suited to sequence-to-sequence transformation — translation, summarization, style transfer — where the input is fully known and the output is new text.
- The deciding question: is the task primarily about understanding an input, generating an output, or transforming one sequence into another?
- Why the pretraining objective, not just the block layout, drives the fit — bidirectional context is exactly what makes BERT strong at classification and unusable for generation.
- Practical considerations: parameter count and serving cost, how much fine-tuning data you have, and whether a prompted decoder-only model now beats a fine-tuned encoder on your task.
asked …