ZZomato·Tech KnowledgeL2DSA Round

LLM Architecture and Training Fundamentals

Problem Explain the basic architecture of large language models and how they are trained.

Be ready to discuss

  • The backbone: the Transformer — stacked blocks of multi-head self-attention plus a position-wise feed-forward network, with residual connections and layer normalization around each sub-layer. Most modern LLMs are decoder-only with causal masking.
  • The input path: tokenization into sub-word units, an embedding lookup, and positional information (learned, sinusoidal, or rotary), since attention on its own is order-blind.
  • What attention computes and why it is the core idea: every token attends over every other, so context is built by weighted mixing rather than fixed-window recurrence.
  • The output path: a final projection to vocabulary logits, softmax, and next-token sampling governed by temperature/top-p at inference time.
  • Stage 1, pre-training: large-scale self-supervised next-token prediction over massive text corpora under cross-entropy loss — the source of both knowledge and fluency, and where nearly all the compute goes.
  • Stage 2, alignment: supervised instruction tuning on demonstration data, then RLHF/DPO against human preferences to make the model helpful and safe rather than merely plausible.
  • Scaling and practicalities: the parameters/data/compute scaling relationship, context-window limits, and KV caching at inference.
  • What's evaluated: fluency across tokenization, embeddings, attention, and the pre-training/fine-tuning distinction, at least at a working level.
asked …
LeaderboardSalaryAccount