ZZomato·Tech KnowledgeL2DSA Round

Transformer Self-Attention Mechanism Explained

Problem Explain the self-attention mechanism in Transformers.

Be ready to discuss

  • Query/Key/Value projections: each token emits a Query (what am I looking for), a Key (what do I contain), and a Value (what information do I pass on); attention scores come from Query-Key dot products.
  • The scaled dot-product formula Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) V, and why the 1/sqrt(d_k) scaling matters — without it large dot products push softmax into saturated regions with vanishing gradients.
  • Multi-head attention: several attention computations run in parallel with different learned projections so heads can specialize (e.g. one tracking syntax, another entity relationships), then outputs are concatenated and projected.
  • Positional encodings: attention is permutation-invariant and has no inherent notion of order, so sine/cosine (or learned) position signals are added to token embeddings.
  • Residual connections and layer normalization around each sub-layer, which keep deep stacks trainable.
  • Complexity: O(n^2 * d) in time and memory over sequence length n, and why that drives long-context variants (sparse, sliding-window, linear, flash attention).
asked …
LeaderboardSalaryAccount