Transformer Self-Attention Mechanism Explained
Problem Explain the self-attention mechanism in Transformers.
Be ready to discuss
- Query/Key/Value projections: each token emits a Query (what am I looking for), a Key (what do I contain), and a Value (what information do I pass on); attention scores come from Query-Key dot products.
- The scaled dot-product formula Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) V, and why the 1/sqrt(d_k) scaling matters — without it large dot products push softmax into saturated regions with vanishing gradients.
- Multi-head attention: several attention computations run in parallel with different learned projections so heads can specialize (e.g. one tracking syntax, another entity relationships), then outputs are concatenated and projected.
- Positional encodings: attention is permutation-invariant and has no inherent notion of order, so sine/cosine (or learned) position signals are added to token embeddings.
- Residual connections and layer normalization around each sub-layer, which keep deep stacks trainable.
- Complexity: O(n^2 * d) in time and memory over sequence length n, and why that drives long-context variants (sparse, sliding-window, linear, flash attention).
asked …