ZZomato·Tech KnowledgeL2DSA Round

Sub-word (BPE) vs Word-level Tokenization

Problem Why do we use sub-word tokenization (like Byte-Pair Encoding) instead of word-level tokenization?

Be ready to discuss

  • The OOV problem: a word-level vocabulary cannot represent a word it never saw — every rare word, typo, name, or product code collapses into a single <UNK> token and its meaning is lost outright.
  • Vocabulary size and cost: word-level needs hundreds of thousands to millions of entries, and both the embedding matrix and the output softmax scale linearly with that.
  • What BPE actually does: start from characters or bytes, then iteratively merge the most frequent adjacent pair into a new token until the vocabulary hits a target size. Frequent words end up as single tokens; rare ones decompose into known sub-word pieces.
  • Why it handles morphology gracefully: prefixes, suffixes, and inflections become reusable units, so "running" and "runner" share structure instead of occupying unrelated vectors.
  • Byte-level BPE and why it guarantees no OOV ever — any input is representable as bytes, which also covers emoji, code, and multilingual text.
  • Alternatives and how they differ: WordPiece (merges by likelihood gain rather than raw frequency), SentencePiece/Unigram (probabilistic, whitespace-agnostic), and character-level (no OOV, but sequences get very long).
  • The underlying trade-off: shorter sequences with a larger vocabulary vs. longer sequences with a smaller one — and why sequence length is expensive under O(n^2) attention.
asked …
LeaderboardSalaryAccount