Sub-word (BPE) vs Word-level Tokenization
Problem Why do we use sub-word tokenization (like Byte-Pair Encoding) instead of word-level tokenization?
Be ready to discuss
- The OOV problem: a word-level vocabulary cannot represent a word it never saw — every rare word, typo, name, or product code collapses into a single <UNK> token and its meaning is lost outright.
- Vocabulary size and cost: word-level needs hundreds of thousands to millions of entries, and both the embedding matrix and the output softmax scale linearly with that.
- What BPE actually does: start from characters or bytes, then iteratively merge the most frequent adjacent pair into a new token until the vocabulary hits a target size. Frequent words end up as single tokens; rare ones decompose into known sub-word pieces.
- Why it handles morphology gracefully: prefixes, suffixes, and inflections become reusable units, so "running" and "runner" share structure instead of occupying unrelated vectors.
- Byte-level BPE and why it guarantees no OOV ever — any input is representable as bytes, which also covers emoji, code, and multilingual text.
- Alternatives and how they differ: WordPiece (merges by likelihood gain rather than raw frequency), SentencePiece/Unigram (probabilistic, whitespace-agnostic), and character-level (no OOV, but sequences get very long).
- The underlying trade-off: shorter sequences with a larger vocabulary vs. longer sequences with a smaller one — and why sequence length is expensive under O(n^2) attention.
asked …