How to Train a Word2Vec Model
Problem How do we train a Word2Vec model?
Be ready to discuss
- The two architectures: Skip-gram predicts surrounding context words from a target word; CBOW predicts the target word from its averaged context. Skip-gram does better on rare words and small corpora, CBOW trains faster.
- The network itself: an embedding lookup (input matrix) and an output matrix, with no hidden non-linearity — the embedding IS the learned input matrix, and the prediction task is only a means to an end.
- Why the naive softmax is intractable: it normalizes over the entire vocabulary at every step, so cost scales with vocab size.
- Negative sampling: reframe as binary classification of real (word, context) pairs against k sampled negatives, drawn from a smoothed unigram distribution (frequency^0.75).
- Hierarchical softmax as the alternative: a Huffman-tree factorization that cuts output cost to O(log V).
- Practical knobs: window size (larger windows capture topical similarity, smaller ones syntactic), embedding dimension, subsampling of frequent words, epochs, and min-count vocabulary pruning.
- Evaluating the result: nearest-neighbour sanity checks, analogy tasks, or downstream task performance — good vectors place semantically similar words close together.
asked …