ZZomato·Tech KnowledgeL3DSA Round

How to Train a Word2Vec Model

Problem How do we train a Word2Vec model?

Be ready to discuss

  • The two architectures: Skip-gram predicts surrounding context words from a target word; CBOW predicts the target word from its averaged context. Skip-gram does better on rare words and small corpora, CBOW trains faster.
  • The network itself: an embedding lookup (input matrix) and an output matrix, with no hidden non-linearity — the embedding IS the learned input matrix, and the prediction task is only a means to an end.
  • Why the naive softmax is intractable: it normalizes over the entire vocabulary at every step, so cost scales with vocab size.
  • Negative sampling: reframe as binary classification of real (word, context) pairs against k sampled negatives, drawn from a smoothed unigram distribution (frequency^0.75).
  • Hierarchical softmax as the alternative: a Huffman-tree factorization that cuts output cost to O(log V).
  • Practical knobs: window size (larger windows capture topical similarity, smaller ones syntactic), embedding dimension, subsampling of frequent words, epochs, and min-count vocabulary pruning.
  • Evaluating the result: nearest-neighbour sanity checks, analogy tasks, or downstream task performance — good vectors place semantically similar words close together.
asked …
LeaderboardSalaryAccount