BERT-based Text Embeddings vs Word2Vec
Problem What is the difference between BERT-based text embeddings and Word2Vec?
Be ready to discuss
- Static vs. contextual: Word2Vec assigns exactly one vector per word regardless of usage, while BERT produces a different vector for the same word depending on surrounding context.
- Polysemy: why "bank" in "river bank" and "bank account" collapses to a single Word2Vec vector but separates cleanly under BERT.
- Training objectives: Word2Vec's shallow CBOW/skip-gram with negative sampling over a local window vs. BERT's deep bidirectional Transformer trained with masked-language-modeling.
- Architecture depth and what it buys: BERT stacks self-attention layers and captures long-range syntactic/semantic structure; Word2Vec is effectively an embedding lookup plus a shallow projection.
- Cost trade-off: Word2Vec is a cheap table lookup — microseconds and megabytes; BERT needs a forward pass per input and an accelerator budget at serving time.
- Getting a sentence vector out of each: averaging Word2Vec vectors vs. CLS-token/mean-pooling from BERT, and why vanilla BERT pooling underperforms sentence-tuned variants (e.g. Sentence-BERT) for similarity search.
asked …