Word2Vec Training, CBOW and Skip-Gram, and Embedding Limitations
Summary
This article reviews Word2Vec as an efficient method for learning word vectors from context, and compares it with matrix-based and clustering approaches to distributed representations. It discusses CBOW and Skip-Gram, emphasizing their different update procedures: CBOW aggregates context words before an update, while Skip-Gram processes target-context pairs individually. Hierarchical Softmax and negative sampling are presented as ways to reduce training cost. The author argues that choice of method depends on the corpus, task, and evaluation criteria.
The notes explain how learned projection weights serve as embeddings and challenge the idea that one-hot inputs are simply inefficient because they are sparse. They also identify limitations: Word2Vec ignores word order within its context window, and nearby or substitutable words may be represented similarly without the model understanding their meaning. The discussion is a personal technical interpretation, partly grounded in code reading and external explanations; it cautions readers to verify claims and experiment. It provides no systematic benchmark comparing methods across tasks.
Key ideas
- Word2Vec learns dense word representations from contextual co-occurrence patterns.
- CBOW aggregates context before updating, whereas Skip-Gram updates using individual target-context pairs.
- Hierarchical Softmax and negative sampling reduce the cost of training output predictions.
- The suitability of Word2Vec, GloVe, FastText, and other representation methods depends on the data and task.
- Word2Vec ignores context word order and captures statistical association rather than human-like semantic understanding.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.