Skip to content
All library documents

Word2Vec Negative Sampling for Faster Training

Article BigQuant

Summary

The article explains why standard stochastic gradient descent can remain inefficient when training Word2Vec on a large vocabulary. Even when an update concerns only a few words, calculating gradients across the full parameter matrices wastes work, and the softmax objective requires scores for every vocabulary item.

Negative sampling replaces that full-vocabulary calculation with a smaller objective: score observed target-context word pairs as positive examples and randomly sampled words as noise. A sigmoid-based loss updates only the vectors involved in those pairs, making each training step cheaper. The article also identifies which word vectors receive gradients, but does not show the promised formulas or derive the gradients. It gives no benchmarks or trading application; its relevance to quantitative research is as a training technique for language models that might be used to analyze financial text.

Key ideas

  • Full-vocabulary gradient calculations can waste computation when only a few word vectors are involved in an update.
  • The softmax objective requires scoring all vocabulary items, which is costly for large vocabularies.
  • Negative sampling contrasts observed word pairs with randomly selected noise words.
  • A sigmoid-based objective limits each update to vectors involved in the sampled examples.
  • The article identifies gradient targets but omits the formulas and provides no empirical comparison.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.