Skip to content
All library documents

How Word2Vec Skip-Gram Learns Word Embeddings from Context

Article BigQuant

Summary

This tutorial explains the skip-gram model as a neural network trained to predict a nearby word from a chosen input word. Training examples are word pairs drawn from a configurable context window. The input word is represented as a one-hot vector, and a single hidden layer maps it to a compact vector; the output layer uses softmax to assign probabilities to vocabulary words.

The hidden-layer weights are the learned word embeddings, so the output layer can be discarded after training. Words that appear in similar contexts tend to acquire similar vectors, allowing the representation to capture both related meanings and similarities such as singular and plural forms. The tutorial also notes that the model does not distinguish context position and that a full softmax over a large vocabulary requires many parameters and substantial computation. It is a conceptual explanation, not a trading application or empirical evaluation; faster training techniques are deferred to a later part.

Key ideas

  • Skip-gram trains on pairs consisting of an input word and a nearby context word.
  • A one-hot input selects a row from the hidden-layer weight matrix, which serves as the word embedding.
  • The softmax output estimates probabilities over the vocabulary for the context word.
  • Words with similar surrounding contexts are encouraged to have similar embeddings.
  • The basic model ignores context position and can be expensive to train over a large vocabulary.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.