How Skip-Gram Uses Word Vectors to Predict Context Words
Summary
This introduction explains the Skip-gram model as a method for predicting nearby context words from a chosen center word. It uses a short example sentence and a window extending two words to either side. The explanation defines one-hot representations, dense word vectors, and two matrices that hold vectors used for center and context words. It then describes how a center word is selected from the one-hot input, compared with vocabulary vectors using dot products, and passed through softmax to produce a probability distribution over possible context words.
The article is conceptual: it focuses on the shapes and intuition of the calculation rather than training the vectors, defining a loss function, or discussing optimization. It presents dot product as a similarity measure and softmax outputs as context probabilities, but gives no experiments or quality metrics. The account is therefore a basic walkthrough of the scoring step, not a full treatment of Word2Vec training or its computational tradeoffs.
Key ideas
- Skip-gram predicts words in a local context window from a selected center word.
- One-hot vectors identify vocabulary entries, while dense vectors represent words in a lower-dimensional space.
- Dot products between a center vector and vocabulary vectors produce similarity scores.
- Softmax converts the scores into a probability distribution over possible context words.
- The walkthrough explains scoring intuition but does not cover model training, optimization, or empirical evaluation.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.