Training Skip-Gram Word Embeddings with Softmax and Stochastic Gradient Descent
Summary
This tutorial walks through a basic Skip-gram word2vec training workflow. A center word is represented by an input vector, and the model predicts each word in its context window using a softmax output layer. The article describes calculating the prediction loss and gradients, summing those quantities across context words, and updating the vectors with stochastic gradient descent. It also shows a simple quadratic objective as a sanity check for the optimizer.
The worked example uses Stanford sentiment text and initializes vectors alongside pretrained GloVe embeddings, then visualizes selected words in two dimensions before and after training. The reported qualitative result is that some words become closer together, consistent with their co-occurrence in the corpus. This is an introductory implementation rather than a financial-language study: it supplies no trading signal, market data, predictive evaluation, or evidence that the learned representations improve investment decisions. Its ideas may nevertheless be useful as background for researchers exploring text features, including sentiment inputs, with the caveat that corpus choice and downstream validation determine whether embeddings carry useful information.
Key ideas
- Skip-Gram learns word representations by predicting context words from a center word.
- A softmax objective produces the loss and gradients for each center-context pair.
- Gradients from all words in the context window are accumulated before updating vectors.
- The tutorial uses stochastic gradient descent and checks it on a simple quadratic function.
- The visualization suggests co-occurring words can cluster, but does not test financial or trading value.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.