Softmax Probability Normalization and Numerical Stabilization
Summary
The document explains softmax as a transformation from a vector of scores to nonnegative values that sum to one, making them interpretable as a probability distribution. It describes its role in the continuous bag-of-words model in Word2Vec, where model scores are converted into probabilities, and notes its broader use for similar normalization tasks.
It also gives an invariance property: adding the same constant to every input leaves the output unchanged. The example implementation uses this property by subtracting the maximum input before exponentiation, reducing the chance of numerical overflow. It handles either a vector or a matrix by normalizing across each row, and demonstrates the resulting distribution for a sample vector. The explanation is introductory; it does not discuss alternatives such as log-softmax or computational tradeoffs for large vocabularies.
Key ideas
- Softmax converts a score vector into values between zero and one that sum to one.
- It is used in Word2Vec's continuous bag-of-words model to turn scores into probabilities.
- Adding a shared constant to all inputs does not change the softmax output.
- Subtracting the maximum input before exponentiation helps reduce numerical overflow.
- The example handles vectors and row-wise matrix inputs.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.