Understanding PyTorch LSTM Tensor Shapes and Sequence Outputs
Summary
The document explains the expected input, initial state, and output shapes for PyTorch’s LSTM layer. With the default layout, inputs are arranged as sequence length, batch size, and feature count; the hidden and cell states also include layer and direction dimensions. It contrasts processing a sequence one item at a time with passing the full sequence in one call, noting that the output contains hidden states across all time steps while the returned final states summarize the sequence end.
Examples illustrate a small LSTM and a tagger that maps token embeddings to labels. The text also sketches preparing indexed inputs and labels, resetting recurrent state between training examples, and using a loss function with gradient updates. Its financial relevance is the mapping from time steps and features to sequential market data. The worked tagger uses a tiny, artificial text sample, and the article’s descriptions of hidden and cell states are simplified; it provides no evidence of trading performance or predictive value.
Key ideas
- PyTorch LSTM inputs use sequence, batch, and feature dimensions unless batch-first layout is selected.
- Initial and final hidden and cell states include dimensions for layers, directions, batches, and hidden features.
- Passing a full sequence returns per-step outputs as well as the final hidden and cell states.
- A token tagger can combine embeddings, an LSTM, and a linear output layer to predict labels.
- Recurrent state should be reset between independent training examples in the illustrated workflow.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.