Preventing Cross-Stock Windows in Time-Series Model Training
Summary
The document raises a data construction question for a rolling-trained deep learning model that predicts a stock’s next-day outcome from a fixed-length history. The dataset concatenates each stock’s chronological observations, then randomly selects training window start points. A window near the boundary between two stocks could therefore contain the tail of one stock’s history and the beginning of another’s, even though those observations are not a continuous sequence for a single security.
The author describes batches containing many such windows, a generator that yields a fixed number of batches per epoch, an unshuffled validation set, and a test process that creates one recent window per stock for portfolio backtesting. The central issue is whether windows should be constrained to one ticker and to valid chronological segments; the document asks this but supplies no answer or results. It also leaves open details such as split boundaries, feature consistency, and how labels are aligned, so it is a problem statement rather than a demonstrated modeling method.
Key ideas
- Randomly sampling windows from concatenated stock histories can create samples that cross ticker boundaries.
- A training sample should preserve the intended chronological relationship between its input window and target label.
- The described setup uses rolling train and test periods, generated training batches, and ordered validation windows.
- The test procedure predicts one next-day value per stock from that stock’s recent history.
- The document asks about window validity but offers no solution or empirical comparison.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.