Turning Time-Series Observations into Supervised ML Examples
Summary
The document addresses how a sequential financial series becomes usable in a supervised machine-learning model. The key idea is to form examples by pairing information available at a decision time with the later outcome being predicted. Each such feature-and-label pair becomes one training observation, so the model can learn recurring relationships between inputs and outcomes. A time series does not need a one-to-one mapping to a single historical observation; the same kind of feature is measured at each example's timestamp.
The question also describes overlapping labels created with a triple-barrier scheme and notes that observations may not be independent. The response focuses on assembling paired inputs and outputs and collecting enough historical examples resembling the current input. It does not explain feature engineering for raw prices, handling temporal order within model inputs, average uniqueness, validation design, or how to avoid leakage. Those decisions remain necessary when applying the basic supervised-learning setup to trading data.
Key ideas
- A supervised-learning observation pairs features available before an outcome with that later outcome as its label.
- Repeat this pairing across historical decision times to create training examples.
- At prediction time, provide features measured at the current decision time to estimate the future outcome.
- The explanation does not specify how to engineer features from raw price sequences or address overlapping-label validation.
Tags
Full text
# How can stationary time series data be used as input in an ML model? # How can stationary time series data be used as input in an ML model? I am halfway through "Advances in Financial Machine Learning" by Marcos Lopez de Prado. I understand that a time series like stock prices can be transformed to make it sufficiently stationary. Pretend a stock series has 100 data points for T=1 to T=100 (one data point per time). I also understand (hopefully) that you can label this data by the three barrier method every so often in the time series. For example, label a training point every 2 units of time where the vertical barrier is 10 units of time out from the start point. So you would have labeled training examples for T=1, 3, 5, ... I understand that there is overlap in the training data and it is not IID and thus, you must use a method such as average uniqueness to counteract this act. Okay. Given that, I still have no idea how this labeled training data can be used as input features in a machine learning model. Like, each data point is just a number. I get that there's some concept of this number holding some concept of "memory" since the series isn't completely differentiated, but how is this even treated in the model? For example, let's pretend I'm gonna train a random forest to predict a stock's performance in the future. I collected training data as described above (so I want to know the stock's performance up to 10 time units in the future). Now I come up with features for my model like let's say the temperature outside, yesterday's NASDAQ price, etc. Those are nice easy features. Now I have this huge time series that I want to use as a feature (or multiple features). There are so many complications. Firstly, it's ordered, if there's any alpha to be collected the order definitely is the thing that will lead you to the alpha. Secondly, there's no one-to-one correspondence between a new testing example's price and a price in the training data. For example, if I wanted to use the NASDAQ price feature on a new testing example I would simply look at yesterday's NASDAQ price (there's a one-to-one correspondence). If I wanted to use yesterday's stock price to predict today's I wouldn't know which "yesterday" to use because any of the data point in the time series could he considered "yesterday". What am I missing here? ## Answer by wildbunny (score 1, accepted) https://quant.stackexchange.com/a/46038 With ML, you're looking to identify patterns in your inputs that result in your output(s). Thus you collect all the outputs you are hoping to be able to identify later, and the inputs which correspond to those outputs (i.e just before the output was generated), collect as many as possible of these relationships, stick them into a ML model and hope you've found an edge. If you want to predict tomorrows stock price using inputs from today, you'd better hope you've collected enough examples of inputs which look like todays in your dataset to to help the ML predict tomorrow.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.