Testing Lagged Time-Series Predictors for Stock Returns
Summary
The document considers how to test whether an arbitrary daily time series predicts a stock’s future movement. The proposed workflow is to split observations into training and test sets, difference both series to work with changes, and examine correlations between the predictor and stock returns at multiple lags. The author suggests validating the selected lag on held-out data and asks how to interpret statistical significance and turn a relationship into a trading signal.
The included answer recommends starting with low-dimensional, binarized features and targets, and considering mutual information or other distribution-based measures for possible causal links. It also mentions full sweeps or Lasso-based heuristics for higher-dimensional causal discovery. No thresholds, validated predictive results, or trading rules are supplied. Lag selection, stationarity, causality, and out-of-sample profitability remain open questions, so correlation and a small p-value alone should not be treated as evidence of a tradable edge.
Key ideas
- The proposed analysis compares differenced predictor and stock-return series across lags.
- Training and test splits can help assess whether a selected lag generalizes.
- The answer suggests low-dimensional binarized features and mutual information as starting points.
- Correlation and statistical significance alone do not establish causality or trading profitability.
- Higher-dimensional causal discovery may use broad searches or Lasso-based heuristics.
Tags
Full text
# Predicting time series based on another # Predicting time series based on another This is more of a generic question, but I'm sure it has a best answer/methodology which is what I'm trying to reach. I'm trying to figure out a solid line of thought when looking at a time series X and seeing if it can predict some stock prices. I've gone through a few threads on this site. The problem statement can be thought of as follows: given the daily closing prices of a stock, let's say AAPL, and an arbitrary predictor time series $X$ that is also given daily at close. Is $X$ a good predictor of AAPL's movement and how far in the future does it predict the price? Here's what I've resolved to doing: - Split data into training/testing. - In order to make the time series stationary, you do first differencing on both time series. (i.e. compute the percent change for each period) - Compute pearson correlation across different lags of $X$, find the highest value. - Find out if you're over-fitting or not by testing that lag with the test dataset. - ???? How do i then use this information. Let's say pearson correlation is `0.45` with a p-value of `2e-9` on the test dataset. Is this good? not good enough? How do I then trade on this information? I've read online about the Granger-Causality test, which sounds like it could help here. But I'm also just not sure about a lot of the assumptions I'm making here. Is percent-change the way to do it? What are the cutoffs for good vs. bad correlations? Also, there are very little posts I could find online that go past this point. I'm not sure what the intuition is here. If they're correlated, and $X$ is found to have the highest lag around 3 days before. Then what do i do? TL;DR - - how do i best test causality/effectiveness of a given predictor series (transforming data+statistical tests)? - how do i find best lag for the test? - how do i then use this knowledge? ## Answer by safetyduck (score 1) https://quant.stackexchange.com/a/51661 Probably the simplest place to start is to pick some binarized features and targets and stay very low dimensional. You can look at mutual information or other distributional estimates to test for causal linkages. I think typically most causal network discovery is either a full sweep or some heuristic basiced on Lasso in higher dimensions. Prado has a lot of heuristics for stationarity but I have always found it a bit unsatisfying from a core learning point of view. Feels like the need for stationarity should arise from out model of uncertainty and locality.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.