Using Lagged Features and Regularization in Stock Price Regression
Summary
The document considers whether a linear regression can predict a stock-related label more effectively by combining current observations with observations from earlier periods. Concatenating these feature vectors can capture temporal information, but it also increases dimensionality and may introduce redundant, correlated predictors. The response suggests Ridge or Elastic Net regularization to address multicollinearity, with hyperparameters selected using forward-walk cross-validation for time series. It also proposes selecting predictors based on Ridge coefficients before fitting a simpler ordinary least squares model.
For longer histories, the response frames the goal as representing a hidden or evolving state and points to recurrent neural networks such as LSTMs. These are suggestions rather than a tested comparison: the document provides no dataset, performance results, or evidence that one approach will outperform another. Feature scaling, leakage controls, target definition, and out-of-sample evaluation would still need careful treatment in an actual forecasting study.
Key ideas
- Adding lagged observations can give a regression access to information from earlier periods.
- Highly correlated current and lagged features can make coefficients unstable and add redundancy.
- Ridge or Elastic Net regularization can help control correlated predictors, with tuning suited to time-ordered validation.
- Longer histories may motivate sequence models such as recurrent neural networks.
- The proposed methods are not compared empirically in the document.
Tags
Full text
# How to use multi-periods and mult-factors to predict stock price by linear regression?
# How to use multi-periods and mult-factors to predict stock price by linear regression?
Give data in $t_n$ denoted by $[x_1^n, x_2^n, ... x_d^n]$ and label $y_n$ to be predicted. We can just train a $d$-dimensional linear regression $y_n=\sum b_ix_i^n$ to make a prediction. However, I think the data in $t_{n-1}$ can also be helpful to predict $y_n$. So my way is to concatenate data in $t_{n-1}$ and $t_n$ by $[x_1^n, x_2^n, ... x_d^n, x_1^{n-1}, x_2^{n-1}, ... x_d^{n-1}]$ (i.e., multi-periods data) and use a $2d$-dimensional linear regression to predict $y_n$.
The problems of my method are: 1st $[x_1^n, x_2^n, ... x_d^n]$ and $[x_1^{n-1}, x_2^{n-1}, ... x_d^{n-1}]$ seems to be much correlated, so feature can be very redundant. 2nd, some normalization may be conducted on $[x_1^n, x_2^n, ... x_d^n]$ and $[x_1^{n-1}, x_2^{n-1}, ... x_d^{n-1}]$ to form a good concatenation, but how to do? 3rd, if we think that data in $t_{n-100}$ is also help to predict $y_n$, then $100d$ dimension is very high.
So could anyone solve my above problems? or have other ways to use multi-periods data.
## Answer by SachaTheBrave (score 1)
https://quant.stackexchange.com/a/54416
To solve your multicollinearity problem I would first perform a regularization technique such as Ridge or Elastic Net. If you choose Ridge for example, once you have tuned your hyperparameter through cross validation (for time series a forward walk approach is preferable) you can fit after your simple OLS by choosing the predictors with the biggest coefficients from your Ridge regression.
However, I am guessing that what you are trying to do is to build some kind of a hidden state model, wanting your model to know "where you are right now". For that purpose, if you wanted to use 100 previous steps as features, I would rather use a recurrent neural network such as LSTM.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.