Estimating HAR Models Separately for Each Stock
Summary
The document clarifies how researchers estimate univariate heterogeneous autoregressive models for realized variance across multiple stocks. The cited study does not pool all stocks into one long dataset to estimate a single common set of coefficients. Instead, it fits a separate regression for each stock and index, producing stock-specific parameter estimates.
The answer points to the study’s web appendix, which reports average in-sample estimates across stocks, and its replication code, which loops through stocks and estimates each regression with a heteroscedasticity-consistent procedure. The paper’s main results summarize averages and forecasting performance rather than displaying every stock’s coefficients. A separate response notes that stacking observations would estimate shared coefficients across assets, which is a different modeling assumption. The document does not provide the study’s data or numerical estimates directly.
Key ideas
- The cited study estimates a separate univariate HAR model for each stock and index.
- Stacking all stocks into one regression would impose common coefficients across assets.
- The study reports average parameter estimates and forecast results rather than every individual estimate.
- Its replication materials are described as evidence for the per-stock estimation procedure.
Tags
Full text
# What is the process for using OLS on time series models (HAR like) # What is the process for using OLS on time series models (HAR like) I am reading about HAR models for realised variance and they all seem to use WLS or OLS to calculate the parameters. Now I understand how that works if you just use say the 10 years of AAPL intraday history. However, if you're doing this for a univariate model on many stocks (like it seems the authors do to create stable parameter estimates) how is this done? Do I just put say TSLA's stock at the bottom of the dataframe containing AAPL's and let it do the error minimisation and continue to just add n more stocks to this dataframe to get more and more robust estimates? Edit re Pleb's request: https://www.sciencedirect.com/science/article/abs/pii/S0304407615002584 Page 7 > We complement our analysis of the aggregate market with additional results for the 27 Dow Jones Constituents as of September 20, 2013 that traded continuously from the start to the end of our sample. Data on these individual stocks comes from the TAQ database. Our sample starts on April 21, 1997, one thousand trading days (the length of our estimation window) before the final decimalization of NASDAQ on April 9, 2001. The sample for the S&P 500 ends on August 30, 2013, while the sample for the individual stocks ends on December 31, 2013, yielding a total of 3096 observations for the S&P 500 and 3202 observations for the DJIA constituents. The first 1000 days are only used to estimate the models, so that the in-sample estimation results and the rolling out-of-sample forecasts are all based on the same samples. To me this reads like they're bundling all of this data together into a 1,000 day training set and then doing the OLS on that. They don't give different model parameters for an S&P model and a DJIA Model, or one for each of the DJIA constituents they analysed. ## Answer by Pleb (score 2, accepted) https://quant.stackexchange.com/a/74623 ## They regress univariately on each individual stock and index: In order to conserve space and be within the scope of the main subject (there are also limitations set by the publisher), the authors only show the in-sample model estimation results for the S&P 500 (see Table 3). However, they do run the univariate (H)ARQ models on all individual constituents as-well as the Dow Jones index, but showing these results would be a "waste of space" since it has no direct value to the main research subject (ie. showing the forecast performance of the HARQ model). If you take a look at the Web appendix and the published Matlab code there's direct evidence that they regress univariately on each individual stock: - You can find the average in-sample parameter estimates across all individual stocks in Appendix C of the Web appendix. It also contains additional tables that is not in the original paper. - Opening up `BPQ2016_Replication_Stocks.m` in the published Matlab code (see number 12) you will find the code snippet for the in-sample parameter estimates provided in Appendix C. In essence, they are simply looping over all stocks in their dataset and estimating the regression models using White's adjusted heteroscedastic consistent Least-squares Regression (`hwhite` function). If you have access to Matlab, I'll advice you to run the code yourself and play around with the results. In conclusion, they run the univariate regression models across all constituents thus providing different parameter estimates for each individual stock. However, they are only providing averages of estimates, performance metrics and out-of-sample statistics throughout the main analysis and in the corresponding web-appendix. I hope this helps. ## Answer by Richard Hardy (score 0) https://quant.stackexchange.com/a/74617 You seem to be interested in a multivariate regression (i.e. multiple equations, one for each stock) with a restriction that the parameter values are the same in each equation. If that is so, then stacking the data for each stock on top of each other to produce tall columns is a way to achieve that. You would have the target variable in the first column and the three regressors in the next three columns. Then just run OLS as we know it.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.