Handling Exchange Holidays and Missing Values in Market Regressions
Summary
The document asks how exchange holidays should be handled when estimating rolling regressions of portfolio returns on recent market returns. Its concern is that non-trading dates or repeated observations could distort regression inputs and time-varying beta estimates. One answer describes a data-provider convention: some Datastream series carry the last available value forward across non-trading days and may continue doing so after a company becomes inactive. A qualifier that leaves null observations unpadded can help expose holiday gaps; a separate option or inactive-date field can help identify delisted firms.
Another answer outlines general choices for missing data, including omission and imputation, and gives model-based prediction as one possible imputation method. These are alternatives, not a tested comparison. The document does not establish which method the paper being replicated used or quantify bias from holiday handling. In practice, researchers should check the provider's date and padding conventions, distinguish holidays from suspensions and delistings, and align returns to actual trading sessions before choosing a treatment.
Key ideas
- Some market-data series repeat the last observed value on holidays or other inactive dates.
- Provider settings that preserve nulls can help identify unpadded missing observations.
- Delisted securities may require a separate inactive-date rule to prevent stale values from persisting.
- Missing data can be handled by omission or imputation, but the choice affects regression inputs.
- Regression dates should be aligned to actual trading sessions and the data provider's conventions.
Tags
Full text
# How to handle Holidays in Time-Series Datasets? # How to handle Holidays in Time-Series Datasets? Im currently analyzing a Dataset of the German Stock market. While Holidays like Christmas or New Year aren't a problem for Return Calculation or Portfolio Performance, im testing some regressions and don't know how to handle these Dates. I'm regressing the Return of my Portfolio, on the Market Returns of the last ten days. Then im adding the betas up, so i can plot the time varying betas of my sample for every point in time. the right side of my regression looks like this: Do these days have to be cancelled? I don't think the guys of the paper im replicating cancelled out each holiday plus the ten days before. However the regression results would be biased if Non-Trading-days are in the sample. ## Answer by skoestlmeier (score 0, accepted) https://quant.stackexchange.com/a/46742 Preliminary: I assume from your previous post, that you are using Thomson Reuters Datastream. There are several additional parameters available on your Datatypes. Let's for example look at the Datatype `MV`, which is the market value of a company: - If you are using `MV`, than you obtain a time-series, where the last available value is repeated, if a stock is (a) not traded, (b) is suspended or (c) on exchange holidays. The last value is also repeated up to your requested date, if the company went dead! - If you are using `MV#S`, the description from Datastream states: The #S qualifier unpads values where the underlying data point is stored as a null value - so displays N/A for null values rather than pad the last real value. You may use `MV#S` to set (incorrect and repeated) values on exchange holidays to `NA` within your request. However, if you are dealing with meanwhile delisted stocks, you may additionally request `MV#T` to set repeated values after delisting to `NA`, or manually search on Data item `TIME` (or Worlscope Item `WC07015`), which > represents the day on which a company became privately held, merged, liquidated, or otherwise became inactive and (manually) set values newer than this inactive date to `NA`. ## Answer by Attack68 (score 1) https://quant.stackexchange.com/a/46731 When you handle data of any type you might have the issue of missing elements. You can generally handle it by global ommission or by imputation. Imputation can be feed forward, feed backward, some averaging or interpolation scheme. One such scheme might be to use your existing data to build a machine learning algorithm that predicts market movements conditioned on other values, as a Bayesian problem. For example FTSE rallies 100 points, S&P rallies 50 points and your model predicts the DAX rallies X points. You then use X as the imputated value for your data. This approach may strengthen your results otherwise.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.