Choosing Trading Dates for Market Data Imputation
Summary
The document discusses how to identify expected trading dates when filling gaps in historical market data. One option is to use a country or exchange calendar; another is to derive the date set from the records already present across the securities in the dataset. The latter approach is illustrated by collecting dates from files, removing duplicates, and optionally converting and sorting them.
An exchange calendar can provide a reference schedule, but it may include dates for which a particular data pipeline has no observations. The response points out that missing records can reflect operational outages or other repeatable disruptions, including interruptions associated with volatile market events. In those cases, imputing values simply because the exchange was scheduled to trade could misrepresent what the system actually captured. The document offers practical guidance rather than a full missing-data procedure: deriving dates from the dataset preserves its observed coverage, while external calendars can help when the intended schedule is the exchange's official one.
Key ideas
- Exchange calendars can supply expected trading dates for a market or country.
- A dataset's own records can be used to build a deduplicated list of observed trading dates.
- Missing observations may reflect data collection failures rather than ordinary calendar gaps.
- Imputation should account for whether the goal is to represent scheduled exchange activity or actual data coverage.
Tags
Full text
# Getting a list of all trading days?
# Getting a list of all trading days?
I have a large dataset (taken from Kaggle: https://www.kaggle.com/borismarjanovic/price-volume-data-for-all-us-stocks-etfs/), and I would like to fill in the missing data.
To do that I can (1) iterate through all the data and get a set of dates that at least one stock was traded on, or (2) get a reference list with dates that trading occurred.
I would like to know from where I can get the data for option (2). Any help?
## Answer by AlexAbrahams (score 3, accepted)
https://quant.stackexchange.com/a/43988
QuantLib provides calendars for given countries and exchanges, see here.
Dates can then be intersected using something like NumPy's `np.intersect1d`, for example `numpy.intersect1d(cal_dates, numpy.array(db_etfs.loc[:, 'date']))`.
## Answer by madilyn (score 1)
https://quant.stackexchange.com/a/43989
@AlexAbrahams's recommended resource is a decent one but here's another (I think better) approach which solves both (1) and (2):
```
#!/usr/bin/env python
# Call from the `Data` directory
import glob
import pandas as pd
import datetime
# Get all traded dates, with possible repetition
dates = []
for f in glob.glob('*/*.us.txt'):
dates += df['Date'].tolist()
# Remove repetitions
dates = set(dates)
# Optional: Convert to `datetime`
dates = map(lambda s: datetime.datetime.strptime(s, '%Y-%m-%d').date(),
dates)
# Optional: Sort
dates = sorted(dates)
```
Two reasons to consider doing it this way instead:
- You don't need an external, and very large, dependency on `QuantLib`.
- In production, there may be practical, structural or systematic reasons why you can't trade on the days that are in `Quantlib` but not in your data. For example, the network that your servers are on may be down, causing you to lose data on those dates. Sometimes this loss is connected with events of significant volatility, e.g. a circuit breaker tripping on the exchange causing a glitch in your own software. It doesn't make sense to impute data on the trading schedule "as though" you would've been able to trade on those dates because there's a very repeatable reason why you wouldn't have been able to. Your own data captures these nuances the best as opposed to a third party library.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.