Skip to content
All library documents

Handling Sparse, Asynchronous Features in Financial Models

Article Quant Q&A · Author: pat

Summary

The document asks how to prepare daily predictive datasets when features arrive at different frequencies. Its example combines infrequent company earnings information with more frequent variables to predict a daily outcome. Because quarterly events leave many daily rows without a new observation, the author considers carrying sparse values forward but is unsure how to combine those data with features that update more often. The question is about practical data preparation before modeling, rather than a particular prediction result.

The answer points to mixed-frequency data sampling, or MIDAS, as a body of research for modeling variables observed at different frequencies. It notes that much of the traditional motivation is forecasting low-frequency outcomes with higher-frequency inputs, such as updating quarterly forecasts as new weekly data arrive. Reverse-MIDAS models are also mentioned, though the answer characterizes their usefulness cautiously. No implementation, imputation comparison, or empirical example is provided, so the document offers a literature direction rather than a complete preprocessing recipe.

Key ideas

  • Financial prediction datasets may combine sparse event data with more frequent features on a daily grid.
  • Carrying the latest sparse observation forward is raised as a possible preprocessing choice, not validated as a method.
  • MIDAS models provide a framework for combining variables observed at different frequencies.
  • Traditional MIDAS applications often use high-frequency inputs to update low-frequency forecasts.
  • The answer offers no direct comparison of imputation methods or worked modeling example.

Tags

Full text
# how are financial data with sparse and asynchronous features imputed in predictive modeling?


# how are financial data with sparse and asynchronous features imputed in predictive modeling?












I watched a presentation from a large quantitative finance firm that spends a lot of effort around predictive modeling. One of the points the presenter emphasized was that they deal with a lot of asynchronous predictive features. So, for example one feature set might revolve around several company quarterly earnings, and it might be used to predict some future daily statistic related to a single company. The data is asynchronous in that (of course) not all data will arrive at the same time, and it is sparse in that the earnings events might occur with a frequency of only four times per year, for example. In this example, let's just assume that the data features matrix is on a daily time scale resolution. He did not, however, discuss exactly how they might aggregate over or impute the missing data.

I could understand that one might just aggregate the data on a quarterly scale, but in the case he described there are other features that occur on a much higher than quarterly frequency, for the same features matrix (e.g. daily time series). My intuition would be that they might just fill down the empty data for the sparse features.

I'm curious to know if anyone builds these kinds of models, and how they might go about cleaning the data set, before continuing to test some kind of model around it. Any literature with pragmatic examples would be great.

## Answer by Igor Pozdeev (score 5, accepted)

https://quant.stackexchange.com/a/40039

There is large literature on MIDAS (mixed-frequency data sampling) models, the leading scholars being Eric Ghysels and Rossen Valkanov — google their research for references. However, the motivation for these models has mostly been to forecast low-frequency stuff with high-frequency variables, updating, say, quarterly GDP predictions as weekly unemployment figures keep appearing.

Recently, reverse-MIDAS models have been introduced as well (link to one such model), but they seem to me like wrappers for good old regressions with lagged values and of limited use.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.