Assessing Whether Market Data Fit a Assumed Data-Generating Process
Summary
The document clarifies a passage about estimating whether a security’s observed returns are compatible with a specified data-generating process, or DGP. The answer frames this as comparing empirical properties with a chosen parametric model to gauge estimation error and the extent to which the model’s assumptions fit actual data. Such assessment matters when interpreting model results as if the assumed distribution applied to the security.
An alternative is to generate synthetic data directly from the specified DGP. Those samples match the model by construction, so researchers can test theoretical behavior without first fitting the model to noisy, limited empirical observations. The response presents empirical and simulated data as complementary settings for robustness checks. It offers a qualitative interpretation of the cited passage, but does not define a formal probability calculation, testing procedure, or criterion for deciding that a security follows a DGP.
Key ideas
- A DGP is a specified model for how observations are generated.
- Empirical returns can be assessed for compatibility with a chosen parametric DGP.
- Synthetic observations generated from a DGP match its assumptions by construction.
- Using both empirical and synthetic data can probe model robustness.
- The explanation does not prescribe a formal fit test or probability estimator.
Tags
Full text
# Tactical Investment Algorithms # Tactical Investment Algorithms I am reading paper "Tactical Investment Algorithms" (link) (NOTE: you can download the paper without registration, just press "Download" and then "Download without registration") by the famous Marcos López de Prado. On the page 8 he writes > In practice, it takes only a few recent observations for the estimated distribution of probability to narrow down the likely DGPs. The reason is, we are comparing two samples, where the synthetic one is comprised of potentially millions of datapoints, and it typically does not take many observations to discard what DGPs are inconsistent with recent observations. Another possibility is to create a basket of securities with a returns distribution that matches the distribution of a given DGP. Under this alternative implementation, rather than estimating the probability that a security follows a DGP (Data Generating Process), we create a synthetic security (as a basket of securities) for which a given algorithm is optimal. What does "estimating the probability that a security follows a DGP" mean? What is the probability that the sample is from the distribution? ## Answer by develarist (score 1) https://quant.stackexchange.com/a/58558 Think of it this way. This has to do with using empirical data or artificial data under the pretense that whichever approach you choose must comply with a pre-specified DGP, installed merely to make the model parametric, for example. There are two possibilities: - Accept the real empirical data for what it is and estimate the probability that the empirical security's "properties" matches the parametric DGP in order to have an idea of how much estimation error might be incurred so that results can be adjusted to nominal levels afterwards, or - Simulate artificial data, that is already designed to match the given parametric DGP, so that you don't need to match the articial security to the parametric DGP since the artificial security already matches it by design He's not saying anything ground-breaking. Researchers are just as much interested in testing their financial models on both artificial data and empirical data to prove the model's robustness. If it's too much work to match (not well-defined small-sample) empirical data to a well-defined continuous parametric DGP, then it is merely more convenient and less of a head-ache to get on with the experiments demonstrating the model's efficacy without trying to force-feed non-compliant empirical data into the model, which would be the path to follow to at least confirm the theoretical model.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.