Designing Transparent Synthetic Market Data from Real Samples
Summary
The document asks whether researchers can create synthetic datasets through explicit, context-aware choices rather than relying on a generative AI model or a uniform library. Its proposed workflow begins with a problem context and a small set of candidate real datasets, identifies the traits the synthetic data should preserve, and uses classical probabilistic distributions or other transparent methods to construct an artificial counterpart. It also suggests statistical tests to assess whether the synthetic sample represents the relevant real data.
The motivation is to reduce reliance on repeated or homogeneous datasets that may contribute to overfitting, while giving domain experts control over assumptions and trade-offs. The example is only illustrative: it mentions matching statistical properties for market data but does not specify which properties, distributions, tests, or acceptance criteria to use. No implementation, empirical comparison, or evidence that this approach reduces overfitting is supplied. The text is therefore a research question and design proposal rather than a reproducible synthetic-data method; faithful market simulation would also need to consider dependence, tails, and changing market regimes.
Key ideas
- The author seeks transparent, non-generative methods for creating synthetic data from contextual knowledge and real dataset candidates.
- A proposed process would select relevant real samples, identify target traits, and build a synthetic counterpart using explicit statistical choices.
- Statistical tests could help assess how well synthetic data represents the selected real datasets.
- The goal is to make dataset construction less homogeneous and expose assumptions and trade-offs to domain experts.
- The document proposes no concrete distributions, tests, or evidence that the approach reduces overfitting.
Tags
Full text
# 85396 # Repeatation of datasets might lead to overfitting.Is synethetic data creation(no genAI) possible given the context and the sample of real datasets? Is there any work done on someone creating synthetic data out of a context and a few candidates of real data.Is there any methodology Edit: By creating,I do not mean to say 'generating' as in using generative AI,rather a more mindful and transparent way which allows us to make the choices ourselves knowing well about the trade-offs and whether or not they are good representatives given the context and a few set of real dataset candidates. One can frame it like this: Anything more transparent(less blackbox-ish) because there might be free synthetic data libraries anyway but not usin them and creating yourown on your own is about independence and not trusting anything. Any method more classical which takes care of the context of the problem,the real datasets that fit there and then capturing several traits of the real dataset that we would like to be present in the synthetic one as well. Maybe even perform statistical tests to see whether the synthetic ones are good representaives of the real dataset candidates. Say,for a very trivial artificial counterpart of market data,if we take help of several probabilistic distributions we gotta match statistical properties given the market and problem at hand. GOAL : Goal of the methods are really simple it would be less homogenous and those who understand the problem and the context would make better datasets,making it anything but a one-step import library and get a uniform-for-everyone process rather something that empowers those who understands the problem and context at hand, and also rewards them with an edge. In the pre-AI days,researchers must have worked on more fundamental ways,isn't it? AI NOTE: The idea is not to be anti-AI,but to be more mindful about the process and once we know how to do that ourselves,then automate it step by step instead of one genAI method which you don't know why it does what it does.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.