Generating Synthetic Classification Data with Informative, Redundant, and Noise Features
Summary
This function creates a synthetic binary classification dataset for studying feature importance and redundancy in an asset-management machine-learning context. It first generates informative and noise features, then constructs additional redundant features by copying randomly selected informative features and optionally adding Gaussian noise. The output is a feature table and corresponding label series.
Parameters control the total feature count, informative and redundant counts, sample count, random seed, and noise level applied to redundant features. Lower noise preserves stronger substitution between a redundant feature and its source. The function is a data-generation utility rather than a trading rule or empirical market study; its synthetic distributions do not establish how real financial features behave or predict returns.
Key ideas
- The generator creates labeled synthetic data with informative, redundant, and noise features.
- Redundant features are derived from randomly selected informative features.
- A noise parameter controls how closely each redundant feature tracks its source.
- The random seed supports repeatable dataset generation.
- Synthetic results from this utility do not establish real-market predictive performance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.