Skip to content
All library documents

GMM Resampling and Brute-Force CatBoost Model Selection for Trading

Article MQL5 articles

Summary

The article addresses weaknesses in randomly sampled training labels for a market classifier, including class imbalance, serial dependence among examples, and overlap between classes. It uses principal component analysis to visualize feature structure and contrasts the original data with a clustering-based illustration. K-means is presented as limited because its cluster shapes are constrained, while a Gaussian Mixture Model can represent probabilistic clusters with more flexible covariance structures.

The proposed workflow fits a GMM to features and labels, samples synthetic examples, converts the generated label dimension back into binary classes, and trains CatBoost models on repeated resamples. A brute-force selection step compares models trained on different generated datasets. The examples use EURUSD hourly data and report visual changes such as fewer apparent loops and reduced feature-label correlation. These observations motivate the approach, but the text provides no detailed quantitative out-of-sample comparison or controls demonstrating improved live trading performance. Synthetic resampling also preserves structure learned from the source data, so it does not by itself resolve weaknesses in the original labels.

Key ideas

  • Random sampling can preserve serial dependence, imbalance, and class overlap in training data.
  • PCA plots can help inspect feature and label structure, though they compress the original dimensions.
  • A Gaussian Mixture Model can generate probabilistic synthetic samples with flexible cluster shapes.
  • Repeated GMM resampling provides alternative training sets for brute-force CatBoost model selection.
  • The reported evidence is primarily visual, and improved plots alone do not prove stronger trading performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.