Training a GBDT for a Virtual Stock Prediction Competition
Summary
This tutorial walks through a machine learning workflow for a virtual stock trend prediction competition using BigQuant. It covers importing competition data, converting CSV files to HDF, inspecting data, selecting input features, extracting derived features, and training a gradient boosted decision tree (GBDT). The model is evaluated on a validation set, then applied to test data to produce a submission file.
For time series data, the tutorial advises splitting by the time-related era field rather than randomly sampling across all rows, which can cause overfitting. Its example trains on earlier eras and validates on later ones. The feature construction is deliberately simple, and the page gives no detailed competition results or complete assessment criteria. It frames feature analysis, model tuning, feature engineering, and cross-validation as ways to improve the demonstration.
Key ideas
- Split time-related data by era to reduce leakage and overfitting risk.
- The example uses GBDT to predict virtual stock trends.
- The workflow covers feature selection, derived features, validation, and test-set prediction.
- Feature quality is presented as a key limit on model performance.
- The tutorial encourages further analysis, feature engineering, and parameter tuning.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.