Using Logistic Regression to Rank CSI 300 Stocks by Predicted Returns
Summary
The article introduces logistic regression as a binary classifier that maps a weighted combination of features to a probability. It outlines fitting the model by maximizing the likelihood of observed labels, equivalently minimizing a loss with gradient descent, then assigning a class based on the predicted probability. The proposed equity strategy uses seven technical and fundamental features for CSI 300 constituents, including market capitalization, valuation measures, OBV, Bollinger Bands, KDJ, and a base-building indicator. Features are standardized using a rolling training window, and labels indicate whether a stock exceeds a specified forward return threshold.
At each scheduled rebalance, the model ranks stocks by the probability of the positive class and equally weights the top three, selling holdings that no longer qualify. The article gives a backtest date range and starting capital, but no reported return, risk, or benchmark comparison. It notes that the return threshold, rebalance interval, and model parameters remain candidates for tuning, and points to cross-validation as a way to select parameters.
Key ideas
- Logistic regression estimates the probability of a binary outcome from a linear combination of features.
- The described model trains on rolling standardized features for CSI 300 stocks and labels based on later returns.
- The strategy equally weights the three stocks with the highest predicted probability at each rebalance.
- The article supplies backtest setup details but no performance statistics or benchmark comparison.
- It identifies label thresholds, rebalance frequency, and model parameters as areas for further validation.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.