Skip to content
All library documents

Feature Selection for One-Month Stock Return Classification

Article Quant Q&A · Author: Matteo

Summary

The document describes a stock-return classification problem with a one-month horizon. The proposed labels distinguish negative and positive cumulative returns, with a threshold separating smaller from larger moves. Candidate predictors include lagged returns, Treasury bill rates, oil and gold prices, Treasury yields, term and default spreads, exchange rates, and broad market index returns. A Random Forest classifier is used, and feature importance is applied to reduce a set of more than one hundred inputs to twenty.

The reported in-sample and out-of-sample accuracy remains around the level expected from random selection among the four classes. The document asks whether literature offers better predictors or strategies, but it supplies no proposed answer, citations, validation details, or evidence that any alternative improves performance. It therefore serves mainly as a problem statement rather than a tested forecasting method. Its results are specific to the described setup; the threshold choice, class balance, feature timing, and validation procedure are not provided, so the reported accuracy cannot be interpreted or generalized confidently.

Key ideas

  • The task classifies one-month stock returns into four signed magnitude categories.
  • The candidate feature set combines price history, rates, spreads, commodities, exchange rates, and index returns.
  • Feature importance reduces more than one hundred candidate inputs to twenty, but accuracy remains near random choice.
  • The document poses a request for better predictors and provides no evidence-backed solution.
  • Missing details about labels, class balance, feature timing, and validation limit interpretation of the reported accuracy.

Tags

Full text
# Optimal predictors for 1-month returns


# Optimal predictors for 1-month returns












I am implementing a Random Forest classifier algorithm on Python for predicting future stock returns (one month). My goal is to foresee whether the cumulative returns in a month will be negative or positive and the magnitude of the movement. The label of the classification is:

- -2 if the cumulative returns are negative and lower than a given threshold

- -1 if the cumulative returns are negative but over the threshold

- 1 if positive but lower than the threshold

- 2 if positive and over the threshold

In doing so, I am using a list of features (predictors) taken from the literature:

- Past $n$ returns

- Some T-Bill rates

- Financial and economical indicators such as Oil price, Gold price, treasury securities yield

- Term and default spreads

- Exchange rate

- Returns on majour indices



At the end I end up with more than 100 features and I apply a feature importance to select the best 20 features. However, both the in-sample and out-of-sample accuracy remain around 25% (as in the case of a random choice).

My question is whether there exist in literature better features or maybe strategies that may improve the predictive power of the model for the horizon I'm considering.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.