Skip to content
All library documents

Feature Selection Under Changing Time-Series Correlations

Article Quant Q&A · Author: Kok Wooi Hew

Summary

The document describes a feature-selection problem in quantitative finance. The author splits a time series chronologically into training, test, and out-of-sample periods, then uses Spearman correlations and a p-value threshold to select features. The selected correlations differ across periods, and some features even change correlation sign. This raises concern that a model validated on the test period may fail on later data.

The account highlights how time variation in observed relationships can undermine feature selection and model validation. It does not give a proposed remedy, report a tested strategy, or provide empirical results beyond the author's observations. Readers should treat it as a problem statement rather than guidance on how to select robust features; it leaves open questions about statistical stability, repeated testing, and suitable time-aware validation methods.

Key ideas

  • Chronological splits can reveal that feature correlations differ across market periods.
  • Some observed feature relationships may reverse sign between training and test data.
  • Selecting features by correlation and p-values in separate samples can yield inconsistent sets.
  • The document raises concerns about out-of-sample performance but does not propose a solution.

Tags

Full text
# Methods for feature selection in quant finance dataset


# Methods for feature selection in quant finance dataset












I want to perform features selection on my dataset. I've split my data into train, test and out-of-sample set. The dataset is time-series based, so the split is sequenced in the order that train set will be taken from the earliest segment of the dataset and the out-of-sample set will be taken from the latest segment of the dataset.

I apply spearman ranking on all 3 sets to obtain corrcoef as well as filter for only features with pvalue < 0.02. I noticed that all 3 sets reported a very different sets of correlated features. And in some features, corrcoef maybe positive in the train set, but on the test set is negative.

Unsurprisingly any model trained using the train set and validated using the test set, most likely will perform miserably on the out-of-sample. I suppose this is a very common problem in a quant finance dataset. How can I solve this?

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.