Why Class Weights May Need to Account for Market Regimes
Summary
The document asks whether a training set spanning different market regimes should use one global class weight or separate weights for each period. Its example is a two-year sample with a bullish first year and a bearish second year: a global adjustment for overall class imbalance might also downweight bullish observations from the already bearish period. The author raises this as an intuition and asks how to formalize it; the document does not propose or evaluate a method.
The issue connects class weighting to changing label distributions over time. Period-specific weights are posed as a possible response, but the discussion gives no empirical evidence that they improve a trading model. Any such approach would need to define periods without relying on future information and be evaluated on time-ordered out-of-sample data. The document leaves open whether regime-specific weighting is appropriate and how to choose the periods or weights.
Key ideas
- A global class weight can obscure changes in class balance across market regimes.
- The example contrasts a bullish period with a later bearish period in the same training set.
- Period-specific class weighting is proposed as an idea, not demonstrated as an effective method.
- Any weighting scheme needs time-ordered out-of-sample evaluation.
Tags
Full text
# Should we split data into several periods before calculating class weight? (Advances in Financial Machine Learning) # Should we split data into several periods before calculating class weight? (Advances in Financial Machine Learning) In the book, section 4.8 class weights, Marcos suggests applying class weight, which I agree because sometimes you have more bullish price action than bearish price action e.g. 52% of the time is bullish. But he doesn't talk about splitting the dataset. My intuition is that if the training set has 2 years of data, and the 1st year is bull market whereas 2nd year is bear market, simply applying the same class weight throughout might not be good enough. Because it implies decreasing the sample weight of bullish samples in 2nd year too. But samples in 2nd year are already in bear market, so it doesn't make much sense to try to make it more bearish. I'm not sure how to formalize this intuition yet. Do you think we need to split dataset into periods to apply different class weight for each period? And how would you go about it?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.