Using Meta-Labels to Filter Trades by Market Regime
Summary
The document discusses how to train a trading model intended to operate only in low-volatility conditions. The question contrasts supplying volatility as a feature with explicitly excluding unsuitable market states. The response recommends retaining triple-barrier outcomes as directional labels and suggests removing neutral observations when they are rare, while treating many neutral outcomes as a possible sign that profit and loss barriers are set too far away.
For regime-sensitive filtering, the response proposes meta-labeling: a secondary model can identify likely false positives and help determine bet sizes. Candidate inputs include volatility, serial correlation, and skew. It also mentions using synthetic data to explore take-profit and stop-loss levels, citing a method intended to determine trading rules without backtesting. The exchange gives suggestions rather than empirical validation; it does not specify how to define low volatility, evaluate the filter, or avoid leakage when constructing labels and secondary-model features.
Key ideas
- Triple-barrier labels encode which outcome barrier is reached first.
- Rare neutral labels may be excluded, while many neutral outcomes can indicate overly distant barriers.
- Meta-labeling can provide a secondary filter for false positives and support bet sizing.
- Volatility, serial correlation, and skew are suggested as secondary-model features.
- Synthetic data is proposed for exploring take-profit and stop-loss settings.
Tags
Full text
# Labeling and excluding specific market conditions # Labeling and excluding specific market conditions I'm going through "Advances in Financial ML" book and got stuck with something which is not covered there (correct me if I'm wrong). Let's assume I labeled data to 0, 1, 2 according to triple barrier method but I know my model will work only in low-volatility market. So one approach would be just to use volatility as a feature and let ML do its magic by figuring everything out but I believe it results in more noise than if I would explicitly excluded specific market conditions from the model (kind of regime switch). How should I adopt my labelling approach to achieve this goal? ## Answer by Jacques Joubert (score 1, accepted) https://quant.stackexchange.com/a/49365 Using the Triple Barrier Labeling you would use the labels [-1, 0, 1] to indicate which barrier was reached first. You should have very few 0 labels and thus you can remove them from the sample. If you have many 0 labels then you have set your take profit and stop loss levels too high. To determine the TP and SL levels you can use synthetic data to determine the optimal trading rules. The following is the paper the technique is based on: Determining Optimal Trading Rules without Backtesting. If your model only works in a low volatility market then you can make use of the Meta-Labeling technique and fit a secondary model to help you filter out false positives and determine optimal bet sizes. This secondary model will rely on features that will be predicitve of false positives so features such as volatility, serial correlation, skew, and so on will be very helpful.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.