Avoiding Future Leakage When Applying Empirical Mode Decomposition
Summary
The document discusses using Empirical Mode Decomposition (EMD) of EUR/USD prices as input to a machine learning model. It addresses a discrepancy between promising test-period results and poor predictions on the latest observation. The proposed explanation is that decomposing the entire series at once lets EMD components near a training or prediction boundary depend on observations beyond that boundary. This makes the derived features unsuitable for a realistic historical evaluation.
The response recommends calculating EMD using only data available within the relevant past window, and suggests Singular Spectrum Analysis as a related technique with a stronger theoretical basis. It does not provide a tested rolling-window procedure, quantify the effect of leakage or establish that EMD will improve forecasts. The broader question of whether noncausal methods can ever be used in machine learning remains unresolved in the source; for trading evaluation, the key concern is whether each feature could have been computed at the time of the prediction.
Key ideas
- Whole-series EMD can use future observations when producing components near a working-window edge.
- That future information can make historical test results misleading.
- Calculate EMD from past-only data available at each prediction point.
- Singular Spectrum Analysis is mentioned as a related method with a stronger theoretical basis.
- The document does not establish that EMD improves predictive performance.
Tags
Full text
# How to perform Empirical Mode Decomposition? # How to perform Empirical Mode Decomposition? I am trying to use the EMD applied to EURUSD open price to train a machine learning algo (RVM). I have run only once the EMD on my training set and once on the training+test set. The results on the test sets only are quite good. However when I apply the algo on the last sample only the predictions are bad. Shall I run the EMD on each sample of my training set using a sliding window ? I understand EMD is non-causal, but can it be used in some ways for training a machine learning algo ? ## Answer by werediver (score 2) https://quant.stackexchange.com/a/10988 If I correctly understood, you have a big training set and EMD calculated over the whole set at once. Then you use a part of training set and the corresponding part of EMD to infer prediction. The problem here is that you peep into the future having EMD on the edge of the working window calculated using information out of the window. Hence, surely you should calculate EMD (and any other derivative of the training data) using only in-window (past) data. EMD is in some way similar to SSA which has more theoretical foundation. Perhaps, it can interest you. Considering usage of non-casual methods in machine learning, I believe it's not a crime. I have neither proof nor disproof though, therefore I too am interested in the related information.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.