Why Empirical Mode Decomposition Repaints in Live Data
Summary
The document describes a practical limitation of Empirical Mode Decomposition (EMD) for streaming financial time series. When new observations arrive and the decomposition is recomputed, the number and values of the intrinsic mode functions can change. The response attributes this repainting partly to spline-based envelope construction and endpoint extrapolation, which require assumptions about how the latest observations continue beyond the available sample.
The author reports trying several EMD variants, including ensemble and noise-assisted approaches, as well as different endpoint treatments. In their experience, these approaches did not make the decomposition stable enough for one-step-ahead forecasting; they view EMD as more suitable for denoising than prediction. They also report difficulties with wavelet methods due to endpoint or phase distortion and mention singular spectrum analysis and sparse coding as possibilities they had not yet evaluated. This is practitioner experience, not a controlled comparison, and the document provides no tested method for preserving old IMF values unchanged as new data arrive.
Key ideas
- Recomputing EMD with new observations can alter earlier IMF values and even the number of modes.
- Spline envelopes and endpoint extrapolation contribute to instability near the end of a time series.
- The response reports that several EMD variants did not solve repainting for one-step-ahead forecasting.
- The author considers EMD more suitable for denoising than for forecasting future IMF behavior.
- The discussion is based on personal experiments and does not establish a generally superior alternative.
Tags
Full text
# What is the best way of updating data while using Empirical Mode Decomposition to analyze # What is the best way of updating data while using Empirical Mode Decomposition to analyze I have a question about EMD updating new data points. For an entire time series, from beginning to the end, the EMD preforms quite good using the cubic spline function. The problem happens when new data points feed in, then after recalculating EMD (including new data) the numbers of output IMFs been change (the IMF data series before and after updating all changed slightly). I suspect that the cubic spline functions do not have memory. What is the best way of avoid this problem, and calculating EMD and keep old results do not change. Best! ## Answer by Philippe A (score 2, accepted) https://quant.stackexchange.com/a/14784 Apologies if it's not a definite answer, but would like to share my experience on this topic. I have been researching EMD for the last 2 years now after reading multiple one-step ahead forecasting papers for financial time series. I have been interested in using each IMF separately to regress them with Relevance Vector Machine and NNs. I have used NI-EMD, EEMD, Statistical-EMD (R package), CEEMDAN, NLMS-EMD and some others. I implemented different versions all with different ways to avoid mode mixing and extrapolating the end point issue. IMF are not only repainting (no memory). Also because of the splines, all artefacts to extrapolate prevent any one-step ahead prediction as we are making assumptions as to which direction last few data are projecting . I even had one researcher working on EMD coding for me this latest implementation Derivative-optimized-EMD, but it failed badly in one-step ahead forecasting. Below is the link to the paper: faculty.nps.edu/pcchu/web_paper/jcam/demd.pdf I have had conversation with researchers like Flandrin and I have been told that the EMD is suitable to denoise but quite unstable for doing prediction on IMFs in the future as modes will repaint and end point is always improper. Hope this help. ## Answer by Philippe A (score 0) https://quant.stackexchange.com/a/15398 Well that is a very good question. Once I realised of the issue with EMD I investigated all wavelets transforms I could access to in Matlab, all seemed to suffer either some end point distorsion or what is called Time phase distorsion for all one-step-ahead forecast where a lag of one sample appear in the forecast at some segments of the test data. I did try an implementation of the Haar A Trou that appeared in dozen of papers, but even with that one I could not get decent one-step ahead prediction. It could be that my learning scheme had some limitation, but I still suspect the multiscale decomposition was the issue, not the machine learning algo (I use Relevance Vector Machine with cross-validation and non-linear kernels). I am yet to try SSA which I think could work, there is an R implementation with a forecasting scheme implemented. I also recently discovered "sparse coding" which I think could be used for time series event though I have not seen implementations. Hope this help and happy to collaborate on algo, just send me message.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.