Causal Smoothing and Trend Detection for Live Market Data
Summary
The document describes a practical signal-processing problem in a trading system that receives live bid and ask data. The author reports that offline backtests can use non-causal smoothing methods with access to future observations, but those methods lag when applied causally in live trading, delaying signals. Several approaches are listed, including moving averages, Savitzky-Golay and Whittaker-Eilers smoothing, PDE-based methods, Kalman filtering, and low-pass filtering. The author found that a least-squares moving average could help but produced overshoots and undershoots during trends.
Incremental curve fitting and segmented linear regression are suggested as possible directions, but neither is developed or evaluated. The document offers no comparative results, implementation details, or recommended solution. Its main practical lesson is that testing must preserve the information constraints of live use: a smoother that uses future data can look effective historically while generating delayed or unrealistic signals when deployed causally.
Key ideas
- A non-causal smoother can use future observations in historical analysis, so its apparent performance may not carry over to live use.
- Causal filters often lag price changes and can delay trading signals.
- The author reports trying several conventional filters but does not provide a controlled comparison of their performance.
- Least-squares moving-average estimation showed promise in the author’s experiments but also produced overshoots and undershoots.
- Incremental curve fitting and segmented linear regression are raised as ideas, not validated solutions.
Tags
Full text
# What would be a good way of smoothing live/real time data or recognizing trends within live data consistently? # What would be a good way of smoothing live/real time data or recognizing trends within live data consistently? I receive live data stream ( ask and bid data ) from a particular data source, I then proceed to save this data in the TimescaleDB and later on process it for both live and historical testing ( forward-testing, backtesting and production ) ( I don't do high frequency trading, only a couple of trades ( ~10 ) per day ). So, after some experiments on historical data I found out that results of my backtesting ( and later on forwardtesting, which are almost identical ) essentially depend on how good I smooth my data/recognize the trends within data. Of course, in an offline setting ( during backtesting and experimenting ) I can apply various non-causal filters/smoothers and this would work perfectly since filter/smoother has information about future data however in an online setting ( during forwardtesting and production ) these filters/smoothers lag and create the effect from the following image: I then utilized different filters and methods in all stages of testing aside from production like the following: - Moving Average ( and different types of moving averages like least square, hull etc... ) - Savitzky-Golay method - Whittaker-Eilers method - Heat equation method - Perona-Malik PDE - Kalman filter - 1 euro filter - Low pass filter And all perform poorly in an online setting since they all create this lag effect. I found some luck with utilizing estimation of least square moving average but that method creates too much over and undershoots in an trend so I assume I will have to combine it with some other specific method ( I don't have knowledge nor experience for something like that but I am still open to that idea ). So, essentially, after some time I realized that these filters and methods are not really an solution in fully causal setting since they create this lag and this essentially creates late signals in my algorithm. I haven't tried the following, but I've got an idea of utilizing incremental curve fitting however the problem is that I need my curve to be consistent with historical data points while being open to introducing new data points and for this I haven't found an solution. Another idea is using segmented linear regression to estimate trends but I haven't explored this idea fully. I am happy to hear different experiences and answers related to this specific issue because I am almost sure that someone had to deal with this while creating trend-following algorithm. And keep in mind that I am new to this and signal processing as well, so be as detailed if you will and it would be great if you had a specific paper related to this, it would be of great help and greatly appreciated!
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.