Robust Sample Statistics and Visual Diagnostics for Time Series
Summary
The article presents methods for estimating basic distribution characteristics from a finite sample, with time-series analysis as the motivating use. It discusses the influence of outliers and describes a center estimate built by sorting five measures: the median, midquartile center, full-sample mean, interquartile mean, and midrange. The middle estimate is then used with dispersion and kurtosis calculations to set empirical bounds; observations outside those bounds are flagged and may be removed.
The accompanying MQL5 examples implement outlier filtering and calculations for descriptive statistics, histograms, and normal-probability plots, with Gnuplot used for visualization. These are tools for inspecting assumptions and sample behavior, not a forecasting or trading strategy. The author cautions that outlier removal is difficult to justify in time series because extreme observations may be genuine, and says the algorithms were implemented directly without stability or accuracy analysis for serious applications.
Key ideas
- The article estimates distribution characteristics from a finite sample when the true process parameters are unknown.
- It combines five center estimates and uses their middle sorted value as a center intended to reduce sensitivity to outliers.
- Empirical bounds based on dispersion and kurtosis are used to flag observations for possible removal.
- MQL5 routines and Gnuplot visualizations demonstrate outlier filtering and descriptive statistical diagnostics.
- Removing extremes from time series requires caution, and the examples lack a formal analysis of numerical stability and accuracy.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.