Skip to content
All library documents

Robust Methods for Handling Outliers in Financial Data

Article MQL5 code base

Summary

This article explains several ways to limit the effect of extreme observations in financial data, where outliers can distort derived features and model training. It covers percentile clipping, which caps values outside chosen quantiles, and a mean-and-standard-deviation rule that flags values several standard deviations from the mean. The article cautions that the mean and standard deviation are themselves sensitive to extremes, so a single pass may fail to reveal obvious outliers.

It presents median absolute deviation (MAD) as a more robust alternative, using the median and MAD to identify distant observations. It also describes a boxplot rule based on quartiles and the interquartile range, noting that ordinary limits can over-flag data with strong positive skew and heavy right tails. A skew-adjusted boxplot using the MedCouple measure is mentioned. These are general preprocessing heuristics; the article gives no comparative tests or guidance for choosing thresholds across different financial distributions.

Key ideas

  • Percentile clipping replaces observations beyond selected quantiles with the corresponding boundary values.
  • Mean-and-standard-deviation rules can be distorted by the same extreme values they aim to detect.
  • MAD uses the median and median absolute deviation to make outlier detection more robust.
  • Boxplot thresholds use quartiles and the interquartile range but may over-flag strongly right-skewed data.
  • A skew-adjusted boxplot can account for asymmetry using a robust skewness measure.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.