Skip to content
All library documents

Choosing and Reviewing Outlier Methods for Financial Time Series

Article Quant Q&A · Author: Malick

Summary

The document compares smoothing based on moving averages with two outlier detection approaches: global winsorization and local outlier factor analysis. Winsorization caps values at selected distribution quantiles, while a local method can flag unusual observations such as a single market event or a data provider error. It also mentions time series preprocessing tools as a practical starting point.

Its main caution is that extreme observations may be genuine market events rather than contamination. Removing major historical moves can materially change the series, and whether a point is an outlier depends partly on the model being used. The suggested methods are not automatic grounds for deletion: thresholds and detections need review in context. The discussion offers no comparative tests or evidence that one method performs best, so method choice depends on the data and modeling purpose.

Key ideas

  • Winsorization replaces values beyond chosen quantiles with the corresponding boundary values.
  • Local outlier methods can identify isolated unusual observations that global distribution rules may miss.
  • A large market move may be important information rather than a data error.
  • Outlier status depends partly on the model and the intended use of the time series.
  • Review flagged observations before removing them.

Tags

Full text
# How to remove outliers in financial times series?


# How to remove outliers in financial times series?












I have a bunch of time series; i need to clean them before modelling. So far I just know the “filtering/smoothing” method : -Ex: moving average methodology (filter the data with a moving average (filter), then obtain a noise (serie minus filter) and remove data points which correspond to a high noise (i.e with a specific threshold) :

(simple) Example of the moving average filter method with three outliers :

Data and filter : Noise and threshold : Cleaned data :

Do you recommend a specific filter ? do you know a better automatic method ?

## Answer by vonjd (score 10, accepted)

https://quant.stackexchange.com/a/10056

Not so fast! I think it is of the utmost importance to first examine whether the data points are real outliers, i.e. noise that is contaminating the data, or perhaps the most important pieces of the time series!

For example when you look at US stock market data of the last 50 years and remove only the ten biggest moves because they are outliers you get a completely different time series!

See page 276 of The Black Swan from Nassim Taleb

So you have to be extremely careful and double check all the data points you remove by whatever available method out there!

In general what you consider an outlier also very much depends on the model you are using. So what seems to be an outlier in one model (e.g. a linear model) is part of the package in a more complex model (e.g. a non-linear model). So it is also a matter of experience how to proceed.

So all in all I think there is no easy answer to your question. A good starting point may be the first chapter of the following new book (2013) which is available online:

Outlier analysis by C. Aggarwal

On a more practical note you can use the forecast-package in R in its new version 5.0 from Rob Hyndman. The new version was just released (27/01/2014) and has upgraded functionality for preprocessing time series and outliers:

http://robjhyndman.com/hyndsight/forecast5/

## Answer by aajajim (score 3)

https://quant.stackexchange.com/a/10055

The is many techniques for Outliers Detection. I separate them into Global and Local techniques.

-One of the Global techniques I usually use is the Winsorization which consiste on replacing the extremes values on the density distribution by the value corresponding to a certain quantile. For example, you replace all the values bellow the 5% quantile by this one, and all the values above the 95% quantile by this one. This technique could be helpful if you want to exclude a certain period, let say a crisis period, from you data to have only the regular period.

-For local techniques, I would recommend the Local Outlier Factor, I discovered it recently, and I believe that it's a good technique to deal with some unexpected events on the market for a singular day, or for data issues from the data providers.

For automatic methods, the first one is very easy to implement, you just need to specify the quantile threshold, the second one is bit more complicated but there is some implementation in R that could help!

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.