Skip to content
All library documents

Cleaning Market and Reference Data for Quantitative Models

Article Quant Q&A · Author: Dylan Kerler

Summary

The document describes why financial institutions clean vendor and market data before using it in quantitative models. It distinguishes high frequency trade and quote records from reference data such as dividends, maturities, coupons, and amortization schedules. Cleaning may involve detecting outliers, correcting timestamp or price errors, checking missing fields with the source, and replacing unreliable observations with missing values.

For intraday equity data, it points to a published realized kernel procedure that filters abnormal trades and records, and reports that realized variance can change substantially as cleaning rules are applied. The accompanying example highlights outliers and unusual premarket and after-hours activity. These procedures are specific to the data type: the suggested rules for high frequency stock records do not automatically apply to alternative data or other datasets. The document also notes that designing the procedures may require quantitative expertise, while routine execution can be operational work.

Key ideas

  • Noisy or incomplete inputs can undermine quantitative models and strategies.
  • High frequency trade and quote data may require filters for outliers, abnormal trades, and timestamp or price errors.
  • Reference data cleaning includes investigating missing or unusual values and correcting security attributes.
  • Cleaning rules should be tailored to the data source and type.
  • Applying different cleaning rules can materially change realized variance estimates.

Tags

Full text
# What kind of data cleansing/scrubbing are hedge funds doing?


# What kind of data cleansing/scrubbing are hedge funds doing?












It's a well-known fact that several hedge funds have a handful of PhDs just doing data cleansing. All day. Every day.

What kind of data cleansing are they actually doing? Is it really that difficult? How much depth is there in such a topic? Why do they need PhDs to do this?

## Answer by Pleb (score 4, accepted)

https://quant.stackexchange.com/a/66336

#### Data cleaning is important for many large institutions:

"It's a well-known fact that several hedge funds have a handful of PhDs just doing data cleansing". Be aware that many large institutions using vast amount of data for their internal models (banks, pension- and hedge-funds, insurance etc.) usually have their own division for data cleaning and gathering. Often, to strengthen internal quantitative models, companies might rely on external data bought from another firm, which needs further cleaning in order to be reliable.

Employing proper data cleaning is an important part of creating a working quantitative model/strategy, since feeding noisy (improperly cleaned) data into a quantitative model will always yield bad results. In my honest opinion, I do not believe you need to be a PhD to do the job. However, there is a large supply of job-seeking quant developers/IT guys wanting to work in a hedge fund. Thus, hedge funds can be selective and hire "the best of the best" for the job, which is usually PhDs.

#### An example of a simple cleaning procedure:

I've provided a quick example of a data-cleaning procedure for better insight.

When you are working with high frequency trades and quotes (TAQ) stock data (ie. intraday stock data), you need to clean it before the data will be useful. A well-known cleaning procedure is described in Barndorff‐Nielsen et al. (2009). Realized kernels in practice: Trades and quotes. (see section 3.1), which gives you the necessary steps to delete outliers, abnormal trades, misrecordings of timestamps and prices in the database, and more. In the paper, they provide a detailed analysis on how the realized variance changes drastically when applying more of their specified data-cleaning rules (see section 4. Data analysis). However, this cleaning procedure only applies to high frequency stock data and it will differ when you need to clean alternative data.

To conclude the answer, I have provided a graphical illustration of cleaned vs raw (noisy) trade data for a single arbitrary day on SPY. The cleaning procedure follows exactly from the rules provided in the above paper (click the picture for better image quality):

We see how the cleaning procedure is able to detect outliers. Also, notice the odd behaviour of trades in the pre- and after-market hours. This is the main reason for the cleaning step, P1.

## Answer by Dimitri Vulis (score 2)

https://quant.stackexchange.com/a/66339

Some years ago I used to work for a large institution that was not a hedge fund and interacted with some folks who worked primarily on data cleanup. I'll share some observations that I hope may help understand what they do.

They focused on two kinds of data: securities indicative data (stock dividends, nond maturities and coupons), and market data (prices, rates). I think these days, "alternative" data (things like the number of people who visited a particular mall on a given date) has also grown more prominent.

Some of the people developing the processes and procedures for cleanup had PhDs. However I'm pretty certain that no one executing the operational procedures did.

Typical examples of indicative data cleanup - Bloomberg is missing local identifier, amortization schedule, exotic coupon formula... all of these need to be amended in the internal databases.

Typical examples of indicative data cleanup - some value is an outlier, accoring to some criteria, then it needs to be investigated with the source (vendor or internal), and possibly replaced with a "missing" value. An unexpectedly "missing" value needs to be investigated with the source and hopefully populated.

As you see, such examples are seldom quantitative, and much of it could be replaced with AI / automation.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.