Skip to content
All library documents

Sanity Checks for Historical Stock Price and Volume Data

Article Quant Q&A · Author: user2278

Summary

The document discusses quality checks for daily stock records containing open, high, low, close, adjusted close, and volume. Beyond basic consistency rules—such as checking whether the close or open falls outside the daily high-low range—it recommends flagging suspicious placeholder prices and reviewing unusually small or large intraday ranges. Large daily percentage changes can also reveal data errors, though genuine market events may produce extreme moves, so outliers need investigation rather than automatic rejection.

Corporate actions and company history create additional pitfalls. Adjusted closing prices may be revised after dividends or splits, so older downloaded records can differ from current data. Name changes and bankruptcy periods may also complicate series continuity. The responses suggest comparing records against another provider, including checks on randomly selected dates. These are practical screening ideas, not a complete validation standard; provider conventions and legitimate market events must be considered before changing or excluding observations.

Key ideas

  • Check that open and close prices fall within the reported daily high-low range.
  • Flag placeholder values and investigate extreme intraday ranges or daily price changes.
  • Treat outliers as candidates for review because real market events can also create large moves.
  • Account for revisions to adjusted close after dividends and stock splits.
  • Cross-check selected observations with another data source and inspect corporate name or bankruptcy changes.

Tags

Full text
# Scanning a stock database for errors/flaws


# Scanning a stock database for errors/flaws












I'm currently working on some matlab code that is supposed to check a stock database for any errors (missing values, wrong values, etc.). The reason for this is that after reading this post I came to the conclusion that I'll probably have to write some data cleaning code in order to get accurate and reliable results when backtesting with this database.

The database has been downloaded from yahoo finance and contains the following columns for each stock: Date, Open, High, Low, Close, Volume, AdjClose.

So far the program scans for the following trivial errors:

- Close > High

- Close < Low

- Open > High

- Open < Low

- High < Low

The program also checks if any of the data columns contains values less than zero or NaN.

What other errors/flaws could I look for in the database?

## Answer by onlyvix.blogspot.com (score 6)

https://quant.stackexchange.com/a/7543

Few points from my experience:

1 Another filters that you that you should consider is for price = 999 or 999.99 that appears in some data providers.

2 Another set of checks is to look at cross-section of e.g. range = (high-low)/close over all names. Check for the smallest range and largest range to see if the values make sense. You can also check daily % change from one day to another. Check all largest moves for errors in the data. Flash crash in the US have created huge ranges, but if you see abnormal ranges on different days, check out the quality of data. Also September 2008 there are many nonsensical values even in very liquid products.

3 You have to be careful using yahoo (and other sources) for companies changing names, or going in / out of bankruptcy.

## Answer by nitin (score 2)

https://quant.stackexchange.com/a/7549

The adjusted close will change after dividends and stock splits. So the old data will have to be replaced by the new. So it is usually a good idea to check for adj close of the downloaded values against current values.

I also like to check for downloaded data against some other source (like Google). I do this by writing a unit test that will randomly pick a date and download the data from Google and check against Yahoo's.

## Answer by Dmitri Nesteruk (score 0)

https://quant.stackexchange.com/a/7542

What you're talking about is sanity checks and if you really have to do those, it's good grounds to classify the data source as unreliable (or, alternatively, your understanding of said data source is inadequate). Typically, when getting from Yahoo, you shouldn't need to do this.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.