Cleaning Datastream Equity Data for Return-Based Research
Summary
The document summarizes screening practices for using Thomson Reuters Datastream in empirical equity research, with particular relevance to momentum strategies. It highlights risks from rounded prices, erroneous extreme returns, nonlocal listings, depositary receipts, and securities that are not common equity. Suggested procedures include excluding low prior-month prices, flagging large returns that reverse quickly, checking geography and security-type fields, and reviewing names for indicators of preferred shares, rights, warrants, units, or other instruments.
It also describes screens for major listings, exceptionally high prices, and extreme monthly returns, drawing on published studies. Stale or zero returns at the end of a sample may require attention, though the document does not prescribe a universal treatment. These are research-specific recommendations, not a complete cleaning pipeline; choices depend on the sample, market, and intended analysis. The discussion notes that alternative data sources can also have serious errors, so source quality and screening should be assessed rather than assumed.
Key ideas
- Datastream prices rounded to small increments can distort returns when share prices are low.
- Large returns followed by sharp reversals may indicate data errors and warrant screening.
- Geography, security type, listing status, and security names help identify unsuitable securities.
- Published studies propose additional thresholds for extreme prices and returns.
- Zero-return observations near sample endpoints may reflect stale or padded data and need scrutiny.
Tags
Full text
# What methods of Data-Screening are necessary before starting an analysis with Thomson Reuters Datastream?
# What methods of Data-Screening are necessary before starting an analysis with Thomson Reuters Datastream?
So im currently focussing my research on Momentum-Trading Strategies. I downloaded Constitutents of different All Share indices (including Price, Return index, Market value and Dividend Yield). For what types of errors do i have to look with regard to Datastream ?
Are there some common methods which has to be used before working with Datastream data? (i'm using R for my analysis)
There are for example a lot of days where the Price/Return doesn't change at all for like 5 days. Do i have to exclude these values?
Since the extraction of large Datasets in my university takes a lot of time, are there any other sources (not necessarily free) for high quality data requests? (Like Prices and MV of all stocks traded on a particular emerging market stockexchange for the last 30 years). I think the PC in my University would explode if i would request 5k stocks at once.
## Answer by skoestlmeier (score 3, accepted)
https://quant.stackexchange.com/a/46815
### Preliminary
Thomson Reuters Datastream is one of the most commonly used and widely accepted data-source for non-US data in empirical finance. Working with financial data is based on many filters prior to any calculations. My answer focuses especially on "data cleaning" methods for Datastream, which are published in academic journals and commonly used in research papers - so i do not step into basic filters like e.g. Winsorizing and other methods.
### Ince/Porter (2006)
This paper analyzes US-data from Datastream with the CRSP database and suggests several data-cleaning methods. The conclude:
> We document important issues of coverage, classification, and data integrity and find that naive use of TDS data can have a large impact on econoic inferences.
Filters:
- TDS internally rounds prices to the nearest penny which can cause nontrivial differences in the calculated returns when prices are small. As a solution, drop (i.e. set to `NA`) observations (at the time of your portfolio formation or variable calculations) where the end-of-previous-month price is less than \$1.00 (or in domestic currency).
- The papers finds many instances of data errors, where prices are far to low and returns therefore are to high. The screen for this type of error by setting to missing / `NA` any return $R_t$ above 300% that is reversed within one month. If $R_t$ or $R_{t-1}$ is greater than 300% and $(1+R_t)(1+R_{t-1})-1$ is less than 50%, they set $R_t$ and $R_{t-1}$ to missing.
- Level 1 screening: Remove all non-local firms where data-type `GEOG` is another one than that of the country you are interested in. Exclude American depository receipts and other non-common equity by eliminating all observations where data-type `TYPE` is not equity ("EQ").
- Level 2 screening: This screening requires more effort than level 1 and is based on analyzing data-type `NAME` and screen for key word or phrases that indicate the security is non common equity (e.g. having "REIT" in the name or a percentage sign which indicates a participatory note).
- As described in my answer here, you may use "unpadded" data or follow this paper an delete all zero returns from the end of the sample until the first non-zero return.
### Schmidt et al. (2011)
Appendix A (especially Table A.2) lists detailed information on data-screening.
The most important screening methods in my opinion are:
- Screen for major listings on stocks (data-type `MAJOR`should equal "Y").
- Screen DS04: Set all returns to missing for which the price is greater than 1,000,000 of the domestic currency.
- Screen DS09: Set all returns to missing, for which the monthly return is greater than 990%.
### Campbell et al. (2010)
In addition to Ince/Porter (2006) level 2 name screening, they suggest the following suspicious word parts: "CV", "CONV", "CVT", "FD", "OPCVM", "PREF", "PF", "PFD", "PFC", "PFCL", "RIGHTS", "RTS", "UNIT", "UNITS", "WTS", "WT", "WARR", "WARRANT", and "WARRANTS".
### Conclusion
Serious research is hard work - especially data cleaning and sample constructions. There are many other (often free) sources like Yahoo Finance. However, there are many fundamental errors in these sources, far worse than the flaws mentioned above in Datastream. I may not recommend these free sources for any serious attempt in getting published in (high quality) financial journals.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.