Cleaning Tick Data Before Backtesting and Analysis
Summary
The document outlines a basic workflow for cleaning tick-level market data before using it in strategies, backtests, or statistical analysis. It identifies duplicate records, implausible prices, out-of-order timestamps, and missing price or volume fields as common problems. The sample process sorts records by timestamp, removes duplicates, drops rows missing key fields, and filters out nonpositive prices and volumes.
The author’s practical point is that data defects can distort results and may be mistaken for unstable strategy behavior. The suggested filters are only a starting point: thresholds and cleaning rules need to reflect the instrument and its trading rules. The discussion offers no measured comparison of cleaned and raw data or validation results, and it briefly favors data sources with clear fields and stable sequencing. Such preferences are anecdotal rather than evidence that a particular source improves trading performance.
Key ideas
- Tick feeds can contain duplicate records, implausible prices, unordered timestamps, and missing fields.
- Sorting by timestamp establishes chronological order before analysis.
- Removing duplicates and rows without valid prices or volumes provides a basic quality screen.
- Filtering nonpositive prices and volumes can remove obvious errors, but rules need instrument-specific tuning.
- Clean input data supports more credible backtests and statistical analysis.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.