Skip to content
All library documents

Market Data Storage Challenges for Intraday Backtesting

Article Quant Q&A · Author: novice

Summary

The document outlines data management issues that can affect backtests using intraday US equity quotes and trades. Beyond missing records, timestamp alignment, and cleaning unusual trades, it highlights changing symbol histories and trading sessions, survivorship bias, canceled trades, and corporate actions. Preserving historical records and transaction history helps researchers revisit decisions about how such events should be represented in a backtest.

It also describes operational pressures: quote and trade volumes can grow quickly, bursts around the open and close can overwhelm persistence systems, and intraday data may not fit models designed for daily bars. A producer-consumer queue and bulk inserts are suggested for handling ingestion backlogs, while flexible data modeling can ease later changes. These are engineering recommendations rather than comparative benchmarks; the actual storage needs depend on symbols, venues, feed granularity, and the intended backtest.

Key ideas

  • Historical symbol availability and trading sessions should be represented to avoid survivorship and calendar bias.
  • Canceled trades and corporate actions require explicit treatment in backtests.
  • Retaining transactional history can preserve options for correcting data handling later.
  • Intraday quote and trade bursts can exceed database write capacity, so queueing and bulk insertion may help.
  • Data models designed for daily data may not suit tick-level records.

Tags

Full text
# Practical challenges in storing and managing market data


# Practical challenges in storing and managing market data












I am trying to store incoming tick data from the US equity markets on a database.

What are the most common problems that come up with respect to storing and managing such a large data set. A few that I already know are: 1) Missing data - you can add some dummy data or not include this in backtest. 2) Timestamp doesn't match. 3) Block trades/Cleaning the data.

Is there anything else I might be missing and if so how can I go about resolving them. The main purpose is to backtest strategies. The data will not be huge - only Level 1 data for around 20 stocks.

## Answer by madilyn (score 8, accepted)

https://quant.stackexchange.com/a/32537

I've designed large data stores (sitting on about $3-10M of hardware). The challenges that @jharonfe brought up are valid, but I don't feel they're the most interesting challenges in your case (20 stocks, L1).

For example, there's no reason why your system, even it's very naively designed and unoptimized, should fall behind on 20 symbols.

You also mentioned missing data - you might have seen that issue brought up by some of the older, publicly available texts on how to store market data. Bear in mind these were written years ago - that's not really an issue nowadays when the exchange networks are highly overprovisioned.

Here's a few that are worth thinking about:

- Non-homogenous trade sessions. For example, IBM still trades today but there was a period when its ticker stopped trading because of anti-trust investigations. People who design their data storage to handle expirations (e.g. options) tend to think about these issues more, but equities have their own peculiarities. Also, if your equities originate from different exchanges, different markets have different holiday schedules.

- Survivorship bias. This is a pretty obvious one. If you know after-the-fact today that you cannot trade XYZ and that's left out from your symbol universe, yet you were collecting data for XYZ 1 year ago, you're incurring survivorship bias by allowing your backtesting engine to be aware of this and be unable to backtest on XYZ 1 year ago or exclude XYZ from your search universe 1 year ago. The reason this gets challenging is that it's easy to describe but it involves interfacing several pieces of software, which people tend to avoid in the first pass of architecting their platform.

- Trade breaks. In certain venues, you can get info after-the-print that a trade has become nullified. This poses issues in deciding how to handle the fill in your backtest but people often ignore this type of edge case in their first pass design.

- Lossiness and transactional history. As you can see in a common theme between my above 3 points, the solution is often just to ignore the problem and worry about it later. This is fine so long as you don't lose the data when you come around to deal with the problem. This is one of the reasons why using a DBMS can be superior to a flat store, because many DBMSes are transactional and have a commit history.

- Corporate actions. Equities go through mergers and acquisitions, stock splits etc.

## Answer by jharonfe (score 2)

https://quant.stackexchange.com/a/32534

A couple engineering issues I have encountered when persisting high frequency tick data:

- The database can grow quickly. For example, a busy symbol like SPY can have a couple hundred thousand trade events on an typical day. If you track all NBBO quote changes, these are about 2 to 4 times more numerous than trade events. If you track all BBO quote changes (which is included in Level 1 data) then multiply the number of events by the number of market data centers you are tracking. This can be a few million records for one symbol for one day. And on high volatility days, this traffic can double, triple, or more.

- There may be times when tick data is arriving faster than you can persist it to the database. Streaming tick data arrives in bursts, and is particularly busy at the market open and close. I have found that a queue processor implementing the producer consumer pattern is the best design to represent this work flow.

- In cases where tick data is arriving too quickly and your processor is falling behind, you will want to have bulk insert methods to help your processor catch up.

- Designing the data model for intraday data can be tricky. Intraday tick data is different than interday data. On the one hand this is obvious. On the other hand, I find most engineers first work with interday data and encounter intraday data problems second (myself included). Hence they bring their data models from what has worked on interday data, which is natural. Although there are similarities, intraday data often violates many of the assumptions that go into interday data models. My suggestion is in the beginning keep this code as agile as possible. Be deliberate about committing to a data model design. Where possible, leave yourself flexibility to refactor.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.