Skip to content
All library documents

Preparing Order Book and Trade Streams for Modeling

Article Quant Q&A · Author: QMath

Summary

The answer describes basic ways to prepare exchange order book updates and trade records for research. It recommends cleaning updates, removing out-of-order events, retaining a chosen number of book levels, and considering latency, slippage, and market impact when working with time-sensitive data. It frames data organization as part of the modeling problem because storage choices affect retrieval and backtest speed.

One approach is to reconstruct and save snapshots at regular intervals, while storing trades by timestamp for alignment. The answer notes that this can consume substantial resources. It suggests storing updates in object-oriented structures as another way to combine trade and book data and support different retrieval methods, while acknowledging that this takes development effort. It does not answer the question about forecast targets, offer a concrete unified representation, or provide evidence comparing the storage approaches; its guidance is practical and high-level.

Key ideas

  • Clean update streams by handling out-of-order events and limiting retained book depth.
  • Reconstructed snapshots can simplify retrieval but may use more storage and slow backtests.
  • Timestamped trades and book data need to be aligned for analysis.
  • An update-oriented object model can support flexible retrieval at the cost of implementation work.

Tags

Full text
# How are order book and trade data consolidated/distilled into a more(?) tractable form for modeling?


# How are order book and trade data consolidated/distilled into a more(?) tractable form for modeling?












Let's say that there's some asset traded on an exchange and that, for this asset, I have access to a snapshot of the limit order book (price level and quantity for bids and offers) and subsequent updates to the order book of this asset allowing me to reconstruct an approximation to the order book at a point in time as well as trades as they occur. For simplicity, let us also assume that there aren't any hidden orders or other complicating factors.

Is there a common way to combine the streams of order book and trade data into something more homogeneous?

Since in one direction, the trades do hold important information about the order book such as allowing us to determine when a change in the quantity at price level of the order book is likely due to orders being matched or order cancellation and I can't think of an example, but I'm assuming the converse holds as well for information content.

Also, on a related note, what are some of examples of the target variable that a high frequency market-taker may try to model/forecast to trade upon? Are these just things such as the mid and micro price + a potential "spread" with any more informative constructions likely being proprietary?

## Answer by quantinho (score 2, accepted)

https://quant.stackexchange.com/a/76516

Usually you need orderbook snapshot and update data to take into account latency, slippage, market impact etc. All these things are time sensitive and data intensive. The orderbook data itself is large, and you would have to retrieve that in milliseconds timestamps and do something with it. Basically, this process becomes very technical.

First step is to clean this data (get rid of out of order updates, take into account N levels and drop the rest, etc.).

Store the data in time series database. The way you store the data will impact how you retrieve it. Your data probably looks like: Snapshot -> Update -> Update -> ... -> Snapshot. One way would be, to construct snapshot from updates for every N millisecond and store only snapshot. You can do the same for trade data and retrieve based on timestamp. However, this approach is resource intensive and backtest would take a lot of time. A more optimal approach would be to use some OOP technics and store the data as classes. This will give you some freedom to store only updates, combine trade data and use various methods when retrieving the data. These approach will require some dev work but will save resources and makes is easier to use.

I did not understand your second question.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.