Skip to content
All library documents

Modeling Latency and Event Ordering in High-Frequency Backtests

Article Quant Q&A · Author: l.m

Summary

The discussion explains why a high-frequency backtest’s latency assumption should reflect both message timing and the strategy’s sensitivity to event ordering. For spread-crossing or arbitrage strategies, it suggests modeling latency as variable, potentially with a Poisson process, and allowing it to change with market activity. For statistical arbitrage, small local perturbations to the sequence of messages may be more informative than a single fixed delay.

A separate approach synchronizes historical instrument streams into a global market state, then shifts timestamps before synchronization to represent feed or execution delays. Latency can be fixed or sampled from a narrow distribution for colocated systems. The advice is conditional: stored data may already include physical delay, and exchange clocks may be misaligned, so blindly adding a delay can distort results. The discussion does not establish a specific latency value for the proposed Eurex setup; feed, venue, and timestamp details need measurement or calibration.

Key ideas

  • Latency assumptions should reflect the strategy’s sensitivity to message ordering.
  • For spread-crossing strategies, latency can vary with message activity and time of day.
  • For statistical arbitrage, perturbing nearby message order may test robustness to timestamp ambiguity.
  • Synchronize instrument streams into a shared market state before applying timestamp shifts.
  • Check whether recorded data already contains latency and whether timestamps are aligned.

Tags

Full text
# What latency should I use for backtesting a high-frequency strategy?


# What latency should I use for backtesting a high-frequency strategy?












We're developing an HFT strategy for highly liquid futures traded at Eurex. We are planning to colocate our server and to use data feed of QuantHouse and execution API of ObjectTrading. Backtesting is performed on tick data bought from QuantHouse, where timestamps have millisecond resolution (BTW, trades and quotes are sorted separately, so if trades and quotes have a same timestamp, they are not sorted). My question is about latency we should use for backtesting. I define latency as: data feed latency + internal processing time (several microseconds) + execution latency. We simulate this by adding X ms to the feed time (T) of the tick that triggered the trade, then if the last tick with the feed time not later than T + X was a quote we use its book for execution, otherwise (if the last tick with the feed time not later than T + X was a trade) we wait for an incoming quote and its book is used for execution. So my questions are:

- Is our execution model reasonable?

- What X should be used for our setup?

- Could you suggest other (not too expensive) setup in order to reduce any kind of latency?

## Answer by lehalle (score 5)

https://quant.stackexchange.com/a/3578

First the kind of strategy you plan to implement is of importance:

If it is scaling (arbitraging spread crosses: buy at one ask one one venue that is cheaper than one bid in another venue), the kind of approach you plan to use is rational. Nevertheless you should take into account:

- the fact that the latency is somehow not deterministic, use a Poisson process for the duration between two messages,

- the latency is not the same at any time during the day, more or less proportional to the activity (U shaped in the US, with an increase when NY is open in Europe).

If it is a stat arb based strategy, your PnL should not be too sensitive tonthe ordering of the messages. Instead of adding latency, you should backtest with local perturbations on the order of the messages (with for instance a radius of 1 to 10 messages).

## Answer by Quinton Pike (score 1)

https://quant.stackexchange.com/a/41736

I would say the average latency for most data providers ( level 1 CTA feed ) is between 100ms-450ms.

This also varies when there are spikes in trade volume, eg: Market Open/Close.

## Answer by Nikolai Zaitsev (score 1)

https://quant.stackexchange.com/a/46658

To make it properly you should have:

- historical data streams per instrument and

- backtest engine is synchronizing those into one GlobalMarket object (instant quote representation after each tick). Synchronization happens by reading timestamp of each tick and then pushing it into GlobalMarket

If you have this setup, then latency can be added as a shift to timestamp of delayed ticks before synchronization. Shift can be fixed or random following some distribution of latency. Even collocated trade-boxes have random latency following very narrow distribution.

However, be aware of the following.

- The data you store might have physical (actual) latency already. So, do not add latency to those. Or at least, subtract expected latency (which is random) and only add the wanted one

- It makes sense to add latency only if you use data stored by exchanges from their matching engine. But even then, there is a chance that exchange clocks (used for timestamps) are misaligned, likely by fixed value. They should use some atomic clock synchronized between them to make perfect match of their timestamps.

There might be more caveats in it. The above might help you to identify them.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.