Skip to content
All library documents

CME Feed Measurements for Designing Low-Latency HFT Receivers

Article arXiv papers · Author: Vincent Maciejewski

Summary

This study examines how CME market-data packets and matching-engine transactions arrive, and uses those measurements to derive design guidance for high-frequency trading receivers. Its evidence comes from more than a year of observations of the Nasdaq-100 E-mini front-month contract, including exchange timestamps for packets and transactions, with a check against a live production receiver.

The authors find that transaction bursts are shaped by the matching engine, while the publisher spaces outgoing packets at roughly 7.5 microsecond intervals. If a receiver services packets within one such interval, arrivals do not create a queue and a single thread is favored. Above that service time, burst timing can produce a long queueing tail; splitting the servicing chain across two threads can reduce that tail if it shortens the slowest stage, though it adds latency to typical handling. Near the publisher interval, the study attributes residual tail latency to multi-message packets and variable service times. These findings concern the measured feed and receiver conditions; they do not establish that the same thread design is optimal for every system.

Key ideas

  • Matching-engine transaction bursts, rather than packet packing, shape the observed arrival clusters.
  • A receiver that processes packets within the publisher interval avoids arrival-driven queues in the measured setting.
  • A two-stage threaded design can reduce queueing tails when it shortens the bottleneck stage.
  • Near the publisher interval, per-message cost and service-time variability matter more than thread count.

Tags

Full text
# Packets, Transactions and Queues: Design Principles for HFT Systems from a Measurement Study of CME Market Data


# Packets, Transactions and Queues: Design Principles for HFT Systems from a Measurement Study of CME Market Data









HFT systems are conventionally built as a single-threaded event loop, on the rule that every thread hop adds latency. We test that rule against a measurement study of more than a year of CME market data for the NQ front-month contract, following every packet and matching-engine transaction through the feed's two exchange timestamps, and checking the results against a live production receiver. Packets arrive in near-critical self-exciting clusters that belong to the matching engine's transactions, not to how the exchange packs them. The engine often processes consecutive transactions within a fraction of a microsecond, while the market-data publisher sends at most one packet per publisher period of about 7.5 microseconds, so a burst reaches the receiver as a train of packets one period apart. This yields design principles for HFT systems. First, a receiver that handles each packet within one publisher period never queues on arrivals, however bursty the market; there one thread is best. Second, above that period a queueing tail appears, driven by the timing of transactions, not by packet rate or size, and two threads can be better than one: splitting the servicing chain into two stages on separate threads removes most of the tail at the cost of one hop on the median. Third, only the slowest stage matters, so a split pays only if it shortens it. Fourth, just under the period, where the production receiver runs, the remaining tail comes from multi-message packets and variable service times, and the levers are cost per message and spread of service, not thread count. An analytic framework, a burst-limit throughput identity and an exact reduction of the tandem to a single bottleneck server, supports these results.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.