跳至正文
返回文库全部文档

CME数据流测量与低延迟HFT接收器设计

文章 arXiv papers · 作者: Vincent Maciejewski

总结

本研究考察CME市场数据包和撮合引擎交易的到达方式,并利用这些测量结果为高频交易接收器提供设计指导。证据来自对Nasdaq-100 E-mini近月合约一年多的观测,包括数据包和交易的交易所时间戳,并通过实盘生产接收器进行核验。

作者发现,交易突发由撮合引擎塑造,而发布方以大约每7.5微秒的间隔发送数据包。如果接收器在一个这样的间隔内处理完数据包,到达的数据不会形成队列,因此单线程更合适。若处理时间超过该间隔,突发的时序可能导致较长的排队尾部;如果将处理链拆分到两个线程能缩短最慢阶段,则可减少这一尾部,但会增加常规处理的延迟。接近发布间隔时,研究将残余尾部延迟归因于多消息数据包和处理时间的变化。这些发现适用于所测量的数据流和接收器条件,并不能证明相同的线程设计适用于所有系统。

核心观点

  • 在观测到的到达簇中,撮合引擎的交易突发比数据包打包方式更关键。
  • 在所测场景中,若接收器在发布间隔内处理完数据包,就不会因数据到达而形成队列。
  • 如果两阶段线程设计能缩短瓶颈阶段,就可以减少排队尾部。
  • 接近发布间隔时,单条消息的处理成本和处理时间变化比线程数量更重要。

标签

全文
# Packets, Transactions and Queues: Design Principles for HFT Systems from a Measurement Study of CME Market Data


# Packets, Transactions and Queues: Design Principles for HFT Systems from a Measurement Study of CME Market Data









HFT systems are conventionally built as a single-threaded event loop, on the rule that every thread hop adds latency. We test that rule against a measurement study of more than a year of CME market data for the NQ front-month contract, following every packet and matching-engine transaction through the feed's two exchange timestamps, and checking the results against a live production receiver. Packets arrive in near-critical self-exciting clusters that belong to the matching engine's transactions, not to how the exchange packs them. The engine often processes consecutive transactions within a fraction of a microsecond, while the market-data publisher sends at most one packet per publisher period of about 7.5 microseconds, so a burst reaches the receiver as a train of packets one period apart. This yields design principles for HFT systems. First, a receiver that handles each packet within one publisher period never queues on arrivals, however bursty the market; there one thread is best. Second, above that period a queueing tail appears, driven by the timing of transactions, not by packet rate or size, and two threads can be better than one: splitting the servicing chain into two stages on separate threads removes most of the tail at the cost of one hop on the median. Third, only the slowest stage matters, so a split pays only if it shortens it. Fourth, just under the period, where the production receiver runs, the remaining tail comes from multi-message packets and variable service times, and the levers are cost per message and spread of service, not thread count. An analytic framework, a burst-limit throughput identity and an exact reduction of the tandem to a single bottleneck server, supports these results.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。