Software Tick-to-Trade Latency and Its Bottlenecks
Summary
This discussion explains why tick-to-trade latency depends on the full market-data and order-routing path, not just the speed of strategy code. Protocol choices, network distance, operating-system and kernel processing, packet parsing, and order transmission all contribute. One response reports an example latency for parsing a UDP update and sending a FIX order, while others describe progressively more demanding software optimizations and latency measurements.
The responses emphasize measuring tail latency, especially the 99th percentile, because a favorable median can hide slower executions. Techniques mentioned for reducing software latency include user-space networking, pinned and isolated cores, lock-free code, profiling, and memory and branch optimizations. The estimates are examples from particular setups rather than universal CPU limits; results depend on the exchange connection, protocols, workload, measurement boundaries, and hardware. The discussion contrasts software with FPGA acceleration but does not provide a controlled benchmark across systems.
Key ideas
- Tick-to-trade time includes networking, protocol handling, operating-system work, and strategy execution.
- The market-data and order-entry protocols affect achievable latency.
- Software optimization can target networking, thread scheduling, locking, branches, and memory behavior.
- Latency should be assessed at the tail, such as the 99th percentile, as well as by the median.
- Reported latency figures are setup-specific and should not be treated as universal limits.
Tags
Full text
# What is the fastest tick-to-trade possible time without FPGAs? # What is the fastest tick-to-trade possible time without FPGAs? I am writing a blackbox model that will react to each market data update (tick) by placing a new order in the market. Without using FPGA, what is the fastest tick-to-trade time that I can expect to achieve with the most modern CPUs? ## Answer by rdalmeida (score 7) https://quant.stackexchange.com/a/24320 That will depend on the protocol you are using for market data (UDP, FAST, MDP, ITCH, etc.) and order routing (FIX, OUCH, etc.). For example, the latency to parse a UDP tick and to place a FIX order was around 8 micros as measured by tcpdump, using CoralFIX and CoralReactor. Disclaimer: I'm one of the developers of CoralFIX ## Answer by experquisite (score 4) https://quant.stackexchange.com/a/42582 It is pretty easy to get down to 50us tick-to-trade measured from start of MD packet on wire at switch to start of trade order at the switch, at the 99th percentile. 10us takes user-space networking, lock-free coding, isolated and pinned cores, profiling and kernel tweaking. (So, still not terribly hard). 5us (99th percentile) takes cache optimization and allocation, branch reduction, and TLB/memory optimization. This is decent, and where many people should stop. 2us (99th) is very hard to do with software, and probably not worth the marginal effort over FPGA (where 2us is, in turn, relatively easy). State of the art FPGA is below 200ns now, probably lower. EDIT: when evaluating vendors, always be sure to ask for 99th percentile numbers. 2us median is pretty easy, but 2us in tails is hard. ## Answer by lehalle (score 1) https://quant.stackexchange.com/a/24316 I do not understant your question. In terms of latency, you have - your distance to the exchange - the network latency - the tcp/ip layer - your operation system - the speed of your code If you just talk about the last step (do not forget you have to count the other steps twice) and you want to go as fast as possible, just write in assembly. In most assembly manuals you will find the execution time for each instruction (for generic considerations about how to count speed, you can have a look at this link). If you write in C, you can have a rough idea of the exec speed if you keep in mind how the compiler translate your code in assembly. But frankly, this last step is from far the fastest of the five steps... Using fpga is about to go fast at tcp/ip and os steps mainly. ## Answer by Ariel Silahian (score 0) https://quant.stackexchange.com/a/24314 I would say something around 1/5 milliseconds. Are you connecting through fix? Any specific engine?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.