Skip to content
All library documents

Latency Engineering for TCP-Based Trading Systems

Article Quant Q&A · Author: Peter Mel

Summary

The document discusses approaches to reducing latency in TCP communication between trading machines. One answer recommends a single-threaded, asynchronous, non-blocking reactor that handles multiple connections on a dedicated, isolated CPU core. It reports benchmark measurements for a Java network library over loopback, while stressing that those measurements exclude network and operating-system delays. It also estimates the transmission time for a small message on a 10-gigabit link and mentions kernel bypass as a way to reduce kernel overhead.

A second answer emphasizes that “latency” must be measured across the full path, including market-data reception, parsing, delivery to the application, trading logic, and order transmission. It suggests FPGA-based processing for systems requiring very low end-to-end delay. The benchmark is a vendor developer’s report and is not a universal comparison; loopback figures cannot establish real network performance. The discussion gives no controlled comparison across hardware, network routes, or complete trading stacks, so its recommendations are starting points rather than guarantees.

Key ideas

  • A pinned reactor thread can handle multiple non-blocking TCP connections with reduced scheduling overhead.
  • The reported software benchmark measures loopback performance and excludes network and kernel latency.
  • Kernel bypass can reduce operating-system overhead in network communication.
  • Trading latency should be measured across market-data, application, analytics, and order paths.
  • FPGA processing is presented as an option for reducing end-to-end delay, with implementation tradeoffs left unspecified.

Tags

Full text
# What is the current lowest possible latency for TCP communication?


# What is the current lowest possible latency for TCP communication?












I have two machines over a 10Gb network that need to communicate with each other through a TCP connection. In terms of technology, what is the current lowest latency possible for this to happen? What are the top players in the market using to bring their TCP latency down to the bare minimum?

## Answer by rdalmeida (score 3, accepted)

https://quant.stackexchange.com/a/17624

For ultra-low-latency network applications it is mandatory to use a single-threaded, asynchronous, non-blocking network library. You can and should handle multiple TCP connections inside the same reactor thread which will be pinned to a dedicated and isolated cpu core.

To give you an idea of TCP latencies you can take a look on our benchmarks using CoralReactor, which is an ultra-low-latency, garbage-free, non-blocking, asynchronous network library implemented in Java.

```
Messages: 1,000,000 (size 256 bytes)
Avg Time: 2.15 micros
Min Time: 1.976 micros
Max Time: 64.432 micros
75% = [avg: 2.12 micros, max: 2.17 micros]
90% = [avg: 2.131 micros, max: 2.204 micros]
99% = [avg: 2.142 micros, max: 2.679 micros]
99.9% = [avg: 2.147 micros, max: 3.022 micros]
99.99% = [avg: 2.149 micros, max: 5.604 micros]
99.999% = [avg: 2.149 micros, max: 7.072 micros]
```

We are not aware of any software-based network library, even in C/C++, that can go below that.

Keep in mind that 2.15 micros is over loopback, so I am not considering network and os/kernel latencies. For a 10Gb network, the over-the-wire latency for a 256-byte size message will be at least 382 nanoseconds from NIC to NIC. If you are using a network card that supports kernel-bypass (Open OnLoad) then the os/kernel latency should be very low.

Disclaimer: I am one of the developers of CoralReactor.

## Answer by lehalle (score 1)

https://quant.stackexchange.com/a/17623

You probably want more than a low "TCP latency". Anyway, what latency is it? The TCP layer to connect to your 10Go (i.e. The implementation of the TCP protocol only)? The implementation of a protocol to read the market data (direct native feed? FIX?)? Or to "move" the content of the feed to a zone of memory shared with your application? And what about the messages you want to send to the market? Don't you want them to go fast too?

There are generic answers in Market Microstructure in Practice (Laruelle, Lasnier, Pelin, Burgot and L).

The rule of thumb is if you are under the millisecond for all the layers I listed, you should be happy. There is not one provider for all of them, you have to investigate on your own. My advice is (from the book): if you really want to go fast, use an FPGA with a physical TCP/IP block, and put on it all the protocols. For simplicity you can keep the state machine of your trading logic outside of the hardware, but keep all the analytics inside it.

On the FPGA for datafeed side, you have firms like novasparks.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.