Skip to content
All library documents

Low-Latency FIX Auditing with Asynchronous Logging and Packet Capture

Article Quant Q&A · Author: Alfred Wilkings

Summary

The document considers how an order gateway can retain sent and received FIX messages for audit without adding substantial latency to trading. One approach is to pass messages to a dedicated logging thread through inter-thread communication, allowing disk writes to happen outside the strategy’s immediate processing path. Other suggestions include caching before later writes and using persistent storage where the audit trail must survive a system outage.

A separate approach is capturing network traffic through a switch mirror port and decoding the captured packets outside the application. The discussion highlights the tradeoff between speed and durability: volatile memory can reduce immediate I/O delays but risks losing records, while persistent media better supports a zero-loss recovery objective. Packet capture can reduce application impact, though decoding depends on the FIX engine and capture setup. The answers are operational suggestions rather than a comparative benchmark, and one response includes a vendor affiliation that should not be treated as independent evidence.

Key ideas

  • A dedicated logging thread can move disk I/O away from the trading strategy’s critical path.
  • The required recovery point objective determines whether volatile buffering is acceptable.
  • Persistent storage can preserve audit records through outages, with latency costs to consider.
  • Switch-based packet capture can record FIX traffic outside the gateway process.
  • Captured packets must be decoded, and the practical effort depends on the FIX engine.

Tags

Full text
# Logging FIX Messages


# Logging FIX Messages












I need to persist every single FIX message received or sent by my order gateway for auditing purposes, however it takes more than 1 millisecond to write the bytes to disk. I tried to write in chunks of 64k but that did not help either. Was wondering what is the best practice to audit their trading gateway without introducing latency into their strategies.

## Answer by rdalmeida (score 4, accepted)

https://quant.stackexchange.com/a/15350

The correct way to do file I/O without introducing latency is to do it asynchronously, in other words, the logger thread just passes the message to another thread that is actually doing the disk I/O.

In the past, it was assumed that it was impossible to do it without creating garbage and lock-contention, but with the emergence of lock-free queues it is now possible to do it using pipelining for inter-thread communication with ultra-low-latency.

For example, CoralLog can easily log and persist messages in less than 100 nanoseconds, with throughput number above 4 million messages per second.

Disclaimer: I am one of the developers of CoralLog.

## Answer by Unknown Coder (score 4)

https://quant.stackexchange.com/a/15302

As suggested you could try a RAM disk or some other form of quick caching and then write to hard disk at intervals or even off-hours.

You could also introduce some multi-threading into your application and create a dedicated thread (or service) to handle the logging side exclusively.

## Answer by Michael Green (score 3)

https://quant.stackexchange.com/a/15339

RAM is super fast but not persistent. If you suffer an outage, such as a power loss, all your FIX records will be gone and your audit trail broken. Solid state disks (SSD) in a RAID configuration give a great combination of throughput and reliability.

In my day job I design database applications for a living. When considering system failure and recovery we have two concepts to keep in mind - recovery time objective (RTO) and recovery point objective (RPO). RTO is the estimated time to get a failed system up and running again. RPO is how much data loss you can accept as a consequence of the outage. From the wording of your question it sounds like your RPO is zero i.e. it is not acceptable for this system to loose any FIX records as a result of a system failure. If this is the case you just cannot use volatile storage (i.e. RAM) for this data. It has to go to persistent media immediately. The trick it to minimise the extra latency inherent in persistent storage. If you can, in fact, accept some small data loss during system failure this, needs to be agreed with your users and management and the system costed and designed accordingly.

There are low latency products specifically for persisting high volume data. RIAK is one that springs to mind, though the search engine of your choice may suggest others.

Another alternative would be to take this function out of the application and ask the network to log all traffic that looks likes a FIX message. There are third part tools and appliances which can do this out of the box.

## Answer by John Greenan (score 3)

https://quant.stackexchange.com/a/16454

Don't write a log file. It's a very old fashioned way to achieve the result you want.

Look into packet capture using a span port on the switch to which your FIX engine is connected.

By running packet capture you grab all of the FIX messages sent and received with zero impact on your application - the packet capture just grabs the messages off the wire.

You then need to decode the captured packets. Depending on your FIX engine vendor this may be something that can be done trivially or it may require some work.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.