Skip to content
All library documents

High-Frequency Trading System Architecture and Data Flow

Article Quant Q&A · Author: May Flower

Summary

The document sketches a high-frequency trading system split between real-time execution and analytical work. The real-time side consumes market data, maintains local limit order books, runs strategies, and creates, updates, or cancels orders. The analytical side uses historical data to build predictive models whose outputs can feed those strategies.

It raises practical design questions about event-driven patterns, in-memory book storage and locking, transactional storage for historical data, and parallel handling of incoming market data for live processing and later analytical normalization. These are questions rather than recommendations: the document does not compare architectures, storage products, latency budgets, or consistency requirements, and it reports no implementation results. Its useful contribution is the separation of low-latency state and order handling from historical model development, alongside recognition that the data pipeline must support both paths. Concrete choices depend on throughput, recovery needs, deployment scale, and the system’s tolerance for contention and delay.

Key ideas

  • The proposed design separates real-time trading from historical analytics.
  • The real-time path processes market data, maintains order books, and manages orders.
  • Predictive models built from historical data can provide inputs to live strategies.
  • The same market data may need both immediate processing and durable analytical storage.
  • Storage and concurrency choices require workload and latency requirements that the document does not specify.

Tags

Full text
# Are there any best practices for designing high frequency trading systems?


# Are there any best practices for designing high frequency trading systems?












I spent some time trying to design some parts of the system, going over the information I found.

At the top-level, the system looks like this

- A "real-time" module that receives market data, processes it, creates a local limit order books and some additional information, there are some "bots" that implement various strategies, and a part that is responsible for creating / updating / canceling orders.

- There is also an "analytical" module, the global goal of which is the creation and use of predictive models, whose predictions are used by some strategies. This part of the system should be able to work with historical data (for the last day, month, etc.)

I have a few questions that are more related to details.

- There are many different design patterns for Event Driven systems. Which of them is preferable for this task?

- What's the best way to store data in a "real-time" module? For example, a local limit order books. Could it just be a data structure in memory? If yes, how can we avoid multiple locks? Can an in-memory database be used, or would it be expensive?

- You also need the ability to access historical data to create predictive models. If I understand correctly, we must first choose an OLTP storage (are there any best practices on this?)

- In this case, the market data flow will, on the one hand, pass through the "real-time" processing module to create / update local limit order books and, at the same time, be stored in some storage with next processing and normalization in OLAP storage ?

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.