Designing a Normalized Market Data Feed for Trading Systems
Summary
The document explains how to design a normalized market data feed for algorithms consuming data from multiple providers. It recommends working backward from downstream strategy and operational needs: identify which fields applications actually use, then normalize those fields rather than attempting to unify every possible asset class and message type at the outset. Examples distinguish strategy inputs such as prices from fields needed for compliance or post-trade analysis.
The answer also highlights implementation and maintenance concerns. Keeping data in memory, reducing allocations and garbage collection, and limiting transfers between threads can reduce processing overhead. Feed schemas need a plan for deprecated fields and version changes as systems evolve. The guidance is architectural rather than a detailed protocol or product comparison; it cautions that supporting multiple feeds takes substantial ongoing effort. It also questions combining an already normalized provider feed with a direct exchange feed without a clear need, since doing so may add overhead.
Key ideas
- Define normalized fields by tracing the actual needs of downstream applications.
- Start with a narrow set of required fields instead of normalizing many asset classes at once.
- Plan for schema changes and deprecated fields throughout the feed's lifecycle.
- Reduce data copying, allocation, garbage collection, and thread handoffs to limit overhead.
- Supporting multiple market data feeds requires sustained maintenance effort.
Tags
Full text
# Answer by madilyn (score 1, accepted) # How to filter and normalize market data obtained from distinct sources (FIX 4.4, bloomberg, etc) in an algorithmic trading system? I'm wondering if some of you known how to resolve this requirement: I have to define the architecture of an algorithmic trading system (but I'm not an architect, so I'm trying to do my best). I have defined an initial architecture, but just now I'm stuck at the data feed handler component. I mean, the system will receive market data from different sources (bloomberg, FIX4.4, etc) and must normalize that data to produce an usable data feed which it's supposed to be consumed by algorithms to make some calculations an create some orders, something like this: mk data providers => data component => normalize data => usable data => algorithms (consume normalized data) So, I'm wondering if you know a good way to make this or maybe you know a good opensource market data feed handler that can receive data from different providers and produce one clean and normalized stream of market data. I will appreciate very much your answer. Thanks in advance. PD: I have been doing some research and I found this: - http://www.openmama.org/ - http://www.stuartreid.co.za/algorithmic-trading-system-architecture-post/ And for now I'm just reviewing those sites... ## Answer by madilyn (score 1, accepted) https://quant.stackexchange.com/a/18679 If I understand your question correctly, you're asking what's a good design for a normalized feed. This is a somewhat trivial question of (i) picking which data fields (e.g. price, volume) to filter out from each feed and (ii) how to keep that in a trading system with minimal computational overhead. ### Regarding (i) I highly recommend you approach this in a test-driven manner. In other words, figure out what data your application is going to use downstream and 'reverse out' what you need to normalize upstream. A trivial example: If your strategy just needs prices, there's probably little use in normalizing the match numbers. If your compliance and post-trade analysis is keyed by transaction times rather than sequence numbers, then you probably don't need that either. Two tips that I can give you are: - Don't overstretch yourself. It's easy to make the mistake of trying to normalize too many from day 1. You don't want to waste time trying to get spot FX, swaps, equity options, equities, exchange-traded futures etc. all onto the same normalized feed. It will probably take you a very long time before you're dealing with all that, and by the time you have to, you probably have someone else redoing this for you anyway. - Think a bit about how you will maintain depreciated fields and version changes you make to your normalized feed. No matter how exhaustive your initial design, I guarantee you will encounter changes over the course of implementation and real-time use. ### Regarding (ii) Without knowing your exact architecture, the best answer we can give you is a collection of obvious software development tips: keep this in memory, minimize allocation and GC overhead, and minimize the number of times you're shuttling the data from thread to thread. ### Important note If you're trading based on more than 1 data feed, this is not something you want to do alone. It's seriously not worth your time. There's a reason why vendor normalized data feeds are costly, they take a lot of time to maintain. The part I don't get about your question is: Bloomberg is a normalized data feed provider whereas FIX is sometimes used by exchanges directly, why are you normalizing across the two? Then you're just introducing unnecessary overhead in the Bloomberg path. I would just use Bloomberg alone if that's the case.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.