High-Throughput Trade Accounting: Logs, Queues, and Data Integrity
Summary
The discussion considers how to record trades at high speed while preserving balanced accounting entries. One proposed design uses an append-only log, fast insertion, and periodic summary transactions; for distributed systems, it raises the possibility of sending summaries between nodes. It also notes that some trading systems log executions and perform simple calculations before reconciling them with an accounting database in batches.
The responses emphasize that accounting has stronger data-integrity needs than many applications commonly assigned to NoSQL systems. One respondent reports handling about 2,000 transactions per second on a mid-range SQL Server and recommends separating front- and back-office work with a message queue. Another view is that operational systems may track and reconcile trades internally while formal accounting uses consolidated totals. The discussion offers personal experience and architectural suggestions, not a comparative benchmark or a validated design for the proposed workload. Its throughput claims and recommendations may depend on system configuration, accounting requirements, and how quickly reconciliations must be available.
Key ideas
- Append-only trade logs and periodic summaries are proposed as ways to separate fast execution capture from formal accounting.
- Accounting systems need reliable data integrity because entries must balance and remain reconcilable.
- A message queue can decouple the trading front end from back-office processing.
- The discussion distinguishes detailed internal trade tracking from consolidated accounting reports.
- The reported SQL Server throughput is an individual example, not a controlled comparison.
Tags
Full text
# Non-SQL methods for high-frequency accounting? # Non-SQL methods for high-frequency accounting? Does anyone know of any prior art for non-SQL data structures for high-frequency accounting, whether client, broker, or exchange-side? I'm thinking specifically of the problem of booking individual trade data into proper transactions, with balanced debits and credits. In my own case, I'll be doing this in or directly adjacent to a fast limit-order book, but I can see other reasons for such a beast existing. And yes, I agree that none of the current raft of popular non-ACID nosql engines are at all right for this job. I'm assuming I'm going to need to write this. A usable answer to this question might be as simple as a link to a paper on the subject of non-SQL or nosql accounting in a high-volume trading context -- I'm obviously using the wrong combinations of search terms, because I'm not finding much yet. What I'm working on is a project that includes a limit-order book and accounting on each node in a distributed grid or fabric. In my case, the traded instruments could best be described as real options or real derivatives, including some mild exotics. The vast majority of the orders would be initiated by machines, and the data rate looks like it could easily hit 60k trades/sec on each node. (Without going into a longer dissertation, it might help to explain that I'm in Silicon Valley these days; this is obviously for a new market, not any existing one.) See http://en.wikipedia.org/wiki/Real_options_valuation if you haven't run across real options before. Partial answers, based in part on feedback to this question so far: - A purpose-built accounting mechanism would probably be log-structured, append-only, probably using a non-SQL API for insertion speed. The engine itself might be a hypergraph database. If running on multiple nodes, it would need a way of providing summary transactions to the other nodes in a peer-to-peer fashion. The more I dig into this, the more it's starting to look like a distributed hypergraph. https://mathoverflow.net/questions/13750/what-are-the-applications-of-hypergraphs http://martin.kleppmann.com/2011/03/07/accounting-for-computer-scientists.html - In the HFT world, it sounds like the standard procedure is still: Log but do not index the trades, do simple arithmetic ignoring debits and credits, and then synthesize balanced summary transactions to the accounting RDBMS periodically. Run MTM in batch. Is there anything anyone can say about how that "simple math and local logging" is done? I know how we did this in the derivatives world 15 years ago, but frankly it and MTM were both slow and ugly, and involved NFS servers, flat files, and shell scripts. Has nothing changed? ;-) - Okay, removing 'accounting' from the search terms just now found me this -- different question at first glance, covering both tick and financial data, but worth reading through -- looks like he had some of the same thoughts: Usage of NoSQL storage in Finance Looks like it would be worth repeating my searches in google, citeseer, etc., substituting "finance" for "accounting". - Complex Event Processing (CEP) tries to solve some of the same problems -- it just occurred to me that including CEP in the same searches might be fruitful. The first thing I found was this (skeptical but humorous) article discussing CEP's slow uptake and some of the nosql hype: http://www.hftreview.com/pg/blog/darkstar/read/32333/whats-wrong-with-complex-event-processing ## Answer by mepuzza (score 4) https://quant.stackexchange.com/a/3102 I know this is probably a naive answer, but when I started doing data analysis for personal trading I looked for something much faster than SQL. I program in C++ and I found that HDF5 was the answer to all my problems http://www.hdfgroup.org/HDF5/ It's not accounting oriented, but the nice thing about it is that you can do almost anything with it and it is very fast. A bit of a learning curve though ## Answer by TomTom (score 3) https://quant.stackexchange.com/a/3071 > I have to think that there are a lot of very fast, very optimized special-purpose accounting engines out there filling this role. Yes and no. I do not think you are high volume at all - you just have a corporate-level server for the database, not a cheap low-end hosting. I do about 2000 transactions per second on a SQL Server with a mid-range database. The core will be: - Decouple front and back with a message queue anyway. - Take trade executions from a FIX backoffice link that reports from clearing / broker. > it seems like a huge waste of data center horsepower when a more modern purpose-built, probably non-SQL accounting engine might be orders of magnitude faster. There is one thing amiss: SQL has data integrity, while NoSql is often written ignoring data integrity requirements. You can get away with a lack of data integrity for a LOT of stuff, but not with accounting. You also miss that accounting is a standardized commodity side. Large companies run something like SAP - and want all their data to be in there, regardless of costs. It is not a waste of time to upgrade the one central system doing your payroll, all invoices for the organization, etc. on top of trade accounting. Also it is a question whether accounting really needs every trade - back office yes, to consolidate and check, but accounting is OK with synthesized balanced summaries. I do not do a lot of trading so far but submit monthly PNL totals with broker statement to my accountant (where it goes straight to my monthly profit / loss and tax calculations). I never will do different , even when volume ramps up - but will consolidate daily or hourly and correlate INTERNALLY, but not for accounting.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.