Skip to content
All library documents

Reducing Order Book Data Capture Overhead in Python

Article Quant Q&A · Author: apt45

Summary

The document addresses capturing a live cryptocurrency order book for later reconstruction and analysis. It recommends recording the incoming network traffic for offline processing when practical, which can avoid much of the overhead of handling each message in an application. If capture must happen in Python, it suggests reducing per-message work: avoid building a pandas DataFrame for every update, defer timestamp parsing, write records through a persistent buffered file handler, and consider a faster JSON decoder.

The motivation is a program that writes snapshots and incremental updates to CSV and becomes CPU intensive during longer runs. These are general engineering recommendations rather than measured benchmarks: the answer reports no comparative performance data and does not provide a complete reconstruction design. Packet capture also shifts work to later processing, while direct application-level logging may be preferable when structured records or selective fields are required. The discussion concerns data collection efficiency, not how to estimate price impact from the recorded book.

Key ideas

  • Per-message DataFrame construction and timestamp parsing can add avoidable overhead.
  • Buffered writes through a persistent file handler can reduce frequent disk calls.
  • A faster JSON decoder may improve message processing throughput.
  • Capturing network traffic for offline processing can reduce application-side work.
  • The document offers no benchmarks or full order book reconstruction method.

Tags

Full text
# Efficient way to store orderbook in Python


# Efficient way to store orderbook in Python












I am using the Coinbase WebSocket API to extract real-time data about the orderbook for BTC-USD.

I am using the following code to store the snapshots of bids and asks and the changes to the orderbook everytime there is an update from the exchange.

```
import websocket,json
import pandas as pd
import numpy as np
from datetime import datetime, timedelta,timezone
from dateutil.parser import parse

pd.DataFrame(columns=['time','side','price','changes']).to_csv("changes.csv")
def on_open(ws):
    print('opened connection')
    
    subscribe_message ={
    "type": "subscribe",
    "channels": [
        {
            "name": "level2",
            "product_ids": [
                "BTC-USD"
            ]
        }
    ]
}
    print(subscribe_message)
    ws.send(json.dumps(subscribe_message))

timeZero = datetime.now(timezone.utc)
timeClose = timeZero+timedelta(seconds=61)
def on_message(ws,message):
    js=json.loads(message)
    #print([js['time'],js['trade_id'],js['last_size'],js['best_ask'],js['best_bid']])
    if js['type']=='snapshot':
        print('Start: ',timeZero)
        pd.DataFrame(js['asks'],columns=['price','size']).to_csv("snapshot_asks.csv")
        pd.DataFrame(js['bids'],columns=['price','size']).to_csv("snapshot_bids.csv")
    elif js['type']=='l2update':
        mydate=parse(js['time'])
        if mydate >= timeClose:
            print('Closing at ', mydate)
            ws.close()
        side = js['changes'][0][0]
        price = js['changes'][0][1]
        change = js['changes'][0][2]
        pd.DataFrame([[js['time'],side, price, change]],columns=['time','side','price','changes']).to_csv("changes.csv",mode='a', header=False)
    
    
    
socket = "wss://ws-feed.exchange.coinbase.com"
ws = websocket.WebSocketApp(socket,on_open=on_open, on_message=on_message)
ws.run_forever()
```

In this way, all the changes are saved in a csv file. This code runs for approximately 1 minute, but I would like to make it run for one day and then reconstruct the orderbook.

Once this is done, I want to analyze the orderbook every second to study what is the price impact of buying (or selling) some specific amount bitcoins.

Of course, this code creates a very huge file 'changes.csv', and if I try to make it run on AWS, the CPU usage reaches 90% after some time and the process gets killed. What is the most efficient way to store the orderbook at every second?

## Answer by databento (score 10, accepted)

https://quant.stackexchange.com/a/72025

The best solution is to change how you're doing it completely:

- Store the data before it even gets to your program and process it later. e.g. Use your program merely as a daemon to subscribe to the data. `tcpdump` everything that's incoming to the interface that you're receiving data on. Post-process it later. Chances are that `tcpdump` will be faster and have less CPU or memory overhead than anything you can write inside this Python application.

But if you have to use this Python program for whatever reason, e.g. out of convenience, then the 3 most obvious optimizations you can do are:

- Reduce casting and function calls. Don't cast every event to a `DataFrame` and don't parse datetime on every message. `pandas` objects come with a lot of overhead that you don't need here if the purpose is simply to write to disk. Function calls are expensive in Python because each call creates a new stack frame. If you have to, cast a large number of records at once. Or better, just write straight to disk without the overhead of marshalling a pandas object. Process the datetime in batch afterwards if you have to.

- Buffered writing. Keep in scope a file handler object, e.g. `f = open(fname, 'w')` and `f.write()` the data inside your message handler. By default, Python will buffer the writes with your OS default buffer size. e.g. Create a class, `LOBWriter` and encapsulate `on_message()` as a class method of it; create the file handler as an instance attribute inside `__init__` when you instantiate `LOBWriter`. Remember to manage this object and close it before your program ends.

- Use a faster JSON decoder. e.g. UltraJSON.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.