Skip to content
All library documents

Designing Tick Data Capture and Storage for Analysis

Article Quant Q&A · Author: Marko

Summary

The document surveys ways to capture, store and later analyze market tick data. It describes collecting Interactive Brokers market updates through an R interface and writing records to files or in-memory time-series objects. For longer-term analysis, contributors favor indexed or memory-mapped access and column-oriented binary formats, which can make queries over large datasets faster than repeatedly reading flat text files.

Storage choices depend on downstream needs: local solid-state disks, databases, serialized blobs, and event-processing systems are all discussed. One example describes querying subsets from a dataset with tens of millions of rows using limited memory, while another routes incoming events to both a datastore and real-time processing. The answers disagree about whether Interactive Brokers supplies true tick data and characterize its updates differently; feed capabilities, API limits, update intervals, and software ecosystems may also change. The recommendations are practical anecdotes rather than a controlled comparison, so format and architecture should be matched to workload and data-feed requirements.

Key ideas

  • Choose a storage format based on how the tick data will be queried and analyzed later.
  • Column-oriented binary storage can improve access speed for large analytical datasets.
  • Memory mapping and indexing can support selective queries without loading all records into RAM.
  • Streaming systems can process incoming events while also persisting data for later research.
  • The answers give conflicting descriptions of Interactive Brokers tick-data availability, so verify current feed capabilities.

Tags

Full text
# How should I store tick data?


# How should I store tick data?












How should I store tick data? For example, if I have an IB trading account, how should I download and store the tick data directly to my computer? Which software should I use?

## Answer by Jeff R (score 25)

https://quant.stackexchange.com/a/497

Using IBrokers from R is going to be the easiest route. A quick example of capturing data to disk would be:

```
library(IBrokers)
tws <- twsConnect()
aapl.csv <- file("AAPL.csv", open="w")

# run an infinite-loop ( <C-c> to break )
reqMktData(tws, twsSTK("AAPL"), 
           eventWrapper=eWrapper.MktData.CSV(1), 
           file=aapl.csv)

close(aapl.csv)
close(tws)
```

This will send CSV style output to disk. Additionally the data can be stored in xts objects within the loop which can be appended to/filled to provide a constant in-memory object to use for analytics. Objects can be shared with many tools - including using the RBerkeley package on CRAN to share objects with other programs with Berkeley DB bindings. This latter approach, if managed intelligently is very, very fast.

Given the symbol limit of IB (100 concurrent more or less) and the 250ms updates - R can typically handle all of this without breaking a sweat (i.e. the JVM running IB's TWS or even IBGateway client is likely to be far surpassing the R/IBrokers process in terms of CPU usage).

You can even extend the syntax above to record more than one symbol by passing a list of Contracts, increasing the number on the eWrapper, and making sure you have a suitable list of files to write to.

In terms of something closer to long-term storage/access, the packages Josh referred to (mmap and indexing) are also very useful. I've given talks with some basic options data examples that are 3-4GB in size without derived columns (12GB total), and I can pull using R-style subsetting syntax any subset I need nearly instantly. e.g. finding 90k+ contracts for AAPL in 2009 (out of 70MM rows) took tens of milliseconds. All without keeping anything in RAM, and all running on a laptop with 2GB of RAM.

I'll likely get some more presentation material for the latter packages put together soon, and will be giving some talk(s) at the upcoming R/Finance conference in Chicago. I am also planning on some public workshops through lemnica related to R and IB for 2011.

## Answer by chrisaycock (score 18)

https://quant.stackexchange.com/a/468

What do you want to do with the tick data later? Run analytics? You can save tick data to a flat file for all the software cares, but that would be really slow to access later.

Instead, you should ideally save the data:

- Column-oriented - all elements in a field are stored contiguously for better caching

- Binary - all elements are ready for immediate use; no lexical casting required

There are number of column-oriented databases, though no production quality ones are open-source at the moment. You can try the non-commercial version of q/kdb+ to see what you think of it, though it's a huge learning curve if you aren't used to it already.

Something else to think about when storing tick data is the physical medium. Ideally you'll want:

- Local storage - fetching across NFS is going to be painful

- Solid state - fetching from disk is also painful

## Answer by Joshua Ulrich (score 13)

https://quant.stackexchange.com/a/493

If you're planning on analyzing the data later in R, you should take a look at the indexing and mmap packages. Though, as @chrisaycock said, you'll need to save the data in a column-oriented, binary format.

If you're downloading the IB data with R, using IBrokers, you can write your own eWrapper to store the data in whatever format you want.

## Answer by Steve Severance (score 4)

https://quant.stackexchange.com/a/476

I am using just a filesystem to store raw tick data. I am using protocol buffers to easily allow multiple languages to consume the data. Part of the reason is that I am moving more stuff onto Amazon's EC2 to use their GPU compute instances and storing data in blobs allows for easy integration with S3. I would love a proper column store put right now this has worked well with low development overhead.

We have also tuned our workloads to get around these constraints. Data is typically worked on for minutes or hours after it is retrieved. We are doing pure algo development and don't have (or need) low latency access in a UI.

## Answer by John Carse (score 2)

https://quant.stackexchange.com/a/1559

I am planning on using a complex event processing platform such as Esper or StreamBase to handle the incoming tick data to generate buy/sell events. The tick data would also be be forwarded through a queue (0MQ or RabbitMQ) to be written to a datastore (file or SQL database). By doing so, I have the data necessary to backtest or analyze in R or MATLAB, while still being able to act on the data "real-time". It seems that, depending on transaction volumes and frequency, that a SQL database may become a bottleneck if you want to act on data near real-time since using a SQL database would require that you pick a trigger or a time period to generate a query.

## Answer by steve8918 (score 2)

https://quant.stackexchange.com/a/1582

IB does not offer tick data. They consolidate their data every 0.3 seconds or so.

If you want to store your data temporarily for import into another system later on, then just store it as a CSV file.

Personally I use IQFeed to download tick data into SQL Server which I then use to run analyses on. When I need to run multiple test runs on the same data, I store the data locally into a file that I just read directly into memory.

## Answer by Andy Flury (score 2)

https://quant.stackexchange.com/a/8224

At AlgoTrader we also use Esper to store all arriving Tick Data in a local Esper Named Window. After a predefined interval, the latest Tick Data snapshot is written to both the filesystem and the database. The actual persistence is done in a separate thread by using Esper Threading. Currently we use MySql but you could just as well use some NoSQL Database (like MongoDB).

Have a look how this is done at our Open Source Page

The Algorithmic Trading Platform AlgoTrader is also available as a Commercial Version

Disclosure: I work for AlgoTrader

## Answer by TraderWerks (score 1)

https://quant.stackexchange.com/a/475

If you have an IB account, you can use their API to request market data and save to a flat file. That being said, IB does not offer true tick data, it is filtered and you may want to consider a different data feed if you need true tick data.

## Answer by HelloWorld (score 0)

https://quant.stackexchange.com/a/34569

I keep seeing that IB does not offer true tick data and I don't know if this is because these comments are old and they DID NOT offer true tick data at the time, but from my research they DO OFFER TRUE TICK DATA. (I called them and asked, because I saw in a few places where people said they do not offer it. You have to sign up for market data to get the true tick data, though.) Just wanted to put this out here for other people that come across it.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.