Skip to content
All library documents

Choosing a Data Layout for OHLCV Backtests

Article Quant Q&A · Author: laughingthunder

Summary

The document compares storing market bars in separate arrays by field with storing each bar as an object. The column-oriented option keeps open, high, low, close, and volume values in separate arrays. One answer recommends this layout for time-series backtests because it can improve cache efficiency and makes the data behave like an in-memory table that can be queried. It also advises planning how symbols and dates will be indexed or partitioned based on the queries the system needs to support.

A second answer recommends an object-oriented design: represent each bar with a data class and keep bars in a list, with a map from symbols to their time series. It argues that performance differences may not matter for many backtests, so ease of use and good object-oriented structure can guide the choice. The document offers design advice rather than benchmark results; the best layout depends on workload, performance needs, and integration with the live-trading API.

Key ideas

  • Separate arrays create a column-oriented layout that can suit time-series queries and improve cache use.
  • An object per bar, stored in a list, can make the model easier to work with.
  • Choose indexing and partitioning for symbols and dates based on expected queries.
  • The performance tradeoff may be small in many backtesting workloads, so benchmark against actual needs.
  • Consider compatibility with the types used by the live-trading API.

Tags

Full text
# For a interdays trading backtest system, should I put day open, close, high, low, volume separately into array?


# For a interdays trading backtest system, should I put day open, close, high, low, volume separately into array?












I think there are two possible ways: 1. day open, close, high, low, volume separately into array, then I have 5 arrays to work with my calculation 2. Put all of these into one array or linklist to do the calculation.

I think the 1. would be more easy to handle all backtest calculations, do you agree?

Background information: I would develop my system with Java and run in either Windows 7 64 bits/ Ubuntu 64 bits. I will eventually connect the live trading version to Interactive Broker Java socket API.

## Answer by chrisaycock (score 4)

https://quant.stackexchange.com/a/8861

Go with the multiple arrays. This would give you a column-oriented store, which is far more cache-efficient when handling time-series data. Specifically, you are describing an "in-memory" database table that can be queried.

You'll also want to think about how to do your look-ups. Will you have a hash table that maps symbols to OHLC tables? Will you partition the tables by date? Think about what kind of queries you're going to run, and then structure your data to make this search easiest.

## Answer by assylias (score 0)

https://quant.stackexchange.com/a/8892

Considering it is backtesting and ultra-performance is not very relevant**, I would suggest following good OOP principles:

- create a `BarData` class which holds the fields you need (date, o, h, l, c, v etc.)

- store the `BarData`s in an `ArrayList<BarData>`

You can then store several time series in a `Map<String, List<BarData>>` or a guava `MultiMap<String, BarData>`.

That should make your life easier.

ps: I don't know the IB API - if they have built-in types that look like my BarData then you sould as well use them directly.

**And note the the difference of performance between arrays and ArrayLists is not that big for most use cases.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.