Skip to content
All library documents

Design Considerations for a Customizable Trading Backtester

Article Quant Q&A · Author: Jake

Summary

The document outlines architectural decisions for building a customizable Python backtesting engine. It highlights market-data storage and serialization, memory limits, handling unstructured feeds, durable output storage, and access to prior run results. It also identifies fill logic and the strategy-facing API as core design choices, including how orders, instruments, prices, and volumes are represented. Dataframes may be convenient for exploration but can add overhead in repeated simulation workloads.

The response recommends examining existing projects while warning that performance and multi-product support vary. It gives an illustrative claim about the speed expected from modern hardware and anecdotes about commercial systems, but provides no benchmark methodology, reproducible measurements, or implementation plan. A separate answer points to another open-source backtesting project as a reference. The advice is an initial checklist rather than a complete specification: realistic fills, event handling, asset-specific mechanics, validation, and the tradeoff between ease of extension and execution speed still require design decisions.

Key ideas

  • Market-data storage must account for feeds that exceed available memory and for irregular formats.
  • Backtest outputs should be persisted so interrupted runs can resume and past results remain accessible.
  • Fill logic and the strategy-facing API are foundational engine design choices.
  • Dataframes support exploration but may impose costs in repeated simulation loops.
  • Claims about framework speed and performance are anecdotal and need independent benchmarking.

Tags

Full text
# Build a customizable trading engine in python


# Build a customizable trading engine in python












I am planning building fully customizable backtesting trading engine in python from scratch as a open source project, the main features i am considering is,

- It should be fully customizable from top to bottom

- Customization is very easy and anyone can customize with a basic knowledge in python

- It have a in built template engine for reports which is also customizable

- Anyone can customize it as per their trading style

So what are the basic things which i have to consider for building a trading engine? Which are the python modules which is useful for this project So anyone know any material regarding this sharing a link will be very helpful....

## Answer by madilyn (score 11, accepted)

https://quant.stackexchange.com/a/15162

Firstly, you'll probably be directed to consider Zipline. It's worth a look but I don't think that it's a good starting point, since:

- Quantopian's developers don't have a financial background and it shows through in the Zipline source code.

- Zipline is dreadfully slow if you compare it to any commercial platform with backtesting functionality in a compiled application, even the low-end retail trading platforms (e.g. NinjaTrader, Sierra, TradeStation).

- Zipline isn't very convenient for trading multiple products. I think the cheapest product that has that level of functionality is Deltix.

A modern processor should be able to backtest a moving average crossover strategy across an entire day of the OPRA feed (all products) without scheduling it overnight. Any less functionality or slower and you have poor developers. (I remember Goldman had 12-14 servers dealing with realtime OPRA in 2007-2008 and 2 persons rewrote the entire thing from scratch to target 128-bit architecture over a weekend. No reason why years of development on Zipline doesn't match up to 2 developers on a weekend before Stack Exchange existed.)

Here are some of the major considerations that you have to make before building your backtesting engine:

- How will you be storing/serializing your market data on disk and in memory? One poor man's approach is to wrap it around a `pandas` dataframe, but this comes at the cost of abstraction and will slow down your backtesting engine. `pandas` is nice for data exploration, but not for a task that you will repeat many times. How will you handle a data source whose size exceeds available memory? How will you deal with unstructured market data?

- How will you be storing your outputs? An obvious, naive problem is that you don't want to restart a backtest that took you 1 night to run if the application crashed midway. Another naive example is that you should be able to access old results from 6 months ago without repeating the backtest loop.

- What's your fill logic?

- What should your API expose? (e.g. Market orders, limit orders, instrument/price/volume queries, changes to fill logic)

## Answer by Raja Pasupuleti (score 0)

https://quant.stackexchange.com/a/15171

There are some other opensource projects at github along with zipline which you can check for some additional inputs for your thoughts. Pyalgotrade is one such project where one can do back testing on their trading strategies

http://gbeced.github.io/pyalgotrade/

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.