Diagnosing and Improving Backtest Performance Without Changing Results
Summary
This engineering guide explains how to locate backtest slowdowns while preserving simulation behavior. It separates startup, historical data loading, strategy computation, and report generation, and recommends first distinguishing cold runs that fetch data from warm runs that should reuse cached data. Repeated downloads point toward data hydration or cache configuration, while slow runs with little downloading call for profiling compute, storage, and artifact costs.
The guide explains that a single strategy advances through one simulated broker clock in sequence because orders, cash, positions, fills, and valuation depend on prior state. Parallelism is therefore more useful for downloads and independent backtests than for arbitrary parallel execution of one strategy. It recommends profiling with YAPPI and inspecting storage, pandas operations, and report generation; options workloads can also be narrowed to needed expirations and strikes. These are diagnostic and implementation recommendations, not a benchmark: no measured speedups are reported, and outcomes depend on the workload and cache setup.
Key ideas
- Separate startup, data loading, computation, and artifact generation when diagnosing a slow backtest.
- Compare cold and warm runs to identify cache coverage or configuration problems.
- A single simulated broker clock processes strategy state serially to preserve execution correctness.
- Parallel downloads and independent backtests can benefit from bounded concurrency.
- Profiling helps locate storage, data transformation, and reporting hotspots.
Tags
Full text
# backtesting.performance .. _backtesting_performance: Backtesting Performance (Speed + Parity) ======================================== .. meta:: :description: This page explains how to make backtests faster without changing strategy correctness. Performance issues are usually dominated by one of:. This page explains how to make backtests faster **without changing strategy correctness**. Performance issues are usually dominated by one of: - **Startup** (python import time, environment loading, first progress update) - **Data hydration** (first run downloads data; warm runs should reuse cache) - **Compute** (strategy logic, pandas transforms, option pricing) - **Artifacts** (tearsheets, plots, indicators) If you are new to backtesting, start with :doc:`backtesting.how_to_backtest`. Warm vs cold runs ----------------- Backtest speed depends heavily on caching: - **Cold run:** the cache is empty, so the backtest must fetch historical data. - **Warm run:** the cache already contains required data, so the backtest should be dramatically faster. If a “warm” run is still slow, the most common causes are: - your cache backend is not configured (or not writable) - your cache namespace changed between runs - a request type is not being cached (so it keeps downloading) For deeper cache semantics (engineering notes), see ``docs/remote_cache.md`` in the repository. Quick diagnosis checklist ------------------------- 1. **Is the backtest downloading a lot of data?** - Look for many “Submitted to queue” log lines (ThetaData) or repeated API calls (Polygon). - If yes, you are hydration-bound: fix request fanout or cache coverage first. 2. **Is the backtest slow even with near-zero downloads?** - If yes, you are compute/IO/artifact-bound: use profiling to attribute time. 3. **Does the backtest look stuck in the UI?** - If data is downloading, progress may not advance unless a heartbeat is enabled. Execution model and realistic parallelism ----------------------------------------- One LumiBot backtest runs a single strategy through a single simulated broker clock. That means the core execution path is intentionally serial: - strategy logic runs, - pending orders are processed, - then the simulated clock advances. This is important for correctness because fills, cash, positions, bracket/OCO behavior, and mark-to-market all depend on prior state. What this means in practice: - **Do not expect arbitrary strategies to become 10x faster just by “making the loop parallel.”** - Generic vectorized or matrix-style execution is not a good fit for the current LumiBot execution model. Where parallelism does help: - provider-side data downloads and prefetch, - running many **independent** backtests at once (for example parameter sweeps), - and deferring expensive report generation until after the simulation is complete. Practical guidance: - If one backtest is slow, first determine whether it is hydration-bound, compute-bound, or artifact-bound. - If you need many backtests, run multiple backtests externally with bounded concurrency rather than expecting one backtest to fan out internally. Profiling (YAPPI) ----------------- To attribute where time is spent (S3 IO vs compute vs artifacts), enable profiling: - Set ``BACKTESTING_PROFILE=yappi`` - Run the backtest - Inspect the produced ``*_profile_yappi.csv`` artifact Common hotspots to look for: - S3 IO (many small objects can be slow even on “warm” runs) - pandas transforms (merge/concat/tz conversions) - artifact generation (tearsheet, indicators, plots) Environment variables --------------------- Many backtesting behaviors are configurable via environment variables. See: - :doc:`environment_variables` (public docs) - ``docs/ENV_VARS.md`` (engineering notes; may include contributor-specific details) Common performance-related flags: - ``LUMIBOT_DISABLE_DOTENV``: disables recursive ``.env`` discovery (reduces startup latency and avoids accidental config overrides) - ``SHOW_TEARSHEET`` / ``SHOW_PLOT`` / ``SHOW_INDICATORS``: disables heavy artifact generation when you only need core results - ``BACKTESTING_PROFILE``: enable profiling (yappi) ThetaData options: common performance pitfalls ---------------------------------------------- Options backtests can be slower than stock backtests because they may need: - option chains (expirations/strikes) - quote history (bid/ask) for realistic pricing - additional mark-to-market logic for illiquid contracts The fastest options backtests are those that: - build **only the chain data they need** (one expiry and a narrow strike neighborhood) - avoid probing hundreds/thousands of strikes when searching for a delta/ATM contract - reuse cached quote history instead of requesting tiny windows repeatedly In practice, the easiest way to get this right is to use :doc:`options_helper` for strike/expiry selection (for example ``OptionsHelper.find_strike_for_delta(...)``) instead of manually scanning chains and calling ``get_greeks()`` per strike. For ThetaData details, see :doc:`backtesting.thetadata`. IBKR history and reuse ---------------------- IBKR equity warmup requests count bars across exchange sessions, including weekends and holidays. A larger indicator lookback can extend the prefetched window; repeated requests for the same completed window reuse it. Routed IBKR daily bars retain their native session-close timestamps. Futures positions use intraday marks rather than the prior daily candle. This may require an initial minute-history fetch even for a daily futures strategy; later marks reuse that series. Missing required history or an unknown valuation gap makes the result incomplete rather than a verified zero-trade result.
Shown in full with attribution under the source's licence. Licence: GPL-3.0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.