MARKET-BENCH: Testing LLMs on Quantitative Trading Backtests
Summary
MARKET-BENCH evaluates whether large language models can turn natural-language strategy descriptions and market assumptions into executable backtesters. Its tasks cover scheduled trading in Microsoft, pairs trading in Coca-Cola and Pepsi, and delta hedging in Microsoft. The benchmark compares generated profit and loss, drawdown, and position paths with a reference implementation, measuring both whether a model’s backtest runs and how closely its results match.
A multi-round evaluation of thirteen models finds that most can execute the simplest strategy reliably, while numerical errors differ substantially across models and tasks. The paper reports strong results from several models, including perfect executability for GPT-5.2, alongside cases where runnable outputs still produce inaccurate paths. The findings suggest that execution success alone does not ensure correct trading logic. Results are limited to the benchmark’s introductory tasks and assumptions, so they do not establish performance on more complex strategies or live markets.
Key ideas
- The benchmark tests code generation for scheduled trading, pairs trading, and delta hedging.
- It separately measures whether backtests execute and how closely their outputs match a reference.
- Thirteen models are evaluated across multiple rounds.
- Runnable backtests can still have inaccurate profit and loss or position paths.
- The results cover introductory benchmark tasks and do not establish live trading ability.
Tags
Full text
# Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics # Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics We introduce MARKET-BENCH, a benchmark that evaluates large language models (LLMs) on introductory quantitative trading tasks by asking them to construct executable backtesters from natural language strategy descriptions and market assumptions. Each instance specifies one of three canonical strategies: scheduled trading on Microsoft (NASDAQ: MSFT), pairs trading on Coca-Cola (NASDAQ: KO) and Pepsi (NASDAQ: PEP), or delta hedging on MSFT. Models must produce code whose profit and loss (P and L), drawdown, and position paths match a verifiable reference implementation. We assess thirteen state-of-the-art models using a multi-round evaluation that separates structural reliability (whether the backtest runs) from numerical accuracy (mean absolute error of the backtest metrics), assigning failed outputs a duplicated-metrics baseline MAE. While most models reliably execute the simplest strategy (average executable passes of 4.08 out of 5 rounds), errors vary by orders of magnitude across models and tasks. Gemini 3 Pro and Claude 4.5 Sonnet combine strong reliability with low error on simpler strategies. GPT-5.2 achieves strong overall performance with perfect executability. GPT-5.1 Codex-Max achieves the lowest best-run error on the easiest task. Qwen3 Max attains perfect executability yet sometimes produces inaccurate profit and loss paths. These results show that current LLMs can scaffold basic trading infrastructure but still struggle to reason robustly about prices, inventory, and risk. We release MARKET-BENCH and a public leaderboard at https://marketbench.ai.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.