MARKET-BENCH:测试大语言模型的量化交易回测能力
文章 arXiv papers · 作者: Abhay Srivastava et al.
总结
MARKET-BENCH 评估大语言模型能否将自然语言策略描述和市场假设转化为可执行的回测程序。任务包括 Microsoft 的定时交易、Coca-Cola 与 Pepsi 的配对交易,以及 Microsoft 的 delta 对冲。该基准将生成的损益、回撤和持仓路径与参考实现进行比较,同时衡量模型的回测能否运行,以及结果与参考实现的吻合程度。
对十三个模型进行的多轮评估发现,大多数模型都能可靠地执行最简单的策略,但不同模型和任务之间的数值误差差异显著。论文报告了若干模型的良好结果,包括 GPT-5.2 完全可执行;同时也有可运行的输出仍产生不准确路径的情况。研究结果表明,仅仅成功运行并不能确保交易逻辑正确。结果仅限于该基准的入门级任务和假设,因此不能证明模型能胜任更复杂策略或实盘市场。
核心观点
- 该基准测试定时交易、配对交易和 delta 对冲的代码生成能力。
- 该基准分别衡量回测能否运行,以及输出与参考结果的吻合程度。
- 研究对十三个模型进行了多轮评估。
- 回测即使可以运行,损益或持仓路径仍可能不准确。
- 结果涵盖入门级基准任务,不能证明模型具备实盘交易能力。
标签
全文
# Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics # Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics We introduce MARKET-BENCH, a benchmark that evaluates large language models (LLMs) on introductory quantitative trading tasks by asking them to construct executable backtesters from natural language strategy descriptions and market assumptions. Each instance specifies one of three canonical strategies: scheduled trading on Microsoft (NASDAQ: MSFT), pairs trading on Coca-Cola (NASDAQ: KO) and Pepsi (NASDAQ: PEP), or delta hedging on MSFT. Models must produce code whose profit and loss (P and L), drawdown, and position paths match a verifiable reference implementation. We assess thirteen state-of-the-art models using a multi-round evaluation that separates structural reliability (whether the backtest runs) from numerical accuracy (mean absolute error of the backtest metrics), assigning failed outputs a duplicated-metrics baseline MAE. While most models reliably execute the simplest strategy (average executable passes of 4.08 out of 5 rounds), errors vary by orders of magnitude across models and tasks. Gemini 3 Pro and Claude 4.5 Sonnet combine strong reliability with low error on simpler strategies. GPT-5.2 achieves strong overall performance with perfect executability. GPT-5.1 Codex-Max achieves the lowest best-run error on the easiest task. Qwen3 Max attains perfect executability yet sometimes produces inaccurate profit and loss paths. These results show that current LLMs can scaffold basic trading infrastructure but still struggle to reason robustly about prices, inventory, and risk. We release MARKET-BENCH and a public leaderboard at https://marketbench.ai.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。