MintEval:测试LLM交易代码是否符合策略规范
文章 arXiv papers · 作者: Siyu Wang et al.
总结
MintEval评估语言模型能否将交易者指令转化为行为符合预期策略的代码。它通过可组合的组件构建参考策略,将其改写为日常语言指令,再要求模型实现。生成的程序和参考程序使用相同市场数据和交易摩擦逐根运行,并比较其动作,因此市场表现差异不会混淆实现忠实度的评估。
初始基准包含800个任务,使用BTCUSDT的15分钟数据,并按执行测得的状态跨度复杂度分组。据报告,模型表现差异很大:成本较低的模型在动作一致性和精确复现方面表现有限;前沿模型在200个任务的子集上表现更好,但仍会出现未显露的失败。即使模型正确识别出所要求的构成要素,许多实现仍会在有交易活动的K线偏离规范。作者还报告称,某个LLM评审器接受了所有此类失败。该基准衡量模型在所定义任务和数据上的行为等价性,并不能证明交易盈利能力,也不能自动推广到其他市场和策略类型。
核心观点
- 该基准比较实际执行的交易动作,而非源代码相似度或利润。
- 参考策略以程序化方式组合,再转写为非正式的交易者指令。
- 两种实现使用相同市场数据和交易摩擦运行,以隔离行为差异。
- 模型表现按执行测得的状态跨度复杂度对任务分组后评估。
- 正确识别策略规范,并不保证生成的代码忠实实现策略。
标签
全文
# 2610.03080 # MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。