MintEval: Testing Whether LLM Trading Code Matches Its Specification
Summary
MintEval evaluates whether language models translate trader instructions into code that behaves like the intended strategy. It builds reference strategies from composable components, turns them into colloquial instructions, and asks models to implement them. The generated and reference programs are run bar by bar on the same market data and with the same frictions; their actions are compared, so market performance differences do not confound implementation fidelity.
The initial benchmark contains 800 tasks using BTCUSDT 15-minute data, grouped by execution-measured state-span complexity. Reported results show that model performance varies substantially: lower-cost models have limited action agreement and exact reproduction, while a frontier model performs better on a 200-task subset but still produces silent failures. Even when models correctly identify the requested building blocks, many implementations diverge from the specification on active bars. The authors also report that an LLM judge accepted all such failures. The benchmark measures behavioral equivalence on its defined tasks and data; it does not establish trading profitability or generalize automatically to other markets and strategy types.
Key ideas
- The benchmark compares executed trading actions rather than source-code similarity or profit.
- Reference strategies are programmatically composed and translated into informal trader instructions.
- Both implementations run on identical market data and trading frictions to isolate behavioral differences.
- Performance is assessed across tasks grouped by execution-based state-span complexity.
- Correctly recognizing a strategy specification does not ensure that generated code implements it faithfully.
Tags
Full text
# 2610.03080 # MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.