본문으로 건너뛰기
라이브러리 문서 전체

MintEval: LLM 매매 코드의 사양 일치 여부 검증

기사 arXiv papers · 저자: Siyu Wang et al.

요약

MintEval은 언어 모델이 트레이더의 지시를 의도한 전략대로 동작하는 코드로 옮기는지 평가합니다. 조합 가능한 구성 요소로 기준 전략을 만들고, 이를 일상적인 표현의 지시로 바꾼 뒤 모델에 구현을 요청합니다. 생성된 프로그램과 기준 프로그램을 같은 시장 데이터와 거래 마찰을 적용해 각 봉마다 실행하고 행동을 비교하므로, 시장 성과 차이가 구현 충실도 비교에 영향을 주지 않습니다.

초기 벤치마크는 800개 과제로 구성되며, BTCUSDT 15분 데이터를 사용하고 실행으로 측정한 상태 구간 복잡도에 따라 분류합니다. 보고된 결과는 모델 성능이 크게 다름을 보여줍니다. 저비용 모델은 행동 일치도와 정확한 재현도가 제한적이며, 최첨단 모델은 200개 과제 하위 집합에서 더 나은 성능을 보이지만 여전히 드러나지 않는 실패를 냅니다. 요청된 구성 요소를 올바르게 파악해도 활성 봉에서 많은 구현이 사양과 달라집니다. 저자들은 LLM 판정자가 이러한 실패를 모두 통과시켰다고도 보고합니다. 이 벤치마크는 정의된 과제와 데이터에서 행동 등가성을 측정합니다. 매매 수익성을 입증하거나 다른 시장과 전략 유형에 자동으로 일반화하지는 않습니다.

핵심 아이디어

  • 소스 코드의 유사성이나 수익이 아니라 실행된 매매 행동을 비교합니다.
  • 기준 전략을 프로그램으로 조합한 뒤 비격식적인 트레이더 지시로 변환합니다.
  • 행동 차이를 분리하기 위해 두 구현을 동일한 시장 데이터와 거래 마찰 조건에서 실행합니다.
  • 실행 기반 상태 구간 복잡도에 따라 분류한 과제 전반에서 성능을 평가합니다.
  • 전략 사양을 올바르게 알아보더라도 생성된 코드가 이를 충실히 구현한다는 보장은 없습니다.

태그

전문
# 2610.03080


# MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code









Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.