MintEval:LLMの売買コードが仕様に合うかを検証
記事 arXiv papers · 著者: Siyu Wang et al.
サマリー
MintEvalは、言語モデルがトレーダーの指示を、意図した戦略どおりに動作するコードへ変換できるかを評価します。組み合わせ可能な部品から参照戦略を作り、口語的な指示に変換して、モデルに実装を求めます。生成されたプログラムと参照プログラムを、同じ市場データとコスト条件で各バーごとに実行し、その行動を比較します。これにより、市場成績の差が実装忠実度の比較に影響しないようにしています。
初期ベンチマークは、BTCUSDTの15分足データを用いた800件のタスクで構成され、実行結果から測定した状態スパンの複雑さで分類されています。報告された結果ではモデル間の性能差が大きく、低コストモデルは行動の一致度や完全再現が限られる一方、最先端モデルは200件のサブセットでより良い成績を示しても、気づきにくい失敗を起こします。モデルが求められた構成要素を正しく特定しても、実際に取引が発生するバーでは多くの実装が仕様から逸脱します。著者らは、LLM判定器がそうした失敗をすべて許容したとも報告しています。このベンチマークが測るのは、定義されたタスクとデータにおける行動の等価性です。売買の収益性を証明するものではなく、他の市場や戦略タイプにそのまま一般化できるものでもありません。
主なアイデア
- このベンチマークは、ソースコードの類似度や利益ではなく、実行された売買行動を比較します。
- 参照戦略をプログラムで組み立て、非形式的なトレーダー向け指示に変換します。
- 行動の差を切り分けるため、両方の実装を同じ市場データと取引コスト条件で実行します。
- 実行に基づく状態スパンの複雑さでタスクを分類し、性能を評価します。
- 戦略仕様を正しく認識しても、生成コードが忠実に実装するとは限りません。
タグ
全文
# 2610.03080 # MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.
出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。