MintEval: בדיקת התאמת קוד מסחר של LLM למפרט
סיכום
MintEval מעריך אם מודלי שפה מתרגמים הוראות סוחר לקוד שמתנהג לפי האסטרטגיה המיועדת. הוא בונה אסטרטגיות ייחוס מרכיבים שניתן להרכיב, מנסח אותן כהוראות בשפה יומיומית ומבקש ממודלים לממש אותן. התוכניות שנוצרו ותוכניות הייחוס מורצות נר אחר נר על אותם נתוני שוק ובאותם חיכוכים; הפעולות שלהן מושוות כך שהבדלי ביצועי השוק אינם מבלבלים את נאמנות המימוש.
מדד ההשוואה הראשוני כולל 800 משימות המשתמשות בנתוני BTCUSDT ברזולוציית 15 דקות, המקובצים לפי מורכבות טווח המצבים הנמדדת בביצוע. התוצאות המדווחות מראות שביצועי המודלים משתנים במידה ניכרת: למודלים הזולים יותר הסכמה מוגבלת על פעולות ושחזור מדויק, בעוד שמודל מתקדם מוביל מציג ביצועים טובים יותר בתת־קבוצה של 200 משימות, אך עדיין נכשל בשקט. גם כאשר המודלים מזהים נכון את אבני הבניין שהתבקשו, מימושים רבים חורגים מהמפרט בנרות שבהם יש פעילות. המחברים מדווחים גם ששופט LLM אישר את כל הכשלים הללו. המדד בוחן שקילות התנהגותית במשימות ובנתונים שהוגדרו; הוא אינו מבסס רווחיות מסחר ואינו מכליל אוטומטית לשווקים ולסוגי אסטרטגיות אחרים.
רעיונות מרכזיים
- מדד ההשוואה בוחן פעולות מסחר שבוצעו, ולא דמיון בקוד המקור או רווח.
- אסטרטגיות ייחוס מורכבות באופן תוכנתי ומתורגמות להוראות לא פורמליות לסוחרים.
- שני המימושים פועלים על נתוני שוק זהים ובאותם חיכוכי מסחר כדי לבודד הבדלים בהתנהגות.
- הביצועים מוערכים בין משימות המקובצות לפי מורכבות טווח מצבים המבוססת על ביצוע.
- זיהוי נכון של מפרט אסטרטגיה אינו מבטיח שהקוד שנוצר יממש אותו בנאמנות.
תגיות
הטקסט המלא
# 2610.03080 # MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.
מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0
הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.