MintEval: اختبار مطابقة كود تداول LLM للمواصفات
الملخص
يقيّم MintEval ما إذا كانت النماذج اللغوية تحول تعليمات المتداول إلى كود يتصرف كما تقصد الاستراتيجية. وينشئ استراتيجيات مرجعية من مكونات قابلة للتركيب، ويحولها إلى تعليمات دارجة، ثم يطلب من النماذج تنفيذها. ويُشغّل البرنامج المولد والمرجعي شمعةً شمعة على بيانات السوق نفسها وبالاحتكاكات نفسها؛ ثم تُقارن أفعالهما كي لا تؤثر اختلافات أداء السوق في قياس أمانة التنفيذ.
يضم المعيار الأولي 800 مهمة باستخدام بيانات BTCUSDT بفاصل 15 دقيقة، مصنفة بحسب تعقيد نطاق الحالة كما يُقاس من تنفيذ الاستراتيجية. وتظهر النتائج المُبلغ عنها تفاوتًا كبيرًا في أداء النماذج: فالنماذج الأقل تكلفة محدودة في توافق الأفعال وإعادة الإنتاج التامة، بينما يحقق نموذج متقدم أداءً أفضل على مجموعة فرعية من 200 مهمة، لكنه يظل ينتج إخفاقات صامتة. وحتى عندما تحدد النماذج مكونات البناء المطلوبة على نحو صحيح، تنحرف تطبيقات كثيرة عن المواصفات في الشموع النشطة. ويفيد المؤلفون أيضًا بأن حكمًا من LLM وافق على كل تلك الإخفاقات. يقيس المعيار التكافؤ السلوكي في المهام والبيانات المحددة، ولا يثبت ربحية التداول أو قابلية التعميم تلقائيًا على أسواق وأنواع استراتيجيات أخرى.
الأفكار الرئيسية
- يقارن المعيار أفعال التداول المنفذة، لا تشابه الكود المصدري أو الأرباح.
- تُركب الاستراتيجيات المرجعية برمجيًا وتُترجم إلى تعليمات غير رسمية للمتداول.
- يعمل التطبيقان على بيانات السوق واحتكاكات التداول نفسها لعزل الفروق السلوكية.
- يُقيّم الأداء عبر مهام مصنفة بحسب تعقيد نطاق الحالة، الذي يُقاس من تنفيذ الاستراتيجية.
- لا يضمن التعرف الصحيح على مواصفات الاستراتيجية أن ينفذها الكود المولد بأمانة.
الوسوم
النص الكامل
# 2610.03080 # MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.
يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0
أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.