الانتقال إلى المحتوى
جميع مستندات المكتبة

MARKET-BENCH: اختبار النماذج اللغوية الكبيرة في الاختبارات التاريخية للتداول الكمي

مقال arXiv papers · المؤلف: Abhay Srivastava et al.

الملخص

يقيّم MARKET-BENCH قدرة النماذج اللغوية الكبيرة على تحويل أوصاف الاستراتيجيات وافتراضات السوق باللغة الطبيعية إلى أدوات قابلة للتنفيذ للاختبار التاريخي. وتشمل مهامه التداول المجدول في مايكروسوفت، وتداول الأزواج في كوكاكولا وبيبسي، والتحوط من دلتا في مايكروسوفت. ويقارن المعيار بين مسارات الربح والخسارة والتراجع والمراكز الناتجة وتطبيق مرجعي، ويقيس ما إذا كان الاختبار التاريخي للنموذج يعمل ومدى تطابق نتائجه.

يجد تقييم متعدد الجولات لثلاثة عشر نموذجًا أن معظمها يستطيع تنفيذ أبسط الاستراتيجيات بموثوقية، بينما تختلف الأخطاء العددية كثيرًا بين النماذج والمهام. وتفيد الورقة بنتائج قوية لعدة نماذج، منها قابلية التنفيذ الكاملة لـGPT-5.2، إلى جانب حالات أنتجت فيها المخرجات القابلة للتشغيل مسارات غير دقيقة. وتشير النتائج إلى أن نجاح التنفيذ وحده لا يضمن صحة منطق التداول. وتقتصر النتائج على المهام والافتراضات التمهيدية للمعيار، لذا لا تثبت الأداء في استراتيجيات أكثر تعقيدًا أو أسواق فعلية.

الأفكار الرئيسية

  • يختبر المعيار توليد الشيفرة للتداول المجدول وتداول الأزواج والتحوط من دلتا.
  • يقيس المعيار على نحو منفصل ما إذا كانت الاختبارات التاريخية تُنفذ ومدى تطابق مخرجاتها مع مرجع.
  • يُقيّم ثلاثة عشر نموذجًا عبر جولات متعددة.
  • قد تنتج الاختبارات التاريخية القابلة للتشغيل مسارات غير دقيقة للربح والخسارة أو المراكز.
  • تغطي النتائج مهام تمهيدية للمعيار، ولا تثبت القدرة على التداول الفعلي.

الوسوم

النص الكامل
# Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics


# Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics









We introduce MARKET-BENCH, a benchmark that evaluates large language models (LLMs) on introductory quantitative trading tasks by asking them to construct executable backtesters from natural language strategy descriptions and market assumptions. Each instance specifies one of three canonical strategies: scheduled trading on Microsoft (NASDAQ: MSFT), pairs trading on Coca-Cola (NASDAQ: KO) and Pepsi (NASDAQ: PEP), or delta hedging on MSFT. Models must produce code whose profit and loss (P and L), drawdown, and position paths match a verifiable reference implementation. We assess thirteen state-of-the-art models using a multi-round evaluation that separates structural reliability (whether the backtest runs) from numerical accuracy (mean absolute error of the backtest metrics), assigning failed outputs a duplicated-metrics baseline MAE. While most models reliably execute the simplest strategy (average executable passes of 4.08 out of 5 rounds), errors vary by orders of magnitude across models and tasks. Gemini 3 Pro and Claude 4.5 Sonnet combine strong reliability with low error on simpler strategies. GPT-5.2 achieves strong overall performance with perfect executability. GPT-5.1 Codex-Max achieves the lowest best-run error on the easiest task. Qwen3 Max attains perfect executability yet sometimes produces inaccurate profit and loss paths. These results show that current LLMs can scaffold basic trading infrastructure but still struggle to reason robustly about prices, inventory, and risk. We release MARKET-BENCH and a public leaderboard at https://marketbench.ai.

يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0

أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.