MARKET-BENCH: مقداری ٹریڈنگ کے بیک ٹیسٹ میں LLMs کی جانچ
خلاصہ
MARKET-BENCH جانچتا ہے کہ آیا بڑے زبان کے ماڈل قدرتی زبان میں لکھی اسٹریٹیجی کی وضاحتوں اور مارکیٹ مفروضوں کو قابلِ عمل بیک ٹیسٹرز میں بدل سکتے ہیں۔ اس کے کاموں میں Microsoft میں مقررہ اوقات پر ٹریڈنگ، Coca-Cola اور Pepsi میں پیئرز ٹریڈنگ، اور Microsoft میں ڈیلٹا ہیجنگ شامل ہیں۔ بینچ مارک تیار کردہ نفع و نقصان، ڈرا ڈاؤن اور پوزیشن کے راستوں کا حوالہ جاتی نفاذ سے موازنہ کرتا ہے، اور یہ بھی ناپتا ہے کہ ماڈل کا بیک ٹیسٹ چلتا ہے یا نہیں اور اس کے نتائج کتنے قریب ہیں۔
تیرہ ماڈلز کی کئی مرحلوں والی جانچ میں زیادہ تر آسان ترین اسٹریٹیجی کو قابلِ اعتماد طور پر چلا لیتے ہیں، جبکہ عددی غلطیاں ماڈلز اور کاموں کے لحاظ سے خاصی مختلف ہیں۔ مقالے میں کئی ماڈلز کے مضبوط نتائج، بشمول GPT-5.2 کے لیے کامل قابلِ اجرا ہونے، کے ساتھ ایسے معاملے بھی درج ہیں جہاں چلنے کے قابل نتائج کے باوجود راستے غلط نکلتے ہیں۔ نتائج بتاتے ہیں کہ صرف کامیاب اجرا درست ٹریڈنگ منطق کی ضمانت نہیں۔ نتائج بینچ مارک کے ابتدائی کاموں اور مفروضوں تک محدود ہیں؛ اس لیے زیادہ پیچیدہ اسٹریٹیجیوں یا لائیو مارکیٹ میں کارکردگی ثابت نہیں کرتے۔
اہم خیالات
- بینچ مارک مقررہ اوقات کی ٹریڈنگ، پیئرز ٹریڈنگ اور ڈیلٹا ہیجنگ کے لیے کوڈ بنانے کی جانچ کرتا ہے۔
- یہ الگ الگ ناپتا ہے کہ بیک ٹیسٹ چلتے ہیں یا نہیں اور ان کے نتائج حوالہ جاتی نتائج سے کتنے ملتے ہیں۔
- تیرہ ماڈلز کی کئی مرحلوں میں جانچ کی گئی۔
- چلنے والے بیک ٹیسٹ میں بھی نفع و نقصان یا پوزیشن کے راستے غلط ہو سکتے ہیں۔
- نتائج ابتدائی بینچ مارک کاموں تک محدود ہیں اور لائیو ٹریڈنگ کی صلاحیت ثابت نہیں کرتے۔
ٹیگز
مکمل متن
# Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics # Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics We introduce MARKET-BENCH, a benchmark that evaluates large language models (LLMs) on introductory quantitative trading tasks by asking them to construct executable backtesters from natural language strategy descriptions and market assumptions. Each instance specifies one of three canonical strategies: scheduled trading on Microsoft (NASDAQ: MSFT), pairs trading on Coca-Cola (NASDAQ: KO) and Pepsi (NASDAQ: PEP), or delta hedging on MSFT. Models must produce code whose profit and loss (P and L), drawdown, and position paths match a verifiable reference implementation. We assess thirteen state-of-the-art models using a multi-round evaluation that separates structural reliability (whether the backtest runs) from numerical accuracy (mean absolute error of the backtest metrics), assigning failed outputs a duplicated-metrics baseline MAE. While most models reliably execute the simplest strategy (average executable passes of 4.08 out of 5 rounds), errors vary by orders of magnitude across models and tasks. Gemini 3 Pro and Claude 4.5 Sonnet combine strong reliability with low error on simpler strategies. GPT-5.2 achieves strong overall performance with perfect executability. GPT-5.1 Codex-Max achieves the lowest best-run error on the easiest task. Qwen3 Max attains perfect executability yet sometimes produces inaccurate profit and loss paths. These results show that current LLMs can scaffold basic trading infrastructure but still struggle to reason robustly about prices, inventory, and risk. We release MARKET-BENCH and a public leaderboard at https://marketbench.ai.
ماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: abstract CC0
یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔