実行可能なアルゴリズム取引コードに向けた言語モデルの特化
記事 arXiv papers · 著者: Alexey Chernysh et al.
サマリー
この研究は、汎用言語モデルを適応させ、Backtraderフレームワーク向けの実行可能な戦略を生成する方法を評価しています。フレームワークのコードによる継続事前学習と、エージェントで検証した指示・コード例による教師あり微調整を組み合わせています。評価には、400件の戦略生成ベンチマークとリポジトリ単位の課題を使用し、判定された正確性、バックテストの成功、修正を重ねる過程でのエージェント性能を測定しています。
主なアイデア
- 取引フレームワークのコードによる継続事前学習は、評価対象モデルの単一ターンでの判定性能を改善します。
- 継続事前学習後の教師あり微調整は、あるモデルでより大きな改善をもたらし、バックテストやエージェントの成功率も高めます。
- 継続事前学習だけでも初回のエージェントによるタスク成功率は上がる一方、修正後の成功率が下がることがあり、指示追従の弱まりが示唆されます。
- 分野特化で構造化されたツール呼び出しの書式が崩れることがあり、復旧のための微調整で書式は回復しますが、リポジトリ単位のエージェント性能は元に戻りません。
- 結果は特定のモデル、課題、Backtraderフレームワークに関するもので、あらゆる取引システムでの性能を示すものではありません。
タグ
全文
# QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code # QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code Large language models are strong general-purpose code generators, but executable algorithmic trading remains a demanding specialization target: a model must translate a natural-language strategy specification into correct program logic for a specialized trading framework, execute on historical data, produce trades, and remain semantically faithful to the request. We study two complementary mechanisms for specializing language models for this setting: continued pretraining on algorithmic-trading framework code and supervised fine-tuning (SFT) on agent-validated request-to-code pairs. Evaluation is centered on QuantCode-Bench, our 400-task benchmark for Backtrader strategy generation, together with a repository-level SWE-bench-like track. Continued pretraining improves single-turn Judge Pass from 41.5% to 47.5% for Qwen3.5-397B-A17B and from 27.8% to 33.0% for Qwen3.6-35B-A3B. SFT applied after continued pretraining yields a larger gain for Qwen3.6-35B-A3B, reaching 58.2% Judge Pass and 83.5% successful backtests; in agentic evaluation it raises first-turn success from 22.3% to 58.3% and final success after up to 10 turns from 47.5% to 79.5%. Continued pretraining alone improves first-turn agentic success but lowers final success after repair from 47.5% to 32.5%, consistent with degraded instruction following, whereas SFT improves both. We also identify a capability-retention failure: domain specialization degrades parser-conformant structured tool calling, and targeted recovery SFT restores tool-call formatting but not the base checkpoint's repository-level agent performance. The results show that framework-oriented pretraining, validated SFT, and explicit capability-retention evaluation address distinct failure modes in domain-specific executable code generation.
出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。