跳至正文
返回文库全部文档

专门化语言模型生成可执行的程序化交易代码

文章 arXiv papers · 作者: Alexey Chernysh et al.

总结

本研究评估调整通用语言模型以生成适用于Backtrader框架的可执行策略的方法。该方法结合了基于框架代码的持续预训练,以及使用经智能体验证的请求到代码示例进行监督微调。评估使用一项包含400个任务的策略生成基准测试和一个代码仓库级赛道,衡量评审认定的正确性、成功回测次数,以及智能体在多轮修复中的表现。

核心观点

  • 在交易框架代码上持续预训练,提高了受评模型单轮评审表现。
  • 持续预训练后再进行监督微调,使一个模型取得更大提升,包括更高的回测和智能体成功率。
  • 仅持续预训练就可能提高智能体首轮成功率,却降低修复后的成功率,这表明其指令遵循能力有所减弱。
  • 领域专门化可能降低结构化工具调用格式的正确性;恢复性微调可以恢复格式,但无法恢复基础模型在代码仓库级任务上的智能体表现。
  • 结果涉及特定模型、任务和Backtrader框架,不能证明其在所有交易系统中的表现。

标签

全文
# QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code


# QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code









Large language models are strong general-purpose code generators, but executable algorithmic trading remains a demanding specialization target: a model must translate a natural-language strategy specification into correct program logic for a specialized trading framework, execute on historical data, produce trades, and remain semantically faithful to the request. We study two complementary mechanisms for specializing language models for this setting: continued pretraining on algorithmic-trading framework code and supervised fine-tuning (SFT) on agent-validated request-to-code pairs. Evaluation is centered on QuantCode-Bench, our 400-task benchmark for Backtrader strategy generation, together with a repository-level SWE-bench-like track. Continued pretraining improves single-turn Judge Pass from 41.5% to 47.5% for Qwen3.5-397B-A17B and from 27.8% to 33.0% for Qwen3.6-35B-A3B. SFT applied after continued pretraining yields a larger gain for Qwen3.6-35B-A3B, reaching 58.2% Judge Pass and 83.5% successful backtests; in agentic evaluation it raises first-turn success from 22.3% to 58.3% and final success after up to 10 turns from 47.5% to 79.5%. Continued pretraining alone improves first-turn agentic success but lowers final success after repair from 47.5% to 32.5%, consistent with degraded instruction following, whereas SFT improves both. We also identify a capability-retention failure: domain specialization degrades parser-conformant structured tool calling, and targeted recovery SFT restores tool-call formatting but not the base checkpoint's repository-level agent performance. The results show that framework-oriented pretraining, validated SFT, and explicit capability-retention evaluation address distinct failure modes in domain-specific executable code generation.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。