Specializing Language Models for Executable Algorithmic Trading Code
Article arXiv papers · Author: Alexey Chernysh et al.
Summary
The study evaluates ways to adapt general language models to generate executable strategies for the Backtrader framework. Its approach combines continued pretraining on framework code with supervised fine-tuning on request-to-code examples validated by agents. Evaluation uses a 400-task strategy-generation benchmark and a repository-level track, measuring judged correctness, successful backtests, and agent performance over repair turns.
Key ideas
- Continued pretraining on trading-framework code improves single-turn judged performance for the evaluated models.
- Supervised fine-tuning after continued pretraining produces larger gains for one model, including higher backtest and agent success rates.
- Continued pretraining alone can help first-turn agentic success while hurting success after repair, suggesting weaker instruction following.
- Domain specialization can degrade structured tool-call formatting, and recovery fine-tuning restores formatting without restoring base repository-level agent performance.
- The results concern particular models, tasks, and the Backtrader framework, so they do not establish performance across all trading systems.
Tags
Full text
# QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code # QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code Large language models are strong general-purpose code generators, but executable algorithmic trading remains a demanding specialization target: a model must translate a natural-language strategy specification into correct program logic for a specialized trading framework, execute on historical data, produce trades, and remain semantically faithful to the request. We study two complementary mechanisms for specializing language models for this setting: continued pretraining on algorithmic-trading framework code and supervised fine-tuning (SFT) on agent-validated request-to-code pairs. Evaluation is centered on QuantCode-Bench, our 400-task benchmark for Backtrader strategy generation, together with a repository-level SWE-bench-like track. Continued pretraining improves single-turn Judge Pass from 41.5% to 47.5% for Qwen3.5-397B-A17B and from 27.8% to 33.0% for Qwen3.6-35B-A3B. SFT applied after continued pretraining yields a larger gain for Qwen3.6-35B-A3B, reaching 58.2% Judge Pass and 83.5% successful backtests; in agentic evaluation it raises first-turn success from 22.3% to 58.3% and final success after up to 10 turns from 47.5% to 79.5%. Continued pretraining alone improves first-turn agentic success but lowers final success after repair from 47.5% to 32.5%, consistent with degraded instruction following, whereas SFT improves both. We also identify a capability-retention failure: domain specialization degrades parser-conformant structured tool calling, and targeted recovery SFT restores tool-call formatting but not the base checkpoint's repository-level agent performance. The results show that framework-oriented pretraining, validated SFT, and explicit capability-retention evaluation address distinct failure modes in domain-specific executable code generation.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.