본문으로 건너뛰기
라이브러리 문서 전체

실행 가능한 알고리즘 매매 코드 생성을 위한 언어 모델 특화

기사 arXiv papers · 저자: Alexey Chernysh et al.

요약

이 연구는 일반 언어 모델을 Backtrader 프레임워크용 실행 가능한 전략을 생성하도록 조정하는 방법을 평가한다. 프레임워크 코드로 추가 사전학습을 하고 에이전트가 검증한 요청-코드 예제로 지도 미세조정을 하는 접근법을 결합한다. 평가는 400개 과제 전략 생성 벤치마크와 저장소 수준 트랙을 사용하며, 판정된 정확도, 성공적인 백테스트, 수정 과정에서의 에이전트 성능을 측정한다.

핵심 아이디어

  • 매매 프레임워크 코드로 추가 사전학습하면 평가 대상 모델의 단일 턴 판정 성능이 향상된다.
  • 추가 사전학습 뒤 지도 미세조정을 하면 한 모델에서 백테스트 및 에이전트 성공률을 포함해 더 큰 향상이 나타난다.
  • 추가 사전학습만으로 첫 시도 에이전트 성공률은 높아질 수 있지만 수정 후 성공률은 낮아질 수 있어, 지시 이행 능력의 약화를 시사한다.
  • 도메인 특화로 구조화된 도구 호출 형식이 저하될 수 있으며, 복구 미세조정은 형식을 회복하지만 기본 저장소 수준 에이전트 성능은 회복시키지 않는다.
  • 결과는 특정 모델, 과제, Backtrader 프레임워크에 관한 것이므로 모든 매매 시스템의 성능을 입증하지 않는다.

태그

전문
# QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code


# QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code









Large language models are strong general-purpose code generators, but executable algorithmic trading remains a demanding specialization target: a model must translate a natural-language strategy specification into correct program logic for a specialized trading framework, execute on historical data, produce trades, and remain semantically faithful to the request. We study two complementary mechanisms for specializing language models for this setting: continued pretraining on algorithmic-trading framework code and supervised fine-tuning (SFT) on agent-validated request-to-code pairs. Evaluation is centered on QuantCode-Bench, our 400-task benchmark for Backtrader strategy generation, together with a repository-level SWE-bench-like track. Continued pretraining improves single-turn Judge Pass from 41.5% to 47.5% for Qwen3.5-397B-A17B and from 27.8% to 33.0% for Qwen3.6-35B-A3B. SFT applied after continued pretraining yields a larger gain for Qwen3.6-35B-A3B, reaching 58.2% Judge Pass and 83.5% successful backtests; in agentic evaluation it raises first-turn success from 22.3% to 58.3% and final success after up to 10 turns from 47.5% to 79.5%. Continued pretraining alone improves first-turn agentic success but lowers final success after repair from 47.5% to 32.5%, consistent with degraded instruction following, whereas SFT improves both. We also identify a capability-retention failure: domain specialization degrades parser-conformant structured tool calling, and targeted recovery SFT restores tool-call formatting but not the base checkpoint's repository-level agent performance. The results show that framework-oriented pretraining, validated SFT, and explicit capability-retention evaluation address distinct failure modes in domain-specific executable code generation.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.