본문으로 건너뛰기
라이브러리 문서 전체

실거래 암호화폐·주식 시장에서 LLM 트레이딩 에이전트 평가

기사 arXiv papers · 저자: Lingfei Qian et al.

요약

이 논문은 암호화폐와 주식 시장에서 대규모 언어 모델 트레이딩 에이전트를 평가하는 연속 실시간 벤치마크인 Agent Market Arena를 소개합니다. 이 프레임워크는 검증된 거래 데이터와 전문가가 확인한 뉴스, 다양한 에이전트 설계를 결합해 실제 시장 상황에서 시스템을 비교하는 것을 목표로 합니다. 기존 평가에서 자산이나 기간이 제한되거나, 완전한 에이전트가 아닌 모델만 시험하거나, 독립적으로 검증되지 않은 데이터를 사용하는 경우가 많았다는 점이 이 연구의 동기입니다.

벤치마크에는 위험 성향이 서로 다른 에이전트와 기억 기반 추론을 사용하는 에이전트가 포함되며, 여러 언어 모델 백본으로 평가합니다. 보고된 실거래 실험에서는 기반 모델 간보다 에이전트 프레임워크 간 행동의 차이가 더 크게 나타났고, 의사결정은 공격적 성향에서 보수적 성향까지 다양했습니다. 초록은 AMA를 지속적이고 재현 가능한 평가를 위한 기반 시설로 제시하지만, 성과 수치, 벤치마크 기간, 전략 비교의 세부 통제 방법은 제공하지 않습니다. 결론은 관찰된 행동 차이에 관한 것이며 어느 에이전트든 꾸준히 수익을 낸다는 점을 입증하지 않습니다.

핵심 아이디어

  • Agent Market Arena는 암호화폐와 주식 시장의 LLM 트레이딩 에이전트를 위한 실시간 벤치마크입니다.
  • 검증된 거래 데이터, 확인된 뉴스, 다양한 에이전트 아키텍처를 결합합니다.
  • 위험 성향과 기억 기반 추론이 서로 다른 에이전트를 평가합니다.
  • 실거래 실험에서는 모델 백본보다 에이전트 프레임워크 간 행동 차이가 더 크게 나타났습니다.
  • 초록은 벤치마킹 기여를 설명하지만 수익성 수치나 평가 통제 세부 사항은 제시하지 않습니다.

태그

전문
# When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents


# When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents









Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies test models instead of agents, cover limited periods and assets, and rely on unverified data. To address these gaps, we introduce Agent Market Arena (AMA), the first lifelong, real-time benchmark for evaluating LLM-based trading agents across multiple markets. AMA integrates verified trading data, expert-checked news, and diverse agent architectures within a unified trading framework, enabling fair and continuous comparison under real conditions. It implements four agents, including InvestorAgent as a single-agent baseline, TradeAgent and HedgeFundAgent with different risk styles, and DeepFundAgent with memory-based reasoning, and evaluates them across GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4, and Gemini-2.0-flash. Live experiments on both cryptocurrency and stock markets demonstrate that agent frameworks display markedly distinct behavioral patterns, spanning from aggressive risk-taking to conservative decision-making, whereas model backbones contribute less to outcome variation. AMA thus establishes a foundation for rigorous, reproducible, and continuously evolving evaluation of financial reasoning and trading intelligence in LLM-based agents.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.