跳至正文
返回文库全部文档

实盘加密货币与股票市场的 LLM 交易智能体基准测试

文章 arXiv papers · 作者: Lingfei Qian et al.

总结

本文介绍 Agent Market Arena,这是一个持续运行的实时基准,用于评估大型语言模型交易智能体在加密货币和股票市场中的表现。该框架结合经过验证的交易数据、经专家核验的新闻和多种智能体设计,旨在实盘条件下比较各系统。其研究动机是,先前评估往往只涵盖有限的资产或时期、测试模型而非完整智能体,或依赖未经独立验证的数据。

该基准纳入了风险风格各异的智能体,其中一个使用基于记忆的推理,并用多个语言模型骨干进行评估。据报告,实盘实验发现,相比底层模型之间的差异,各智能体框架在行为上的差异更大,决策从激进到保守不等。摘要将 AMA 描述为持续、可复现评估的基础设施;但未提供表现数据、基准测试时长或比较策略的详细控制方法。其结论涉及观察到的行为差异,并未证明任何智能体能够持续盈利。

核心观点

  • Agent Market Arena 是用于评估加密货币和股票市场 LLM 交易智能体的实时基准。
  • 该框架结合经过验证的交易数据、经核验的新闻和多样的智能体架构。
  • 评估纳入了风险风格各异及采用基于记忆推理的智能体。
  • 实盘实验报告称,智能体框架之间的行为差异大于模型骨干之间的差异。
  • 摘要介绍了基准测试方面的贡献,但没有盈利数据或详细的评估控制方法。

标签

全文
# When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents


# When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents









Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies test models instead of agents, cover limited periods and assets, and rely on unverified data. To address these gaps, we introduce Agent Market Arena (AMA), the first lifelong, real-time benchmark for evaluating LLM-based trading agents across multiple markets. AMA integrates verified trading data, expert-checked news, and diverse agent architectures within a unified trading framework, enabling fair and continuous comparison under real conditions. It implements four agents, including InvestorAgent as a single-agent baseline, TradeAgent and HedgeFundAgent with different risk styles, and DeepFundAgent with memory-based reasoning, and evaluates them across GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4, and Gemini-2.0-flash. Live experiments on both cryptocurrency and stock markets demonstrate that agent frameworks display markedly distinct behavioral patterns, spanning from aggressive risk-taking to conservative decision-making, whereas model backbones contribute less to outcome variation. AMA thus establishes a foundation for rigorous, reproducible, and continuously evolving evaluation of financial reasoning and trading intelligence in LLM-based agents.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。