暗号資産・株式市場で実運用されるLLM取引エージェントのベンチマーク
記事 arXiv papers · 著者: Lingfei Qian et al.
サマリー
本論文は、暗号資産市場と株式市場における大規模言語モデルの取引エージェントを評価する、継続的なリアルタイムベンチマーク「Agent Market Arena」を紹介します。この枠組みは、検証済みの取引データ、専門家が確認したニュース、多様なエージェント設計を組み合わせ、実運用環境でシステムを比較することを目指します。背景には、従来の評価が対象資産や期間を限定しがちであること、完全なエージェントではなくモデルを検証すること、独立検証されていないデータに依存することがあります。
このベンチマークには、異なるリスクスタイルを持つエージェントや、記憶に基づく推論を使うエージェントが含まれ、複数の言語モデル基盤で評価されます。報告されたライブ実験では、基盤となるモデル間よりもエージェントの枠組み間のほうが、積極的な判断から保守的な判断まで行動の差が大きく見られました。要旨はAMAを継続的で再現可能な評価の基盤として示していますが、成績数値、ベンチマーク期間、戦略比較の詳細な統制条件は示していません。結論は観察された行動の違いに関するもので、いずれかのエージェントが一貫して利益を上げることを示すものではありません。
主なアイデア
- Agent Market Arenaは、暗号資産・株式市場におけるLLM取引エージェント向けのリアルタイムベンチマークです。
- 検証済みの取引データ、確認済みニュース、多様なエージェント構成を組み合わせています。
- 評価対象には、異なるリスクスタイルや記憶に基づく推論を持つエージェントが含まれます。
- ライブ実験では、モデル基盤間よりエージェントの枠組み間で行動の差が大きいと報告されています。
- 要旨はベンチマーク構築への貢献を述べていますが、収益性の数値や詳細な評価統制条件は示していません。
タグ
全文
# When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents # When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies test models instead of agents, cover limited periods and assets, and rely on unverified data. To address these gaps, we introduce Agent Market Arena (AMA), the first lifelong, real-time benchmark for evaluating LLM-based trading agents across multiple markets. AMA integrates verified trading data, expert-checked news, and diverse agent architectures within a unified trading framework, enabling fair and continuous comparison under real conditions. It implements four agents, including InvestorAgent as a single-agent baseline, TradeAgent and HedgeFundAgent with different risk styles, and DeepFundAgent with memory-based reasoning, and evaluates them across GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4, and Gemini-2.0-flash. Live experiments on both cryptocurrency and stock markets demonstrate that agent frameworks display markedly distinct behavioral patterns, spanning from aggressive risk-taking to conservative decision-making, whereas model backbones contribute less to outcome variation. AMA thus establishes a foundation for rigorous, reproducible, and continuously evolving evaluation of financial reasoning and trading intelligence in LLM-based agents.
出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。