Skip to content
All library documents

Benchmarking LLM Trading Agents in Live Crypto and Stock Markets

Article arXiv papers · Author: Lingfei Qian et al.

Summary

This paper introduces Agent Market Arena, a continuous real-time benchmark for evaluating large language model trading agents across cryptocurrency and stock markets. The framework combines verified trading data with expert-checked news and multiple agent designs, aiming to compare systems under live conditions. Its motivation is that prior evaluations often use limited assets or periods, test models rather than complete agents, or rely on data that has not been independently verified.

The benchmark includes agents with different risk styles and one that uses memory-based reasoning, and evaluates them with several language model backbones. The reported live experiments find more variation in behavior across agent frameworks, from aggressive to conservative decisions, than across the underlying models. The abstract presents AMA as infrastructure for ongoing, reproducible evaluation; it does not give performance figures, benchmark duration, or detailed controls for comparing strategies. Its conclusions concern observed behavioral differences and do not establish that any agent is consistently profitable.

Key ideas

  • Agent Market Arena is a real-time benchmark for LLM trading agents across crypto and stock markets.
  • The framework combines verified trading data, checked news, and diverse agent architectures.
  • The evaluation includes agents with different risk styles and memory-based reasoning.
  • Live experiments report larger behavioral differences between agent frameworks than between model backbones.
  • The abstract describes a benchmarking contribution but gives no profitability figures or detailed evaluation controls.

Tags

Full text
# When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents


# When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents









Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies test models instead of agents, cover limited periods and assets, and rely on unverified data. To address these gaps, we introduce Agent Market Arena (AMA), the first lifelong, real-time benchmark for evaluating LLM-based trading agents across multiple markets. AMA integrates verified trading data, expert-checked news, and diverse agent architectures within a unified trading framework, enabling fair and continuous comparison under real conditions. It implements four agents, including InvestorAgent as a single-agent baseline, TradeAgent and HedgeFundAgent with different risk styles, and DeepFundAgent with memory-based reasoning, and evaluates them across GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4, and Gemini-2.0-flash. Live experiments on both cryptocurrency and stock markets demonstrate that agent frameworks display markedly distinct behavioral patterns, spanning from aggressive risk-taking to conservative decision-making, whereas model backbones contribute less to outcome variation. AMA thus establishes a foundation for rigorous, reproducible, and continuously evolving evaluation of financial reasoning and trading intelligence in LLM-based agents.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.