Skip to content
All library documents

Backtesting Trading Agents with Prediction Market Order Book Replay

Article arXiv papers · Author: Avi Arora et al.

Summary

The document presents a benchmark for evaluating algorithmic and language-model trading agents in prediction markets. It builds historical episodes from order book updates, trades, contract lifecycle events, and settlements, then replays them deterministically so strategies face a consistent sequence of market conditions. An agent interface lets both conventional algorithms and tool-using language models interact with the simulation.

The simulator models maker and taker execution along with fees, making it possible to assess outcomes under trading frictions and settlement risk. Four episodes based on cryptocurrency, weather, and sports markets are described. Baseline findings indicate that simple agents may lose to transaction costs or settlement outcomes, while strategies that account for fees can stay competitive in volatile episodes. The evidence is limited to those four episodes and the stated baselines; the summary provides no detailed performance figures or broader claims about results across other markets.

Key ideas

  • Historical market streams can be replayed deterministically to compare agents under consistent conditions.
  • The benchmark combines book data, trades, contract lifecycle events, and settlements into episodes.
  • Its simulator models maker and taker behavior as well as fees.
  • Naive trading can be weakened by costs and settlement losses.
  • Fee-aware algorithms are reported as competitive in volatile episodes, with evidence limited to four markets.

Tags

Full text
# PredictionMarketBench: A SWE-bench-Style Framework for Backtesting Trading Agents on Prediction Markets


# PredictionMarketBench: A SWE-bench-Style Framework for Backtesting Trading Agents on Prediction Markets









Prediction markets offer a natural testbed for trading agents: contracts have binary payoffs, prices can be interpreted as probabilities, and realized performance depends critically on market microstructure, fees, and settlement risk. We introduce PredictionMarketBench, a SWE-bench-style benchmark for evaluating algorithmic and LLM-based trading agents on prediction markets via deterministic, event-driven replay of historical limit-order-book and trade data. PredictionMarketBench standardizes (i) episode construction from raw exchange streams (orderbooks, trades, lifecycle, settlement), (ii) an execution-realistic simulator with maker/taker semantics and fee modeling, and (iii) a tool-based agent interface that supports both classical strategies and tool-calling LLM agents with reproducible trajectories. We release four Kalshi-based episodes spanning cryptocurrency, weather, and sports. Baseline results show that naive trading agents can underperform due to transaction costs and settlement losses, while fee-aware algorithmic strategies remain competitive in volatile episodes.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.