基于预测市场订单簿回放的交易智能体回测
文章 arXiv papers · 作者: Avi Arora et al.
总结
该文提出了一套基准,用于评估预测市场中的算法交易智能体和语言模型交易智能体。研究根据订单簿更新、交易、合约生命周期事件和结算信息构建历史片段,再以确定性方式回放,使策略面对一致的市场条件序列。智能体接口允许传统算法和使用工具的语言模型与模拟环境交互。
模拟器对做市方和吃单方的执行行为以及费用进行建模,从而可以评估交易摩擦和结算风险下的结果。文中介绍了四个基于加密货币、天气和体育市场的片段。基线结果显示,简单智能体可能因交易成本或结算结果而亏损;而考虑费用的策略在波动较大的片段中仍能保持竞争力。证据仅限于这四个片段和所述基线;摘要没有提供详细表现数据,也未对其他市场的结果作更广泛的断言。
核心观点
- 历史市场数据流可以通过确定性回放,让不同智能体在一致条件下进行比较。
- 该基准将订单簿数据、交易、合约生命周期事件和结算信息组合为多个片段。
- 模拟器对做市方和吃单方的行为以及费用进行建模。
- 成本和结算损失可能削弱简单交易策略的表现。
- 据报告,考虑费用的算法在波动较大的片段中具有竞争力,但证据仅限于四个市场。
标签
全文
# PredictionMarketBench: A SWE-bench-Style Framework for Backtesting Trading Agents on Prediction Markets # PredictionMarketBench: A SWE-bench-Style Framework for Backtesting Trading Agents on Prediction Markets Prediction markets offer a natural testbed for trading agents: contracts have binary payoffs, prices can be interpreted as probabilities, and realized performance depends critically on market microstructure, fees, and settlement risk. We introduce PredictionMarketBench, a SWE-bench-style benchmark for evaluating algorithmic and LLM-based trading agents on prediction markets via deterministic, event-driven replay of historical limit-order-book and trade data. PredictionMarketBench standardizes (i) episode construction from raw exchange streams (orderbooks, trades, lifecycle, settlement), (ii) an execution-realistic simulator with maker/taker semantics and fee modeling, and (iii) a tool-based agent interface that supports both classical strategies and tool-calling LLM agents with reproducible trajectories. We release four Kalshi-based episodes spanning cryptocurrency, weather, and sports. Baseline results show that naive trading agents can underperform due to transaction costs and settlement losses, while fee-aware algorithmic strategies remain competitive in volatile episodes.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。