针对 LLM 智能体提出因子的任意时点有效检验
文章 arXiv papers · 作者: Bo Qu et al.
总结
本研究将因子发现与因子批准分开。智能体可以提出候选因子并创建诊断性探针,而冻结的统计裁判只使用提案提交后观察到的市场结果来评估每项提案。无论智能体如何选择提案,该裁判都使用基于投注的检验,旨在控制任意停止时点的虚假发现。
研究人员将脚本、老虎机算法和语言模型分别与冻结裁判以及三个刻意设置为信息泄漏的裁判进行比较。证据来自植入真实因子的合成环境、探针编写环境,以及在 CSI 500 上进行的十年滚动研究。使用脚本提案器时,冻结裁判接受的低于阈值因子少得多;更换提案器也无法消除这一差距。语言模型产生的因子多于脚本,与老虎机算法相当,并能编写诊断性探针。代价是等待时间:真实因子需要约 500 个交易日才能通过资格检验,获认证投资组合的夏普比率也低于未设门槛的组合。结果取决于研究设计,不能证明所有因子或市场都会呈现相同表现。
核心观点
- 冻结裁判可以根据提案提交后揭晓的结果评估候选因子。
- 基于投注的检验旨在确保在任意停止时点都能控制虚假发现。
- 裁判的选择会显著影响被接受的弱因子数量。
- 语言模型可以通过提出因子和编写诊断性探针发挥作用。
- 认证可能延迟真实因子的采用,并使投资组合夏普比率低于未设门槛的筛选结果。
标签
全文
# Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors # Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。