Anytime-Valid Testing for Factors Proposed by LLM Agents
Summary
The study separates factor discovery from factor approval. An agent may propose candidates and create diagnostic probes, while a frozen statistical referee evaluates each proposal using market outcomes observed only after submission. The referee uses betting-based tests designed to control false discoveries at any stopping time, regardless of how the agent chooses its proposals.
The researchers compare a script, a bandit, and a language model with the frozen referee and with three deliberately leaky referees. Evidence comes from a synthetic setting with planted true factors, a probe-authoring environment, and a ten-year walk-forward study on the CSI 500. With the scripted proposer, the frozen referee admits far fewer sub-threshold factors, and changing the proposer does not remove that gap. The language model yields more factors than the script, matches the bandit, and can author diagnostic probes. The tradeoff is delay: true factors take about 500 trading days to qualify, and the certified portfolio has a lower Sharpe ratio than an ungated one. Results depend on the study designs and do not establish that every factor or market will behave similarly.
Key ideas
- A frozen referee can evaluate candidates on outcomes revealed after submission.
- Betting-based tests aim to preserve false-discovery control at any stopping time.
- The choice of referee strongly affects how many weak factors are admitted.
- Language models can contribute by proposing factors and authoring diagnostic probes.
- Certification can delay adoption of true factors and reduce portfolio Sharpe relative to ungated selection.
Tags
Full text
# Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors # Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.