본문으로 건너뛰기
라이브러리 문서 전체

LLM 에이전트가 제안한 요인의 언제든 유효한 검정

기사 arXiv papers · 저자: Bo Qu et al.

요약

이 연구는 요인 발견과 요인 승인을 분리합니다. 에이전트는 후보를 제안하고 진단용 탐침을 만들 수 있으며, 고정된 통계 심판은 제안 이후에만 관측된 시장 결과를 사용해 각 제안을 평가합니다. 심판은 에이전트가 어떤 방식으로 제안하든 정지 시점에 관계없이 거짓 발견을 통제하도록 설계된 베팅 기반 검정을 사용합니다.

연구진은 고정 심판을 두고 스크립트, 밴딧, 언어 모델을 비교하며, 의도적으로 정보 누출이 있는 심판 세 가지와도 비교합니다. 근거는 참 요인을 심어둔 합성 환경, 탐침 작성 환경, CSI 500에서 진행한 10년 워크포워드 연구에서 나옵니다. 스크립트 제안자를 사용하면 고정 심판은 기준 미달 요인을 훨씬 적게 승인하며, 제안자를 바꿔도 그 차이는 사라지지 않습니다. 언어 모델은 스크립트보다 더 많은 요인을 산출하고 밴딧과 비슷한 수준을 보이며 진단용 탐침도 작성할 수 있습니다. 단점은 지연입니다. 참 요인이 승인되기까지 약 500 거래일이 걸리고, 인증된 포트폴리오의 샤프 비율은 필터링하지 않은 포트폴리오보다 낮습니다. 결과는 연구 설계에 따라 달라지며 모든 요인이나 시장에서 비슷하게 작동한다는 점을 입증하지 않습니다.

핵심 아이디어

  • 고정 심판은 제안 이후 공개된 결과로 후보를 평가할 수 있습니다.
  • 베팅 기반 검정은 정지 시점에 관계없이 거짓 발견 통제를 유지하는 것을 목표로 합니다.
  • 심판 선택에 따라 승인되는 약한 요인의 수가 크게 달라집니다.
  • 언어 모델은 요인 제안과 진단용 탐침 작성에 기여할 수 있습니다.
  • 인증 절차는 참 요인의 채택을 늦추고 필터링하지 않은 선택보다 포트폴리오 샤프 비율을 낮출 수 있습니다.

태그

전문
# Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors


# Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors









Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.