コンテンツへスキップ
ライブラリの全資料

LLMエージェントが提案するファクターの随時有効な検定

記事 arXiv papers · 著者: Bo Qu et al.

サマリー

本研究は、ファクターの発見と承認を分けています。エージェントが候補を提案して診断用プローブを作成し、固定された統計的審査役が、提案の提出後にのみ観測された市場結果を使って各候補を評価します。審査役は、エージェントが提案をどのように選ぶかにかかわらず、任意の停止時点で偽発見を制御するよう設計された、賭けに基づく検定を使います。

研究者らは、スクリプト、バンディット、言語モデルを、固定された審査役および意図的に情報漏洩を含ませた3つの審査役と比較しています。証拠は、真の因子を仕込んだ合成環境、プローブ作成環境、そしてCSI 500を対象とする10年間のウォークフォワード研究から得られています。スクリプトによる提案者を使った場合、固定された審査役が採用する基準未満の因子は大幅に少なく、提案者を変えてもこの差は解消されません。言語モデルはスクリプトより多くの因子を提示し、バンディットと同程度の結果を示すほか、診断用プローブも作成できます。トレードオフは時間です。真の因子が認定されるまで約500取引日を要し、認定済みポートフォリオのシャープレシオはゲートを設けないポートフォリオより低くなっています。結果は研究設計に依存しており、すべての因子や市場で同様の結果が得られることを示すものではありません。

主なアイデア

  • 固定された審査役は、提案後に明らかになる市場結果を使って候補を評価できます。
  • 賭けに基づく検定は、任意の停止時点で偽発見の制御を保つことを目指します。
  • 審査役の選択によって、承認される弱いファクターの数は大きく変わります。
  • 言語モデルはファクターの提案に加え、診断用プローブの作成にも活用できます。
  • 認定によって真のファクターの採用が遅れ、ゲートなしの選択に比べてポートフォリオのシャープレシオが下がる場合があります。

タグ

全文
# Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors


# Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors









Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。