בדיקה תקפה בכל עת לגורמים שמציעים סוכני LLM
סיכום
המחקר מפריד בין גילוי גורמים לבין אישורם. סוכן עשוי להציע מועמדים וליצור בדיקות אבחון, בעוד בוחן סטטיסטי מקובע מעריך כל הצעה לפי תוצאות שוק שנצפו רק לאחר הגשתה. הבוחן משתמש במבחנים המבוססים על הימורים, שנועדו לשלוט בתגליות שווא בכל מועד עצירה, ללא תלות באופן שבו הסוכן בוחר את הצעותיו.
החוקרים משווים תסריט, אלגוריתם שודד ומודל שפה לשופט הקפוא ולשלושה שופטים דולפים במכוון. הראיות מגיעות מסביבה סינתטית שבה הוטמעו גורמים אמיתיים, מסביבת יצירת בדיקות וממחקר ווק-פורוורד בן עשר שנים על מדד CSI 500. עם מציע המבוסס על תסריט, השופט הקפוא מאשר הרבה פחות גורמים שמתחת לסף, והחלפת המציע אינה מבטלת את הפער. מודל השפה מניב יותר גורמים מהתסריט, משתווה לאלגוריתם השודד ויכול ליצור בדיקות אבחון. המחיר הוא עיכוב: נדרשים כ־500 ימי מסחר כדי לאשר גורמים אמיתיים, ויחס שארפ של התיק המאושר נמוך מזה של תיק ללא שער סינון. התוצאות תלויות בתכנון המחקרים ואינן מבססות שכל גורם או שוק יתנהגו באופן דומה.
רעיונות מרכזיים
- בוחן מקובע יכול להעריך מועמדים על סמך תוצאות שמתגלות לאחר הגשתם.
- מבחנים המבוססים על הימורים נועדו לשמור על שליטה בתגליות שווא בכל מועד עצירה.
- בחירת השופט משפיעה מאוד על מספר הגורמים החלשים שמאושרים.
- מודלי שפה יכולים לתרום באמצעות הצעת גורמים ויצירת בדיקות אבחון.
- אישור עלול לעכב אימוץ של גורמים אמיתיים ולהפחית את יחס השארפ של התיק לעומת בחירה ללא סינון.
תגיות
הטקסט המלא
# Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors # Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.
מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0
הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.