اختبارات صالحة في أي وقت للعوامل التي يقترحها وكلاء LLM
الملخص
تفصل الدراسة بين اكتشاف العوامل واعتمادها. يمكن لوكيل اقتراح مرشحين وإنشاء مجسات تشخيصية، بينما يقيّم محكّم إحصائي ثابت كل مقترح باستخدام نتائج السوق التي لا تُرصد إلا بعد تقديمه. ويستخدم المحكّم اختبارات قائمة على الرهان صُممت لضبط الاكتشافات الزائفة عند أي وقت للتوقف، بغض النظر عن كيفية اختيار الوكيل لمقترحاته.
يقارن الباحثون بين برنامج نصي وخوارزمية متعددة الأذرع ونموذج لغوي، باستخدام المحكّم الثابت وثلاثة محكّمين متساهلين عمداً. تأتي الأدلة من بيئة اصطناعية زُرعت فيها عوامل حقيقية، وبيئة لإنشاء الاختبارات، ودراسة باختبار التقدم المتحرك لمدة عشر سنوات على CSI 500. ومع الجهة المقترحة المبرمجة، يعتمد المحكّم الثابت عدداً أقل بكثير من العوامل دون العتبة، ولا يزيل تغيير الجهة المقترحة هذه الفجوة. وينتج النموذج اللغوي عوامل أكثر من البرنامج النصي، ويضاهي الخوارزمية متعددة الأذرع، ويمكنه إنشاء اختبارات تشخيصية. والمقابل هو التأخير: تستغرق العوامل الحقيقية نحو 500 يوم تداول حتى تُعتمد، وتكون نسبة شارب للمحفظة المعتمدة أقل منها لمحفظة بلا بوابة. وتعتمد النتائج على تصاميم الدراسة ولا تثبت أن كل عامل أو سوق سيتصرف على نحو مماثل.
الأفكار الرئيسية
- يمكن لمحكّم ثابت تقييم المرشحين بناءً على نتائج تُكشف بعد تقديمهم.
- تهدف الاختبارات القائمة على الرهان إلى الحفاظ على ضبط الاكتشافات الزائفة عند أي وقت للتوقف.
- يؤثر اختيار المحكّم بشدة في عدد العوامل الضعيفة التي تُعتمد.
- يمكن للنماذج اللغوية الإسهام باقتراح العوامل وإنشاء الاختبارات التشخيصية.
- قد يؤخر الاعتماد تبني العوامل الحقيقية ويخفض نسبة شارب للمحفظة مقارنة بالاختيار دون بوابة.
الوسوم
النص الكامل
# Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors # Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.
يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0
أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.