تقييم القدرات المتطورة ذاتيًا في أبحاث ألفا طويلة الأفق
الملخص
EverMine إطار لاختبار ما إذا كان وكلاء الأبحاث يتحسنون مع تراكم المهارات والأدوات والقواعد القابلة لإعادة الاستخدام خلال اكتشاف ألفا طويل الأفق. ويمثل حالة البحث عبر سجل الأبحاث ومحفظة العوامل الحالية والقدرات المتراكمة. وتقارن التجارب ذات الموارد المتطابقة بين القدرات الثابتة والمتطورة؛ كما تبدل تجارب الفروع القدرات مع تثبيت سجل البحث وحالة المحفظة. ويُستخدم تكرار الحالة التاريخية لفحص أثر اختيارات المرشحين المدفوعة بالخبرة في نتائج المحفظة.
عبر 18 مسارًا، لا تظهر القدرات المتطورة مكسبًا متسقًا من البداية إلى النهاية، وعبر 48 فرعًا استمراريًا لا تتفوق القدرات المتراكمة باستمرار على المجموعة الأولية. ومع ذلك، قد يفيد ضبط هياكل العوامل القائمة. وفي إعادة تشغيل استكشافية لدفعتين من الفرز في مسار متطور واحد، خفضت المرشحات المرفوضة الواعدة كل على حدة قليلًا معامل المعلومات النهائي للمحفظة عند إضافتها بالتتابع. وتؤكد النتائج أن قيمة المرشح تعتمد على حالة المحفظة وترتيب التقديم، لكن أدلة إعادة التشغيل محدودة ولا تبين أن كل أسلوب لتطوير القدرات سيعمل بالطريقة نفسها.
الأفكار الرئيسية
- يفصل الإطار بين سجل الأبحاث ومحفظة العوامل الحالية والقدرات القابلة لإعادة الاستخدام.
- تقدر التجارب ذات الموارد المتطابقة وتبديلات القدرات المضبوطة بالحالة قيمة القدرات من منظورات مختلفة.
- لم يؤد تطور القدرات إلى مكاسب متسقة في المسارات أو الفروع الاستمرارية المذكورة.
- تعتمد القيمة الحدية للمرشح على المحفظة الحالية وعلى ترتيب إضافة المرشحين.
- إعادة تشغيل المرشحات استكشافية وتستند إلى مسار متطور واحد فقط.
الوسوم
النص الكامل
# EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research # EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the portfolio, the predictive information already covered changes, so the value of the same candidate or experience may change over time. We introduce EverMine, an empirical framework for studying self-evolving research capabilities in long-horizon alpha discovery. EverMine decomposes the research state into history (Hist), the current factor portfolio (Frontier), and reusable capabilities (Cap). Under matched resource limits, we compare complete runs with fixed or evolving Cap, and replace Cap while holding Hist and Frontier fixed to estimate the conditional value of accumulated capabilities. We also combine full trajectories with historical-state replay to examine how experience-based decisions affect candidate selection and portfolio outcomes. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from Cap evolution. Across 48 continuation branches from shared Hist and Frontier states, accumulated Cap also does not consistently outperform the initial Cap. Parameter tuning of existing factor structures can still improve the portfolio. In an exploratory replay of two screening batches from one Evolving trajectory, some screened-out candidates have positive marginal value at the original state, yet submitting all screened-out candidates sequentially slightly lowers final portfolio IC in both batches. These results show that candidate value depends on the evolving portfolio and submission order, and motivate evaluating self-evolving research capabilities through end-to-end outcomes, conditional capability value, and the consequences of experience-based decisions.
يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0
أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.