הערכת יכולות מתפתחות במחקר אלפא לטווח ארוך
סיכום
EverMine היא מסגרת לבדיקת השאלה אם סוכני מחקר משתפרים ככל שהם צוברים מיומנויות, כלים וכללים לשימוש חוזר במהלך גילוי אלפא לטווח ארוך. היא מייצגת את מצב המחקר באמצעות היסטוריית המחקר, תיק הגורמים הנוכחי והיכולות שנצברו. הרצות עם משאבים תואמים משוות בין יכולות קבועות למתפתחות; ניסויי הסתעפות נוספים מחליפים יכולות תוך שמירה על היסטוריה ומצב תיק קבועים. השמעה חוזרת של מצבים היסטוריים משמשת לבדיקת האופן שבו בחירת מועמדים המבוססת על ניסיון משפיעה על תוצאות התיק.
לאורך 18 מסלולים, היכולות המתפתחות אינן מציגות רווח עקבי מתחילת התהליך ועד סופו, וב־48 ענפי המשך היכולות שנצברו אינן גוברות בעקביות על קבוצת היכולות הראשונית. כוונון של מבני גורמים קיימים עדיין עשוי לעזור. בהשמעה חוזרת חקרנית של שני מקבצי סינון מהרצה מתפתחת אחת, מועמדים שנדחו, אך נמצאו מבטיחים בהערכה פרטנית, מפחיתים מעט את מקדם המידע של התיק הסופי כאשר מוסיפים אותם בזה אחר זה. התוצאות מדגישות שערכם של המועמדים תלוי במצב התיק ובסדר ההגשה, אך ראיות ההשמעה החוזרת מוגבלות ואינן מראות שכל שיטה להתפתחות יכולות תפעל באותו אופן.
רעיונות מרכזיים
- המסגרת מפרידה בין היסטוריית המחקר, תיק הגורמים הנוכחי ויכולות לשימוש חוזר.
- הרצות עם משאבים תואמים והחלפות יכולות תחת בקרת מצב אומדות את ערך היכולות מנקודות מבט שונות.
- התפתחות היכולות לא הניבה רווחים עקביים במסלולים או בענפי ההמשך שדווחו.
- הערך השולי של מועמד תלוי בתיק הנוכחי ובסדר שבו מוסיפים מועמדים.
- ההשמעה החוזרת של מועמדים שעברו סינון היא חקרנית ומבוססת על מסלול מתפתח אחד בלבד.
תגיות
הטקסט המלא
# EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research # EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the portfolio, the predictive information already covered changes, so the value of the same candidate or experience may change over time. We introduce EverMine, an empirical framework for studying self-evolving research capabilities in long-horizon alpha discovery. EverMine decomposes the research state into history (Hist), the current factor portfolio (Frontier), and reusable capabilities (Cap). Under matched resource limits, we compare complete runs with fixed or evolving Cap, and replace Cap while holding Hist and Frontier fixed to estimate the conditional value of accumulated capabilities. We also combine full trajectories with historical-state replay to examine how experience-based decisions affect candidate selection and portfolio outcomes. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from Cap evolution. Across 48 continuation branches from shared Hist and Frontier states, accumulated Cap also does not consistently outperform the initial Cap. Parameter tuning of existing factor structures can still improve the portfolio. In an exploratory replay of two screening batches from one Evolving trajectory, some screened-out candidates have positive marginal value at the original state, yet submitting all screened-out candidates sequentially slightly lowers final portfolio IC in both batches. These results show that candidate value depends on the evolving portfolio and submission order, and motivate evaluating self-evolving research capabilities through end-to-end outcomes, conditional capability value, and the consequences of experience-based decisions.
מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0
הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.