评估长周期阿尔法研究中的自我进化能力
文章 arXiv papers · 作者: Siyuan Li et al.
总结
EverMine 是一个框架,用于测试研究智能体在长期阿尔法发现过程中累积可复用技能、工具和规则后是否会有所改进。它通过研究历史、当前因子投资组合和累积能力来表示研究状态。资源匹配实验比较固定能力与不断演化的能力;其他分支实验则在保持研究历史和投资组合状态不变的同时替换能力。研究还通过历史状态回放,考察基于经验的候选选择如何影响投资组合结果。
在18条轨迹中,能力演化没有带来一致的端到端收益;在48个延续分支中,累积能力也未能持续胜过初始能力集。调整现有因子结构仍可能有所帮助。在对一次演化运行的两批筛选结果进行探索性回放时,依次加入单独看来有前景但被拒绝的候选项,略微降低了最终投资组合的信息系数。结果强调,候选项的价值取决于投资组合状态和提交顺序;但回放证据有限,也未表明所有能力演化方法都会有相同表现。
核心观点
- 该框架将研究历史、当前因子投资组合和可复用能力区分开来。
- 资源匹配实验和控制状态的能力替换实验从不同角度估算能力的价值。
- 报告中的轨迹和延续分支未显示能力演化带来一致收益。
- 候选项的边际价值取决于当前投资组合及候选项加入顺序。
- 被筛选候选项的回放属于探索性分析,且只来自一条演化轨迹。
标签
全文
# EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research # EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the portfolio, the predictive information already covered changes, so the value of the same candidate or experience may change over time. We introduce EverMine, an empirical framework for studying self-evolving research capabilities in long-horizon alpha discovery. EverMine decomposes the research state into history (Hist), the current factor portfolio (Frontier), and reusable capabilities (Cap). Under matched resource limits, we compare complete runs with fixed or evolving Cap, and replace Cap while holding Hist and Frontier fixed to estimate the conditional value of accumulated capabilities. We also combine full trajectories with historical-state replay to examine how experience-based decisions affect candidate selection and portfolio outcomes. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from Cap evolution. Across 48 continuation branches from shared Hist and Frontier states, accumulated Cap also does not consistently outperform the initial Cap. Parameter tuning of existing factor structures can still improve the portfolio. In an exploratory replay of two screening batches from one Evolving trajectory, some screened-out candidates have positive marginal value at the original state, yet submitting all screened-out candidates sequentially slightly lowers final portfolio IC in both batches. These results show that candidate value depends on the evolving portfolio and submission order, and motivate evaluating self-evolving research capabilities through end-to-end outcomes, conditional capability value, and the consequences of experience-based decisions.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。