Skip to content
All library documents

Evaluating Self-Evolving Capabilities in Long-Horizon Alpha Research

Article arXiv papers · Author: Siyuan Li et al.

Summary

EverMine is a framework for testing whether research agents improve as they accumulate reusable skills, tools, and rules during long-horizon alpha discovery. It represents research state through the research history, current factor portfolio, and accumulated capabilities. Matched-resource runs compare fixed and evolving capabilities; additional branch experiments swap capabilities while keeping history and portfolio state constant. Historical-state replay is used to examine how experience-driven candidate choices affect portfolio outcomes.

Across 18 trajectories, evolving capabilities show no consistent end-to-end gain, and across 48 continuation branches accumulated capabilities do not consistently beat the initial set. Tuning existing factor structures can still help. In an exploratory replay of two screening batches from one evolving run, individually promising rejected candidates slightly reduce final portfolio information coefficient when added sequentially. The results emphasize that candidate value depends on portfolio state and submission order, but the replay evidence is limited and does not show that every capability-evolution method will behave the same way.

Key ideas

  • The framework separates research history, current factor portfolio, and reusable capabilities.
  • Matched-resource runs and state-controlled capability swaps estimate capability value from different perspectives.
  • Capability evolution did not produce consistent gains in the reported trajectories or continuation branches.
  • Candidate marginal value depends on the current portfolio and on the order in which candidates are added.
  • The screened-candidate replay is exploratory and comes from only one evolving trajectory.

Tags

Full text
# EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research


# EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research









Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the portfolio, the predictive information already covered changes, so the value of the same candidate or experience may change over time. We introduce EverMine, an empirical framework for studying self-evolving research capabilities in long-horizon alpha discovery. EverMine decomposes the research state into history (Hist), the current factor portfolio (Frontier), and reusable capabilities (Cap). Under matched resource limits, we compare complete runs with fixed or evolving Cap, and replace Cap while holding Hist and Frontier fixed to estimate the conditional value of accumulated capabilities. We also combine full trajectories with historical-state replay to examine how experience-based decisions affect candidate selection and portfolio outcomes. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from Cap evolution. Across 48 continuation branches from shared Hist and Frontier states, accumulated Cap also does not consistently outperform the initial Cap. Parameter tuning of existing factor structures can still improve the portfolio. In an exploratory replay of two screening batches from one Evolving trajectory, some screened-out candidates have positive marginal value at the original state, yet submitting all screened-out candidates sequentially slightly lowers final portfolio IC in both batches. These results show that candidate value depends on the evolving portfolio and submission order, and motivate evaluating self-evolving research capabilities through end-to-end outcomes, conditional capability value, and the consequences of experience-based decisions.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.