본문으로 건너뛰기
라이브러리 문서 전체

장기 알파 연구에서 진화 역량 평가

기사 arXiv papers · 저자: Siyuan Li et al.

요약

EverMine은 연구 에이전트가 장기 알파 발굴 과정에서 재사용 가능한 기술, 도구, 규칙을 축적하며 개선되는지 검증하는 프레임워크입니다. 연구 이력, 현재 요인 포트폴리오, 축적된 역량으로 연구 상태를 표현합니다. 동일한 리소스를 맞춘 실행에서는 고정 역량과 진화 역량을 비교하고, 추가 분기 실험에서는 이력과 포트폴리오 상태를 고정한 채 역량을 교체합니다. 경험에 따른 후보 선택이 포트폴리오 결과에 미치는 영향을 살펴보기 위해 과거 상태를 재생합니다.

18개 궤적에서는 진화 역량이 처음부터 끝까지 일관된 개선을 보이지 않았고, 48개 연속 분기에서도 축적된 역량이 초기 집합을 꾸준히 앞서지 못했습니다. 기존 요인 구조를 조정하면 여전히 도움이 될 수 있습니다. 진화 실행 하나에서 두 차례 선별 과정을 탐색적으로 재생한 결과, 개별적으로는 유망했던 탈락 후보를 순차적으로 추가했을 때 최종 포트폴리오 정보계수가 소폭 낮아졌습니다. 결과는 후보의 가치가 포트폴리오 상태와 제출 순서에 따라 달라진다는 점을 강조하지만, 재생 근거는 제한적이며 모든 역량 진화 방법이 같은 양상을 보인다는 점을 입증하지 않습니다.

핵심 아이디어

  • 이 프레임워크는 연구 이력, 현재 요인 포트폴리오, 재사용 가능한 역량을 구분합니다.
  • 동일 리소스 실행과 상태를 통제한 역량 교체로 서로 다른 관점에서 역량의 가치를 추정합니다.
  • 보고된 궤적과 연속 분기에서는 역량 진화가 일관된 개선을 만들지 못했습니다.
  • 후보의 한계 가치는 현재 포트폴리오와 후보를 추가하는 순서에 따라 달라집니다.
  • 선별 후보 재생은 탐색적이며 진화 궤적 하나만을 대상으로 합니다.

태그

전문
# EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research


# EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research









Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the portfolio, the predictive information already covered changes, so the value of the same candidate or experience may change over time. We introduce EverMine, an empirical framework for studying self-evolving research capabilities in long-horizon alpha discovery. EverMine decomposes the research state into history (Hist), the current factor portfolio (Frontier), and reusable capabilities (Cap). Under matched resource limits, we compare complete runs with fixed or evolving Cap, and replace Cap while holding Hist and Frontier fixed to estimate the conditional value of accumulated capabilities. We also combine full trajectories with historical-state replay to examine how experience-based decisions affect candidate selection and portfolio outcomes. Across 18 long-horizon trajectories, end-to-end comparisons show no consistent gain from Cap evolution. Across 48 continuation branches from shared Hist and Frontier states, accumulated Cap also does not consistently outperform the initial Cap. Parameter tuning of existing factor structures can still improve the portfolio. In an exploratory replay of two screening batches from one Evolving trajectory, some screened-out candidates have positive marginal value at the original state, yet submitting all screened-out candidates sequentially slightly lowers final portfolio IC in both batches. These results show that candidate value depends on the evolving portfolio and submission order, and motivate evaluating self-evolving research capabilities through end-to-end outcomes, conditional capability value, and the consequences of experience-based decisions.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.