본문으로 건너뛰기
라이브러리 문서 전체

자산 가격 요인을 위한 문맥 기반 뉴스 주제

기사 arXiv papers · 저자: Kevin Foley et al.

요약

이 연구는 문장 수준의 맥락이 체계적 자산가격결정 요인 구축에 쓰이는 뉴스 토픽 모델을 개선하는지 검증합니다. 동일한 394,661개 기사와 후속 포트폴리오 구성 절차를 사용해 잠재 디리클레 할당과 고정된 문장 변환기 뒤에 k-평균 군집화를 적용한 방식을 비교합니다. 두 접근법은 사용하는 기사 텍스트의 양과 토픽 용어의 순위를 매기는 방식에서도 차이가 있습니다.

변환기 방식은 NPMI로 측정한 토픽 응집도가 더 높고 포트폴리오 샤프 지수도 높았지만, 현재 검증으로는 해당 방식이 LDA보다 우수하다고 입증되지 않았습니다. 구면 군집화와 여러 투자 기간의 익스포저를 적용한 탐색적 버전은 초과수익 샤프 지수 1.03의 통합 모델을 만들었습니다. 연구 결과는 맥락을 고려한 표현이 뉴스에서 재무적으로 유용한 신호를 추출하는 데 도움이 될 수 있음을 시사합니다. 다만 확신에는 한계가 있습니다. 연구는 각 시점에 이용 가능했던 정보만 입력으로 제한하는 더 엄격한 검증과 더 폭넓은 데이터셋 평가가 필요하다고 밝혔습니다.

핵심 아이디어

  • 이 연구는 LDA 토픽과 문장 변환기 임베딩을 k-평균으로 군집화해 얻은 토픽을 비교합니다.
  • 보고된 분석에서 변환기 접근법은 주제 응집도와 포트폴리오 샤프 지수가 모두 더 높았습니다.
  • 두 방법은 같은 기사 모음과 후속 포트폴리오 구성 절차를 쓰지만 텍스트 입력과 용어 순위는 다릅니다.
  • 탐색적 구면 군집화와 다중 기간 익스포저를 적용한 통합 모델은 초과수익 샤프 지수 1.03을 기록했습니다.
  • 근거는 결정적이지 않으며 시점별 정보만 사용하는 테스트와 더 폭넓은 데이터셋이 필요합니다.

태그

전문
# From Word Counts to Context: Topic Models for Asset Pricing


# From Word Counts to Context: Topic Models for Asset Pricing









News may reveal systematic risk, but whether its context enhances the construction of systematic risk factors is still unclear. We seek to test whether utilizing a sentence transformer represents an improvement over techniques such as Latent Dirichlet Allocation (LDA) in the coherence of topic term lists generated from unstructured text data. To test this, the same collection of unstructured text data comprising of 394,661 articles and the same downstream financial portfolio construction pipeline were applied with the text layer differing, including the length of article text each model used and how topic terms were ranked: we benchmark LDA against a frozen sentence transformer with k-means clustering. We find that the sentence transformer branch had higher observed scores both in terms of coherence (measured by NPMI) as well as financial performance (measured by Sharpe), although the available tests do not establish outperformance. Further exploratory specifications such as utilizing spherical clustering and multi-horizon exposures had an observed excess-return Sharpe of 1.03 for the combined model. We believe that there is some promise in applying context-aware techniques on unstructured news text, but stricter tests using only information available at each date and broader datasets may be required to enhance the confidence in the observed performance.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.