コンテンツへスキップ
ライブラリの全資料

資産価格ファクター向けの文脈対応型ニューストピック

記事 arXiv papers · 著者: Kevin Foley et al.

サマリー

本研究は、文単位の文脈情報によって、システマティックな資産価格ファクターの構築に使うニューストピックモデルが改善するか検証します。同一の394,661記事と後続のポートフォリオ構築プロセスを用いて、潜在的ディリクレ配分法と、固定した文埋め込みモデルにk-meansクラスタリングを組み合わせた手法を比較します。両手法では、使用する記事本文の量とトピック語の順位付け方法も異なります。

トランスフォーマー手法では、NPMIで測定したトピックの一貫性とポートフォリオのシャープレシオがより高かったものの、利用可能な検証からはLDAを上回るとは確認できません。球面クラスタリングと複数の期間にわたるエクスポージャーを用いた探索的な手法では、超過リターンのシャープレシオが1.03の統合モデルが得られました。これらの結果は、文脈を考慮した表現がニュースから金融上有用なシグナルを抽出するのに役立つ可能性を示唆しています。ただし確度には限界があります。この研究では、各時点で利用可能な情報に入力を限定する、より厳密な検証と、より幅広いデータセットでの評価が必要だとしています。

主なアイデア

  • LDAのトピックと、文埋め込みをk-meansでクラスタリングしたトピックを比較します。
  • 報告された分析では、トランスフォーマーモデルの手法でトピックの一貫性とポートフォリオのシャープレシオが高くなりました。
  • 記事コレクションと後続のポートフォリオ構築工程は共通ですが、入力する本文量と用語の順位付けは異なります。
  • 探索的な球面クラスタリングと複数期間のエクスポージャーから、統合モデルで超過リターンのシャープレシオ1.03が得られました。
  • 証拠は決定的ではなく、時点情報に基づく検証とより広範なデータセットが必要です。

タグ

全文
# From Word Counts to Context: Topic Models for Asset Pricing


# From Word Counts to Context: Topic Models for Asset Pricing









News may reveal systematic risk, but whether its context enhances the construction of systematic risk factors is still unclear. We seek to test whether utilizing a sentence transformer represents an improvement over techniques such as Latent Dirichlet Allocation (LDA) in the coherence of topic term lists generated from unstructured text data. To test this, the same collection of unstructured text data comprising of 394,661 articles and the same downstream financial portfolio construction pipeline were applied with the text layer differing, including the length of article text each model used and how topic terms were ranked: we benchmark LDA against a frozen sentence transformer with k-means clustering. We find that the sentence transformer branch had higher observed scores both in terms of coherence (measured by NPMI) as well as financial performance (measured by Sharpe), although the available tests do not establish outperformance. Further exploratory specifications such as utilizing spherical clustering and multi-horizon exposures had an observed excess-return Sharpe of 1.03 for the combined model. We believe that there is some promise in applying context-aware techniques on unstructured news text, but stricter tests using only information available at each date and broader datasets may be required to enhance the confidence in the observed performance.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。