Skip to content
All library documents

Context-Aware News Topics for Asset Pricing Factors

Article arXiv papers · Author: Kevin Foley et al.

Summary

This study tests whether sentence-level context improves news topic models used to build systematic asset-pricing factors. It compares Latent Dirichlet Allocation with a frozen sentence transformer followed by k-means clustering, using the same collection of 394,661 articles and downstream portfolio construction process. The approaches also differ in how much article text they use and how they rank topic terms.

The transformer approach recorded higher topic coherence, measured by NPMI, and higher portfolio Sharpe scores, but the available tests do not establish that it outperforms LDA. Exploratory versions using spherical clustering and exposures across multiple horizons produced a combined model with an excess-return Sharpe of 1.03. The findings suggest that context-aware representations may help extract financially useful signals from news. Confidence remains limited: the study calls for stricter tests that restrict inputs to information available at each date and for evaluation on broader datasets.

Key ideas

  • The study compares LDA topics with topics from sentence-transformer embeddings clustered by k-means.
  • Both topic coherence and portfolio Sharpe scores were higher for the transformer approach in the reported analysis.
  • The comparison uses the same article collection and downstream portfolio construction pipeline, though the text input and term ranking differ.
  • Exploratory spherical clustering and multi-horizon exposures yielded an excess-return Sharpe of 1.03 for the combined model.
  • The evidence is not conclusive, and point-in-time tests and broader datasets are needed.

Tags

Full text
# From Word Counts to Context: Topic Models for Asset Pricing


# From Word Counts to Context: Topic Models for Asset Pricing









News may reveal systematic risk, but whether its context enhances the construction of systematic risk factors is still unclear. We seek to test whether utilizing a sentence transformer represents an improvement over techniques such as Latent Dirichlet Allocation (LDA) in the coherence of topic term lists generated from unstructured text data. To test this, the same collection of unstructured text data comprising of 394,661 articles and the same downstream financial portfolio construction pipeline were applied with the text layer differing, including the length of article text each model used and how topic terms were ranked: we benchmark LDA against a frozen sentence transformer with k-means clustering. We find that the sentence transformer branch had higher observed scores both in terms of coherence (measured by NPMI) as well as financial performance (measured by Sharpe), although the available tests do not establish outperformance. Further exploratory specifications such as utilizing spherical clustering and multi-horizon exposures had an observed excess-return Sharpe of 1.03 for the combined model. We believe that there is some promise in applying context-aware techniques on unstructured news text, but stricter tests using only information available at each date and broader datasets may be required to enhance the confidence in the observed performance.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.