Skip to content
All library documents

Text-Based Earnings Surprise Signals for Post-Earnings Drift

Article BigQuant

Summary

This study builds a text-derived earnings surprise factor, SUE.txt, to capture post-earnings announcement drift in Chinese equities. It uses analyst report titles and summaries related to earnings forecasts, converts them into word-frequency features, and trains rolling Logistic and XGBoost classifiers. The training labels classify each forecast's two-day excess return relative to the CSI 500 as rising, flat, or falling. The factor is based on the difference between predicted rise and fall log odds, with exponential decay applied to older observations.

Monthly quintile tests over 2013–2021 reportedly show stronger long-side and sorting performance for XGBoost than Logistic regression, and the text factor outperforms a decayed two-day return signal. The authors also examine word and paragraph importance, finding intuitive links between positive or negative earnings language and factor direction. A further combination with selected research factors produces a reported enhanced portfolio. These are historical backtest claims; the summary gives limited detail on costs, turnover, implementation, or out-of-sample validation.

Key ideas

  • The factor extracts earnings-related signals from analyst report text about company earnings forecasts.
  • Word frequencies from titles and summaries serve as features for rolling Logistic and XGBoost models.
  • Model labels use the two-day excess return around an earnings forecast, grouped into three directions.
  • The study reports that XGBoost sorts stocks more effectively than Logistic regression in its historical tests.
  • Words associated with upgrades and profit growth tend to contribute positively, while loss-related terms contribute negatively.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.