Building a Text-Based Earnings Surprise Factor for PEAD Stock Selection
Summary
This study summary describes a text-based factor intended to capture post-earnings-announcement drift in Chinese equities. It uses analyst report titles and abstracts about earnings forecasts, converts selected words into frequency features, and trains Logistic and XGBoost classifiers to categorize relative returns around announcements. The difference between predicted upside and downside log-odds, decayed over time, forms the SUE.txt factor. Monthly portfolio sorts use announcements from the preceding quarter; the summary reports stronger long-side and stratification results for XGBoost than Logistic regression, and says the text factor outperformed a simple announcement-window return measure.
An enhancement stage combines the strongest factor signals within the top SUE.txt group and selects 30 stocks. The reported backtest covers 2013 through 2021 and includes annualized return, excess return, and Sharpe statistics, but these are historical results rather than evidence of future performance. The source warns that machine-learning signals can fail and that model interpretability is limited; the full paper is referenced but not included here.
Key ideas
- The method models analyst commentary on earnings forecasts to estimate a text-based earnings surprise signal.
- Word-frequency features from report titles and abstracts train classifiers against announcement-window relative returns.
- The factor compares predicted upside and downside probabilities and applies time decay.
- The summary reports that XGBoost performed better than Logistic regression in its portfolio sorts.
- Historical backtest results do not establish future performance, and the source flags model failure and interpretability risks.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.