Using Search Query Volumes to Predict S&P 100 Stock Returns
Summary
The paper tests whether Google search activity can help predict which S&P 100 stocks will outperform the index median the next day. It combines lagged financial variables with search query volumes and trains gradient boosted decision trees to classify those outcomes. The study covers 2005 to 2017 and reports average ROC areas from 54.2% to 56.7%, indicating predictive power above chance. Models using combinations of data sources rank best in the reported portfolio exercise.
The authors also form daily portfolios of ten stocks using a simple statistical arbitrage approach and report annual performance above 57% before transaction costs. Those figures are backtest results, not evidence of realized, net trading returns. Transaction costs are excluded, and the summary provides no details on robustness, portfolio construction, or implementation constraints. The findings suggest search behavior may add information for short-horizon stock forecasting, while leaving practical profitability and implications for market efficiency uncertain.
Key ideas
- Search query volumes are combined with historical financial data to classify next-day relative stock returns.
- Gradient boosted decision trees produce forecasts for S&P 100 stocks.
- The reported classification performance exceeds random guessing over the study period.
- A ten-stock portfolio exercise reports high pre-cost returns, but excludes transaction costs.
- The results raise questions about market efficiency without establishing realizable net profits.
Tags
Full text
# 2205.15853 # Predicting Day-Ahead Stock Returns using Search Engine Query Volumes: An Application of Gradient Boosted Decision Trees to the S&P 100 The internet has changed the way we live, work and take decisions. As it is the major modern resource for research, detailed data on internet usage exhibits vast amounts of behavioral information. This paper aims to answer the question whether this information can be facilitated to predict future returns of stocks on financial capital markets. In an empirical analysis it implements gradient boosted decision trees to learn relationships between abnormal returns of stocks within the S&P 100 index and lagged predictors derived from historical financial data, as well as search term query volumes on the internet search engine Google. Models predict the occurrence of day-ahead stock returns in excess of the index median. On a time frame from 2005 to 2017, all disparate datasets exhibit valuable information. Evaluated models have average areas under the receiver operating characteristic between 54.2% and 56.7%, clearly indicating a classification better than random guessing. Implementing a simple statistical arbitrage strategy, models are used to create daily trading portfolios of ten stocks and result in annual performances of more than 57% before transaction costs. With ensembles of different data sets topping up the performance ranking, the results further question the weak form and semi-strong form efficiency of modern financial capital markets. Even though transaction costs are not included, the approach adds to the existing literature. It gives guidance on how to use and transform data on internet usage behavior for financial and economic modeling and forecasting.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.