跳至正文
返回文库全部文档

搜索量预测标普100股票收益

文章 arXiv papers · 作者: Christopher Bockel-Rickermann

总结

论文检验 Google 搜索活动能否帮助预测哪些标普100股票次日表现将超过指数中位数。研究将滞后金融变量与搜索查询量结合,并训练梯度提升决策树来分类这些结果。研究覆盖2005至2017,报告的平均ROC面积为54.2%至56.7%,表明预测能力高于随机水平。在报告的投资组合测试中,使用多种数据源组合的模型排名最高。

作者还采用简单的统计套利方法构建每日十只股票的投资组合,并报告交易成本前的年化表现高于57%。这些数字是回测结果,不能证明实际实现了扣除成本后的交易收益。研究未计入交易成本,摘要也没有说明稳健性、投资组合构建或实施限制的细节。研究结果表明,搜索行为或可为短期股票预测提供信息,但其实际盈利能力及对市场效率的影响仍不确定。

核心观点

  • 研究将搜索查询量与历史金融数据结合,用于分类股票次日的相对收益。
  • 梯度提升决策树用于预测标普100股票。
  • 报告的分类表现表明,在研究期间预测效果优于随机猜测。
  • 十只股票的投资组合测试报告了较高的成本前收益,但未计入交易成本。
  • 结果引发了有关市场效率的问题,但并未证明能够实现净利润。

标签

全文
# 2205.15853


# Predicting Day-Ahead Stock Returns using Search Engine Query Volumes: An Application of Gradient Boosted Decision Trees to the S&P 100









The internet has changed the way we live, work and take decisions. As it is the major modern resource for research, detailed data on internet usage exhibits vast amounts of behavioral information. This paper aims to answer the question whether this information can be facilitated to predict future returns of stocks on financial capital markets. In an empirical analysis it implements gradient boosted decision trees to learn relationships between abnormal returns of stocks within the S&P 100 index and lagged predictors derived from historical financial data, as well as search term query volumes on the internet search engine Google. Models predict the occurrence of day-ahead stock returns in excess of the index median. On a time frame from 2005 to 2017, all disparate datasets exhibit valuable information. Evaluated models have average areas under the receiver operating characteristic between 54.2% and 56.7%, clearly indicating a classification better than random guessing. Implementing a simple statistical arbitrage strategy, models are used to create daily trading portfolios of ten stocks and result in annual performances of more than 57% before transaction costs. With ensembles of different data sets topping up the performance ranking, the results further question the weak form and semi-strong form efficiency of modern financial capital markets. Even though transaction costs are not included, the approach adds to the existing literature. It gives guidance on how to use and transform data on internet usage behavior for financial and economic modeling and forecasting.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。