コンテンツへスキップ
ライブラリの全資料

検索クエリ量によるS&P 100株リターン予測

記事 arXiv papers · 著者: Christopher Bockel-Rickermann

サマリー

Google検索の動向が、翌日にS&P 100のどの銘柄が指数中央値を上回るかの予測に役立つかを検証しています。過去の金融変数に検索クエリ量を加え、勾配ブースティング決定木で結果を分類します。対象期間は2005から2017で、平均ROC面積は54.2%から56.7%と報告されており、偶然を上回る予測力を示しています。報告されたポートフォリオ分析では、複数のデータソースを組み合わせたモデルの成績が最良です。

著者らは単純な統計的アービトラージ手法を使って10銘柄の日次ポートフォリオも構築し、取引コスト控除前の年率成績が57%を上回ると報告しています。これらの数値はバックテストの結果であり、実現したコスト控除後の売買リターンを示すものではありません。取引コストは含まれておらず、要約にも頑健性、ポートフォリオ構築、実装上の制約に関する詳細はありません。検索行動が短期の株価予測に情報を加える可能性を示唆する一方、実務的な収益性や市場効率性への示唆は不確かなままです。

主なアイデア

  • 検索クエリ量を過去の金融データと組み合わせ、翌日の相対的な株式リターンを分類します。
  • 勾配ブースティング決定木を使い、S&P 100銘柄のリターンを予測します。
  • 対象期間における分類成績は、報告上、ランダムな予測を上回っています。
  • 10銘柄のポートフォリオ分析では高いコスト控除前リターンが報告されていますが、取引コストは含まれていません。
  • 結果は市場効率性への疑問を提起しますが、実現可能なコスト控除後利益を立証するものではありません。

タグ

全文
# 2205.15853


# Predicting Day-Ahead Stock Returns using Search Engine Query Volumes: An Application of Gradient Boosted Decision Trees to the S&P 100









The internet has changed the way we live, work and take decisions. As it is the major modern resource for research, detailed data on internet usage exhibits vast amounts of behavioral information. This paper aims to answer the question whether this information can be facilitated to predict future returns of stocks on financial capital markets. In an empirical analysis it implements gradient boosted decision trees to learn relationships between abnormal returns of stocks within the S&P 100 index and lagged predictors derived from historical financial data, as well as search term query volumes on the internet search engine Google. Models predict the occurrence of day-ahead stock returns in excess of the index median. On a time frame from 2005 to 2017, all disparate datasets exhibit valuable information. Evaluated models have average areas under the receiver operating characteristic between 54.2% and 56.7%, clearly indicating a classification better than random guessing. Implementing a simple statistical arbitrage strategy, models are used to create daily trading portfolios of ten stocks and result in annual performances of more than 57% before transaction costs. With ensembles of different data sets topping up the performance ranking, the results further question the weak form and semi-strong form efficiency of modern financial capital markets. Even though transaction costs are not included, the approach adds to the existing literature. It gives guidance on how to use and transform data on internet usage behavior for financial and economic modeling and forecasting.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。