رفتن به محتوا
همه اسناد کتابخانه

پیش‌بینی بازده سهام S&P 100 با حجم جست‌وجو

مقاله arXiv papers · نویسنده: Christopher Bockel-Rickermann

خلاصه

این مقاله می‌آزماید که آیا فعالیت جست‌وجوی گوگل می‌تواند به پیش‌بینی سهام S&P 100 که روز بعد از میانه شاخص عملکرد بهتری خواهند داشت کمک کند. متغیرهای مالی با وقفه را با حجم جست‌وجو ترکیب می‌کند و درخت‌های تصمیم تقویت‌شده با گرادیان را برای طبقه‌بندی این پیامدها آموزش می‌دهد. پژوهش بازه 2005 تا 2017 را پوشش می‌دهد و میانگین مقادیر مساحت زیر منحنی ROC را در بازه 54.2% تا 56.7% گزارش می‌کند که نشان‌دهنده قدرت پیش‌بینی بالاتر از شانس است. در تمرین سبد گزارش‌شده، مدل‌های ترکیبیِ منابع داده رتبه بهتری دارند.

نویسندگان همچنین با رویکردی ساده برای آربیتراژ آماری، سبدهای روزانه‌ای از ده سهم می‌سازند و عملکرد سالانه بیش از 57% را پیش از هزینه‌های تراکنش گزارش می‌کنند. این ارقام نتایج بک‌تست‌اند، نه شواهدی از بازده معاملاتی خالصِ تحقق‌یافته. هزینه‌های تراکنش لحاظ نشده‌اند و خلاصه جزئیاتی درباره استحکام، ساخت سبد یا محدودیت‌های اجرا ارائه نمی‌کند. یافته‌ها نشان می‌دهند رفتار جست‌وجو شاید اطلاعاتی به پیش‌بینی سهام در افق کوتاه بیفزاید، اما سودآوری عملی و پیامدهای آن برای کارایی بازار همچنان نامشخص‌اند.

ایده‌های کلیدی

  • حجم جست‌وجو با داده‌های مالی تاریخی ترکیب می‌شود تا بازده نسبی سهام در روز بعد طبقه‌بندی شود.
  • درخت‌های تصمیم تقویت‌شده با گرادیان برای سهام S&P 100 پیش‌بینی تولید می‌کنند.
  • عملکرد طبقه‌بندی گزارش‌شده در دوره پژوهش از حدس تصادفی بهتر است.
  • تمرین سبدی با ده سهم، بازده بالای پیش از هزینه را گزارش می‌کند، اما هزینه‌های تراکنش را لحاظ نمی‌کند.
  • نتایج پرسش‌هایی درباره کارایی بازار مطرح می‌کنند، اما سود خالصِ قابل‌تحقق را اثبات نمی‌کنند.

برچسب‌ها

متن کامل
# 2205.15853


# Predicting Day-Ahead Stock Returns using Search Engine Query Volumes: An Application of Gradient Boosted Decision Trees to the S&P 100









The internet has changed the way we live, work and take decisions. As it is the major modern resource for research, detailed data on internet usage exhibits vast amounts of behavioral information. This paper aims to answer the question whether this information can be facilitated to predict future returns of stocks on financial capital markets. In an empirical analysis it implements gradient boosted decision trees to learn relationships between abnormal returns of stocks within the S&P 100 index and lagged predictors derived from historical financial data, as well as search term query volumes on the internet search engine Google. Models predict the occurrence of day-ahead stock returns in excess of the index median. On a time frame from 2005 to 2017, all disparate datasets exhibit valuable information. Evaluated models have average areas under the receiver operating characteristic between 54.2% and 56.7%, clearly indicating a classification better than random guessing. Implementing a simple statistical arbitrage strategy, models are used to create daily trading portfolios of ten stocks and result in annual performances of more than 57% before transaction costs. With ensembles of different data sets topping up the performance ranking, the results further question the weak form and semi-strong form efficiency of modern financial capital markets. Even though transaction costs are not included, the approach adds to the existing literature. It gives guidance on how to use and transform data on internet usage behavior for financial and economic modeling and forecasting.

با ذکر منبع و مطابق مجوز اثر، به‌طور کامل نمایش داده می‌شود. مجوز: abstract CC0

این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخه‌ای از اثر منبع نیست.