עבור לתוכן
כל מסמכי הספרייה

שימוש בנפחי חיפוש לחיזוי תשואות מניות S&P 100

מאמר arXiv papers · מחבר: Christopher Bockel-Rickermann

סיכום

המאמר בוחן אם פעילות החיפוש בגוגל יכולה לסייע בחיזוי אילו מניות ב־S&P 100 יניבו ביצועים טובים מחציון המדד ביום הבא. הוא משלב משתנים פיננסיים בפיגור עם נפחי שאילתות חיפוש ומאמן עצי החלטה מוגברים בגרדיאנט כדי לסווג את התוצאות האלה. המחקר מכסה את התקופה מ־2005 עד 2017 ומדווח על שטחים ממוצעים מתחת לעקומת ROC בטווח שבין 54.2% ל־56.7%, המעידים על יכולת ניבוי העולה על אקראיות. המודלים המשלבים מקורות נתונים שונים דורגו במקום הגבוה ביותר בתרגיל התיקים המדווח.

החוקרים בונים גם תיקים יומיים של עשר מניות בגישת ארביטראז׳ סטטיסטי פשוטה ומדווחים על ביצועים שנתיים מעל 57% לפני עלויות עסקה. אלה תוצאות בקטסט, ולא ראיה לתשואות מסחר ממומשות נטו. עלויות העסקה אינן נכללות, והסיכום אינו מפרט עמידות, בניית תיק או מגבלות יישום. הממצאים מרמזים שהתנהגות חיפוש עשויה להוסיף מידע לחיזוי מניות בטווח קצר, אך הרווחיות המעשית וההשלכות על יעילות השוק נותרות לא ודאיות.

רעיונות מרכזיים

  • נפחי שאילתות חיפוש משולבים בנתונים פיננסיים היסטוריים לסיווג תשואות יחסיות של מניות ביום הבא.
  • עצי החלטה מוגברים בגרדיאנט מפיקים תחזיות למניות S&P 100.
  • ביצועי הסיווג המדווחים עולים על ניחוש אקראי בתקופת המחקר.
  • תרגיל תיק של עשר מניות מדווח על תשואות גבוהות לפני עלויות, אך אינו כולל עלויות עסקה.
  • התוצאות מעלות שאלות על יעילות השוק, אך אינן מוכיחות רווחים נטו שניתן לממש.

תגיות

הטקסט המלא
# 2205.15853


# Predicting Day-Ahead Stock Returns using Search Engine Query Volumes: An Application of Gradient Boosted Decision Trees to the S&P 100









The internet has changed the way we live, work and take decisions. As it is the major modern resource for research, detailed data on internet usage exhibits vast amounts of behavioral information. This paper aims to answer the question whether this information can be facilitated to predict future returns of stocks on financial capital markets. In an empirical analysis it implements gradient boosted decision trees to learn relationships between abnormal returns of stocks within the S&P 100 index and lagged predictors derived from historical financial data, as well as search term query volumes on the internet search engine Google. Models predict the occurrence of day-ahead stock returns in excess of the index median. On a time frame from 2005 to 2017, all disparate datasets exhibit valuable information. Evaluated models have average areas under the receiver operating characteristic between 54.2% and 56.7%, clearly indicating a classification better than random guessing. Implementing a simple statistical arbitrage strategy, models are used to create daily trading portfolios of ten stocks and result in annual performances of more than 57% before transaction costs. With ensembles of different data sets topping up the performance ranking, the results further question the weak form and semi-strong form efficiency of modern financial capital markets. Even though transaction costs are not included, the approach adds to the existing literature. It gives guidance on how to use and transform data on internet usage behavior for financial and economic modeling and forecasting.

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.