استخدام أحجام البحث للتنبؤ بعوائد أسهم S&P 100
الملخص
تختبر الورقة ما إذا كان نشاط البحث في Google يساعد في التنبؤ بأسهم S&P 100 التي ستتجاوز وسيط المؤشر في اليوم التالي. وتجمع بين متغيرات مالية متأخرة وأحجام استعلامات البحث، وتدرّب أشجار القرار المعززة بالتدرج لتصنيف تلك النتائج. تغطي الدراسة الفترة من 2005 إلى 2017، وتورد متوسط مساحات ROC بين 54.2% و56.7%، بما يشير إلى قدرة تنبؤية أعلى من المصادفة. وجاء ترتيب النماذج التي تجمع مصادر البيانات في الصدارة ضمن تجربة المحفظة المذكورة.
كما كوّن المؤلفون محافظ يومية من عشرة أسهم باستخدام نهج بسيط للمراجحة الإحصائية، وأفادوا بأداء سنوي يزيد على 57% قبل تكاليف المعاملات. هذه الأرقام ناتجة عن اختبار تاريخي، وليست دليلًا على عوائد تداول محققة وصافية. وتُستثنى تكاليف المعاملات، ولا يقدم الملخص تفاصيل عن المتانة أو بناء المحفظة أو قيود التنفيذ. وتشير النتائج إلى أن سلوك البحث قد يضيف معلومات للتنبؤ بالأسهم على أفق قصير، مع بقاء الربحية العملية والآثار على كفاءة السوق غير مؤكدة.
الأفكار الرئيسية
- تُدمج أحجام استعلامات البحث مع البيانات المالية التاريخية لتصنيف العوائد النسبية للأسهم في اليوم التالي.
- تنتج أشجار القرار المعززة بالتدرج توقعات لأسهم S&P 100.
- يتجاوز أداء التصنيف المعلن التخمين العشوائي خلال فترة الدراسة.
- تعرض تجربة محفظة من عشرة أسهم عوائد مرتفعة قبل التكاليف، لكنها تستبعد تكاليف المعاملات.
- تثير النتائج تساؤلات حول كفاءة السوق من دون إثبات أرباح صافية قابلة للتحقق.
الوسوم
النص الكامل
# 2205.15853 # Predicting Day-Ahead Stock Returns using Search Engine Query Volumes: An Application of Gradient Boosted Decision Trees to the S&P 100 The internet has changed the way we live, work and take decisions. As it is the major modern resource for research, detailed data on internet usage exhibits vast amounts of behavioral information. This paper aims to answer the question whether this information can be facilitated to predict future returns of stocks on financial capital markets. In an empirical analysis it implements gradient boosted decision trees to learn relationships between abnormal returns of stocks within the S&P 100 index and lagged predictors derived from historical financial data, as well as search term query volumes on the internet search engine Google. Models predict the occurrence of day-ahead stock returns in excess of the index median. On a time frame from 2005 to 2017, all disparate datasets exhibit valuable information. Evaluated models have average areas under the receiver operating characteristic between 54.2% and 56.7%, clearly indicating a classification better than random guessing. Implementing a simple statistical arbitrage strategy, models are used to create daily trading portfolios of ten stocks and result in annual performances of more than 57% before transaction costs. With ensembles of different data sets topping up the performance ranking, the results further question the weak form and semi-strong form efficiency of modern financial capital markets. Even though transaction costs are not included, the approach adds to the existing literature. It gives guidance on how to use and transform data on internet usage behavior for financial and economic modeling and forecasting.
يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0
أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.