검색량으로 S&P 100 주식 수익률 예측
기사 arXiv papers · 저자: Christopher Bockel-Rickermann
요약
이 논문은 구글 검색 활동이 다음 날 지수 중앙값을 웃도는 S&P 100 종목을 예측하는 데 도움이 되는지 검증합니다. 시차를 둔 금융 변수와 검색어 검색량을 결합하고 그래디언트 부스팅 의사결정나무를 학습시켜 결과를 분류합니다. 연구 기간은 2005년부터 2017년까지이며, 평균 ROC 면적은 54.2%에서 56.7%까지로 보고되어 우연보다 높은 예측력을 나타냅니다. 보고된 포트폴리오 실험에서는 여러 데이터 출처를 결합한 모델의 순위가 가장 높습니다.
저자들은 단순한 통계적 차익거래 접근법으로 매일 10개 종목 포트폴리오를 구성하고 거래 비용 차감 전 연간 성과가 57%를 넘는다고 보고합니다. 이 수치는 백테스트 결과이지 실현된 순 트레이딩 수익의 증거가 아닙니다. 거래 비용은 제외되어 있으며, 요약에는 견고성, 포트폴리오 구성, 구현 제약의 세부 사항이 없습니다. 검색 행동이 단기 주가 예측에 추가 정보를 제공할 수 있음을 시사하지만, 실제 수익성과 시장 효율성에 대한 함의는 불확실합니다.
핵심 아이디어
- 검색어 검색량을 과거 금융 데이터와 결합해 다음 날 상대 주식 수익률을 분류합니다.
- 그래디언트 부스팅 의사결정나무로 S&P 100 종목의 수익률을 예측합니다.
- 보고된 분류 성능은 연구 기간 동안 무작위 추측을 넘어섰습니다.
- 10개 종목 포트폴리오 실험에서 높은 비용 차감 전 수익률을 보고했지만 거래 비용은 제외했습니다.
- 결과는 시장 효율성에 의문을 제기하지만 실현 가능한 순수익을 입증하지는 않습니다.
태그
전문
# 2205.15853 # Predicting Day-Ahead Stock Returns using Search Engine Query Volumes: An Application of Gradient Boosted Decision Trees to the S&P 100 The internet has changed the way we live, work and take decisions. As it is the major modern resource for research, detailed data on internet usage exhibits vast amounts of behavioral information. This paper aims to answer the question whether this information can be facilitated to predict future returns of stocks on financial capital markets. In an empirical analysis it implements gradient boosted decision trees to learn relationships between abnormal returns of stocks within the S&P 100 index and lagged predictors derived from historical financial data, as well as search term query volumes on the internet search engine Google. Models predict the occurrence of day-ahead stock returns in excess of the index median. On a time frame from 2005 to 2017, all disparate datasets exhibit valuable information. Evaluated models have average areas under the receiver operating characteristic between 54.2% and 56.7%, clearly indicating a classification better than random guessing. Implementing a simple statistical arbitrage strategy, models are used to create daily trading portfolios of ten stocks and result in annual performances of more than 57% before transaction costs. With ensembles of different data sets topping up the performance ranking, the results further question the weak form and semi-strong form efficiency of modern financial capital markets. Even though transaction costs are not included, the approach adds to the existing literature. It gives guidance on how to use and transform data on internet usage behavior for financial and economic modeling and forecasting.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.