Skip to content
All library documents

Using K-Nearest Neighbors for Multi-Factor Stock Selection

Article SuperMind

Summary

The article introduces K-nearest neighbors (KNN) classification using Euclidean distance: a sample is assigned the class most common among its nearest training examples. It then outlines a stock-selection application in which financial and technical factors are standardized, labeled by subsequent returns, and used to estimate which stocks are more likely to meet a return threshold. The proposed features include valuation, market capitalization, OBV, Bollinger Bands, KDJ, and a base-building indicator.

In its example, the universe is the CSI 300. Training labels distinguish stocks whose returns over the following 20 days exceed a 2% threshold from those that do not. Each day, the method ranks stocks by predicted probability of the positive class, selects three, sells holdings outside that group, and equal-weights the selected stocks. The document gives a backtest period and initial capital, but no resulting return, risk, or benchmark statistics. It also leaves important design details unclear, including how training and test samples are separated and how feature scaling avoids using future information.

Key ideas

  • KNN classifies observations according to the labels of their nearest neighbors in feature space.
  • The proposed stock screen standardizes seven technical and fundamental factors before classification.
  • Training labels use subsequent 20-day returns and a 2% threshold to define the positive class.
  • The portfolio selects three stocks with the highest predicted positive-class probabilities and equal-weights them.
  • The document gives backtest setup details but does not report performance statistics or explain safeguards against look-ahead bias.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.