Skip to content
All library documents

Support Vector Machines for Stock Selection: Kernels, Tuning, and Backtesting

Article SuperMind

Summary

This tutorial explains how support vector machines classify data by finding a boundary with a wide margin, and how slack variables allow some classification errors in noisy data. It introduces kernel methods as a way to handle nonlinear boundaries by computing relationships in a higher-dimensional feature space, and discusses linear, polynomial, sigmoid, and Gaussian or RBF kernels. The article also describes how the penalty parameter C and kernel parameter gamma affect fit and overfitting, while noting the trade-off between model flexibility and computational cost.

Its stock-selection example uses A-share factor features, labels the top 5% of future five-day returns as positive, preprocesses missing and extreme values, and compares kernel choices in a weekly equal-weight portfolio backtest. The article reports that linear and RBF kernels outperformed polynomial and sigmoid variants in that experiment, with RBF described as more stable and having the smallest maximum drawdown at 8.7%. These findings are specific to the example’s data and setup; the page says it concerns an older platform version and notes that feature selection, labels, portfolio rules, and SVM parameters remain open to improvement.

Key ideas

  • Linear SVMs seek a wide-margin decision boundary, while slack variables permit some violations when data are noisy.
  • Kernel functions let SVMs model nonlinear classification boundaries without explicitly computing all higher-dimensional features.
  • The choice of kernel and the values of C and gamma affect model complexity, fit, and overfitting risk.
  • The stock-selection example labels the highest-returning 5% of stocks over the following five days as positive.
  • In the reported comparison, linear and RBF kernels performed better than polynomial and sigmoid kernels, but the result is limited to that backtest.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.