Skip to content
All library documents

Support Vector Machines: Margins, Soft Constraints, and Kernels

Article QuantStart

Summary

This guide explains support vector machines as supervised binary classifiers. It builds from a separating hyperplane to the maximal margin classifier, which chooses a boundary with the greatest distance from nearby training points. Because real data often overlap, it introduces the support vector classifier, which permits some classification errors while balancing them against a wider margin. The points that constrain the boundary are support vectors.

The article then describes how kernel functions extend the linear method to nonlinear boundaries. Polynomial kernels represent feature interactions, while radial kernels give nearby training observations more influence than distant ones. The discussion is theoretical and uses examples such as spam classification to illustrate feature spaces; it does not present a trading test or empirical performance results. It notes that SVMs can work in high-dimensional settings and store only a subset of training observations for decisions, but may perform poorly when features greatly outnumber samples. Their classifications also do not directly provide probabilities, and the guide leaves practical implementation for a later article.

Key ideas

  • An SVM classifies observations by locating them on either side of a separating hyperplane.
  • A maximal margin classifier chooses a boundary with the greatest separation from nearby training observations.
  • Soft constraints allow a support vector classifier to handle overlapping classes and imperfect separation.
  • Kernel functions enable nonlinear decision boundaries while using inner products to represent feature similarity.
  • SVM outputs are not inherently probabilistic, and performance can weaken when features greatly outnumber training samples.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.