Skip to content
All library documents

Ten Introductory Machine Learning Algorithms and When to Try Them

Article BigQuant

Summary

This overview introduces ten supervised learning approaches: linear and logistic regression, linear discriminant analysis, decision trees, naive Bayes, k-nearest neighbors, learning vector quantization, support vector machines, bagging and random forests, and boosting with AdaBoost. It sketches how each method represents or learns predictions, covering regression, classification, distance-based prediction, margins, and ensemble aggregation.

The central guidance is that no single algorithm is best for every dataset. Practitioners should match the method to the task and compare candidates on held-out test data. The article notes practical considerations such as correlated or noisy inputs, distribution assumptions, feature scaling, memory use, high dimensionality, and outliers. It is a broad beginner survey rather than a rigorous comparison: it offers no benchmark results, detailed tuning procedures, or trading-specific applications, and its simplified descriptions do not cover all variants or assumptions.

Key ideas

  • Supervised algorithms learn a mapping from input features to outcomes, but performance depends on the task and data.
  • A held-out test set helps compare candidate algorithms without assuming one method is universally best.
  • Linear models, probabilistic classifiers, trees, distance methods, and margin-based methods make predictions in different ways.
  • Bagging and random forests aggregate models, while boosting builds models sequentially to address earlier errors.
  • Preprocessing needs vary, including feature scaling, removal of irrelevant or redundant inputs, and attention to outliers or high dimensionality.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.