A Reference Map of Machine Learning and Data Mining Methods
Summary
This document is a broad glossary of machine learning and data mining terms, organized by task. It surveys sampling methods, clustering, classification and regression, model evaluation, probabilistic graphical models, neural networks, deep learning, dimensionality reduction, text mining, association discovery, recommendation, similarity measures, feature selection, outlier detection, and learning to rank. Examples range from common statistical models and tree ensembles to specialized algorithms such as MCMC, DBSCAN, conditional random fields, and pairwise ranking methods.
The material is an inventory rather than a tutorial: it names methods and gives expansions or brief translations, but does not explain their assumptions, implementation, comparative performance, or suitability for trading. It provides no empirical results or worked examples. Some abbreviations recur with different meanings across fields, and the list includes methods that are related but not interchangeable. Readers can use it to orient themselves in the vocabulary, but need additional sources to understand model selection, validation, data leakage, and practical application to financial data.
Key ideas
- The glossary groups algorithms by analytical task, including clustering, prediction, evaluation, and dimensionality reduction.
- It covers both conventional statistical learning and neural or deep learning approaches.
- Text mining, recommendation, association mining, and learning to rank are included alongside general-purpose methods.
- Distance measures and feature selection methods are listed as tools for representing and preparing data.
- The entries provide terminology but little guidance on assumptions, implementation, or performance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.