Using Support Vector Machines to Classify Data with Noisy Labels
Summary
This example introduces support vector machines through a hypothetical animal-classification task. It defines seven numeric features and labels an observation positive when every feature falls within a specified range. Randomly generated observations and their rule-based labels form the training set, which is passed to an SVM tool for training and classification.
The script then corrupts some training examples by replacing their features and labels with random values, trains a second model, and compares both models on newly generated observations labeled by the same rule. This illustrates classification under label and feature noise, with accuracy as the evaluation measure. The evidence is limited to a synthetic dataset whose target is generated by simple thresholds; the text gives no reported accuracy values or details about financial data, out-of-sample market testing, feature scaling, or model selection. The example motivates possible use in market-trend assessment but does not establish trading performance.
Key ideas
- The example trains an SVM to classify synthetic observations from seven input features.
- A threshold rule supplies the target labels for both training and evaluation data.
- One model is trained on clean examples and another on data altered with random errors.
- Accuracy is compared on new synthetic examples labeled by the same rule.
- Results on this constructed task do not demonstrate performance on real market data.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.