Reuse Training-Set Scaling Parameters for Live Model Inputs
Summary
The document addresses how to normalize real-time observations for a support vector machine trained with automatic scaling. Its central guidance is to apply the mean and standard deviation estimated from the training data to out-of-sample observations. Recomputing these values from each live batch would change the feature scale from the one the model learned and can lead to incorrect predictions.
If live inputs have shifted substantially from the training distribution, the answer recommends addressing that shift with more representative training data. It does not describe a drift-detection method, retraining schedule, or safeguards for feature changes, so it is practical preprocessing guidance rather than a complete procedure for handling nonstationarity. The same principle applies to any fitted transformation: estimate it on training data, then carry those fixed parameters into validation and deployment.
Key ideas
- Use the training set’s fitted mean and standard deviation to scale live observations.
- Recomputing scaling parameters on out-of-sample data changes the representation learned by the model.
- A substantial distribution shift may call for collecting more representative training data.
- The guidance does not specify how to detect drift or when to retrain.
Tags
Full text
# Normalized data # Normalized data I am new to this. I trained and tested my data using SVM in Matlab with the autoscale option true => the data would be normalized with unit SD. Let's say the training data have the price around 200. During real time, let's say the price is around 300. I am wondering what is the practise in obtaining the mean and standard deviation in real time so that I could normalize the data for the model I built previously? Thanks. ## Answer by Louis Marascio (score 2) https://quant.stackexchange.com/a/3642 You must use the same scaling parameters out of sample as you did during training. Do otherwise will give you incorrect results. If the distribution of your out of sample observations different significantly from your training data then you need to address that by gathering more training data.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.