Fractional Differentiation for Financial Machine Learning Features
Summary
The document asks how to apply a feature transformation from financial machine learning: cumulatively sum a series, test fractionally differentiated versions over a range of differentiation orders, and choose the least order that passes an augmented Dickey–Fuller stationarity test at a stated significance threshold. The resulting transformed series is then used as a predictive feature. The question is whether this procedure applies to every feature and whether transformed features should be scaled before model fitting.
The excerpt presents the procedure and the practitioner’s uncertainties, but contains no answers, experiments, or model comparisons. It therefore motivates two feature-engineering decisions without establishing a universal rule. The appropriate differentiation order and need for scaling can depend on the feature’s properties and the model; the passage itself does not discuss those contingencies or how to avoid using future information when selecting transformations. It is a useful prompt about stationarity and preprocessing, rather than a validated prescription for all financial inputs.
Key ideas
- The proposed procedure selects a fractional differentiation order using an augmented Dickey–Fuller stationarity test.
- The transformed series is intended to retain predictive information while improving stationarity.
- The document asks whether this procedure should be applied to every feature.
- It also raises feature scaling as a modeling choice but provides no answer or empirical comparison.
Tags
Full text
# Do we need to fractionally differentiated all features in ML prediction for finance time series? # Do we need to fractionally differentiated all features in ML prediction for finance time series? I am reading Prof. Marcos Lopez de Prado's book Advances in Financial Machine Learning, and have a question on feature engineering. On page 88, he says: In practice, I suggest you experiment with the following transformation of your features: First, compute a cumulative sum of the time series. This guarantees that some order of differentiation is needed. Second, compute the FFD(d) series for various d ∈ [0, 1]. Third, determine the minimum d such that the p-value of the ADF statistic on FFD(d) falls below 5%. Fourth, use the FFD(d) series as your predictive feature. I wonder: - Should we scale the result predictive features when fitting a model? The book seems didn't discuss the topic about scaling, is it needless to scale? - Should we follow this procedure for all features?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.