Monitoring StockRanker Validation NDCG to Tune Tree Count
Summary
This guide explains how to monitor a StockRanker model’s validation performance during training. It recommends preparing a validation dataset, either from a separate data module or by reusing the training data, and connecting it to the model’s validation input. The module calculates NDCG, a ranking metric, even though the newer interface does not plot it directly; the guide shows how to read and chart the stored metric.
The example curve indicates that performance begins to converge around the eighth tree. The guide suggests considering a smaller tree count, such as eight or ten, to reduce runtime and potentially improve generalization. This is an example rather than a universal rule: the suitable count depends on dataset size and model complexity, including leaf count. Reusing training data for validation also does not provide an independent assessment of generalization, so a separate validation set is preferable when available. The document gives no broader benchmark or evidence that the example setting works across datasets.
Key ideas
- Connect a validation dataset to StockRanker’s validation input to calculate validation NDCG.
- The newer module calculates NDCG but may require a separate charting step to inspect its progression.
- The example curve suggests considering fewer trees after the metric begins to converge.
- Choose tree count based on data size and model complexity rather than treating the example as a fixed rule.
- Using training data as validation does not provide an independent generalization check.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.