Skip to content
All library documents

How Tree Count and Feature Count Affect Random Forest Coverage

Article Quant Q&A · Author: Jacques Joubert

Summary

The document discusses how the number of trees and the number of predictors may affect random forest performance. Because forests use bootstrap samples of observations and random subsets of features, too few trees may leave some observations with little or no representation across the ensemble. With many predictors and few trees, some features could also be omitted from the sampled subsets, though the answer notes that feature selection occurs at each node, making complete omission less likely.

To explore the relationship, the author describes a synthetic classification experiment with informative and irrelevant features, varying the tree count and the maximum features available to each tree, then measuring test-set accuracy in a heat map. The document ends before presenting the heat map or its results, so it provides no empirical conclusion about an optimal tree-to-feature ratio. Its discussion is a limited synthetic example, not evidence that directly establishes performance in financial data.

Key ideas

  • Random forests use samples of observations and random subsets of features when building trees.
  • With too few trees, some observations may receive little or no coverage across the ensemble.
  • A small forest may also fail to sample some predictors, although features are selected repeatedly at individual nodes.
  • The author describes a synthetic experiment that varies tree count and maximum features and evaluates test accuracy.
  • The reported material omits the heat map and findings, so it does not establish an optimal relationship.

Tags

Full text
# Random Forests - Trees vs Predictors


# Random Forests - Trees vs Predictors












This question relates to the use of random forests in finance and the relationship between the number of features, the observations, and the number of trees.

Consider the relation between an RF, the number of trees it is composed of, and the number of features utilized:

- Could you envision a relation between the minimum number of trees needed in an RF and the number of features utilized?

- Could the number of trees be too small for the number of features used?

- Could the number of trees be too high for the number of observations available?

Question sourced from (AFML de Prado 2018)

## Answer by Jacques Joubert (score 1, accepted)

https://quant.stackexchange.com/a/46669

The following post on cross-validated has quite a good answer:

"Random forest uses bagging (picking a sample of observations rather than all of them) and random subspace method (picking a sample of features rather than all of them, in other words - attribute bagging) to grow a tree. If the number of observations is large, but the number of trees is too small, then some observations will be predicted only once or even not at all. If the number of predictors is large but the number of trees is too small, then some features can (theoretically) be missed in all subspaces used. Both cases results in the decrease of random forest predictive power. But the last is a rather extreme case, since the selection of subspace is performed at each node."

I ran some empirical tests to validate this. First I created synthetic data using the following:

```
# Create data
X, y = make_classification(n_samples=20000, n_features=50,
                            n_informative=10, n_redundant=0,
                            random_state=42, shuffle=True, n_classes=2, class_sep=1.0)

# Split data
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42, shuffle=True, stratify=None)
```

There are 50 features, with only 10 of them being informative. There is a clear class separation and 20000 observations.

Next, I fit a random forest which is limited by the number of trees, and the number of features it may use (n_estimators, max_features).

```
max_trees = 100
max_feat_used = 50

store = []
for num_trees in range(2, max_trees, 2):
    print(num_trees)
    for num_feat in range(1, max_feat_used, 2):
        rnd_clf = RandomForestClassifier(criterion='entropy', n_estimators=num_trees, max_features=num_feat, n_jobs=-1)
        rnd_clf.fit(X_train, y_train) 
        y_pred_rf = rnd_clf.predict(X_test)
        
        store.append([num_trees, num_feat, accuracy_score(y_test, y_pred_rf)])

# Pivot and save results
results = pd.DataFrame(store, columns=['N', 'F', 'Score'])
pivot_results = results.pivot(index='N', columns='F', values='Score')
pivot_results = pivot_results.sort_index(ascending=False)
```

Finally, we can observe the relationship in a heat map:

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.