Using Random Forests to Classify Stock Price Direction
Summary
The author describes a binary stock-direction classifier built with a random forest. Historical open, high, low, close, and volume data are used to derive technical features, including momentum, oscillator, volume, and moving-average indicators. The target labels whether price rose or fell. The model is trained on one portion of the data and evaluated on a held-out portion; the author reports mean accuracy of 85 percent and asks how to apply the fitted model to future observations.
The excerpt explains the feature and target setup and the distinction between testing predictions and forecasting beyond the available sample, but it does not provide an answer to the deployment question. It also does not describe the forecast horizon, how future feature values would be obtained, or whether the split respects time order. The reported test accuracy alone does not establish out-of-sample trading performance, and the discussion gives no transaction-cost or risk evaluation.
Key ideas
- The model predicts a binary up-or-down stock direction from historical price and volume features.
- The described features include several technical indicators derived from market data.
- The author reports 85 percent mean accuracy on testing data but does not explain the evaluation setup in depth.
- Using a fitted model for future predictions requires feature inputs available at the forecast date.
- The excerpt does not resolve whether the data split or reported accuracy supports realistic future performance.
Tags
Full text
# stock price trend classification using Random Forest in sklearn
# stock price trend classification using Random Forest in sklearn
I have created a random forest classification model in skicit-learn, but I am unsure how to finalise my forecast.
I have built the model and it is showing good results on the testing data. I get a mean accuracy of 85%. Predicting whether the stock price will go up or down. I used data from Yahoo finance consisting of open, high, low, close and volume. From these I worked out some technical indicators such as the RSI, ROC, stochastic oscillators (fast and slow), macd, on balance volume and the 200 day moving average and used these as features (independent variables) in the random forest classifier. I created another column, showing 1 when the price went up and 0 when the price went down. This column was used as the dependent variable. (the thing I want to predict)
The thing I am trying to find out now is how can I run the forecast into the unknown future? For now, I have split my data into training and testing, trained the model on the training dataset, and then used the predict function on the testing dataset. The model performs well and after a little more tweaking it can be used.
But how? I can't seem to find anywhere in the sklearn random forest documentation about how to actually run the forecast for the future (not on the testing data). I hope you understand what I mean. Below is my code.
```
X_train2, X_test2, y_train2, y_test2 =
train_test_split(data2.drop('prediction',axis=1),data2.prediction,test_size=0.02)
from sklearn.ensemble import RandomForestClassifier
model1 = RandomForestClassifier(random_state=13)
model1.fit(X_train2,y_train2)
predicted = model1.predict(X_test2)
model1.score(X_test2, y_test2)
from sklearn.metrics import roc_auc_score
probabilities = model1.predict_proba(X_test)
probabilities
roc_auc_score(y_test2, probabilities[:,1])
from sklearn.metrics import confusion_matrix
confusion_matrix(y_test2, predicted
```Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.