Diagnosing Lagged Neural Network Forecasts in Financial Time Series
Summary
The document discusses a neural network trained on exogenous variables to forecast Bitcoin prices ahead and examines why its predictions appear delayed despite having no past-price inputs. The answer suggests that the model may effectively predict the latest price with noise, which can make forecasts trail a moving target. It recommends predicting relative returns instead of raw price and first checking whether simpler linear or logistic models extract useful signal from the features.
The response also emphasizes time-aware validation: split observations into non-overlapping chronological segments with gaps to reduce leakage, and avoid shuffling time-series examples. Feature selection is proposed because irrelevant inputs can impede learning. These are diagnostic suggestions, not a demonstrated explanation of the plotted lag or a validated forecasting recipe. The source offers no independent performance evidence, and its claim about shuffling is presented generally rather than tied to details of the model’s evaluation setup.
Key ideas
- Predicting returns can be more informative than forecasting raw price levels.
- Compare neural networks with simpler models before adding complexity.
- Use chronological validation segments with gaps to limit information leakage.
- Avoid shuffling observations when evaluating a time-series forecasting model.
- Feature selection may help when inputs contain irrelevant variables.
Tags
Full text
# Why are my Neural Network predictions “correct”, but offset from true value? Not using any past lagged values
# Why are my Neural Network predictions “correct”, but offset from true value? Not using any past lagged values
Please bear with me through the whole question - I just want to make it very clear what I've done so far and why I'm so perplexed.
I am working with a neural network with the Keras package in R, trying to predict hourly Bitcoin price, 24 hours ahead. Here is the code for my model:
```
batch_size = 2
model <- keras_model_sequential()
model%>%
layer_dense(units=13,
batch_input_shape = c(batch_size, 1, 13), use_bias = TRUE) %>%
layer_dense(units=17, batch_input_shape = c(batch_size, 1, 13)) %>%
layer_dense(units=1)
model %>% compile(
loss = 'mean_squared_error',
optimizer = optimizer_adam(lr= 0.000025, decay = 0.0000015),
metrics = c('mean_squared_error')
)
summary(model)
Epochs <- 25
for (i in 1:Epochs){
print(i)
model %>% fit(x_train, y_train, epochs=1, batch_size=batch_size, verbose=1, shuffle=TRUE)
#model %>% reset_states()
}
```
You may notice that I am working on a time-series problem, but not using LSTM. This is because none of my inputs are time-series values. They are all exogenous variables. You'll also notice that I commented out the line "model %>% reset_states()". I'm not sure if that's the right thing to do here, but from what I read, that line is for LSTM models and since I'm not using one anymore, I commented it out.
Again, because I have no time-series inputs, I also set "shuffle=" to TRUE. So, below are the predictions in blue vs. true value in red:
You can see that the predictions very often (but not always) lag behind the true value. Additionally, this lag is not constant. Let me again emphasize that all of the variables are exogenous. There is not past-price input that the model could be using in order to generate these late predictions. AND don't forget I set shuffle= TRUE, which confuses me even further as to how the model could be giving such results if there's no way (that I know of) that it can "see" past values in order to replicate them. Here is the graph of the training data fit:
It's harder to tell, but the lag exists in the training data as well. I'll also say that if I change how far ahead I'm trying to predict, the apparent prediction lag changes too. If I try to predict 0 hours ahead (so I'm "predicting" current price given current conditions), there is no lag in prediction.
I checked to make sure the tables/columns were set up right so that the model is trained on "current" conditions predicting price 24 hours ahead. I've also played around with the network architecture and batch size. The only thing that seems to affect this lag is how far ahead I'm trying to predict - that's to say, how many rows I shifted my Bitcoin Price column by so that "future" price rows are matched with past predictor rows.
Something else weird I've noticed is that with this non-LSTM model, it has a rather high error when training (MSE=0.07) which it reaches after only 5-9 Epochs and then doesn't go any lower. I don't think this is relevant because the LSTM model I used before achieved MSE=0.005 and still had the same lag issue, but I figured I'd mention it.
Any advice, tips or links would be enormously appreciated. I can't for the life of me figure out what's going on.
## Answer by alexprice (score 3)
https://quant.stackexchange.com/a/55352
Some tips:
- Do not predict price directly, set your target variable to be (relative) return. as you mentioned your model trains really fast (just 5-9 epochs) and then error does not get lower. I suspect that your neural network just predicts last price + some noise. (you can also see it from your graphs as prediction just lags the true value)
- Test your features first on simpler model such as linear regression (or logistic, if you try to predict price direction). Compare results with Neural Network.
- run cross validation (split your time series into non overlapping chunks, with some gap in between to insure you do not leak information from the future)
- run proper feature selection.(univariate and backward feature selection, for example) Complex models such as neural networks with many hidden layers depend heavily on input feature quality.If you are inputting a lot of non-relevant features, often neural network will not be able to learn.
- Set shuffle to false. It is time-series prediction model, shuffling is just leaking future information, making the model over optimistic (i.e. in live set-up prediction of your model will be much worse than back-tested)Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.