Skip to content
All library documents

Using TRIX and Williams Percent Range Patterns with a DQN

Article MQL5 articles

Summary

This article explains how a Deep Q Network can evaluate three trading patterns built from TRIX, a smoothed momentum indicator, and Williams Percent Range (WPR), an overbought and oversold measure. It outlines the value based reinforcement learning setup: indicator features form the state, long, short, and neutral are possible actions, and rewards reflect trading outcomes. Replay buffers, a target network, and exploratory action selection are described as parts of DQN training. The article also compares value based learning with policy and actor critic approaches and discusses exporting a trained quantile regression network for use in MetaTrader.

The reported forward walk covers a year after training and optimization on earlier data. Of the three tested patterns, only pattern 1 was profitable in that period; patterns 4 and 5 struggled. The author cautions that the networks may have been weak or insufficiently adapted, and that deployed ONNX models run as fixed inference engines rather than learning online. The results are therefore exploratory and do not establish that the approach will generalize or be profitable in other markets or periods.

Key ideas

  • A DQN estimates long term action values for indicator states instead of relying only on fixed thresholds.
  • TRIX and WPR features are used to evaluate three patterns with long, short, and neutral actions.
  • Replay sampling and a target network are described as ways to stabilize DQN training.
  • Only pattern 1 was profitable in the reported forward walk, while patterns 4 and 5 struggled.
  • The exported model operates as an inference engine and cannot update its weights during deployment.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.