Skip to content
All library documents

Using Q-Learning and Markov Chains to Guide an MLP Trading Model

Article MQL5 articles

Summary

The article describes adding reinforcement learning to a multilayer perceptron (MLP) in an MQL5 trading system. Q-learning is used to guide model training rather than act as the direct signal generator. The market environment is represented by combinations of short- and longer-horizon bullish, bearish, or flat conditions. For each state, a Q-map scores the choices to buy, sell, or remain inactive; a critic converts the resulting price change into a normalized reward or penalty. A Markov transition-probability matrix is added as a supplementary weight to the learning process.

The article compares strategy-tester runs with and without Markov chains. It cautions that the settings were not thoroughly optimized and that the tests were not walk-forward evaluations, so they do not establish that the Markov component improves results. The author also describes the implementation as complex and parameter-sensitive. The method is therefore an experimental design for integrating Q-learning, state transitions, and an MLP, rather than validated evidence of a robust trading strategy.

Key ideas

  • Q-learning is used to guide an MLP’s training rather than serve as the system’s raw trading signal.
  • The market state is encoded from short- and longer-horizon bullish, bearish, or flat conditions.
  • A Q-map assigns action scores to buying, selling, or taking no action in each state.
  • A critic maps price changes after actions to normalized rewards or penalties.
  • A Markov transition matrix is tested as an additional weight, but the reported tests lack optimization and walk-forward validation.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.