Multi-Pair Q-Learning with Correlation-Adjusted Trading Actions
Summary
The article describes an MQL5 Expert Advisor that combines Q-learning, cross-pair correlation adjustments, and market features drawn from three timeframes. It represents market observations as an indexed state and scores a set of actions that includes opening positions, adding to them, and closing profitable positions. A Nash-inspired adjustment modifies each pair’s action scores using other pairs’ Q-values when their measured correlation passes a stated threshold. Indicator caching is presented as a way to reduce repeated calculations across symbols and timeframes.
The text also describes an opportunity-cost mechanism that increases the learning rate after a poor decision when an alternative action could have earned a reward. The evidence is primarily implementation sketches and explanations; the article does not provide backtest results, out-of-sample evaluation, or risk-adjusted performance. Its claims about adaptability and profitability therefore remain unverified. Selectively closing only winners can also leave losing positions open, while hashing a large feature set into a fixed state range may merge distinct market conditions.
Key ideas
- The system scores actions using Q-values and adjusts them with information from correlated currency pairs.
- It encodes price, moving-average, momentum, RSI, and stochastic readings across multiple timeframes into a state index.
- Available actions include adding to positions and selectively closing profitable trades.
- Indicator and price caches are used to limit repeated calculations.
- The article explains implementation concepts but supplies no empirical performance validation.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.