MASA: Combining Reinforcement Learning and Risk Optimization for Portfolios
Summary
The article describes the Multi-Agent Self-Adaptive (MASA) framework for dynamic portfolio management. It assigns return generation to a TD3 reinforcement-learning agent, risk adjustment to a separate optimization-based controller, and market analysis to an observer. The observer estimates trends and supplies auxiliary market features, including a risk boundary and market vector, to inform the other agents. An entropy-based divergence measure encourages diversity in proposed actions.
For its observer, the article outlines an implementation combining piecewise-linear trend representation, attention with relative positional encoding, and an MLP forecast. It reports that the framework was tested on ten years of data from the CSI 300, Dow Jones Industrial Average, and S&P 500, and says it outperformed traditional RL approaches. However, this installment focuses on the method and implementing separate agent modules in MQL5; full system integration and evaluation of that implementation are deferred. The article also says further analysis on more complex datasets and in other domains is needed.
Key ideas
- MASA separates portfolio return optimization, risk adjustment, and market observation into interacting agents.
- The return agent uses TD3, while a separate optimizer revises its portfolio weights to address risk.
- The observer combines trend representation, attention, and an MLP forecast to provide auxiliary market information.
- The framework was evaluated on ten years of data from the CSI 300, DJIA, and S&P 500, with reported gains over traditional RL approaches.
- The MQL5 implementation and its long-term generality still require full-system evaluation and broader testing.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.