MASA Trading Networks: Combining RL Decisions with Risk Control
Summary
This article’s final part assembles the MASA framework from a market observer, a reinforcement learning agent, and a controller agent. The observer forecasts market sequences, the RL agent proposes portfolio actions, and the controller adjusts those actions using the forecasts and its interpretation of trade directions. A reverse-normalization layer and a second gradient input support an alternative training path for the observer. The controller uses a Transformer decoder, whose internal residual connections are presented as carrying information from the RL proposal into the adjusted result.
The document focuses on architecture and implementation in MQL5, describing initialization, data flow, output activation synchronization, and the start of gradient handling. It reports that testing produced a favorable balance trend when average winning trades exceeded twice the average losing trade, but the excerpt omits most experimental details. The conclusion describes results as promising and says further training on representative data and extensive testing across conditions are needed before live use.
Key ideas
- MASA assigns market forecasting, action selection, and risk adjustment to separate interacting agents.
- The controller receives both the RL agent’s proposed actions and the market observer’s forecasts.
- The controller outputs the final action tensor, with its Transformer residual connections intended to preserve information from the RL proposal.
- Reverse normalization and an extra gradient input support a separate training route for the market observer.
- The reported test outcome is preliminary and the document calls for broader training and evaluation before deployment.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.