Implementing a Segmented-Attention PSformer Encoder for Trading Policies
Summary
This article describes a trading-oriented implementation of the PSformer time-series architecture. It reviews parameter-sharing blocks and spatial-temporal segmented attention, then explains how the encoder handles sequential market data. Because each bar already contains multiple features, the author reduces the original patching and transformation step to a tensor transposition. The design uses three parameter-sharing blocks and two relative-attention modules, with SAM optimization incorporated into the implementation.
The encoder is trained alongside an Actor policy using normalized inputs, without the output mapping and reverse normalization components needed for a separate forecasting stage. The article cites the source study’s comparison of encoder depths on hourly and minute-level electrical-transformer datasets to motivate explicit, relatively shallow encoder layers. It also reports promising out-of-sample historical testing, but the excerpt provides no detailed performance figures or evidence of live profitability. The described architecture is the author’s adaptation of PSformer, so its departures from the original design and the limited reported trading evidence matter when assessing its results.
Key ideas
- The implementation reduces patch formation to transposing the sequential bar-data tensor.
- The encoder combines parameter-sharing blocks with relative-attention modules and SAM optimization.
- The author trains the encoder jointly with an Actor policy using normalized inputs.
- The source study’s results inform a design favoring a small, explicitly specified number of encoder layers.
- Reported historical tests are described as promising, but the excerpt does not establish live trading performance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.