Adam-mini for Lower-Memory Neural Network Training
Summary
The article explains Adam-mini, an Adam optimizer variant designed to reduce the memory used for per-parameter second-moment estimates. It groups model parameters into blocks and tracks a moving average of squared gradients for each block, using the block mean to set learning rates. The proposed grouping treats attention query and key parameters by head, other parameters by layer, and embedding layers separately with classic Adam. The article also describes an MQL5 and OpenCL implementation for neural network layers, including a fully connected layer approach that factors output gradients from averages of squared inputs to reduce repeated global-memory access.
The cited research reports comparable or better training performance than Adam in experiments on a small Transformer, while the article says reduced memory can allow larger GPU batches and less CPU transfer. These benefits depend on model composition and available hardware. The block average is an approximation, the memory reduction varies with the share of non-embedding parameters, and the article provides limited detail about its own test results. Its focus is neural network training implementation, rather than a trading strategy or market forecast.
Key ideas
- Adam-mini reduces optimizer state by sharing second-moment estimates across parameter blocks.
- The method uses moving averages of block-level mean squared gradients to guide parameter updates.
- The article suggests grouping Transformer query and key parameters by attention head and other parameters by layer.
- Embedding layers are kept on classic Adam because their gradient distributions differ.
- Lower optimizer memory may support larger batches and reduce CPU-to-GPU data movement, though savings vary by model.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.