Skip to content
All library documents

Reinforcement Learning for Cryptocurrency Portfolio Management

Article arXiv papers · Author: Kamal Paykan

Summary

This paper proposes using Soft Actor-Critic (SAC) and Deep Deterministic Policy Gradient (DDPG) to manage cryptocurrency portfolios. The agents learn continuous trading actions in a simulated environment using historical market data, adjusting portfolio weights to seek cumulative returns while accounting for downside risk and transaction costs. SAC uses an entropy-regularized objective, which the paper associates with greater stability in noisy conditions.

Experiments across multiple cryptocurrencies report that both agents outperform equal-weight and mean-variance portfolio baselines, with SAC more stable and robust than DDPG. The summary does not specify the assets, sample period, market assumptions, transaction-cost model, or numerical results. As the evidence comes from experimental evaluation using historical data and a simulated environment, it does not by itself show that the methods will perform similarly in live trading or under different market conditions.

Key ideas

  • SAC and DDPG agents learn continuous portfolio actions from historical market data in a simulated environment.
  • The agents adjust portfolio weights while considering returns, downside risk, and transaction costs.
  • The paper reports better performance than equal-weight and mean-variance baselines across multiple cryptocurrencies.
  • SAC is reported to be more stable than DDPG in noisy market conditions.
  • The summary does not provide the experimental setup or numerical performance details.

Tags

Full text
# 2511.20678


# Cryptocurrency Portfolio Management with Reinforcement Learning: Soft Actor--Critic and Deep Deterministic Policy Gradient Algorithms









This paper proposes a reinforcement learning--based framework for cryptocurrency portfolio management using the Soft Actor--Critic (SAC) and Deep Deterministic Policy Gradient (DDPG) algorithms. Traditional portfolio optimization methods often struggle to adapt to the highly volatile and nonlinear dynamics of cryptocurrency markets. To address this, we design an agent that learns continuous trading actions directly from historical market data through interaction with a simulated trading environment. The agent optimizes portfolio weights to maximize cumulative returns while minimizing downside risk and transaction costs. Experimental evaluations on multiple cryptocurrencies demonstrate that the SAC and DDPG agents outperform baseline strategies such as equal-weighted and mean--variance portfolios. The SAC algorithm, with its entropy-regularized objective, shows greater stability and robustness in noisy market conditions compared to DDPG. These results highlight the potential of deep reinforcement learning for adaptive and data-driven portfolio management in cryptocurrency markets.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.