Skip to content
All library documents

Continuous-Action TD3 for Stock and Bitcoin Trading

Article arXiv papers · Author: Naseh Majidi et al.

Summary

This study applies Twin-Delayed Deep Deterministic Policy Gradient (TD3), a continuous-action reinforcement learning method, to daily closing prices for Amazon stock and Bitcoin. The trading action represents both the position to take and the number of shares to trade, extending approaches that use discrete actions. The paper compares the resulting strategies with technical-analysis methods, other reinforcement learning approaches, and stochastic and deterministic strategies.

Performance is assessed using return and the Sharpe ratio. The reported results suggest that selecting both position and trade size can improve performance on those measures. The document does not provide the evaluation details, specific results, or evidence about robustness across other assets or market periods. Its conclusions therefore describe the markets and comparison methods studied, rather than establishing that continuous-action TD3 will perform better in other settings.

Key ideas

  • TD3 is used to generate trading actions from daily closing prices for Amazon and Bitcoin.
  • The action space represents both the trading position and the number of shares traded.
  • The study compares TD3 with technical-analysis, reinforcement learning, stochastic, and deterministic strategies.
  • Return and Sharpe ratio are the stated performance measures.
  • The reported comparison favors combining position selection with trade sizing, within the study's tested settings.

Tags

Full text
# Algorithmic Trading Using Continuous Action Space Deep Reinforcement Learning


# Algorithmic Trading Using Continuous Action Space Deep Reinforcement Learning









Price movement prediction has always been one of the traders' concerns in financial market trading. In order to increase their profit, they can analyze the historical data and predict the price movement. The large size of the data and complex relations between them lead us to use algorithmic trading and artificial intelligence. This paper aims to offer an approach using Twin-Delayed DDPG (TD3) and the daily close price in order to achieve a trading strategy in the stock and cryptocurrency markets. Unlike previous studies using a discrete action space reinforcement learning algorithm, the TD3 is continuous, offering both position and the number of trading shares. Both the stock (Amazon) and cryptocurrency (Bitcoin) markets are addressed in this research to evaluate the performance of the proposed algorithm. The achieved strategy using the TD3 is compared with some algorithms using technical analysis, reinforcement learning, stochastic, and deterministic strategies through two standard metrics, Return and Sharpe ratio. The results indicate that employing both position and the number of trading shares can improve the performance of a trading system based on the mentioned metrics.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.