Skip to content
All library documents

Applying Deep Reinforcement Learning to Trading Decisions

Article Quant Q&A · Author: GoFaster

Summary

The document asks whether an agent trained to play Atari games from screen images and scores could learn to trade by observing market data and choosing actions. It introduces deep Q-networks as an example of reinforcement learning, where the agent improves through repeated interaction and reward feedback, and points to a deep reinforcement learning framework for financial portfolio management as a related approach. A cryptocurrency portfolio example is also mentioned as an application of that framework after a market crash.

The discussion highlights a central difficulty: unlike a game, market behavior can change in response to trading, so historical or simulated experiences may not be repeatable under identical conditions. This complicates the training process and limits direct comparisons with video-game learning. The document presents no performance results or implementation details for the cited trading examples, so it motivates experimentation rather than establishing that reinforcement learning can reliably predict markets or produce profitable trades.

Key ideas

  • Deep Q-networks can learn actions from observations and reward signals through repeated training.
  • A market trading agent could frame observations as states, trades as actions, and financial outcomes as rewards.
  • Market behavior may change in response to trading, making repeated experiments less reliable than in a fixed game.
  • A cited research framework applies deep reinforcement learning to financial portfolio management.
  • The document offers no evidence that the proposed approach consistently predicts markets or earns profits.

Tags

Full text
# Predict the financial markets in the fashion of a video game?


# Predict the financial markets in the fashion of a video game?












DeepMind have demonstrated amazing capabilities of a reinforcement machine learning agent to competently play Atari video games. It is most astounding that that during training nothing more than the image frames of the game and the score were provided to the `deep Q-network` (DQN). The agent learned appropriate actions to accurately play a game and operate competently without any specific adaptions to the source code or network hyper-parameters. It simply needs training on a large number of game sequences to learn a new game.

Could this technology feasibly be adapted to permit a machine learning agent to take competent & appropriate trading actions in the financial markets? Just like playing Pong, but with the markets? A high score would be quite agreeable!

Does anyone have experience to articulate or advice on how this could be practically experimented upon?

## Answer by M. Jeunesse (score 4)

https://quant.stackexchange.com/a/28215

This is an interesting question.

I would reformulate a little bit your question and try an attempt of answer of why using neural networks is not a good idea for predicting market direction.

IMHO, one main reason would be that it is not possible to experiment a strategy without modifying the market behavior and thus it is impossible to repeat the same experiment again and again, which is a prerequisite of training.

To make an analogy with the video-game machine learning, it is as if when you start the level again, the traps where you have lost have been replaced by new ones you never thought about it before.

I encourage other quant.stackexchange members to share their view on the subject.

## Answer by sonaam1234 (score 3)

https://quant.stackexchange.com/a/42633

Z. Jiang, D. Xu, J. Liang, in A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem. demonstrate a Deep RL framework for Trading. The approach is based on Tensor flow and uses the ideas similar to the Open AI Gym used by Deepmind for video games!

In my blog Optimizing a Portfolio of Cryptocurrencies with Deep Reinforcement Learning I pick up their framework and see how the same strategy performs after the crypto crash.

Please have a look!

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.