Skip to content
All library documents

Practical Risks of Reinforcement Learning for Algorithmic Trading

Article Quant Q&A · Author: QMath

Summary

The document outlines a proposed reinforcement-learning setup for trading a single asset pair. The suggested observation includes time, holdings in both assets, and market information such as quotes, order-book data, or automated-market-maker reserves. Trading either asset or taking no action would form the action set, with the market treated as a partially observable decision process.

The author proposes learning from live interaction with a small amount of capital, hoping that real executions would capture details such as queue priority and other market behavior. The document asks what practical problems this approach may face, but supplies no answers, experiments, or performance evidence. It raises an important design question about learning through market interaction; the suitability of the state representation, exploration, execution modeling, and live deployment remains unresolved in the text.

Key ideas

  • A proposed trading agent observes time, current holdings, and actionable market data.
  • The setup frames trading as a partially observable decision process with trade and hold actions.
  • Live learning is proposed as a way to encounter execution details absent from a simplified model.
  • The document asks about practical obstacles but does not identify or evaluate specific solutions.

Tags

Full text
# Potential problems with trying to apply reinforcement learning to algorithmic trading


# Potential problems with trying to apply reinforcement learning to algorithmic trading












I have been attempting to develop an algorithmic trading agent for a single asset pair and upon researching, it seems as if, in theory, reinforcement learning would be a natural way to approach this problem.

My idea was to have our observations be defined as follows:

$$o_i = (t_i, v_i, w_i, b_i)$$

where

- $t_i$: time of observation

- $v_i$: amount held by agent of asset 0

- $w_i$: amount held by agent of asset 1

- $b_i$: some actionable representation of the market such as quote data or order book data for an order book market, asset reserve levels for an automated market maker decentralized exchange marker, etc.

and to assume these observations are observations from a partially-observable markov decision process (described here), where the actions are defined as trading some of either asset 0 or 1 or doing nothing, so that the literature on those can be applied.

This would allow us to then implicitly model many real-world intricacies such as orders being executed before ours is executed given that we received a new observation by running our agent live and having it learn from the actual market by giving it a small amount of actual capital to work with so that it may explore and learn through actions based on the observations, perhaps pointing us more towards other insights about the market.

I am looking for what problems this approach may run into as I know I am very likely not the first to approach trading this way, e.g. this may sound good in theory, but the practical implementation is a lot more intricate, etc.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.