Skip to content
All library documents

Challenges in Designing Reinforcement Learning for Stock Trading

Article Quant Q&A · Author: Mircea

Summary

The document asks how to normalize stock data and define rewards for a reinforcement learning trading system. The author describes transforming price changes into percentage returns, applying min-max scaling, and using the Sortino ratio as the reward. The response does not recommend a particular alternative; it emphasizes that feature selection and reward design require substantial experimentation.

The answer treats sound statistical practice as central to developing and evaluating predictive signals. It recommends building foundations in statistics, hypothesis testing, linear algebra, and calculus, and suggests studying machine learning and financial modeling. These are broad learning recommendations rather than a tested workflow. No experiments compare the proposed normalization or reward against alternatives, and the discussion gives no evidence that a particular model, feature set, or reward function will work. Its main practical lesson is that there is no universal recipe and that avoiding misleading conclusions requires careful statistical reasoning.

Key ideas

  • The question proposes percentage price changes followed by min-max scaling as model inputs.
  • The proposed reward is the Sortino ratio, but the response does not validate it or suggest a replacement.
  • Feature selection and reward design in reinforcement learning require substantial experimentation.
  • Statistical grounding and hypothesis testing help researchers avoid misleading apparent signals.
  • The document provides study guidance rather than a tested trading model or method.

Tags

Full text
# What data should I use for a machine learning model


# What data should I use for a machine learning model












I would like to ask you for an advice of any of you could help me with this information it would be really helpful.

I am trying to build a reinforcement learning trading bot that based on the current and past stock data it will try to predict when its the perfect moment for buying or selling a stock. But my problem is how could I normalize the data in a way that it could help my model to learn faster and also what reward should I use?

At the current stage my normalizing strategy is consist from 2 parts

- Converting the price into percentage of how much or less did the price increase or decreased from the previous one

- Use MinMax scale

Also in terms of reward I'm using sortino ratio.

Is there any other better alternative for this?

## Answer by R110 (score 2, accepted)

https://quant.stackexchange.com/a/63033

If you are asking these kinds of questions, it is very unlikely you will be able to successfully train any kind of ML or AI model on stock data with your current level of skill.

In fact, each of the questions you ask usually involves weeks, months and -- in some cases -- years of experimentation to discover the right mix of potential features. In reinforcement learning, just figuring out the correct reward function is a task that is much, much harder than it seems.

That said, I would encourage you to research widely. Start with a good book like Marcos Lopez de Prado's "Advances in Financial Machine Learning". If you are starting right at the beginning, begin with Andrew Trask's "Grokking Deep Learning" book and then move on to The Deep Learning Book https://www.deeplearningbook.org/

Make sure you know the fundamentals of statistics in depth. Really, really, really in depth. Don't skip over this. Learn and practice foundational statistics skills like your life depends on it. Do Kaggle challenges to cut your teeth with them. You need to know how to not fool yourself and a solid grounding in statistical methods and hypothesis testing will be your only guide on this lonely, long road. Be sharp in linear algebra and matrix math. Have some solid calculus footing.

There are tons of courses and tutorials across the web - particularly on YouTube. Udacity also has good courses on AI in general and even a specific Machine Learning for Trading course.

Naturally, none of these will point you to "the answer". But they will get you to start asking the right questions and, at the end of it all, begin to develop some intuition for how you might build your own form of ML / AI.

Good luck -- it's tremendously difficult; one of the hardest technical challenges on earth. But with the right attitude and work perhaps you will find some good predictive signals and come back here to set someone else on the same journey you went on!

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.