Skip to content
All library documents

PPO Configuration for Reinforcement Learning Order Execution

Code Qlib

Summary

This configuration specifies a reinforcement learning setup for order execution using Proximal Policy Optimization. It defines a categorical action interpreter, a recurrent network, and a full-history state representation built from intraday and prior-day price and volume features. The simulator uses five-minute data over a 240-minute session, with the reward evaluated over most of that session. The environment is configured for parallel episodes, and the training section sets learning, collection, validation, checkpoint, and stopping parameters.

The file is an experiment configuration rather than a description of the reward formula, execution objective, order data, or evaluation results. It identifies the PPO reward class but does not reveal how reward is calculated, so the intended trade-off between execution quality and other costs cannot be assessed from this document alone. It also sets CUDA off and uses dummy parallel mode; runtime performance and generalization would depend on the implementation, data, and validation procedure, none of which are reported here.

Key ideas

  • The setup trains a recurrent PPO policy for an order-execution task.
  • The state uses intraday and previous-day price and volume features at five-minute granularity.
  • A categorical action interpreter defines discrete execution actions and a bounded step parameter.
  • The configuration specifies training, validation, and checkpoint settings but no empirical results.
  • The reward class is named without its formula, limiting assessment of the execution objective.

Tags

Full text
# train_ppo.yml


```yml
simulator:
  data_granularity: 5
  time_per_step: 30
  vol_limit: null
env:
  concurrency: 32
  parallel_mode: dummy
action_interpreter:
  class: CategoricalActionInterpreter
  kwargs:
    values: 4
    max_step: 8
  module_path: qlib.rl.order_execution.interpreter
state_interpreter:
  class: FullHistoryStateInterpreter
  kwargs:
    data_dim: 5
    data_ticks: 48  # 48 = 240 min / 5 min
    max_step: 8
    processed_data_provider:
      class: HandlerProcessedDataProvider
      kwargs:
        data_dir: ./data/pickle/
        feature_columns_today: ["$high", "$low", "$open", "$close", "$volume"]
        feature_columns_yesterday: ["$high_1", "$low_1", "$open_1", "$close_1", "$volume_1"]
        backtest: false
      module_path: qlib.rl.data.native
  module_path: qlib.rl.order_execution.interpreter
reward:
  class: PPOReward
  kwargs:
    max_step: 8
    start_time_index: 0
    end_time_index: 46  # 46 = (240 - 5) min / 5 min - 1
  module_path: qlib.rl.order_execution.reward
data:
  source:
    order_dir: ./data/orders
    feature_root_dir: ./data/pickle/
    feature_columns_today: ["$close0", "$volume0"]
    feature_columns_yesterday: []
    total_time: 240
    default_start_time_index: 0
    default_end_time_index: 235
    proc_data_dim: 5
  num_workers: 0
  queue_size: 20
network:
  class: Recurrent
  module_path: qlib.rl.order_execution.network
policy:
  class: PPO  # PPO, DQN
  kwargs:
    lr: 0.0001
  module_path: qlib.rl.order_execution.policy
runtime:
  seed: 42
  use_cuda: false
trainer:
  max_epoch: 500
  repeat_per_collect: 25
  earlystop_patience: 50
  episode_per_collect: 10000
  batch_size: 1024
  val_every_n_epoch: 4
  checkpoint_path: ./outputs/ppo
  checkpoint_every_n_iters: 1

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.