PPO Configuration for Reinforcement Learning Order Execution
Summary
This configuration specifies a reinforcement learning setup for order execution using Proximal Policy Optimization. It defines a categorical action interpreter, a recurrent network, and a full-history state representation built from intraday and prior-day price and volume features. The simulator uses five-minute data over a 240-minute session, with the reward evaluated over most of that session. The environment is configured for parallel episodes, and the training section sets learning, collection, validation, checkpoint, and stopping parameters.
The file is an experiment configuration rather than a description of the reward formula, execution objective, order data, or evaluation results. It identifies the PPO reward class but does not reveal how reward is calculated, so the intended trade-off between execution quality and other costs cannot be assessed from this document alone. It also sets CUDA off and uses dummy parallel mode; runtime performance and generalization would depend on the implementation, data, and validation procedure, none of which are reported here.
Key ideas
- The setup trains a recurrent PPO policy for an order-execution task.
- The state uses intraday and previous-day price and volume features at five-minute granularity.
- A categorical action interpreter defines discrete execution actions and a bounded step parameter.
- The configuration specifies training, validation, and checkpoint settings but no empirical results.
- The reward class is named without its formula, limiting assessment of the execution objective.
Tags
Full text
# train_ppo.yml
```yml
simulator:
data_granularity: 5
time_per_step: 30
vol_limit: null
env:
concurrency: 32
parallel_mode: dummy
action_interpreter:
class: CategoricalActionInterpreter
kwargs:
values: 4
max_step: 8
module_path: qlib.rl.order_execution.interpreter
state_interpreter:
class: FullHistoryStateInterpreter
kwargs:
data_dim: 5
data_ticks: 48 # 48 = 240 min / 5 min
max_step: 8
processed_data_provider:
class: HandlerProcessedDataProvider
kwargs:
data_dir: ./data/pickle/
feature_columns_today: ["$high", "$low", "$open", "$close", "$volume"]
feature_columns_yesterday: ["$high_1", "$low_1", "$open_1", "$close_1", "$volume_1"]
backtest: false
module_path: qlib.rl.data.native
module_path: qlib.rl.order_execution.interpreter
reward:
class: PPOReward
kwargs:
max_step: 8
start_time_index: 0
end_time_index: 46 # 46 = (240 - 5) min / 5 min - 1
module_path: qlib.rl.order_execution.reward
data:
source:
order_dir: ./data/orders
feature_root_dir: ./data/pickle/
feature_columns_today: ["$close0", "$volume0"]
feature_columns_yesterday: []
total_time: 240
default_start_time_index: 0
default_end_time_index: 235
proc_data_dim: 5
num_workers: 0
queue_size: 20
network:
class: Recurrent
module_path: qlib.rl.order_execution.network
policy:
class: PPO # PPO, DQN
kwargs:
lr: 0.0001
module_path: qlib.rl.order_execution.policy
runtime:
seed: 42
use_cuda: false
trainer:
max_epoch: 500
repeat_per_collect: 25
earlystop_patience: 50
episode_per_collect: 10000
batch_size: 1024
val_every_n_epoch: 4
checkpoint_path: ./outputs/ppo
checkpoint_every_n_iters: 1
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.