Configuring PPO and TWAP Strategies for Order-Execution Backtests
Summary
This configuration describes an order-execution backtest using five-minute market data and an order file. The main strategy uses a recurrent network with a PPO policy, a categorical action interpreter, and a state interpreter that supplies recent intraday history and selected price and volume features from the current and previous day. Its settings include a learning rate and action-related parameters. A separate 30-minute strategy is configured as TWAP, alongside exchange settings for deal prices and optional limits.
The file shows how components and data sources are wired together, but it is configuration rather than a full description of the training process or execution objective. It does not report results, costs, slippage, or comparison against a benchmark, and the referenced paths and checkpoints depend on the surrounding project. The setup therefore provides a reproducible starting structure only for an environment with compatible data and Qlib components; it does not establish that PPO outperforms TWAP or another execution method.
Key ideas
- The configuration assigns a PPO policy and recurrent network to a five-minute order-execution strategy.
- The state uses recent price and volume features from the current and previous day.
- A separate TWAP strategy is configured with a 30-minute data granularity.
- The file specifies wiring and parameters but gives no execution results or benchmark comparison.
Tags
Full text
# backtest_ppo.yml
```yml
order_file: ./data/orders/test_orders.pkl
start_time: "9:30"
end_time: "14:54"
data_granularity: "5min"
qlib:
provider_uri_5min: ./data/bin/
exchange:
limit_threshold: null
deal_price: ["$close", "$close"]
volume_threshold: null
strategies:
1day:
class: SAOEIntStrategy
kwargs:
data_granularity: 5
action_interpreter:
class: CategoricalActionInterpreter
kwargs:
max_step: 8
values: 4
module_path: qlib.rl.order_execution.interpreter
network:
class: Recurrent
kwargs: {}
module_path: qlib.rl.order_execution.network
policy:
class: PPO # PPO, DQN
kwargs:
lr: 0.0001
# Restore `weight_file` once the training workflow finishes. You can change the checkpoint file you want to use.
# weight_file: outputs/ppo/checkpoints/latest.pth
module_path: qlib.rl.order_execution.policy
state_interpreter:
class: FullHistoryStateInterpreter
kwargs:
data_dim: 5
data_ticks: 48
max_step: 8
processed_data_provider:
class: HandlerProcessedDataProvider
kwargs:
data_dir: ./data/pickle/
feature_columns_today: ["$high", "$low", "$open", "$close", "$volume"]
feature_columns_yesterday: ["$high_1", "$low_1", "$open_1", "$close_1", "$volume_1"]
module_path: qlib.rl.data.native
module_path: qlib.rl.order_execution.interpreter
module_path: qlib.rl.order_execution.strategy
30min:
class: TWAPStrategy
kwargs: {}
module_path: qlib.contrib.strategy.rule_strategy
concurrency: 16
output_dir: outputs/ppo/
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.