Skip to content
All library documents

Configuring PPO and TWAP Strategies for Order-Execution Backtests

Code Qlib

Summary

This configuration describes an order-execution backtest using five-minute market data and an order file. The main strategy uses a recurrent network with a PPO policy, a categorical action interpreter, and a state interpreter that supplies recent intraday history and selected price and volume features from the current and previous day. Its settings include a learning rate and action-related parameters. A separate 30-minute strategy is configured as TWAP, alongside exchange settings for deal prices and optional limits.

The file shows how components and data sources are wired together, but it is configuration rather than a full description of the training process or execution objective. It does not report results, costs, slippage, or comparison against a benchmark, and the referenced paths and checkpoints depend on the surrounding project. The setup therefore provides a reproducible starting structure only for an environment with compatible data and Qlib components; it does not establish that PPO outperforms TWAP or another execution method.

Key ideas

  • The configuration assigns a PPO policy and recurrent network to a five-minute order-execution strategy.
  • The state uses recent price and volume features from the current and previous day.
  • A separate TWAP strategy is configured with a 30-minute data granularity.
  • The file specifies wiring and parameters but gives no execution results or benchmark comparison.

Tags

Full text
# backtest_ppo.yml


```yml
order_file: ./data/orders/test_orders.pkl
start_time: "9:30"
end_time: "14:54"
data_granularity: "5min"
qlib:
  provider_uri_5min: ./data/bin/
exchange:
  limit_threshold: null
  deal_price: ["$close", "$close"]
  volume_threshold: null
strategies:
  1day:
    class: SAOEIntStrategy
    kwargs:
      data_granularity: 5
      action_interpreter:
        class: CategoricalActionInterpreter
        kwargs:
          max_step: 8
          values: 4
        module_path: qlib.rl.order_execution.interpreter
      network:
        class: Recurrent
        kwargs: {}
        module_path: qlib.rl.order_execution.network
      policy:
        class: PPO  # PPO, DQN
        kwargs:
          lr: 0.0001
          # Restore `weight_file` once the training workflow finishes. You can change the checkpoint file you want to use.
          # weight_file: outputs/ppo/checkpoints/latest.pth
        module_path: qlib.rl.order_execution.policy
      state_interpreter:
        class: FullHistoryStateInterpreter
        kwargs:
          data_dim: 5
          data_ticks: 48
          max_step: 8
          processed_data_provider:
            class: HandlerProcessedDataProvider
            kwargs:
              data_dir: ./data/pickle/
              feature_columns_today: ["$high", "$low", "$open", "$close", "$volume"]
              feature_columns_yesterday: ["$high_1", "$low_1", "$open_1", "$close_1", "$volume_1"]
            module_path: qlib.rl.data.native
        module_path: qlib.rl.order_execution.interpreter
    module_path: qlib.rl.order_execution.strategy
  30min:
    class: TWAPStrategy
    kwargs: {}
    module_path: qlib.contrib.strategy.rule_strategy
concurrency: 16
output_dir: outputs/ppo/

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.