Skip to content
All library documents

Configuring Qlib TRA with Alpha158 for CSI 300 Ranking

Code Qlib

Summary

This configuration defines a Qlib experiment using Alpha158 features for the CSI 300 and the Shanghai 300 index as benchmark. It normalizes features with robust z-scores, fills missing values, filters a selected subset of columns, and applies cross-sectional rank normalization to labels. The prediction target is a forward close-to-close return, and the dataset uses sequences of 60 observations.

The model is a recurrent TRA setup using an LSTM architecture with attention, three latent states, and a router transport method. Training, validation, and test periods are separated, with the portfolio analysis covering the test period. The portfolio uses a top-k dropout strategy with specified transaction costs, price assumptions, and a limit threshold. These are experimental settings, not reported findings: the document supplies no metrics, comparison, or validation of predictive performance. Results would depend on the data and framework implementation, including processing and execution assumptions.

Key ideas

  • The configuration applies Alpha158 features to a CSI 300 universe with a Shanghai 300 benchmark.
  • Feature processing includes column filtering, robust normalization, and missing-value filling.
  • Labels use a forward close-to-close return and are cross-sectionally rank-normalized.
  • The model combines an LSTM with attention and a three-state TRA setup.
  • Portfolio analysis uses top-k dropout and specified trading cost and price assumptions.
  • No model performance results or comparisons are included.

Tags

Full text
# workflow_config_tra_Alpha158.yaml


```yaml
qlib_init:
  provider_uri: "~/.qlib/qlib_data/cn_data"
  region: cn

market: &market csi300
benchmark: &benchmark SH000300

data_handler_config: &data_handler_config
  start_time: 2008-01-01
  end_time: 2020-08-01
  fit_start_time: 2008-01-01
  fit_end_time: 2014-12-31
  instruments: *market
  infer_processors:
    - class: FilterCol
      kwargs:
        fields_group: feature
        col_list: ["RESI5", "WVMA5", "RSQR5", "KLEN", "RSQR10", "CORR5", "CORD5", "CORR10",
                   "ROC60", "RESI10", "VSTD5", "RSQR60", "CORR60", "WVMA60", "STD5",
                   "RSQR20", "CORD60", "CORD10", "CORR20", "KLOW"]
    - class: RobustZScoreNorm
      kwargs:
        fields_group: feature
        clip_outlier: true
    - class: Fillna
      kwargs:
        fields_group: feature
  learn_processors:
    - class: CSRankNorm
      kwargs:
        fields_group: label
  label: ["Ref($close, -2) / Ref($close, -1) - 1"]

num_states: &num_states 3

memory_mode: &memory_mode sample

tra_config: &tra_config
  num_states: *num_states
  rnn_arch: LSTM
  hidden_size: 32
  num_layers: 1
  dropout: 0.0
  tau: 1.0
  src_info: LR_TPE

model_config: &model_config
  input_size: 20
  hidden_size: 64
  num_layers: 2
  rnn_arch: LSTM
  use_attn: True
  dropout: 0.0

port_analysis_config: &port_analysis_config
    strategy:
        class: TopkDropoutStrategy
        module_path: qlib.contrib.strategy
        kwargs:
            signal: <PRED>
            topk: 50
            n_drop: 5
    backtest:
        start_time: 2017-01-01
        end_time: 2020-08-01
        account: 100000000
        benchmark: *benchmark
        exchange_kwargs:
            limit_threshold: 0.095
            deal_price: close
            open_cost: 0.0005
            close_cost: 0.0015
            min_cost: 5

task:
  model:
    class: TRAModel
    module_path: qlib.contrib.model.pytorch_tra
    kwargs:
      tra_config: *tra_config
      model_config: *model_config
      model_type: RNN
      lr: 1e-3
      n_epochs: 100
      max_steps_per_epoch:
      early_stop: 20
      logdir: output/Alpha158
      seed: 0
      lamb: 1.0
      rho: 0.99
      alpha: 0.5
      transport_method: router
      memory_mode: *memory_mode
      eval_train: False
      eval_test: True
      pretrain: True
      init_state:
      freeze_model: False
      freeze_predictors: False
  dataset:
    class: MTSDatasetH
    module_path: qlib.contrib.data.dataset
    kwargs:
      handler:
        class: Alpha158
        module_path: qlib.contrib.data.handler
        kwargs: *data_handler_config
      segments:
        train: [2008-01-01, 2014-12-31]
        valid: [2015-01-01, 2016-12-31]
        test: [2017-01-01, 2020-08-01]
      seq_len: 60
      input_size:
      num_states: *num_states
      batch_size: 1024
      n_samples:
      memory_mode: *memory_mode
      drop_last: True
  record:
    - class: SignalRecord
      module_path: qlib.workflow.record_temp
      kwargs: 
        model: <MODEL>
        dataset: <DATASET>
    - class: SigAnaRecord
      module_path: qlib.workflow.record_temp
      kwargs: 
        ana_long_short: False
        ann_scaler: 252
    - class: PortAnaRecord
      module_path: qlib.workflow.record_temp
      kwargs: 
        config: *port_analysis_config

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.