Skip to content
All library documents

Preparing Daily Equity Risk Model Estimates with Structured Covariance

Code Qlib

Summary

This document describes a data preparation workflow for daily risk estimates on China A-shares. For each date, it selects the CSI 300 constituents, gathers a rolling window of closing prices, calculates returns, and clips extreme returns at the cross-sectional 2.5th and 97.5th percentiles. A structured covariance estimator then decomposes the risk into factor exposures, factor covariance, and specific variance. The workflow saves these components by date for downstream use.

The example explains how to construct and store model inputs, but it does not report predictive performance or show how to use the estimates in portfolio construction. Results depend on the input data, the chosen rolling window, universe membership, and estimator assumptions. The saved specific-risk measure is expressed as volatility by taking the square root of the estimated specific variance. The code is a preparation example rather than a complete trading or validation procedure.

Key ideas

  • The workflow estimates risk components separately for each trading date.
  • It uses a rolling window of closing prices to calculate returns for current CSI 300 constituents.
  • Extreme returns are clipped using cross-sectional percentile thresholds before model fitting.
  • The structured covariance model outputs factor exposures, factor covariance, and specific variance.
  • Specific variance is converted to volatility before it is saved.

Tags

Full text
# prepare_riskdata.py


```py
# Copyright (c) Microsoft Corporation.
# Licensed under the MIT License.
import os
import numpy as np
import pandas as pd

from qlib.data import D
from qlib.model.riskmodel import StructuredCovEstimator


def prepare_data(riskdata_root="./riskdata", T=240, start_time="2016-01-01"):
    universe = D.features(D.instruments("csi300"), ["$close"], start_time=start_time).swaplevel().sort_index()

    price_all = (
        D.features(D.instruments("all"), ["$close"], start_time=start_time).squeeze().unstack(level="instrument")
    )

    # StructuredCovEstimator is a statistical risk model
    riskmodel = StructuredCovEstimator()

    for i in range(T - 1, len(price_all)):
        date = price_all.index[i]
        ref_date = price_all.index[i - T + 1]

        print(date)

        codes = universe.loc[date].index
        price = price_all.loc[ref_date:date, codes]

        # calculate return and remove extreme return
        ret = price.pct_change()
        ret.clip(ret.quantile(0.025), ret.quantile(0.975), axis=1, inplace=True)

        # run risk model
        F, cov_b, var_u = riskmodel.predict(ret, is_price=False, return_decomposed_components=True)

        # save risk data
        root = riskdata_root + "/" + date.strftime("%Y%m%d")
        os.makedirs(root, exist_ok=True)

        pd.DataFrame(F, index=codes).to_pickle(root + "/factor_exp.pkl")
        pd.DataFrame(cov_b).to_pickle(root + "/factor_cov.pkl")
        # for specific_risk we follow the convention to save volatility
        pd.Series(np.sqrt(var_u), index=codes).to_pickle(root + "/specific_risk.pkl")


if __name__ == "__main__":
    import qlib

    qlib.init(provider_uri="~/.qlib/qlib_data/cn_data")

    prepare_data()

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.