Preparing Daily Equity Risk Model Estimates with Structured Covariance
Summary
This document describes a data preparation workflow for daily risk estimates on China A-shares. For each date, it selects the CSI 300 constituents, gathers a rolling window of closing prices, calculates returns, and clips extreme returns at the cross-sectional 2.5th and 97.5th percentiles. A structured covariance estimator then decomposes the risk into factor exposures, factor covariance, and specific variance. The workflow saves these components by date for downstream use.
The example explains how to construct and store model inputs, but it does not report predictive performance or show how to use the estimates in portfolio construction. Results depend on the input data, the chosen rolling window, universe membership, and estimator assumptions. The saved specific-risk measure is expressed as volatility by taking the square root of the estimated specific variance. The code is a preparation example rather than a complete trading or validation procedure.
Key ideas
- The workflow estimates risk components separately for each trading date.
- It uses a rolling window of closing prices to calculate returns for current CSI 300 constituents.
- Extreme returns are clipped using cross-sectional percentile thresholds before model fitting.
- The structured covariance model outputs factor exposures, factor covariance, and specific variance.
- Specific variance is converted to volatility before it is saved.
Tags
Full text
# prepare_riskdata.py
```py
# Copyright (c) Microsoft Corporation.
# Licensed under the MIT License.
import os
import numpy as np
import pandas as pd
from qlib.data import D
from qlib.model.riskmodel import StructuredCovEstimator
def prepare_data(riskdata_root="./riskdata", T=240, start_time="2016-01-01"):
universe = D.features(D.instruments("csi300"), ["$close"], start_time=start_time).swaplevel().sort_index()
price_all = (
D.features(D.instruments("all"), ["$close"], start_time=start_time).squeeze().unstack(level="instrument")
)
# StructuredCovEstimator is a statistical risk model
riskmodel = StructuredCovEstimator()
for i in range(T - 1, len(price_all)):
date = price_all.index[i]
ref_date = price_all.index[i - T + 1]
print(date)
codes = universe.loc[date].index
price = price_all.loc[ref_date:date, codes]
# calculate return and remove extreme return
ret = price.pct_change()
ret.clip(ret.quantile(0.025), ret.quantile(0.975), axis=1, inplace=True)
# run risk model
F, cov_b, var_u = riskmodel.predict(ret, is_price=False, return_decomposed_components=True)
# save risk data
root = riskdata_root + "/" + date.strftime("%Y%m%d")
os.makedirs(root, exist_ok=True)
pd.DataFrame(F, index=codes).to_pickle(root + "/factor_exp.pkl")
pd.DataFrame(cov_b).to_pickle(root + "/factor_cov.pkl")
# for specific_risk we follow the convention to save volatility
pd.Series(np.sqrt(var_u), index=codes).to_pickle(root + "/specific_risk.pkl")
if __name__ == "__main__":
import qlib
qlib.init(provider_uri="~/.qlib/qlib_data/cn_data")
prepare_data()
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.