Managing Rolling Model Training and Strategy Updates with Qlib
Summary
This Qlib example shows an online manager coordinating rolling model training and strategy updates. Its workflow begins by training an initial strategy set, then runs a routine that updates online predictions, prepares new tasks and models, and generates signals. It also demonstrates adding another strategy configuration between routine runs, after which the manager updates the expanded set. The manager’s state is saved between calls so a later routine can resume from the stored state.
The example uses rolling task generation and configurable trainers, including trainer options that can delegate work through a worker process. It prints collected results and signals, and includes a reset operation that removes task and experiment records before starting over. The sample is implementation guidance rather than a trading study: it reports no predictive accuracy, portfolio returns, or evidence that rolling updates improve performance. Its configurations and assumptions, including how strategy identifiers are derived, may need adjustment for other workflows.
Key ideas
- The example separates initial model training from recurring model and prediction updates.
- Rolling tasks are generated for each configured strategy and managed through Qlib’s online manager.
- New strategies can be added after the initial run and incorporated in later routines.
- Saved manager state allows subsequent routines to resume the workflow.
- The code demonstrates orchestration and signal collection, but gives no evidence of trading performance.
Tags
Full text
# rolling_online_management.py
```py
# Copyright (c) Microsoft Corporation.
# Licensed under the MIT License.
"""
This example shows how OnlineManager works with rolling tasks.
There are four parts including first train, routine 1, add strategy and routine 2.
Firstly, the OnlineManager will finish the first training and set trained models to `online` models.
Next, the OnlineManager will finish a routine process, including update online prediction -> prepare tasks -> prepare new models -> prepare signals
Then, we will add some new strategies to the OnlineManager. This will finish first training of new strategies.
Finally, the OnlineManager will finish second routine and update all strategies.
"""
import os
import fire
import qlib
from qlib.model.trainer import DelayTrainerR, DelayTrainerRM, TrainerR, TrainerRM, end_task_train, task_train
from qlib.utils.pickle_utils import validate_trusted
from qlib.workflow import R
from qlib.workflow.online.strategy import RollingStrategy
from qlib.workflow.task.gen import RollingGen
from qlib.workflow.online.manager import OnlineManager
from qlib.tests.config import CSI100_RECORD_XGBOOST_TASK_CONFIG_ROLLING, CSI100_RECORD_LGB_TASK_CONFIG_ROLLING
from qlib.workflow.task.manage import TaskManager
class RollingOnlineExample:
trusted = False
def __init__(
self,
provider_uri="~/.qlib/qlib_data/cn_data",
region="cn",
trainer=None, # defaults to DelayTrainerRM; a supplied trainer keeps its own trust policy
task_url="mongodb://10.0.0.4:27017/", # not necessary when using TrainerR or DelayTrainerR
task_db_name="rolling_db", # not necessary when using TrainerR or DelayTrainerR
rolling_step=550,
tasks=None,
add_tasks=None,
*,
trusted=False,
):
if add_tasks is None:
add_tasks = [CSI100_RECORD_LGB_TASK_CONFIG_ROLLING]
if tasks is None:
tasks = [CSI100_RECORD_XGBOOST_TASK_CONFIG_ROLLING]
mongo_conf = {
"task_url": task_url, # your MongoDB url
"task_db_name": task_db_name, # database name
}
qlib.init(provider_uri=provider_uri, region=region, mongo=mongo_conf)
self.tasks = tasks
self.add_tasks = add_tasks
self.rolling_step = rolling_step
self.trusted = validate_trusted(trusted)
strategies = []
for task in tasks:
name_id = task["model"]["class"] # NOTE: Assumption: The model class can specify only one strategy
strategies.append(
RollingStrategy(
name_id,
task,
RollingGen(step=rolling_step, rtype=RollingGen.ROLL_SD),
trusted=self.trusted,
)
)
self.trainer = DelayTrainerRM(trusted=trusted) if trainer is None else trainer
self.rolling_online_manager = OnlineManager(strategies, trainer=self.trainer)
_ROLLING_MANAGER_PATH = (
".RollingOnlineExample" # the OnlineManager will dump to this file, for it can be loaded when calling routine.
)
def worker(self):
# train tasks by other progress or machines for multiprocessing
print("========== worker ==========")
if isinstance(self.trainer, TrainerRM):
for task in self.tasks + self.add_tasks:
name_id = task["model"]["class"]
self.trainer.worker(experiment_name=name_id)
else:
print(f"{type(self.trainer)} is not supported for worker.")
# Reset all things to the first status, be careful to save important data
def reset(self):
for task in self.tasks + self.add_tasks:
name_id = task["model"]["class"]
TaskManager(task_pool=name_id).remove()
exp = R.get_exp(experiment_name=name_id)
for rid in exp.list_recorders():
exp.delete_recorder(rid)
if os.path.exists(self._ROLLING_MANAGER_PATH):
os.remove(self._ROLLING_MANAGER_PATH)
def first_run(self):
print("========== reset ==========")
self.reset()
print("========== first_run ==========")
self.rolling_online_manager.first_train()
print("========== collect results ==========")
print(self.rolling_online_manager.get_collector()())
print("========== dump ==========")
self.rolling_online_manager.to_pickle(self._ROLLING_MANAGER_PATH)
def routine(self):
print("========== load ==========")
self.rolling_online_manager = OnlineManager.load(self._ROLLING_MANAGER_PATH)
print("========== routine ==========")
self.rolling_online_manager.routine()
print("========== collect results ==========")
print(self.rolling_online_manager.get_collector()())
print("========== signals ==========")
print(self.rolling_online_manager.get_signals())
print("========== dump ==========")
self.rolling_online_manager.to_pickle(self._ROLLING_MANAGER_PATH)
def add_strategy(self):
print("========== load ==========")
self.rolling_online_manager = OnlineManager.load(self._ROLLING_MANAGER_PATH)
print("========== add strategy ==========")
strategies = []
for task in self.add_tasks:
name_id = task["model"]["class"] # NOTE: Assumption: The model class can specify only one strategy
strategies.append(
RollingStrategy(
name_id,
task,
RollingGen(step=self.rolling_step, rtype=RollingGen.ROLL_SD),
trusted=self.trusted,
)
)
self.rolling_online_manager.add_strategy(strategies=strategies)
print("========== dump ==========")
self.rolling_online_manager.to_pickle(self._ROLLING_MANAGER_PATH)
def main(self):
self.first_run()
self.routine()
self.add_strategy()
self.routine()
if __name__ == "__main__":
####### to train the first version's models, use the command below
# Only opt in for artifacts whose writer and store you trust. first_run resets the experiments.
# python rolling_online_management.py --trusted=True first_run
####### to update the models and predictions after the trading time, use the command below
# The saved manager is a separately trusted local pickle and retains its original trust settings.
# python rolling_online_management.py routine
####### to give newly added strategies the same explicit consent
# python rolling_online_management.py --trusted=True add_strategy
####### to define your own parameters, use `--`
# python rolling_online_management.py --trusted=True --rolling_step=40 first_run
fire.Fire(RollingOnlineExample)
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.