High-Frequency Dataset Handling and a Trend Prediction Benchmark
Summary
This document introduces two high-frequency trading examples: handling a dataset for reinforcement learning and predicting price trends. It explains that the dataset is represented by a Qlib DatasetH object, which can be serialized to disk and reloaded. After loading, users can reinitialize dataset or data-handler settings, including instruments, date boundaries, and segments, to generate data for a changed configuration.
The document also reports a LightGBM benchmark for trend prediction on a dataset labeled Alpha158. It provides information coefficient, rank information coefficient, their information ratios, long and short precision, and long-short average return and Sharpe. These are reported as point estimates with zero displayed variation. The excerpt does not define the target, sampling interval, evaluation period, benchmark procedure, or transaction costs, so the figures alone do not establish expected live trading performance. Its main value is as a brief example of dataset persistence and a compact model-evaluation report.
Key ideas
- A Qlib DatasetH high-frequency dataset can be serialized and later reloaded.
- Reloaded dataset settings can be reset to generate data for different instruments, dates, or segments.
- The document reports LightGBM trend-prediction metrics, including rank-based measures and long-short performance.
- The benchmark excerpt omits evaluation details needed to assess robustness or live tradability.
Tags
Full text
# Introduction
# Introduction
This folder contains 2 examples
- A high-frequency dataset example
- An example of predicting the price trend in high-frequency data
## High-Frequency Dataset
This dataset is an example for RL high frequency trading.
### Get High-Frequency Data
Get high-frequency data by running the following command:
```bash
python workflow.py get_data
```
### Dump & Reload & Reinitialize the Dataset
The High-Frequency Dataset is implemented as `qlib.data.dataset.DatasetH` in the `workflow.py`. `DatatsetH` is the subclass of [`qlib.utils.serial.Serializable`](https://qlib.readthedocs.io/en/latest/advanced/serial.html), whose state can be dumped in or loaded from disk in `pickle` format.
### About Reinitialization
After reloading `Dataset` from disk, `Qlib` also support reinitializing the dataset. It means that users can reset some states of `Dataset` or `DataHandler` such as `instruments`, `start_time`, `end_time` and `segments`, etc., and generate new data according to the states.
The example is given in `workflow.py`, users can run the code as follows.
### Run the Code
Run the example by running the following command:
```bash
python workflow.py dump_and_load_dataset
```
## Benchmarks Performance (predicting the price trend in high-frequency data)
Here are the results of models for predicting the price trend in high-frequency data. We will keep updating benchmark models in future.
| Model Name | Dataset | IC | ICIR | Rank IC | Rank ICIR | Long precision| Short Precision | Long-Short Average Return | Long-Short Average Sharpe |
|---|---|---|---|---|---|---|---|---|---|
| LightGBM | Alpha158 | 0.0349±0.00 | 0.3805±0.00| 0.0435±0.00 | 0.4724±0.00 | 0.5111±0.00 | 0.5428±0.00 | 0.000074±0.00 | 0.2677±0.00 |Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.