Preparing External Market Data for QTPyLib Backtests
Summary
This guide explains how to bring market data from an outside provider into QTPyLib for backtesting. The workflow module’s preparation step converts a data frame into the library’s expected format and can write the result as a CSV file. The example uses historical equity data, but the core method is to supply an instrument identifier and source data to the conversion function, then choose an output location. The instrument identifier must follow the format QTPyLib expects for strategy instruments.
Converted data can also be stored in the Blotter’s MySQL database, provided a Blotter is running. For backtests that use CSV files, the guide describes passing the data directory as a command-line option. It gives an example workflow rather than performance evidence or a comparison of data sources. It does not cover data validation, adjustment policies, timestamp alignment, or the risks of differences between external data and live captured data, so those details need separate consideration.
Key ideas
- External market data must be converted to QTPyLib’s expected format before strategies can consume it.
- The preparation workflow can save converted bars as CSV files.
- Converted data can optionally be stored in the Blotter’s MySQL database.
- CSV data can be selected for a backtest through a command-line data-directory option.
- The instrument identifier supplied for conversion must use a valid QTPyLib instrument format.
Tags
Full text
# load the workflow module
Data Workflow
=============
QTPyLib's new ``workflow`` module includes some handy methods for
working with external data sources when backtesting.
Working with External Data
--------------------------
Sometimes, you'd want to backtest your strategies using market data
you already have from sources other than the ``Blotter``.
Before you can use market data from external data sources,
you'll need to convert it into a QTPyLib-compatible data format.
Once the data is converted, it can be read by your strategies as CSV files.
You can also save the converted data in your ``Blotter``'s MySQL database
so it can be used just like any other data captured by your ``Blotter``.
**Code Example:**
.. code:: python
# load the workflow module
from qtpylib import workflow as wf
# Load some data from Quandl
import quandl
aapl = quandl.get("WIKI/AAPL", authtoken="your token here")
# convert the data into a QTPyLib-compatible
# data will be saved in ~/Desktop/AAPL.BAR.csv
df = wf.prepare_data("AAPL", data=aapl, output_path="~/Desktop/")
# store converted bar data in MySQL
# optional, requires a running Blotter
wf.store_data(df, kind="BAR")
.. note::
The first argument in ``prepare_data()`` must be a **valid string as IB tuple**
(just like the those specified in your strategy's ``instruments`` parameter).
For a complete list of available methods and parameters for each
method, please refer to the `Workflow API <./api.html#workflow-api>`_
for a full list of available parameters for each method.
Using CSV files when Backtesting
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Once you have your CSV files in a QTPyLib-compatible format,
you can backtest using this data using the ``--data`` flag when
running your backtests, for example:
.. code:: bash
$ python strategy.py --backtest --start 2015-01-01 --end 2015-12-31 --data ~/mycsvdata/ --output ~/portfolio.pkl
Please refer to `Back-Testing Using QTPyLib <./algo.html#back-testing-using-qtpylib>`_
for more information about back-testing using QTPyLib.Shown in full with attribution under the source's licence. Licence: Apache-2.0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.