Skip to content
All library documents

Scaling High-Frequency Minute Data into Daily Features

Article BigQuant

Summary

This overview explains why high-frequency research systems process tick snapshots, trades, and order data with distributed computing. It describes the challenge of loading very large datasets and the time required to clean and analyze them on a single machine. Horizontal scaling and cluster computation are presented as ways to make this research more practical and support the extraction of high-frequency factors for machine-learning work.

The document introduces a module for converting minute-level observations into daily features and notes that its input data includes 143 fields. It also mentions that the order-count fields for the first ten bid and ask levels were added after a specified date. The article does not provide the feature definitions, extraction procedure, empirical results, or evidence that derived factors predict returns. Its discussion is therefore an introduction to the research infrastructure and data scope, rather than a reproducible strategy or validation study.

Key ideas

  • High-frequency research can use tick snapshots, trade records, and order data to construct factors.
  • Large datasets can exceed single-machine memory and processing capacity.
  • Distributed cluster computing can reduce the time needed for data cleaning and analysis.
  • The described module converts minute-level data into daily features and includes 143 fields.
  • The overview does not explain the extraction method or validate any resulting factors.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.