Skip to content
All library documents

Local Transformer Training with Data Format Alignment for Cloud Inference

Article BigQuant

Summary

This guide describes a workflow for training an end-to-end stock model locally on compressed minute and multi-minute bar data, then submitting its weights and training code for cloud prediction. It explains that local files store prices and amounts as scaled integers, replace stock symbols with integer identifiers, and retain only three order-book levels. A canonical preprocessing step is proposed to restore units, mark missing OHLC values, align available fields, and create a common instrument key for local training and cloud inference.

A sample Transformer predicts a next-day return from sequences of bar features, with training-set normalization statistics saved alongside model weights for reuse at inference. The guide also covers monthly Feather partitions, field mappings, and submission artifacts. Its example is explicitly a minimal demonstration, and the document presents no predictive performance results. Reliable reproduction depends on matching preprocessing across environments and checking the handling of missing values, features, and labels.

Key ideas

  • Local compressed prices and amounts must be rescaled before feature construction.
  • Missing OHLC values and instrument identifiers need consistent handling across local and cloud data.
  • A shared canonical preprocessing path helps reduce train and inference mismatches.
  • The example Transformer uses historical bar sequences to predict a future daily return.
  • Normalization statistics and training code are saved with the model for reuse and review.
  • The example is a workflow demonstration and does not establish model performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.