Skip to content
All library documents

Cross-Sectional Factor Preparation and Rolling Linear Regression Evaluation

Article BigQuant

Summary

This short project note outlines a workflow for preparing data and fitting a predictive model. It divides SQL data construction into a base dataset, a normalized dataset, and a final feature-and-label stage. The note mentions cross-sectional factor ranking and a step intended to avoid look-ahead bias, then describes fitting linear regression with 50 days of training data and 10 days of test data. The author favors linear regression because an XGBoost run takes too long and says to print the model’s R-squared.

The document is a brief description of an assignment, not a full research report. It gives no feature definitions, target construction, market universe, dates, or reported R-squared, and provides no comparison with a benchmark or trading strategy. The stated train/test window alone does not establish that the evaluation is robust or free from leakage. Readers can learn the rough sequence of data preparation and model fitting, but would need the linked project material and further validation details to reproduce or assess it.

Key ideas

  • The proposed SQL workflow separates raw fields, normalized data, and feature and label definitions.
  • Cross-sectional factor ranking is included in the feature preparation process.
  • The note mentions a step intended to prevent look-ahead bias.
  • The model uses a 50-day training window and a 10-day test window with linear regression.
  • No R-squared value or trading performance result is reported.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.