Testing Financial Data Dependence with Chi-Square and Correlation Ratios
Summary
This article explains how to test whether two variables are independent using Pearson’s chi-square test on a grouped contingency table. For continuous data, observations are binned, actual cell counts are compared with counts expected under independence, and the resulting statistic is evaluated against a chi-square distribution. It also describes Cramer’s contingency coefficient as a normalized measure of association, and the correlation ratio as a way to detect nonlinear dependence that a linear correlation measure can miss.
The MQL5 tools described apply these methods to adjacent price increments, pairs of instruments, and model series such as autoregressive, ARCH, and logistic-map data. The article recommends standardizing data and trimming heavy tails when selecting bins, while noting that sparse expected counts complicate inference. Grouping loses information, and conclusions depend on binning, sample size, and significance level. The examples illustrate how nonlinear dependence can evade a linear measure, but do not establish a profitable trading signal.
Key ideas
- The chi-square independence test compares observed contingency-table counts with expected counts under independence.
- Continuous observations must be grouped into intervals before applying the test.
- Cramer’s contingency coefficient summarizes association strength, while the correlation ratio can indicate nonlinear dependence.
- Standardization and tail trimming can help create usable bins, but sparse expected counts remain a concern.
- Statistical dependence in price data does not by itself imply a tradable edge.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.