Testing Hourly FX Returns with Multiple Comparisons and Market Controls
Summary
The discussion considers whether average hourly foreign-exchange log returns differ from zero, based on separate one-sample t-tests for 24 hour-of-day buckets. A central statistical issue is multiplicity: running many tests at a nominal significance level increases the chance of finding apparently significant results when all null hypotheses hold. Bonferroni adjustment is suggested as a conservative correction, with confidence intervals offered as an interpretable companion to p-values.
The replies also caution that hourly returns can reflect more than time-of-day effects. News, trading volume, and market opens and closes may confound comparisons, so observed differences should not automatically be attributed to the hour. Normality and sample selection also deserve attention, although one reply discusses equities despite the question's later clarification that the asset class is FX. These are general cautions rather than a complete analysis plan; the thread does not specify a model for controlling news or other intraday factors.
Key ideas
- Testing each of 24 hourly buckets separately raises a multiple-comparisons problem.
- Bonferroni adjustment reduces the family-wise false-positive risk by tightening each test's significance threshold.
- Confidence intervals can show plausible mean returns and whether zero lies within the interval.
- News, volume, and market opening or closing activity may confound apparent time-of-day return effects.
- The thread does not provide a complete method for controlling intraday FX-specific factors.
Tags
Full text
# Hourly Returns Statistical test
# Hourly Returns Statistical test
I am trying to do an analysis on time zones effect on intraday returns.
As a first step, I collected hourly log returns for the past 3 years and bucketed them by hour (so that I have 24 buckets with around 700 data points)
I am now trying to see if for some hours of the day, the average log return is significantly different than 0. To do that, I performed a 1sample tstat test on each of the buckets.
Are there any additional tests that I need to do to make sure the analysis is valid? (For each bucket, data looks normally distributed and there seems to be little autocorrelation)
Thanks for the help
Edit: the asset class is FX
## Answer by Comp_Warrior (score 2)
https://quant.stackexchange.com/a/32079
You should consider adjusting your p-values for multiplicity. Otherwise you would expect 5% of your tests to come out significant even if the null hypothesis is true (assuming you use 5% as the significance level).
## Answer by Brumder (score 2)
https://quant.stackexchange.com/a/32080
You would have to include a lot of variables here to actually isolate timing's effect on an index's returns. Given the most-influential variable on intraday returns is market news, this would be very, very difficult.
Assuming you could correct for news, remember that equity returns as a whole are leptokurtic and therefore not perfectly normally-distributed. Also, the benchmarks you select for this will have to be carefully chosen for specific reasons and compared to each other just make sure you are working with an appropriate sample to begin with. You don't mention what they are or what the asset-class is, so I thought this was worth mentioning. Volume is another huge variable you're leaving out here. Markets opening and closing will also greatly affect returns. There are other things to look for as well within the data, but the fact that it's misspecified to begin with makes it a moot point, really.
## Answer by David Kozak (score 2)
https://quant.stackexchange.com/a/32097
Agreed with @compwarrior. The Bonferroni correction, while conservative, is a reasonable way to test multiple hypotheses. Confidence regions are (to me) a more intuitive way of evaluating plausible values, and have a direct relationship with p-values. If what follows doesn't make sense, just consider that if zero is not in your interval than your returns are significantly different from zero
Where the (one-sided interval) t-statistic to test that the returns are greater than 0 is:
$\bar{x} \pm t_{n-1}(\alpha) \sqrt{(s^2/n)}$
The Bonferroni correction for your case (since you are testing your hypothesis on 24 data sets) would be :
$\bar{x} \pm t_{n-1}(\alpha/24) \sqrt{(s^2/n)}$
Here, $\bar{x}$ are your returns, $\alpha$ is your significance level, $t_{n-1}(\alpha)$ is the t-statistic of level alpha with $n-1$ degrees of freedom, and $s$ is your sample standard deviation.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.