Skip to content
All library documents

Interpreting Peaks and Zero Returns in One-Minute XAUUSD Data

Article Quant Q&A · Author: Michael Teguh Laksana

Summary

The document describes an exploratory analysis of one-minute XAUUSD log returns over a long historical sample, compared with fitted normal and Student-t distributions. It raises questions about irregular density peaks, a pronounced concentration near zero, unchanged bars, and sparse histogram bins, and whether these patterns reflect price increments, market microstructure, time aggregation, volatility changes, or data issues.

It provides observations rather than answers: most returns fall within a narrow range around zero, and a small fraction are exactly zero. No statistical tests, alternative modeling method, or causal evidence are supplied. The author is uncertain whether tiny moves should be removed or modeled separately. The key practical implication is that a simple continuous distribution may not capture discretization and zero-return mass in high-frequency data; the observations need validation against the data construction and market mechanics before they can guide a model. The document does not establish which explanation is correct or recommend a specific distribution.

Key ideas

  • The author compares one-minute XAUUSD log-return densities with fitted normal and Student-t distributions.
  • The observed return histogram has smaller peaks and uneven decay away from its center.
  • A large concentration of returns lies near zero, including bars with no open-to-close price change.
  • The document asks whether price increments, time aggregation, changing volatility, or data problems explain the histogram patterns.
  • It offers no test or conclusion for choosing how to model or filter small and zero returns.

Tags

Full text
# Peaks and gaps in log-return of XAUUSD 1-minute log-return density


# Peaks and gaps in log-return of XAUUSD 1-minute log-return density












I'm tinkering around a 1-minute XAUUSD data from March 2009-December 2023 to see if I can model it with a log-normal or log-t distribution and I happen to notice some interesting properties in the log-return ($log(\frac{close}{open})$). The zoomed-in density histogram below is the data log-return (blue) and the fitted t-distribution of the data:

and the one vs a fitted normal distribution:

I'm relatively new with financial data and I don't quite understand why these observations occur:

- I noticed the log-return doesn't smoothly decay as it goes further from the peak, but instead have some smaller peaks of data. It's quite visible especially in the second plot where the blue distribution has periods where the density decays more slowly. I don't quite understand why this happen. Does this mean certain amount of price change is more likely than the others and price is more likely to move ins a certain step sizes? Will this affect the model and if so, how should I integrate this into the model?

- I noticed a very large peak of low or no return (which I guess is to be expected for smaller timeframe data?) near the center of the distribution. About 99% of the log-return occur between log return of -0.001 and 0.001 (about 0.1% of price change) and 4% of the total data shows a 0 log-return (no-price change between open and close in 1 minute). I'm not really sure how to handle this data since it really mess up the distributions I'm trying to fit. Any model recommendations? Should I include the no-change in the data? Coming from a non-finance background, we usually treat these smaller movements as noises. While I cannot eliminate 99% of the data, is it alright to eliminate the even smaller price movements?

- I also noticed the log-return of data nearer to 0 is more sparsely distributed from the bins in the histogram. Prices are concentrated in some bins while the adjacent bins are relatively lower. While it does make sense to me that price difference that are too small are less likely than a bigger one from minimum lot sizing and the chosen timeframe of 1 minute masks the smaller price movements, etc., shouldn't this just strictly mean lower frequency for lower price changes and not a more sparsely distributed data? Is this normal or an artifact of the computation or bad data or its because gold prices are less volatile in 2009 than 2023 or something else entirely?

Thank you in advance.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.