Interpreting ADF Results on Repeated-Value Spread Data
Summary
This exchange examines why an augmented Dickey–Fuller test can produce surprising results for a constructed spread. The question compares a cumulative random series with minute-by-minute spread data and reports that the spread's p-value changes when its time-series object is converted to a numeric vector. The response points to a visible feature of the spread: many consecutive observations repeat. That creates many zero first differences, so the test's behavior need not match the visual impression that the series resembles a random walk.
The example is a reminder to inspect data structure and sampling before interpreting a unit-root test. Repeated or stale prices can affect the test, and changing the object representation may alter how the data are handled. The exchange does not establish exactly why the conversion changes the reported p-value, nor does it provide a broader diagnostic or a definitive cointegration analysis. It concerns one dataset and old software versions, so it should motivate checking the observations and test inputs rather than serve as a general rule about spread stationarity.
Key ideas
- Consecutive repeated observations create many zero first differences in a time series.
- A visual resemblance to a random walk does not establish that the series has a unit root.
- Inspect sampling and repeated values before interpreting an augmented Dickey–Fuller result.
- The exchange reports a p-value change after numeric conversion but does not explain its cause fully.
Tags
Full text
# ADF test in R yielding perfect cointegration. How is this possible? # ADF test in R yielding perfect cointegration. How is this possible? I am using the famous conintegrated pairs tutorial to just different stocks for cointegration. The adf.test yeilds perfect cointegration, which I feel must be incorrect. Here is why: When I run adf.test() on a cumsum of a random series, the plot looks like this: And it yields the following adf.test output: Augmented Dickey-Fuller Test data: sp Dickey-Fuller = -2.8333, Lag order = 4, p-value = 0.2314 alternative hypothesis: stationary Here is a spread I constructed, notice how it looks similar to the random walk: Which yields the following adf.test() output: Augmented Dickey-Fuller Test data: sprd3 Dickey-Fuller = 3.719, Lag order = 7, p-value = 0.99 alternative hypothesis: stationary Warning message: In adf.test(sprd3) : p-value greater than printed p-value Any ideas what could be going on here? Why is the p-value extremely different between the two cases? I have a hard time believing that the spread I constructed in the graph is has a p-value of .99... Thanks. UPDATE I have looked into this problem some more and have revealed a little more that may help us get to the bottom of the .99 p-value. Here is another spread I created: The spread looks a little more stable than the previous one I posted. I ran the adf.test() on this spread two different ways. The first was adf.test(sprd1). This came up with a p-value of .99, similar to what I have been experiencing. However, when I use as.numeric() on the spread, the result is quite different. Executing adf.test(as.numeric(sprd1)) gives me a p-value of .07 Interesting. A little more info, the sprd1 data is an xts object with minute-by-minute data and no missing values. xts version: 0.8-8 zoo version: 1.7-9 R version: 2.14 Maybe older packages are causing the problem? ## Answer by Joshua Ulrich (score 3) https://quant.stackexchange.com/a/9211 Your spread does not look similar to the random walk. Many of the observations are the same as the previous observation. This means most of the first differences are zero, which is why the test indicates your series has a unit-root. The current value is very good at explaining what the next value will be.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.