Assessing Correlation with Skewed Social Media Activity Data
Summary
The question considers whether cross-correlation is informative when comparing security price changes with social media discussion counts. The price series have fat tails, while the discussion measure is strongly right-skewed, with many smaller values and some much larger observations. The observed contemporaneous and short-lag correlations are positive but small, and the question assumes a linear relationship between the series.
The answer suggests using Pearson and Spearman correlation estimates with asymptotic p-values, as available through a statistical package. Pearson measures linear association, while Spearman assesses rank-based monotonic association and can be less sensitive to extreme values. The response does not explain how to interpret the cross-correlation matrix itself or address dependence over time. Its significance estimates should therefore be treated cautiously for stationary time series, where serial correlation can affect standard errors; the short answer offers a possible check, not a complete inference procedure.
Key ideas
- Pearson correlation measures linear association, while Spearman correlation measures association between ranked observations.
- Highly skewed discussion counts make it useful to compare more than one correlation measure.
- Small positive correlations at short lags do not by themselves establish statistical significance.
- Asymptotic p-values may need adjustment when observations have serial dependence.
Tags
Full text
# Interpretation of cross-correlation matrix when one sample distribution is not normal
# Interpretation of cross-correlation matrix when one sample distribution is not normal
I am looking at the variance of (log) price changes in securities vs. the amount of social media discussion about them. I'm not interested in building a model. I'm just looking to see if there is a significant correlation.
Suppose the social media is represented by a numeric variable “sm”. All the series I'm working with are weakly stationary. The distribution of the price data is as one would expect: normal with fat tails. The basic stats of a typical set of “sm” observations, however, are:
```
nobs 240.000000
NAs 0.000000
Minimum 0.000000
Maximum 725.000000
1. Quartile 52.000000
3. Quartile 119.250000
Mean 99.245833
Median 82.000000
Sum 23819.000000
SE Mean 5.573789
LCL Mean 88.265806
UCL Mean 110.225861
Variance 7456.110861
Stdev 86.348775
Skewness 3.428570
Kurtosis 17.793173
```
For price vs. “sm” contemporaneous, lag(1), and sometimes lag(2), the correlation is positive but small, about what I would expect. Because the distribution is not normal, I'm wondering if the cross-correlation matrix (ccf() function in R) provides a reasonable assessment of cross correlation (assuming linearity). I welcome any comments regarding how to interpret these results as well as any comments on best-practices.
## Answer by Walter (score 0, accepted)
https://quant.stackexchange.com/a/9439
You could use
```
rcorr(x, y, type=c("pearson","spearman"))
```
e.g.
```
# Correlations with significance levels
library(Hmisc)
rcorr(x, type="pearson") # type can be pearson or spearman
```
from the Hmisc package. It gives asymptotic p-values.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.