Skip to content
All library documents

Assessing Correlation with Skewed Social Media Activity Data

Article Quant Q&A · Author: SCallan

Summary

The question considers whether cross-correlation is informative when comparing security price changes with social media discussion counts. The price series have fat tails, while the discussion measure is strongly right-skewed, with many smaller values and some much larger observations. The observed contemporaneous and short-lag correlations are positive but small, and the question assumes a linear relationship between the series.

The answer suggests using Pearson and Spearman correlation estimates with asymptotic p-values, as available through a statistical package. Pearson measures linear association, while Spearman assesses rank-based monotonic association and can be less sensitive to extreme values. The response does not explain how to interpret the cross-correlation matrix itself or address dependence over time. Its significance estimates should therefore be treated cautiously for stationary time series, where serial correlation can affect standard errors; the short answer offers a possible check, not a complete inference procedure.

Key ideas

  • Pearson correlation measures linear association, while Spearman correlation measures association between ranked observations.
  • Highly skewed discussion counts make it useful to compare more than one correlation measure.
  • Small positive correlations at short lags do not by themselves establish statistical significance.
  • Asymptotic p-values may need adjustment when observations have serial dependence.

Tags

Full text
# Interpretation of cross-correlation matrix when one sample distribution is not normal


# Interpretation of cross-correlation matrix when one sample distribution is not normal












I am looking at the variance of (log) price changes in securities vs. the amount of social media discussion about them. I'm not interested in building a model. I'm just looking to see if there is a significant correlation.

Suppose the social media is represented by a numeric variable “sm”. All the series I'm working with are weakly stationary. The distribution of the price data is as one would expect: normal with fat tails. The basic stats of a typical set of “sm” observations, however, are:

```
nobs          240.000000
NAs             0.000000
Minimum         0.000000
Maximum       725.000000
1. Quartile    52.000000
3. Quartile   119.250000
Mean           99.245833
Median         82.000000
Sum         23819.000000
SE Mean         5.573789
LCL Mean       88.265806
UCL Mean      110.225861
Variance     7456.110861
Stdev          86.348775
Skewness        3.428570
Kurtosis       17.793173
```

For price vs. “sm” contemporaneous, lag(1), and sometimes lag(2), the correlation is positive but small, about what I would expect. Because the distribution is not normal, I'm wondering if the cross-correlation matrix (ccf() function in R) provides a reasonable assessment of cross correlation (assuming linearity). I welcome any comments regarding how to interpret these results as well as any comments on best-practices.

## Answer by Walter (score 0, accepted)

https://quant.stackexchange.com/a/9439

You could use

```
rcorr(x, y, type=c("pearson","spearman"))
```

e.g.

```
# Correlations with significance levels
library(Hmisc)
rcorr(x, type="pearson") # type can be pearson or spearman
```

from the Hmisc package. It gives asymptotic p-values.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.