How Log Transformations Affect Statistical Trends and Estimates
Summary
The document asks how trends in cumulative data change when viewed on a logarithmic rather than linear scale. Its answer emphasizes that the relationship depends on the estimation method, so a trend fitted to raw observations cannot always be converted directly into an equivalent trend on logged observations. Logarithms compress values for easier comparison, but they also change the statistical problem being modeled.
The discussion distinguishes several approaches. Classical Pearson and Neyman frequentist estimators are described as not invariant to logging; maximum likelihood estimates are described as invariant under transformation, though potentially biased; Bayesian results may vary with the transformation and depend on the prior and generative model. The answer recommends choosing the scale based on assumptions about the process that generated the data, rather than convenience alone. It provides conceptual guidance but no worked example or empirical comparison, and its broad claims should be applied with care because invariance and bias depend on the estimator, model, and transformation details.
Key ideas
- A logarithmic scale compresses data and can make comparisons easier, but it changes the modeling question.
- Whether estimates remain consistent under logging depends on the estimation method.
- Maximum likelihood estimates are described as transformation invariant, though not necessarily unbiased.
- Bayesian estimates depend on the generative model and prior, so changing scale can change the result.
- The choice between linear and logarithmic data should reflect assumptions about how the observations were generated.
Tags
Full text
# What is the relation between the trend of log and linear in accumulative data? # What is the relation between the trend of log and linear in accumulative data? From this discussion, I know when to use the log, but now I am wondering how to guess the trends of log graph of accumulative data based on linear accumulative data graph? For example, this table is linear data of the cumulative number of deaths per capita And this table is about the log of data of the cumulative number of deaths per capita ## Answer by Dave Harris (score 1, accepted) https://quant.stackexchange.com/a/68628 The answer will depend on the estimation method. Classical Pearson and Neyman Frequentist statistics are not invariant to the log transformation. The estimator under the logarithmic transform, when taken as the power of Euler's number, will not match the estimator under the raw data. That is an artifact of the log function itself. The maximum likelihood estimator will be invariant under the logarithmic transformation but it will generally not be unbiased. The Bayesian estimator may or may not be invariant under the logarithmic transformation. However, if you did a good, professional job in building the prior distribution, it will not be invariant and taking the log will give you a different answer than if you left it with raw data. Nonetheless, if you were careful in constructing your prior distribution, it is quite possible they will be close enough that you do not care. Because Bayesian models are generative instead of sampling based, the choice of using the log or not would depend upon your model and not on what would be convenient or mentally helpful. If you believe that nature uses logs, then you use logs. If you believe that nature does not use logs, then you do not use logs. The use of logarithms to make it mentally easier to understand is different from asking how to project it mathematically. The logarithm compresses data allowing ease of comparison. It can do something other than what you planned it to do in some statistical paradigms.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.