Skip to content
All library documents

Scaling Histogram Frequencies to Estimate a Probability Density

Article Quant Q&A · Author: naz

Summary

The document explains how to place a histogram and a probability density curve on a comparable scale. It connects the cumulative distribution function to its density: probability accumulated over a small interval is approximately the density at that location multiplied by the interval width.

To convert histogram counts into density heights, divide each class frequency by the total number of observations to obtain the class probability, then divide that probability by the bin width. The reverse conversion follows by multiplying density by bin width and sample size. The approximation depends on bin width, so the result is sensitive to the histogram's interval choice; the response does not discuss alternative estimators or finite-sample uncertainty.

Key ideas

  • A density describes probability per unit of the variable, while histogram counts describe observations per bin.
  • Convert each bin count to a probability by dividing by the total sample size.
  • Divide the bin probability by its width to obtain a histogram density height.
  • The approximation depends on the selected bin width.

Tags

Full text
# Scaling of probability mass function


# Scaling of probability mass function












Given a histogram and the probability mass function values for each observation, when plotting the histogram and the curve (this is bell curve since the data is assumed to be normal) on the same figure, the two will not be of the same scaling. How does one actually scale the probability mass/density function so that it is 'equivalent' to the histogram?

## Answer by Neeraj (score 2, accepted)

https://quant.stackexchange.com/a/22883

This is pretty simple. Let's assume that $F_X(x)$ is cummulative density function, such that $$F_X(x) = \int_{-\infty}^{x} f_X(x) dx \quad \cdots \cdots (1)$$ where $f_X(x)$ is density function. Differentiating on the both side (ignoring subscript X), we get $$\frac{d}{dx}\, F(x) = f(x) $$ It can be written as; $$dF(x)=f(x) \,d(x)$$ $$F(x+h) - F(x) = f(x)dx$$ so, more precisely your density $f(x)$ is: $$f(x) = \frac{F(x+h)-F(x)}{dx}\quad \cdots \cdots (2)$$

Now, you can use identity 2, to scale your histogram into density. But your degree of accuracy depend upon the width of class interval. To scale your histogram into density:

- Convert your frequency for each class into probability by dividing total number of observations{ensure that your class interval is sufficiently small}. This represent your $dF(x)$.

- Now, divide your probability from $dx$ is size of class interval and you will get $f(x)$.

You can reverse the procedure to get the frequency from the density and vice-versa.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.