Probability Integral Transform for Mapping Weibull Data to Normal Scores
Summary
The document describes an attempt to transform Weibull-distributed observations into standard normal scores using the probability integral transform, with the transformed data intended as neural-network input. The author reports that the resulting histogram does not look normal and contains a gap, and asks whether an alternative procedure could produce a closer match. The author also questions a normalization calculation involving the inverse-CDF density and an approximate sum of one-half.
No response, derivation, plots, or numerical details are included, so the source does not establish what caused the gap or whether the density calculation is valid. In particular, it presents questions rather than a tested transformation or corrective method. It is useful as a statement of a statistical preprocessing problem, but readers should not treat it as evidence that the probability integral transform failed or as guidance on how to normalize the data.
Key ideas
- The author applies a probability integral transform to map Weibull data toward a standard normal distribution.
- The reported transformed histogram has a gap and does not appear normally shaped.
- The author asks about alternative transformations and a normalization calculation involving an inverse CDF.
- The document gives no answer or evidence identifying the cause of the observed histogram.
Tags
Full text
# Probability Integral Transform: Standardisation # Probability Integral Transform: Standardisation I've been applying the probability integral transform as shown here to standardise date for input into a neural network: https://math.stackexchange.com/questions/592076/mapping-cdfs-to-each-other?noredirect=1&lq=1 I want to standardise the data that is wiebull distibuted by mapping it to standard normal: The original histogram is shown below: When I transform to the standard Normal I get the histogram shown below, which doesn't look very Normal. I don't want the big gap I want the transformed histogram to look as much like a standard normal as possible. Is there anything else I can do? Here is the pdf of the inverse CDF it is approximately flat, When I normalize by the data set size the resultant sum is approximately 0.5? I would expect it to be approximately one? Should I be normalising the data by some other value other than the size of the data set?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.