Interpreting Probabilities from Logistic Regression
Summary
This short note raises a basic question about what a probability estimated by logistic regression represents. It contrasts interpreting a prediction as the chance that one particular person succeeds with interpreting it as the expected success rate among many people who share the same observed characteristics.
The author argues that an individual probability would refer to repeated attempts by the same person under identical conditions, while ordinary observational data may instead consist of different individuals with similar measured features. The note introduces this distinction through an interview anecdote and mentions diagrams, but the diagrams and further explanation are absent from the supplied text. It therefore serves as a prompt about the assumptions behind probability interpretation rather than a complete treatment of logistic regression, calibration, or the data-generating process. Quant researchers can use the distinction to be careful about what population a model’s predicted probabilities describe.
Key ideas
- A logistic regression probability needs interpretation in relation to the model’s data-generating assumptions.
- An individual’s repeated-trial success probability differs conceptually from a success rate across people with similar observed features.
- Observational data often compare different individuals rather than repeated outcomes for the same individual.
- The supplied note raises this issue but omits the diagrams and does not fully resolve it.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.