How Grouping Choices Affect Regression on Panel Data
Summary
The document describes a panel regression problem in which a large dataset is reduced to fewer observations by selecting a shorter time window and aggregating firms into groups. The author reports that estimated significance changes depending on how firms are sorted before forming groups: sorting by the dependent variable produces less significant results than a default ordering. The underlying concern is how to aggregate data without making regression conclusions depend on arbitrary grouping choices.
No solution or statistical analysis is provided, so the document does not establish which averaging method is valid. It highlights a methodological risk: sorting on the outcome before aggregation can affect the resulting group means and regression estimates. The example is limited to the author’s described setup and offers no robustness checks, alternative estimators, or evidence about the source of the significance change. It motivates careful aggregation design and sensitivity analysis rather than prescribing a specific procedure.
Key ideas
- Reducing panel data by grouping firms and averaging their observations can change regression results.
- Sorting firms by the dependent variable before aggregation may affect estimated significance.
- The document raises an aggregation-method question but does not supply a definitive method.
- Researchers should treat sorting and grouping choices as consequential modeling decisions.
Tags
Full text
# How to properly take averages to reduce data in regression/panel data analysis # How to properly take averages to reduce data in regression/panel data analysis I'm trying to do a regression on my panel data. Say I have T=3500 days of data and N=125 firms. Since Matlab get's major memory issues (which I try to prevent by the usual solutions as seen on the Mathworks site), because my panel was too big, I decided to only look at 700 days of data and to take averages of a number of firms to get N=5. I discovered that the way I take these averages/medians heavily influences my regression results. If I first sort the firms on the size of the dependent variable and then make 5 groups to average them, I'll get less significant results than when I used another (default) sorting (almost alphabetic). So what is the way to go in these kind of situations?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.