Handling Incomplete Windows in Moving Average Calculations
Summary
The document considers how to calculate a long moving average when a price history contains fewer observations than the target window. A strict approach omits early rows until the full window is available; a dynamic approach computes an average from whatever history exists, producing values immediately but with changing sample sizes. The question asks whether one approach is more accurate and whether comparisons with an index or different bar frequencies can guide the choice.
The response describes using the available observations up to the target window as a pragmatic way to keep a routine returning values, while signaling that early values are incomplete and can be disregarded. It treats the choice as a programming interface concern rather than a settled quantitative-finance rule, and does not offer a statistical accuracy test or evidence that dynamic windows are more representative. Researchers should distinguish partial-window values from fully formed moving averages when using them in analysis.
Key ideas
- A strict moving average starts only after the full lookback window is present.
- A dynamic moving average uses fewer observations at the start of a price history.
- Partial-window values can be marked as incomplete so downstream analysis can ignore them.
- The source presents dynamic windows as a pragmatic programming choice, not a definitive quant method.
- Comparisons with an index or across sampling frequencies are raised as questions but not evaluated.
Tags
Full text
# How a chose between strict vs dynamic measurement of moving average
# How a chose between strict vs dynamic measurement of moving average
I am building a simple stock model with Pandas and part of that is calculating moving averages.
I would like to understand what are and how to measure accuracy implications of using strict vs dynamic time window size when calculating long-term moving averages.
Here is an example what I mean:
I have a csv file with TLSA stock history between 1.1.2000 - 28.2.2017
```
df = pd.read_csv('TSLA.csv', index_col=0, parse_dates=True)
# add new col to dataframe consisting of a 100 day moving avarage
df['100ma'] = df['Adj Close'].rolling(window=100, min_periods=0).mean()
```
min_periods=0 allows me to calculate moving average for the first 100 days where don't have 100 days using first 0, 1, 2, 3... moving average.
My alternative to this is call `df.dropna()` it would drop first 100 rows that don't have a 100-day moving average. Downside I would have exact average but fewer data points.
I am building a tool using a large number of moving averages. What would be a correct method to figure which data frame to use? Correlation to a moving average of a representing stock index? Also which technique is usually applied in the quant world?
Intuitively I think dynamic version is more representing but is that true also with minute and shorter timeframe data?
## Answer by nbbo2 (score 1, accepted)
https://quant.stackexchange.com/a/32998
What I do in my code is to calculate the MA for "whatever number of days is available or 100 whichever is smaller". But that is just a pragmatic hack to make sure my program produces some value instead of dying of unexplained causes, and matches the "correct" values once T>= 100. I would not claim that this is the definitive solution by any means. The routine should probably return a Warning Code that the full MA cannot yet be computed, so the caller can disregard those fake values if he wants. Of course before the date of TSLA IPO I return a Hard Error code saying "can't do it". This is programming style rather than Quant Finance.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.