Why Bar High-Low Position Does Not Give a Hitting Probability
Summary
The document asks why a simulated price series with independent, normally distributed increments appears to contradict a proposed probability for reaching an upper barrier before a lower one. It compares a bar’s close-to-high fraction with whether the bar’s high or low was reached first, then sorts and groups observations to estimate conditional frequencies.
The central issue is that a bar’s close, high, and low do not reveal the order in which its extremes occurred. A close-to-high ratio is a location within the observed range, not by itself a probability of which barrier will be hit first. The hitting-probability formula cited in the question applies to a specified process and barrier setup; the described sampling, bar construction, and rolling calculation do not directly estimate that quantity. The document presents the simulation and the question but contains no answer or reported numerical results, so it does not establish a corrected estimate.
Key ideas
- A bar’s close-to-high fraction describes the close’s position within the observed range, not a barrier-hitting probability.
- Bar high and low values alone do not show which extreme occurred first.
- A first-passage probability depends on the process and the definition of the two barriers.
- The described rolling statistic does not directly test the stated theoretical hitting probability.
- The document provides a question and simulation procedure but no resolution or numerical findings.
Tags
Full text
# Inconsistency between simulation and the probability of a "stock" hitting take profit before stop loss
# Inconsistency between simulation and the probability of a "stock" hitting take profit before stop loss
Let's assume a stock at time $t$ is worth $X(t)$. If the returns of $X(t)$ are i.i.d. and normally distributed,the probability of $X(t)$ hitting a value $H>X(t)$ before $L<X(t)$ is $\frac{H-X(t)}{H-L}$ (not considering exchange fees and the cost of shorting).
Now, I ran a python simulation on a stochastic process with normal i.i.d returns (price diffs) and obtained results that seem inconsistent with my assumptions about the probabilities mentioned above. I took the folling steps in the simulation:
- obtain a simulated time series of prices using the "fake_stock" function.
- upsample the price data to 5 minute bars.
- obtain the high, low, and close of each 5 minute bar.
- calculate the close to high percentage (ch) of each 5 minute bar. $ch = \frac{high-close}{high-low}, 0 <=ch <= 1 $. In theory, $ch$ denotes the probability of $X(t)$ hitting the high after the low, correct? This is where the inconsistency lies.
- run 'high_after_low' to see if $H$ was hit after $L$.
- based on step five, I obtain a dataframe of close_to_highs and a boolean showing if they hit their respective highs before lows.
```
def fake_stock():
'''returns a time series of prices based on gaussian returns'''
dti = pd.date_range(start="2018-01-01 17:00", end="2019-01-01 16:00", freq="S")
returns = np.random.normal(0, 1, len(dti))
df = pd.DataFrame(index=dti, columns=["price"])
df["price"] = returns
df = df.between_time("17:00", "16:00")
df["price"] = df["price"].cumsum()
df["price"] = df["price"] - df["price"].min()
return df
def upsample(df, freq="5T"):
'''upsample second bars to minutes. obtain close, high, and low of each bar'''
upsampled_df = pd.DataFrame()
upsampled_df.loc[:, "high"] = df["price"].resample(freq, label="left", closed="left").max()
upsampled_df.loc[:, "low"] = df["price"].resample(freq, label="left", closed="left").min()
upsampled_df.loc[:, "close"] = df["price"].resample(freq, label="left", closed="left").last()
upsampled_df.loc[:, "close_to_high"] = (upsampled_df["high"]-upsampled_df["close"])/(upsampled_df["high"]-upsampled_df["low])
return upsampled
def high_after_low_apply(row, minute_df):
'''returns True if high was reached after low, False otherwise'''
minute_df_after = minute_df.loc[row.end_idx:]
first_highs = (minute_df_after.ge(row.high))
first_lows = (minute_df_after.le(row.low))
if ((len(first_highs) == 0) & (len(first_lows) == 0)):
return None
elif (len(first_highs) == 0):
return True
elif (len(first_lows) == 0):
return False
return first_highs.idxmax() > first_lows.idxmax()
```
Now, we are ready to run the simulation.
```
df = fake_stock()
upsampled_df = upsample(df)
upsampled_df.loc[:,"end_idx"]=upsampled_df.index+pd.Timedelta(minutes=5)
upsampled_df.loc[:,"high_after_low"] = upsampled_df.apply(high_after_low_apply, axis=1, args=(df,))
upsampled_df = pd.sort_values(upsampled_df, by="close_to_high")
```
Now we have the upsampled_df dataframe sorted by close_to_high values. Let's drop close_to_high values that are 0 or 1.
```
upsampled_df = upsampled_df.loc[(upsampled_df["high_after_low"]!=0)&(upsampled_df["high_after_low"]!=1)]
```
Now, I can calculate the historic probabilities of hitting high after low using a rolling window.
```
probs = (upsampled_df["close_to_high"]*upsampled_df["high_after_low]).rolling(1000).sum()/upsampled_df["close_to_high"].rolling(1000).sum()
```
when I compare probs to the average values of high_after_low, I get inconsistent results. Namely, for small close_to_high values the probability of hitting the high after the low is larger than theory suggests and vice versa.
What is the explanation for this inconsistency?Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.