티커 재배정 전후 시점 기준 가격 연속성 유지
코드 Machine Learning for Trading
요약
이 모듈은 티커 기호가 시간에 따라 서로 다른 증권을 가리킬 수 있을 때 과거 S&P 500 가격을 준비하는 방법을 설명합니다. 기업 활동 조정 계수를 적용하고 증권 식별자를 추적하며 정체성 경계를 표시합니다. 수익률 계산은 정체성이 유지되는 증권 구간 안에서만 수행합니다. 따라서 티커가 재배정된 시점에는 잘못된 시장 움직임을 기록하지 않고 수익률을 정의하지 않습니다. 연속된 가격 수준이 필요한 백테스트 패널에는 후속 구간을 별도로 재조정해 이전 수준과 연결하고, 각 티커의 마지막 조정 가격을 공시 종가에 맞춥니다.
가격과 거래량의 곱이 관측된 거래대금과 일치하도록 거래량에도 역수 배율을 적용합니다. 안정적인 배율을 계산하려면 전체 봉 이력이 필요합니다. 날짜별 구간에서 따로 계산하면 구간 경계에 인위적인 불연속이 생길 수 있습니다. 검증 점검은 조정 수익률 계산을 확인하고 연결할 수 없는 경계를 거부합니다. 이 방식은 일관된 레이블과 백테스트 패널을 지원하지만 조정 후 주식 수가 가격 배율에 따라 바뀌므로 변환된 패널의 주당 거래 비용은 실거래 수수료와 다를 수 있습니다.
핵심 아이디어
- 기초 상품의 변경을 찾을 때 티커와 함께 증권 식별자도 사용하세요.
- 증권 정체성이 유지되는 구간 안에서만 조정 수익률을 계산하세요.
- 연속적인 백테스트 가격 수준을 만들려면 재배정된 티커 구간을 연결하고 결과를 데이터에 표시된 최신 종가에 맞추세요.
- 전체 이력으로 조정 배율을 계산해 별도로 불러온 구간도 일치하게 하세요.
- 거래대금을 유지하려면 거래량을 가격의 역수로 조정하되, 주식 수 기준 비용도 변환된다는 점을 고려하세요.
태그
전문
# sp500_price_lineage.py
```py
"""Point-in-time S&P 500 prices that respect security identity boundaries.
``load_sp500_daily_bars`` returns the close as it printed plus ``adj_factor``, a
cumulative price factor. ``close * adj_factor`` is the series in which a split, a
reverse split or a cash dividend is no longer a price move, and it is what
``sp500_equity_option_analytics/02_labels`` builds every label from.
The factor restarts at 1.0 when the exchange reassigns a ticker to a new
security, so a series taken across that boundary carries a jump that is neither a
price move nor a corporate action on the security being held. Both consumers of
these bars need that boundary, and they need it in different shapes: a return
series can leave it null, a backtest price panel cannot. Both shapes are here so
the boundary is defined once.
"""
from __future__ import annotations
import polars as pl
PRICE_COLS = ("open", "high", "low", "close")
def validate_reconciled_returns(frame: pl.DataFrame) -> None:
"""Fail if a return crosses a security identity or violates adjusted-price arithmetic."""
required = {
"sec_id",
"adjusted_close",
"clean_log_return",
"identity_boundary",
}
missing = required - set(frame.columns)
if missing:
raise ValueError(f"Reconciled returns missing columns: {sorted(missing)}")
if frame.filter(pl.col("identity_boundary") & pl.col("clean_log_return").is_not_null()).height:
raise ValueError("A return crosses a security identity boundary")
expected = pl.col("adjusted_close").log().diff().over(["symbol", "sec_id"])
checked = frame.with_columns(expected.alias("expected_log_return"))
violations = checked.filter(
~(
pl.col("clean_log_return").eq_missing(pl.col("expected_log_return"))
| ((pl.col("clean_log_return") - pl.col("expected_log_return")).abs() <= 1e-12)
)
)
if not violations.is_empty():
raise ValueError(f"Adjusted-return identity violations: {violations.height}")
def _validated(prices: pl.DataFrame) -> pl.DataFrame:
"""Check the bar columns both shapes depend on, and mark each identity boundary."""
required = {"timestamp", "symbol", "sec_id", "close", "adj_factor"}
missing = required - set(prices.columns)
if missing:
raise ValueError(f"Underlying bars missing columns: {sorted(missing)}")
if prices.select(pl.struct("timestamp", "symbol").is_duplicated().any()).item():
raise ValueError("Underlying bars contain duplicate timestamp-symbol keys")
if prices["sec_id"].null_count():
raise ValueError("Underlying bars contain null sec_id values")
invalid_levels = prices.filter(
((pl.col("close").is_not_null()) & (pl.col("close") <= 0))
| ((pl.col("adj_factor").is_not_null()) & (pl.col("adj_factor") <= 0))
)
if not invalid_levels.is_empty():
raise ValueError(
f"Underlying bars contain nonpositive price/factor rows: {invalid_levels.height}"
)
return prices.sort(["symbol", "timestamp"]).with_columns(
(pl.col("close") * pl.col("adj_factor")).alias("adjusted_close"),
(
pl.col("sec_id").shift(1).over("symbol").is_not_null()
& (pl.col("sec_id") != pl.col("sec_id").shift(1).over("symbol"))
).alias("identity_boundary"),
)
def reconcile_underlying_log_returns(prices: pl.DataFrame) -> pl.DataFrame:
"""Compute adjusted daily log returns within stable ``sec_id`` segments."""
frame = _validated(prices).with_columns(
pl.col("adjusted_close").log().diff().over(["symbol", "sec_id"]).alias("clean_log_return")
)
validate_reconciled_returns(frame)
return frame
def adjustment_scale(bars: pl.DataFrame) -> pl.DataFrame:
"""Return the per ``(symbol, timestamp)`` multiplier that back-adjusts a printed price.
Two things go into it. ``adj_factor`` removes the corporate action. Then each
``sec_id`` segment after a ticker's first is rescaled to meet the level the
previous segment closed at, so the two securities keep their own returns and
the changeover contributes zero rather than the jump a factor restarting at
1.0 would produce - the treatment ``cme_futures`` gives a contract roll,
applied to a ticker reassignment. A backtest feed reads a level, so unlike
:func:`reconcile_underlying_log_returns` it cannot say that a return does not
exist.
Finally the series is anchored so each ticker's **last** row equals the close
that printed, which is what makes it a back-adjusted price rather than an
index. The level is not cosmetic: this case study sizes positions in whole
shares against a fixed cash budget, and on the factor's own scale AAPL sits
near 4000 instead of near 130, which turns integer rounding into a material
allocation error.
**Pass the complete bar history.** Both halves depend on every segment a
ticker has and on which row is its last, so a scale derived from a
date-filtered frame is a function of the window as well as the session, and
two windows would disagree about the same date. That is not hypothetical:
the holdout path concatenates a validation load and a holdout load to give
the rolling-volatility allocators their burn-in, and a window-dependent
scale puts a fabricated return on the seam - 8x for GE, whose 1-for-8 falls
inside the holdout year.
"""
frame = _validated(bars).with_columns(
pl.col("identity_boundary").cum_sum().over("symbol").alias("_seg")
)
splice = (
frame.group_by("symbol", "_seg")
.agg(
pl.col("adjusted_close").first().alias("_open"),
pl.col("adjusted_close").last().alias("_close"),
)
.sort("symbol", "_seg")
.with_columns((pl.col("_close").shift(1) / pl.col("_open")).over("symbol").alias("_step"))
)
# The null `_step` belongs to a ticker's first segment, where there is nothing
# to splice onto. A null anywhere else means a null close reached the ratio -
# `_validated` tolerates those - and filling it would silently drop the splice
# and carry a wrong offset into every later segment, visible only as P&L.
unspliceable = splice.filter((pl.col("_seg") > 0) & pl.col("_step").is_null())
if not unspliceable.is_empty():
raise ValueError(
"cannot splice a security identity boundary whose adjusted close is null: "
f"{unspliceable.select('symbol', '_seg').rows()}"
)
splice = (
splice.with_columns(pl.col("_step").fill_null(1.0))
.with_columns(pl.col("_step").cum_prod().over("symbol").alias("_splice"))
.select("symbol", "_seg", "_splice")
)
return (
frame.join(splice, on=["symbol", "_seg"], how="inner")
.with_columns((pl.col("adj_factor") * pl.col("_splice")).alias("price_scale"))
.with_columns(
(
pl.col("price_scale")
/ pl.col("price_scale").last().over("symbol", order_by="timestamp")
).alias("price_scale")
)
.select("symbol", "timestamp", "price_scale")
)
def continuous_adjusted_panel(
prices: pl.DataFrame,
*,
scale: pl.DataFrame,
price_cols: tuple[str, ...] = PRICE_COLS,
volume_col: str | None = "volume",
) -> pl.DataFrame:
"""Apply a full-history :func:`adjustment_scale` to a panel that may be one window of it.
Keeping the scale a separate argument is what makes the result a function of
``(symbol, timestamp)`` alone: the caller resolves it once over the whole
series, and every window of that series then agrees about every date it
contains.
``volume_col`` is divided by the same factor its row's price is multiplied
by, so ``price * volume`` stays the dollar volume that printed. Note that the
share **count** moves with the level, so a per-share cost schedule evaluated
on this panel is charged on adjusted share counts and is not comparable to a
live per-share commission.
"""
columns = [column for column in price_cols if column in prices.columns]
if "close" not in columns:
raise ValueError("the adjusted panel requires a close column")
if "price_scale" in prices.columns:
raise ValueError("the panel already carries a price_scale column")
frame = prices.join(scale, on=["symbol", "timestamp"], how="left")
unscaled = frame.filter(pl.col("price_scale").is_null())
if not unscaled.is_empty():
raise ValueError(
"the adjustment scale does not cover every panel row; it must be resolved "
f"over the complete bar history. First uncovered: {unscaled.head(3).rows()}"
)
volume = (
[(pl.col(volume_col) / pl.col("price_scale")).alias(volume_col)]
if _has(frame, volume_col)
else []
)
return frame.with_columns(
[(pl.col(column) * pl.col("price_scale")).alias(column) for column in columns] + volume
).drop("price_scale")
def _has(frame: pl.DataFrame, column: str | None) -> bool:
return bool(column) and column in frame.columns
```출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.