Skip to content
All library documents

Delisted Stock Data, Identifier Changes, and Survivorship Bias

Article Quant Q&A · Author: Rocky the Owl

Summary

The document explains why a current Yahoo Finance or yfinance lookup may fail for companies that have gone bankrupt, merged, or retired their tickers. It treats this as a historical data coverage problem rather than a query syntax issue. Relying on currently available ticker symbols can leave older research with only surviving companies, creating survivorship bias and potentially overstating strategy performance.

For longer historical studies, the answer recommends a curated or commercial archive that retains delisted securities, and suggests checking whether academic access to a research database is available. It also advises mapping permanent identifiers to the ticker valid on each date, constructing the investment universe from historical membership, and preserving downloaded source files so results can be reproduced. These recommendations address data coverage and identifier management; the document does not compare vendors or provide a dataset evaluation. Its discussion is centered on US equity examples, and access, adjustment policies, and identifier coverage vary by source.

Key ideas

  • Current ticker lists and live price services may omit securities that have been delisted or renamed.
  • A survivor-only historical universe can bias backtests by excluding failed or acquired companies.
  • Historical research should use a data source that retains delisted securities and point-in-time identifiers.
  • Map permanent security identifiers to the ticker in use on each date.
  • Build universes from historical membership and preserve source files to support reproducible analysis.

Tags

Full text
# How to obtain yahoo finance ticker data for stocks that are no longer listed (etc. merge, bankruptcy, etc.)?


# How to obtain yahoo finance ticker data for stocks that are no longer listed (etc. merge, bankruptcy, etc.)?












I am trying to use the Python library `yfinance` to obtain stock data for some companies on a list provided to me (just a fairly generic list). However, amongst this list are companies which are no longer listed on the exchanges for a variety of reasons (e.g. bankruptcy, merged, etc.).

Question: How can I obtain the stock price data for these companies during an earlier time period? One earlier post () suggests that I would need to pay to get the data for securities which are no longer listed - is this still true?

Some examples are:

- Lehman Brothers (the former bank)

- Yahoo

- United Technologies (merged with Raytheon)

- and a few others

Attempt: My current use of the module uses the following code (for use within a Google Colab environment):

```
import pandas as pd
import numpy as np

# pip install yfinance
!pip install yfinance

import yfinance as yf

start_date_string = "2006-01-01" # some made up dates
end_date_string = "2006-12-31" 

d = yf.download("[insert string of the ticker]", start=start_date_string, end=end_date_string)
```

However, the tickers for some of these companies don't work so I don't know how to get that data.

## Answer by s teve (score 0)

https://quant.stackexchange.com/a/85855

Yahoo Finance, and therefore `yfinance`, generally does not serve usable OHLCV for names that are gone, including bankruptcies, acquisitions, and ticker retirements. That is a data coverage limitation of the source, not something a different `yf.download` calling convention can fix. Comments on this question already note the same issue with commercial vendors.

What that means for your examples (Lehman, pre-merger UTX, etc.):

- You cannot recover those series from Yahoo. Empty frames or "symbol may be delisted" messages are expected.

- Any universe built only from today's Yahoo tickers is survivors-only. Over long windows that bias is large. Strategies that look great on survivors often fail once dead names are restored.

- Paid or curated archives (CRSP via WRDS, Norgate, Algoseek, etc.) keep point-in-time history with proper identifiers. University finance departments sometimes have WRDS access for students, so it is worth asking before paying personally.

If you need something downloadable for Python without a WRDS seat, look for vendors that ship day-partitioned files that still contain symbols on their last trading day, rather than live Yahoo scrapes. I maintain one such Parquet archive (US stocks, ETFs, and futures, split-adjusted, with delisted symbols retained). Overview: Delisted stock data for survivorship-bias-free backtesting. Disclosure: I am the founder.

Practical workflow regardless of vendor:

- Prefer a permanent ID (FIGI, CIK, or an internal ID) mapped to the ticker as of each date. Do not base research on today's ticker string alone, since names can change like UTX to RTX.

- Build the universe from as-of listing membership, then join prices. Do not start from `yf.Tickers(sp500_today)`.

- Keep raw files immutable so reruns do not drift.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.