What is survivorship bias in financial data?
Survivorship bias is what happens when a dataset contains only the companies that still exist. Firms that were delisted, acquired, or went bankrupt are missing, so any historical study run on that universe silently excludes most of the ways an investment can go wrong.
A universe built from today's index membership is the common case. Screening "the S&P 500 over the last twenty years" against today's constituents does not test a strategy on the S&P 500 of 2006 — it tests it on the companies that were good enough to still be in the index in 2026. The failures were removed from the sample by the very outcome the study is trying to measure.
The effect is large and always flattering. Strategies look more profitable, drawdowns look shallower, and bankruptcy risk approaches zero, because bankrupt companies are not in the file. Nothing in the output signals this; the numbers are internally consistent and completely wrong.
Avoiding it means keeping the dead companies — retaining entities after they stop filing, and reconstructing index membership as it stood on each historical date rather than as it stands now. That is expensive to build and impossible to retrofit, which is why it is worth asking a provider about before anything else.