Most quantitative work below the institutional tier is tested against data that did not exist on the decision date. Two errors do most of the damage. The first is survivorship: a universe built from today's index members excludes every company that was delisted, acquired, or demoted along the way, and those are disproportionately the ones that performed badly. The second is restatement: fundamental data that has been revised after the fact is used as if the original figure had never been published.
Both errors flatter the result in the same direction, and neither is visible in the output. A clean-looking backtest with a high Sharpe ratio is exactly what these errors produce. An examiner does not need to understand factor models to ask the question that exposes them: what data was available on the date this decision was made, and can you show it?
The fix is not a better model. It is a data layer that preserves as-reported values with their original publication timestamps and keeps monthly snapshots of index membership including the names that later disappeared. That layer is expensive to build and boring to maintain, which is why it is rarely maintained below the enterprise price tier and why Marnello describes it before describing any factor. Our demo builds, generated on free data, disclose the survivorship problem on every page rather than hide it; the build-status page states which parts of that layer exist today.