Every quantitative shop guards against the idea that looked good by luck. Far fewer notice the mirror hazard — an idea that looked dead because the test could never have seen it alive — and unlike a false positive, a false negative has nothing downstream to catch it.
Defences against the false positive
Test enough ideas and some will look brilliant by luck alone. This is the central problem of quantitative research and the profession has built a real defence against it: hold out data the idea was not built on, penalise the evidence bar for the number of ideas tried, deflate the headline statistic for the size of the search, track the survivors forward in real time, and let somebody adversarial try to break the result. Harvey, Liu and Zhu made the argument unignorable in 2016; Bailey and Lopez de Prado made the search-size correction concrete. Any shop that skips all of this is not doing research.
Every one of those defences catches a false positive: out-of-sample testing catches a fluke that will not repeat, deflation a fluke found by searching, forward tracking a fluke that repeats for a while and then does not, adversarial review a fluke created by a bug. The entire apparatus points in one direction.
The false negative
The opposite failure begins with a real effect and a test built in a way that could never have resolved something that small. It returns a negative, and the negative is recorded, dated, indexed, and from that day forward cited — correctly, by the shop's own procedures — as a reason not to look there again.
Nothing catches that. There is no out-of-sample check for a null, no deflation and no forward tracking, and nobody re-runs a rejection. Rejections are the cheap, virtuous, disciplined outcome, what a rigorous shop is supposed to produce in bulk, and each one quietly removes a piece of territory from the map. A shop with strong positive-result hygiene and no negative-result hygiene will, over a few years, convince itself that a great deal of ground is barren without ever having looked at it properly.
An underpowered null is not evidence of absence. It is absence of evidence wearing the same label.
The smallest detectable effect
Every statistical test has a smallest effect it could detect: the evidence bar you have committed to, multiplied by the measurement noise of your design. Compute that number, take from published work how big the effect actually is for things of this kind, and compare the two. The check is simple enough that there is no excuse for skipping it, which is precisely why skipping it is so common.
A design that appears entirely respectable can fail that comparison. Suppose you want to know whether some characteristic of a company predicts which shares do better than others, and you measure it as a correlation between the characteristic and next period's returns across the names you hold, over two dozen names and forty observation dates — a real universe and a serious span of work.
What that design could resolve
| Quantity | Value |
|---|---|
| The design, as it would be frozen | 24 names × 40 dates |
| Smallest correlation it can detect, at the conventional two-standard-error bar | about 0.07 |
| Smallest correlation it can detect once the bar is raised for search size | about 0.10 |
| Correlations that published stock characteristics actually have | about 0.02 to 0.06 |
| Verdict | needs an effect larger than any on record |
The detection floor is the evidence bar divided by the square root of the number of independent observations in the panel — here the number of names less one, multiplied by the number of dates. Harvey, Liu and Zhu argue that the conventional bar is far too low once you count how many ideas the profession has already tried; raising it, as they recommend, widens this gap rather than closing it.
That test cannot succeed, not because the idea is bad or the market efficient but because of arithmetic that was knowable before a single price was loaded. Run it anyway and you will get a negative result, and the negative result will be true of the test rather than of the world.
The repair is cheap, and nothing about the idea, the estimator or the evidence bar has to change: widen the cross-section to forty-five names, run it across a year of dates, and the floor falls to roughly 0.03, still at the raised bar — inside the documented range, so the test can now return a meaningful answer in either direction. Breadth is the cheapest power there is, bought with a bigger universe rather than a better model. The design that could not have worked and the design that can differ only in how many things are being ranked against each other at each point in time.
Detection floors larger than the mechanism
A sharper form of the same mistake, common once you know to look for it, is a smallest detectable effect larger than the entire quantity the study set out to measure.
Our own files hold two de-identified examples. In the first, the question was whether a change to how a portfolio trades would reduce its trading costs; the study's implicit detection floor was larger than the whole cost it measured, so eliminating one hundred percent of that cost — a perfect, impossible outcome — would still have fallen short of the evidence bar. In the second, the question was how much return a rebalancing discipline harvests from volatility, a quantity to which theory gives a ceiling, and the detection floor sat above that ceiling.
In both cases the study was competently executed, honestly reported and completely uninformative. The verdict was determined at the moment the design was frozen, and no amount of data, no cleverer estimator and no better market conditions could have changed it.
Headroom
A further check, which we now run first because it costs a single line of arithmetic, applies whenever the thing being proposed is an estimator — a model, a smoothing, a structural surface, a filter — whose purpose is to reduce noise in some quantity.
Take the noise the mechanism exists to remove and subtract the error the mechanism itself introduces, since it is an approximation of a reality that does not match it exactly. If the difference is negative, the mechanism is worse than the problem, and no sample size rescues it: you are replacing sampling noise with a larger, systematic misfit.
We killed a proposal at the design stage on exactly this, using two numbers its own preliminary work had already published: the model's own shape error exceeded, by a wide margin, the sampling noise it existed to remove. A synthetic test confirmed the diagnosis, since even in a simulated world where the model's assumed structure was exactly true the fitted version still lost to a plain textbook alternative. Both numbers had been sitting in the file for weeks, never placed next to each other.
Audit of our own negative results
An argument of this kind is worth little without its own record attached, so we went through every negative result our register holds — every idea we had tested and rejected — and computed, retrospectively, the smallest effect each test could have seen, against a plausible effect for that class of idea fixed in advance so we could not tune it afterwards. Ties were resolved in favour of the existing verdict, so the count below is a floor.
- The closed-vein population was 86, of which 81 were audited for power (the remaining five were already labelled measurement-limited and were assessed separately). 66 of the 81 — 81% — held up: the design could have seen a realistic effect and did not, so those are real evidence and the ground really is barren.
- 14 — 17% — could not have seen a plausible effect at all, and were reclassified from settled to unresolved. One more was indeterminate: the record was too thin to compute power from. The honest qualifier matters as much as the headline: 9 of those 14 carried their substantive conclusion on a second, adequately-powered bar, so for those the reclassification is a labelling correction rather than a reopened question. Five are genuinely open again. The number worth quoting is five of eighty-one, not fourteen.
- Six of the fourteen failed in a way that no sample size can fix: the bar demanded an effect larger than the entire quantity the study set out to measure. One needed 4.41% per year against a cost drag of 4.21% it was built to address; another needed 337 basis points a year against its own theoretical gamma of 314; a third needed 4.43 basis points a day against its own measured concession of 2.71. Those were unpassable on the day they were frozen, whatever the market did.
- Not one record contained a power calculation at the time it was frozen, and more than a dozen described themselves as well-powered with no arithmetic behind the phrase. That is the finding that stings, and it is why the check is now mandatory before any design can be frozen rather than optional afterwards.
Our adopted rule has three numbers in it, and a design cannot be frozen without all three written down: the smallest effect this test can detect; a plausible effect for this class of idea, with the published source it came from; and the size of the mechanism the study itself is measuring. If the first exceeds the third, the test cannot pass even when the theory is exactly right, and it does not get run.
Design shapes that are almost always underpowered
Auditing found the failures clustered, which is useful, because the hazard can then be spotted from the shape of a proposal before any arithmetic is done at all.
- Monthly predictive regressions over about twenty years almost never resolve: there are only a couple of hundred observations, and the explanatory power needed to clear a serious bar is roughly double that of the best-documented economic predictors.
- Event studies with modest event counts: below roughly a hundred and thirty events a design of this kind cannot resolve an effect the size of the best-known post-announcement drift, and many event ideas have fewer events than that available in the whole history.
- Narrow cross-sections of a few dozen names lack the breadth for a cross-sectional test, no matter how long you run it, because the noise is dominated by how many things you are ranking against each other at each point in time.
Designs that came out of the audit best were the ones whose headline check was a rearrangement test rather than a ratio: shuffle the labels many times, rebuild the result on the shuffled data, and ask where the real result sits in that distribution. Such a test is scale free, makes almost no distributional assumption, and is typically far better powered than the ratio it replaces.
Publication of negative results
The general reason for publishing them is old. Robert Rosenthal named the file-drawer problem in 1979: when only successes are published, the published record is a distorted sample of what was tried, and everyone downstream re-tests the same dead ideas because nobody wrote down that they were dead. Finance has an unusually severe version, since a firm's negative results are also its competitors' saved effort, so nobody shares them, and the consequence is an entire industry repeatedly rediscovering the same nothing.
A more specific reason is sharper. Publishing your negative results is a good start; publishing an audit of them — which of your own rejections you now think you should not trust — is the part that costs something, and the part that tells a reader whether the rest of the record was built honestly. A firm that will tell you which of its conclusions are shaky is telling you something real about the ones it stands behind.
Limits of the claim
- None of this licenses reviving dead ideas, since an underpowered test is a reason to say unresolved rather than promising. Reviving one requires re-deriving its original numbers first, from the original artefacts, because a revival built on a misremembered result is worse than no revival at all; we have caught two such misreadings, one of which had a sign the wrong way round.
- Nor is it an argument for weaker evidence bars: the fix is better designs, more breadth, longer spans and scale-free tests rather than a lower bar, because lowering the bar trades a false-negative problem for a false-positive one, and the false-positive one costs money.
- None of the statistics is novel. Power analysis is taught in every introductory course and is mandatory in clinical trials; the observation is that quantitative finance, which is otherwise obsessed with statistical hygiene, has almost entirely skipped it, and that the reason is structural, because power protects against the error nobody is embarrassed by.
- The job is not finished: the audit is retrospective and the mandatory check is new. Ask us again in a year.
Further reading
- Robert Rosenthal, “The file drawer problem and tolerance for null results”, Psychological Bulletin, 1979.
- Campbell R. Harvey, Yan Liu and Heqing Zhu, “. . . and the cross-section of expected returns”, Review of Financial Studies, 2016.
- David H. Bailey and Marcos Lopez de Prado, “The deflated Sharpe ratio”, Journal of Portfolio Management, 2014.
- Jacob Cohen, Statistical Power Analysis for the Behavioral Sciences, 1988 — the standard reference, and the one this corner of finance never opened.
What we have tested, and what we rejected
Our validation policy and a plain-language record of the ideas that did not survive it.
Research Integrity More Research