From the desk · Research design

The negative result nobody checks

Every quantitative shop guards against the idea that looked good by luck. Far fewer notice the mirror hazard — an idea that looked dead because the test could never have seen it alive — and unlike a false positive, a false negative has nothing downstream to catch it.

What the test could have seen Same idea, same data source, two designs — only one of them can return a meaningful negative EFFECTS THAT ACTUALLY OCCUR ADEQUATE DESIGN everything from here rightwards is visible the band of real effects is inside it A negative result here means the effect is absent. UNDERPOWERED DESIGN visible only from here rightwards nothing that could plausibly be true is in range A negative result here means nothing at all — and gets filed as if it did. size of the effect being looked for a false positive is caught downstream · a false negative is caught by nothing
The gold band is the range of effect sizes that signals of a given class actually have, taken from published evidence. The pink region is what each design is capable of resolving. In the lower design the two never overlap, so the answer was fixed before the data arrived.

Defences against the false positive

Test enough ideas and some will look brilliant by luck alone. This is the central problem of quantitative research and the profession has built a real defence against it: hold out data the idea was not built on, penalise the evidence bar for the number of ideas tried, deflate the headline statistic for the size of the search, track the survivors forward in real time, and let somebody adversarial try to break the result. Harvey, Liu and Zhu made the argument unignorable in 2016; Bailey and Lopez de Prado made the search-size correction concrete. Any shop that skips all of this is not doing research.

Every one of those defences catches a false positive: out-of-sample testing catches a fluke that will not repeat, deflation a fluke found by searching, forward tracking a fluke that repeats for a while and then does not, adversarial review a fluke created by a bug. The entire apparatus points in one direction.

The false negative

The opposite failure begins with a real effect and a test built in a way that could never have resolved something that small. It returns a negative, and the negative is recorded, dated, indexed, and from that day forward cited — correctly, by the shop's own procedures — as a reason not to look there again.

Nothing catches that. There is no out-of-sample check for a null, no deflation and no forward tracking, and nobody re-runs a rejection. Rejections are the cheap, virtuous, disciplined outcome, what a rigorous shop is supposed to produce in bulk, and each one quietly removes a piece of territory from the map. A shop with strong positive-result hygiene and no negative-result hygiene will, over a few years, convince itself that a great deal of ground is barren without ever having looked at it properly.

An underpowered null is not evidence of absence. It is absence of evidence wearing the same label.

The smallest detectable effect

Every statistical test has a smallest effect it could detect: the evidence bar you have committed to, multiplied by the measurement noise of your design. Compute that number, take from published work how big the effect actually is for things of this kind, and compare the two. The check is simple enough that there is no excuse for skipping it, which is precisely why skipping it is so common.

A design that appears entirely respectable can fail that comparison. Suppose you want to know whether some characteristic of a company predicts which shares do better than others, and you measure it as a correlation between the characteristic and next period's returns across the names you hold, over two dozen names and forty observation dates — a real universe and a serious span of work.

What that design could resolve

QuantityValue
The design, as it would be frozen24 names × 40 dates
Smallest correlation it can detect, at the conventional two-standard-error barabout 0.07
Smallest correlation it can detect once the bar is raised for search sizeabout 0.10
Correlations that published stock characteristics actually haveabout 0.02 to 0.06
Verdictneeds an effect larger than any on record

The detection floor is the evidence bar divided by the square root of the number of independent observations in the panel — here the number of names less one, multiplied by the number of dates. Harvey, Liu and Zhu argue that the conventional bar is far too low once you count how many ideas the profession has already tried; raising it, as they recommend, widens this gap rather than closing it.

That test cannot succeed, not because the idea is bad or the market efficient but because of arithmetic that was knowable before a single price was loaded. Run it anyway and you will get a negative result, and the negative result will be true of the test rather than of the world.

The repair is cheap, and nothing about the idea, the estimator or the evidence bar has to change: widen the cross-section to forty-five names, run it across a year of dates, and the floor falls to roughly 0.03, still at the raised bar — inside the documented range, so the test can now return a meaningful answer in either direction. Breadth is the cheapest power there is, bought with a bigger universe rather than a better model. The design that could not have worked and the design that can differ only in how many things are being ranked against each other at each point in time.

Detection floors larger than the mechanism

A sharper form of the same mistake, common once you know to look for it, is a smallest detectable effect larger than the entire quantity the study set out to measure.

Our own files hold two de-identified examples. In the first, the question was whether a change to how a portfolio trades would reduce its trading costs; the study's implicit detection floor was larger than the whole cost it measured, so eliminating one hundred percent of that cost — a perfect, impossible outcome — would still have fallen short of the evidence bar. In the second, the question was how much return a rebalancing discipline harvests from volatility, a quantity to which theory gives a ceiling, and the detection floor sat above that ceiling.

In both cases the study was competently executed, honestly reported and completely uninformative. The verdict was determined at the moment the design was frozen, and no amount of data, no cleverer estimator and no better market conditions could have changed it.

Headroom

A further check, which we now run first because it costs a single line of arithmetic, applies whenever the thing being proposed is an estimator — a model, a smoothing, a structural surface, a filter — whose purpose is to reduce noise in some quantity.

Take the noise the mechanism exists to remove and subtract the error the mechanism itself introduces, since it is an approximation of a reality that does not match it exactly. If the difference is negative, the mechanism is worse than the problem, and no sample size rescues it: you are replacing sampling noise with a larger, systematic misfit.

We killed a proposal at the design stage on exactly this, using two numbers its own preliminary work had already published: the model's own shape error exceeded, by a wide margin, the sampling noise it existed to remove. A synthetic test confirmed the diagnosis, since even in a simulated world where the model's assumed structure was exactly true the fitted version still lost to a plain textbook alternative. Both numbers had been sitting in the file for weeks, never placed next to each other.

The best experiment you run this quarter may be the one you cancel on the strength of two numbers you already had.

Audit of our own negative results

An argument of this kind is worth little without its own record attached, so we went through every negative result our register holds — every idea we had tested and rejected — and computed, retrospectively, the smallest effect each test could have seen, against a plausible effect for that class of idea fixed in advance so we could not tune it afterwards. Ties were resolved in favour of the existing verdict, so the count below is a floor.

Our adopted rule has three numbers in it, and a design cannot be frozen without all three written down: the smallest effect this test can detect; a plausible effect for this class of idea, with the published source it came from; and the size of the mechanism the study itself is measuring. If the first exceeds the third, the test cannot pass even when the theory is exactly right, and it does not get run.

Design shapes that are almost always underpowered

Auditing found the failures clustered, which is useful, because the hazard can then be spotted from the shape of a proposal before any arithmetic is done at all.

Designs that came out of the audit best were the ones whose headline check was a rearrangement test rather than a ratio: shuffle the labels many times, rebuild the result on the shuffled data, and ask where the real result sits in that distribution. Such a test is scale free, makes almost no distributional assumption, and is typically far better powered than the ratio it replaces.

Publication of negative results

The general reason for publishing them is old. Robert Rosenthal named the file-drawer problem in 1979: when only successes are published, the published record is a distorted sample of what was tried, and everyone downstream re-tests the same dead ideas because nobody wrote down that they were dead. Finance has an unusually severe version, since a firm's negative results are also its competitors' saved effort, so nobody shares them, and the consequence is an entire industry repeatedly rediscovering the same nothing.

A more specific reason is sharper. Publishing your negative results is a good start; publishing an audit of them — which of your own rejections you now think you should not trust — is the part that costs something, and the part that tells a reader whether the rest of the record was built honestly. A firm that will tell you which of its conclusions are shaky is telling you something real about the ones it stands behind.

Limits of the claim

Further reading

What we have tested, and what we rejected

Our validation policy and a plain-language record of the ideas that did not survive it.

Research Integrity More Research