A retrospective census of eighty-one settled negative results in our register found fourteen whose design could not have detected a plausible effect, five of them because the evidence bar demanded a larger effect than the entire quantity the study set out to measure. Whether a bar is reachable is arithmetic, knowable on the day it is frozen. When the smallest effect a design can resolve exceeds anything the mechanism under test could produce, the study returns a negative result whatever the truth, which is then indexed as evidence the question is closed.
Low power leaves a real effect a poor chance of detection; a detection floor above the mechanism's ceiling leaves it none.
The detection floor
A test's minimum detectable effect, the smallest true effect it could report as a finding, is the evidence bar in standard errors multiplied by the standard error of the estimator. At the Harvey, Liu and Zhu (2016) threshold of |t| > 3, no true effect below 3 × SE can be reported as a positive finding, however carefully it is estimated.
Retrospectively, the standard error is most reliably backed out of the record rather than rebuilt from a formula. Where a completed study reports an effect E and its t-statistic t, SE = |E| / |t| and
Back-out estimator for the detection floor
MDE = bar × SE = |E| × bar / |t|
The estimate inherits the autocorrelation, overlapping-window and clustering corrections already applied to the reported t-statistic, which no closed-form substitute does, and was computable for the large majority of the units audited below. Where no effect-and-t pair was recorded, the audit used the standard forms: the cross-sectional information coefficient as bar / √((N−1)T), the Lo (2002) standard error for a Sharpe level, bar × σevent / √nevents for an event study, and the null distribution's own upper percentile for a permutation bar.
The design that prompted this work ranked twenty-four names against one another on a flow-derived characteristic, with a quorum of forty observation dates, judged at |t| > 3. Its detection floor is 3 / √(23 × 40) = 0.099 in correlation units, against published cross-sectional characteristic correlations of about 0.02 to 0.06, so it could have confirmed only an effect roughly twice the strongest on record and would have returned a null irrespective of the market. That null would then have been cited as settled against every later proposal in the same territory, although the calculation that predicts it is one line long.
Four quantities to declare at freeze
A preregistration should state four quantities before it is frozen, in this order.
- The minimum detectable effect at the frozen bar, with the formula and arithmetic that produced it, rather than an assertion that the design is adequately powered.
- A plausible effect for this class of strategy, with its published source, fixed in advance from a common reference table so that it cannot be adjusted after the result is known.
- The size of the mechanism the study itself measures, such as the cost being removed, the premium being harvested or the concession being recovered: the ceiling on what a perfect version of the idea could deliver.
- Headroom, wherever the proposed mechanism is an estimator (defined below).
The binding comparison is the first quantity against the third. If the detection floor exceeds the size of the mechanism, the design cannot pass even when the theory is exactly correct at full strength, and it should not be registered.
Floors above the mechanism
Six studies in the census failed that comparison. Each had measured, or theory had bounded, the quantity being pursued, and the frozen bar demanded more of it than existed.
At the lowest ratio, a study of whether a slower adjustment schedule would reduce the trading costs of a market-neutral book was frozen with a detection floor of 4.41 percent a year against a measured cost drag of 4.21 percent a year. Removing every basis point of cost the book incurred would have produced t = 2.87, short of the bar, so the competent null the study reported was fixed when the design was written down.
The other five fail the same way at the wider margins shown in the chart, and in every case the ceiling was already in the study's own record.
A bar the mechanism cannot reach at full strength tests the arithmetic of the design rather than the mechanism, and its verdict is fixed before any data arrives.
Headroom
When the proposal is an estimator meant to reduce noise, such as a structural model, a shrinkage, a fitted surface or a filter, it buys lower sampling error at the price of approximation error, since it imposes a structure the world does not exactly satisfy. Define
Headroom
headroom = noise the mechanism removes − the mechanism's own approximation error
Negative headroom proves impossibility. Misspecification does not shrink with the sample, so no quantity of data, refinement of the fit or change of bar can rescue the proposal.
One proposal was rejected at the design stage on this check alone: a two-parameter structural correlation surface to replace ordinary sample correlations between volatility contracts of different maturities. Its approximation error over the full sample, reported weeks earlier by its own preliminary study, was 0.0113; the sampling noise it was meant to remove, at the sixty-observation window planned as the primary specification, was 0.00827. Its approximation error exceeded the noise it was meant to remove by 37 percent before a single parameter was estimated.
In a synthetic control where the assumed structure is exactly true, the fitted estimator still lost to a standard shrinkage estimator by 0.0193 and to raw sample correlation by 0.0075, on a loss scale where lower is better, so the hypothesis lost even in its best case.
Census of settled nulls
We computed the detection floor for every settled negative result we hold, with plausible effect sizes fixed in advance in a nine-row reference table citing a source per row. Classification ties were resolved in favour of the existing verdict, so every count of a failure is an undercount.
Retrospective power census of the settled nulls
| Class | Count | Share of audited |
|---|---|---|
| Settled negative results | 86 | — |
| Audited for power | 81 | — |
| Could have seen a plausible effect and did not | 66 | 81% |
| of which carried by a wrong-signed estimate | 21 | 26% |
| Could not have seen a plausible effect at all | 14 | 17% |
| Record too thin to compute power from | 1 | 1% |
| Already labelled measurement-limited, audited separately | 5 | — |
| Records containing a power calculation at freeze time | 0 of 152 | — |
Fifteen register entries described themselves as well-powered with no arithmetic behind the phrase. Nine of the fourteen underpowered units carry their substantive conclusion on a second, adequately-powered bar (a permutation control, a deterministic threshold count or a sign-decisive point estimate), so for those the reclassification corrects a label without reopening the question.
Four fifths of the settled negatives survived, and 21 of the 66 survivors are carried by an estimate significantly of the wrong sign, which is stronger evidence than a failure to reject. The damage is concentrated and identifiable unit by unit.
Detection floors by design
The failures clustered by design shape rather than by subject matter, so the hazard can be recognised from the outline of a proposal before any arithmetic is done.
Detection floors at the conventional strict bar
| Design | Detection floor | Plausible effect |
|---|---|---|
| Monthly predictive regression, 20 years (n = 244) | correlation 0.192, R² 3.7% | best macro predictors, R² 0.25–2% |
| Monthly Sharpe, 10 years (n = 120) | annualised Sharpe 0.95 | single-asset-class sleeve, 0.30–0.80 |
| Daily Sharpe, 20 years (n = 5,040) | annualised Sharpe 0.67 | as above |
| Event study, 48 events at 12% event volatility | abnormal return 5.2% | corporate-event effects, 1–5% |
| Event study, 237 events | abnormal return 2.3% | as above |
| Cross-section, 24 names × 40 dates | correlation 0.099 | characteristic correlations, 0.02–0.06 |
| Cross-section, 45 names × 245 dates | correlation 0.029 | as above |
Sharpe floors use the Lo (2002) standard error; the cross-sectional floors use 3 / √((N−1)T); the event-study floors use 3σ/√n at a sixty-three-day idiosyncratic volatility of 12%. Plausible ranges are the reference table's, with Welch and Goyal (2008) supplying the predictive-regression row.
The two cross-sectional designs share an estimator, a bar and a class of signal, and differ only in how many names are ranked and over how many dates. The narrower cannot resolve any documented effect; the wider resolves most of the plausible range. Breadth, bought with a wider universe and a longer panel rather than a better model, moves the detection floor faster than anything else available.
Designs whose declared primary test was a permutation test rather than a ratio came through the census in the best condition. Such a test is scale free, imposes almost no distributional assumption, and in this register was consistently better powered than the t-statistic it accompanied.
Limitations
- Plausible effect sizes are judgement calls. Nine reference rows covered most units and roughly twenty needed a bespoke figure, each stated and flagged and, where possible, anchored on the unit's own measured quantities rather than outside literature, which is conservative in the direction that matters.
- Sharpe-difference floors between correlated arms are upper bounds: where no paired t-statistic was recorded, an independent-arms standard error overstates the floor, and three units carry high ratios for that reason alone and are classified as adequately powered on other evidence.
- The audit's own error rate is not low. Two of its incidental per-row findings, since checked against source artefacts, were both wrong. One alleged a citation error in a register entry that carried the correct figure verbatim. The other inverted a sign, recording a covariance forecaster as beating its controls when it lost to all three in twelve of twelve combinations of window and control; under the audit's rule that a significantly wrong-signed estimate settles a question regardless of power, that unit is adequately powered, so the census counts above, which predate this correction, overstate the underpowered class by one (thirteen rather than fourteen, with sixty-seven survivors of which twenty-two are wrong-signed), and the count of genuinely reopened questions falls from five to four.
- An underpowered test warrants the label unresolved rather than promising and does not license reviving an idea. Any revival must first re-derive the original headline numbers from the original artefacts with the sign convention stated explicitly; in the case above that re-derivation ended the revival before any data was pulled.
- Lowering the threshold exchanges a false-negative problem for a false-positive one, which is the more expensive error. The remedies are breadth, higher sampling frequency, longer panels, paired constructions and scale-free primary tests.
Implications
Ioannidis (2005) established that the credibility of a published finding depends on the power of the studies producing it. Low power also inflates the share of negative claims that are uninformative, and where negative results are the routine output of a disciplined process, those are the larger population. Harvey, Liu and Zhu's case for a stricter bar in asset pricing is correct, and because a stricter bar raises the detection floor proportionally, it carries an obligation to check that the floor still sits below anything worth finding.
- A settled negative result whose power was never computed should be treated as inconclusive until it is; the back-out estimator above needs two numbers any competent record contains.
- The detection floor is best compared with the size of the mechanism, usually measured inside the same study, rather than with a literature figure.
- Where the proposal is an estimator, headroom should be checked before power, design or data acquisition, since it is the only check in this set that can return a proof rather than a probability.
A companion piece, Underpowered null results, presents the same census in non-technical terms and discusses why the incentive structure of the field leaves this class of error unguarded.
References
- Campbell R. Harvey, Yan Liu and Heqing Zhu, “. . . and the cross-section of expected returns”, Review of Financial Studies 29(1), 2016.
- John P. A. Ioannidis, “Why most published research findings are false”, PLoS Medicine 2(8), 2005.
- Andrew W. Lo, “The statistics of Sharpe ratios”, Financial Analysts Journal 58(4), 2002.
- Ivo Welch and Amit Goyal, “A comprehensive look at the empirical performance of equity premium prediction”, Review of Financial Studies 21(4), 2008.
- Jacob Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd edition, 1988.
What we have tested, and what we rejected
Our validation policy and a plain-language record of the ideas that did not survive it.
Research Integrity More Research