Research note · Research design

Minimum detectable effects at design time

A retrospective census of eighty-one settled negative results in our register found fourteen whose design could not have detected a plausible effect, five of them because the evidence bar demanded a larger effect than the entire quantity the study set out to measure. Whether a bar is reachable is arithmetic, knowable on the day it is frozen. When the smallest effect a design can resolve exceeds anything the mechanism under test could produce, the study returns a negative result whatever the truth, which is then indexed as evidence the question is closed.

Low power leaves a real effect a poor chance of detection; a detection floor above the mechanism's ceiling leaves it none.

The detection floor

A test's minimum detectable effect, the smallest true effect it could report as a finding, is the evidence bar in standard errors multiplied by the standard error of the estimator. At the Harvey, Liu and Zhu (2016) threshold of |t| > 3, no true effect below 3 × SE can be reported as a positive finding, however carefully it is estimated.

Retrospectively, the standard error is most reliably backed out of the record rather than rebuilt from a formula. Where a completed study reports an effect E and its t-statistic t, SE = |E| / |t| and

Back-out estimator for the detection floor

MDE  =  bar × SE  =  |E| × bar / |t|

The estimate inherits the autocorrelation, overlapping-window and clustering corrections already applied to the reported t-statistic, which no closed-form substitute does, and was computable for the large majority of the units audited below. Where no effect-and-t pair was recorded, the audit used the standard forms: the cross-sectional information coefficient as bar / √((N−1)T), the Lo (2002) standard error for a Sharpe level, bar × σevent / √nevents for an event study, and the null distribution's own upper percentile for a permutation bar.

The design that prompted this work ranked twenty-four names against one another on a flow-derived characteristic, with a quorum of forty observation dates, judged at |t| > 3. Its detection floor is 3 / √(23 × 40) = 0.099 in correlation units, against published cross-sectional characteristic correlations of about 0.02 to 0.06, so it could have confirmed only an effect roughly twice the strongest on record and would have returned a null irrespective of the market. That null would then have been cited as settled against every later proposal in the same territory, although the calculation that predicts it is one line long.

Four quantities to declare at freeze

A preregistration should state four quantities before it is frozen, in this order.

  1. The minimum detectable effect at the frozen bar, with the formula and arithmetic that produced it, rather than an assertion that the design is adequately powered.
  2. A plausible effect for this class of strategy, with its published source, fixed in advance from a common reference table so that it cannot be adjusted after the result is known.
  3. The size of the mechanism the study itself measures, such as the cost being removed, the premium being harvested or the concession being recovered: the ceiling on what a perfect version of the idea could deliver.
  4. Headroom, wherever the proposed mechanism is an estimator (defined below).

The binding comparison is the first quantity against the third. If the detection floor exceeds the size of the mechanism, the design cannot pass even when the theory is exactly correct at full strength, and it should not be registered.

Floors above the mechanism

Six studies in the census failed that comparison. Each had measured, or theory had bounded, the quantity being pursued, and the frozen bar demanded more of it than existed.

Six bars that demanded more than the whole mechanism detection floor at the frozen bar, divided by the quantity the study set out to measure FLOOR = WHOLE MECHANISM sizing rule on a growth book 39 pp/yr needed · 8 pp/yr declared 4.88× gated harvest of an excess-growth term 11.05 %/yr needed · 4.3 %/yr measured 2.57× auction-cycle rebound in government bonds 4.43 bp/day needed · 2.71 bp/day measured 1.63× rebalancing premium, cross-asset basket 183 bps/yr needed · 138 bps/yr in theory 1.33× rebalancing premium, single-asset-class 337 bps/yr needed · 314 bps/yr in theory 1.07× cost reduction in a market-neutral book 4.41 %/yr needed · 4.21 %/yr of cost existed 1.05× 0 1 2 3 4 5 Anything to the right of the gold line was unpassable on the day it was frozen.
Each bar divides the study's detection floor by the quantity it was built to capture, taken from the study's own measurement or from the theoretical ceiling it cites. Above one, capturing the whole available effect still falls short of the evidence bar.

At the lowest ratio, a study of whether a slower adjustment schedule would reduce the trading costs of a market-neutral book was frozen with a detection floor of 4.41 percent a year against a measured cost drag of 4.21 percent a year. Removing every basis point of cost the book incurred would have produced t = 2.87, short of the bar, so the competent null the study reported was fixed when the design was written down.

The other five fail the same way at the wider margins shown in the chart, and in every case the ceiling was already in the study's own record.

A bar the mechanism cannot reach at full strength tests the arithmetic of the design rather than the mechanism, and its verdict is fixed before any data arrives.

Headroom

When the proposal is an estimator meant to reduce noise, such as a structural model, a shrinkage, a fitted surface or a filter, it buys lower sampling error at the price of approximation error, since it imposes a structure the world does not exactly satisfy. Define

Headroom

headroom  =  noise the mechanism removes  −  the mechanism's own approximation error

Negative headroom proves impossibility. Misspecification does not shrink with the sample, so no quantity of data, refinement of the fit or change of bar can rescue the proposal.

One proposal was rejected at the design stage on this check alone: a two-parameter structural correlation surface to replace ordinary sample correlations between volatility contracts of different maturities. Its approximation error over the full sample, reported weeks earlier by its own preliminary study, was 0.0113; the sampling noise it was meant to remove, at the sixty-observation window planned as the primary specification, was 0.00827. Its approximation error exceeded the noise it was meant to remove by 37 percent before a single parameter was estimated.

In a synthetic control where the assumed structure is exactly true, the fitted estimator still lost to a standard shrinkage estimator by 0.0193 and to raw sample correlation by 0.0075, on a loss scale where lower is better, so the hypothesis lost even in its best case.

Census of settled nulls

We computed the detection floor for every settled negative result we hold, with plausible effect sizes fixed in advance in a nine-row reference table citing a source per row. Classification ties were resolved in favour of the existing verdict, so every count of a failure is an undercount.

Retrospective power census of the settled nulls

ClassCountShare of audited
Settled negative results86
Audited for power81
Could have seen a plausible effect and did not6681%
  of which carried by a wrong-signed estimate2126%
Could not have seen a plausible effect at all1417%
Record too thin to compute power from11%
Already labelled measurement-limited, audited separately5
Records containing a power calculation at freeze time0 of 152

Fifteen register entries described themselves as well-powered with no arithmetic behind the phrase. Nine of the fourteen underpowered units carry their substantive conclusion on a second, adequately-powered bar (a permutation control, a deterministic threshold count or a sign-decisive point estimate), so for those the reclassification corrects a label without reopening the question.

Four fifths of the settled negatives survived, and 21 of the 66 survivors are carried by an estimate significantly of the wrong sign, which is stronger evidence than a failure to reject. The damage is concentrated and identifiable unit by unit.

Detection floors by design

The failures clustered by design shape rather than by subject matter, so the hazard can be recognised from the outline of a proposal before any arithmetic is done.

Detection floors at the conventional strict bar

DesignDetection floorPlausible effect
Monthly predictive regression, 20 years (n = 244)correlation 0.192, R² 3.7%best macro predictors, R² 0.25–2%
Monthly Sharpe, 10 years (n = 120)annualised Sharpe 0.95single-asset-class sleeve, 0.30–0.80
Daily Sharpe, 20 years (n = 5,040)annualised Sharpe 0.67as above
Event study, 48 events at 12% event volatilityabnormal return 5.2%corporate-event effects, 1–5%
Event study, 237 eventsabnormal return 2.3%as above
Cross-section, 24 names × 40 datescorrelation 0.099characteristic correlations, 0.02–0.06
Cross-section, 45 names × 245 datescorrelation 0.029as above

Sharpe floors use the Lo (2002) standard error; the cross-sectional floors use 3 / √((N−1)T); the event-study floors use 3σ/√n at a sixty-three-day idiosyncratic volatility of 12%. Plausible ranges are the reference table's, with Welch and Goyal (2008) supplying the predictive-regression row.

The two cross-sectional designs share an estimator, a bar and a class of signal, and differ only in how many names are ranked and over how many dates. The narrower cannot resolve any documented effect; the wider resolves most of the plausible range. Breadth, bought with a wider universe and a longer panel rather than a better model, moves the detection floor faster than anything else available.

Designs whose declared primary test was a permutation test rather than a ratio came through the census in the best condition. Such a test is scale free, imposes almost no distributional assumption, and in this register was consistently better powered than the t-statistic it accompanied.

Limitations

Implications

Ioannidis (2005) established that the credibility of a published finding depends on the power of the studies producing it. Low power also inflates the share of negative claims that are uninformative, and where negative results are the routine output of a disciplined process, those are the larger population. Harvey, Liu and Zhu's case for a stricter bar in asset pricing is correct, and because a stricter bar raises the detection floor proportionally, it carries an obligation to check that the floor still sits below anything worth finding.

A companion piece, Underpowered null results, presents the same census in non-technical terms and discusses why the incentive structure of the field leaves this class of error unguarded.

References

What we have tested, and what we rejected

Our validation policy and a plain-language record of the ideas that did not survive it.

Research Integrity More Research