A year or two of daily prices pins down how much an asset moves to within a few percent of itself. Pinning down what the same asset earns, to within one point a year, takes four centuries of the same data. That asymmetry is arithmetic rather than opinion, and almost every argument about track records ignores it.
Two estimation problems
Two questions can be put to the same price history: how much the asset bounces around, and how much it earns per year, on average, going forward. Both look like one kind of problem — take the data, compute a summary number, report it with an error bar — but they differ by enough to reorganise how a research firm ought to behave.
The bounce is the easy one. Because adding observations — daily instead of weekly, hourly instead of daily — tightens the estimate of how much an asset moves, a year of daily closes already pins volatility down to within a few percent of itself. That is why risk models work at all, why option pricing has a century of practical success behind it, and why a firm can say something honest about the shape of a portfolio's risk after a fairly short observation window.
The average behaves differently. Robert Merton set the problem out precisely in 1980, and the result has never been seriously contested: the precision of an estimated average return depends only on how long a span of calendar time you observe, not on how often you sample within it. Chopping the same decade into minute bars gives you millions of observations and does not improve the estimate of the mean at all, since every extra observation adds one more draw of the noise along with one more slice of the drift, and the two cancel.
Measurement error on an average return
The measurement error on an estimated annual average return is the asset's annual volatility divided by the square root of the number of years observed, and nothing else enters. For an asset that moves about twenty points a year, roughly a broad equity index:
Measurement error on an estimated average annual return
| Calendar span observed | Error, points per year |
|---|---|
| 1 year | ±20.0 |
| 10 years | ±6.3 |
| 25 years | ±4.0 |
| 100 years | ±2.0 |
| 400 years | ±1.0 |
Assumes an asset whose annual volatility is twenty points and whose behaviour is stable over the whole span, both of them generous assumptions. Sampling more frequently than daily changes none of these figures.
The bottom row is the binding one. Pinning an average annual return down to within one point — a precision most allocators would consider the bare minimum for a decision — requires four centuries of a stationary asset; the oldest continuously quoted equity indices are not a quarter of that, and nobody believes the underlying economy held still for the part we do have.
More data is not the answer, because what is missing is not observations but calendar time, and calendar time arrives at one year per year.
Evidence in a track record
The same arithmetic governs the judgement people actually care about — whether this manager, or this strategy, is any good — which is usually summarised as the ratio of return to volatility, whose measurement error has a known form.
Suppose a strategy genuinely has a ratio of half, earning half a point of return for every point of risk taken; a great many respectable institutional programmes sit near that decent and entirely realistic figure. The question is how many standard errors of evidence a track record of that strategy accumulates. A standard error here just means one unit of measurement noise: an estimate two standard errors away from zero is conventionally called suggestive, and anyone correcting properly for the number of ideas that were tried before this one demands considerably more than two.
Evidence accumulated by a genuinely good strategy
| Length of live record | Standard errors of evidence |
|---|---|
| 5 years | 1.1 |
| 10 years | 1.5 |
| 20 years | 2.1 |
| 40 years | 3.0 |
For a strategy whose true return-to-risk ratio is one half, using the standard measurement error for that ratio. Shorter records are not weak evidence of a good strategy; they are no evidence either way.
A decade of a genuinely good strategy does not reach the level of evidence a careful reader would call suggestive, and two decades barely does. Four decades gets you to a bar that a properly corrected analysis would take seriously, by which point the market, the fee structure, the venue, and probably the manager have all changed.
Because this cuts in both directions, a five-year record is weak evidence that a strategy works and equally weak evidence that it fails; the second direction is the one that gets forgotten. Most of what passes for evaluation in this industry is reading tea leaves out of samples that could not have settled the question either way.
Differences rather than levels
The arithmetic above is brutal because the noise in a raw return series is enormous relative to its drift, so the way out is to compare two things that share most of that noise: the shared part cancels, only the difference remains, and the difference is a much quieter series. This is the reason serious quantitative work looks the way it does.
Measuring whether a portfolio earned more than a matched benchmark is a far easier statistical problem than measuring what either of them earned. If the two move together closely, the difference between them might wobble by a handful of points a year rather than twenty, and the four-century requirement collapses to something closer to a couple of decades — still slow, but possible within a career.
The rest follows from that: disciplined shops obsess over benchmark matching, paired comparisons and difference-in-difference designs beat before-and-after ones, and a test that asks “is version B better than version A on identical inputs” is worth an order of magnitude more than a test that asks “does version B make money”. The same logic explains why cost, turnover, tracking error and drawdown are quoted so much more confidently than expected return: those are all quantities the data can actually resolve. Cost is the clearest case, since a few hundred trades settle it, which is why measuring it properly repays the effort far faster than any attempt to sharpen a forecast of return.
Consequences for how we work
Three consequences shape how we work, and we would rather state them than leave them to be inferred.
- We do not publish forecasts of return levels, not because forecasting is disreputable but because the honest error bar on any such forecast is wider than the forecast itself. Anyone quoting an expected annual return to one decimal place is quoting an assumption rather than a measurement, and the polite thing is to say which.
- We design tests as comparisons wherever the data allows it, running a candidate against an incumbent on identical inputs, identical dates, identical costs, and examining the difference. This is the single highest-leverage design decision available, and it is free.
- We treat “we could not tell” as a real and common outcome, since a large fraction of investment questions are not open or closed but unresolvable at the sample size available. Filing those as failures is as wrong as filing them as successes, a trap we have written about separately.
Scope of the claim
- Expected returns plainly exist; assets have risk premia and those premia are real. The claim here is narrower and harder to escape: you cannot measure one precisely from its own price history in any relevant amount of time.
- None of this argues against quantitative investing; it argues about which quantities deserve confidence. Risk, correlation, cost, capacity, turnover and relative behaviour are all estimable on human timescales, which makes the expected level of return the outlier rather than the rule.
- Nor is it a reason to prefer judgement to arithmetic, since a discretionary manager reading the same five-year record faces exactly the same measurement problem and typically does not compute the error bar at all.
- The arithmetic is not the whole difficulty either. Real markets are not stationary, so the long spans it demands are precisely the spans over which the underlying quantity has changed, which makes the situation worse than the table above and never better.
Further reading
- Robert C. Merton, “On estimating the expected return on the market: an exploratory investigation”, Journal of Financial Economics, 1980 — the original statement of the sampling-frequency result.
- Andrew W. Lo, “The statistics of Sharpe ratios”, Financial Analysts Journal, 2002 — measurement error for return-to-risk ratios, including the serial-correlation corrections this article skipped.
- Campbell R. Harvey, Yan Liu and Heqing Zhu, “. . . and the cross-section of expected returns”, Review of Financial Studies, 2016 — why the conventional evidence bar is too low once you count how many ideas were tried.
How we hold ourselves to this
Our validation policy, the bar every idea has to clear, and the record of what we have rejected.
Research Integrity More Research