From the desk · Execution

When trading costs more than the edge is worth

An edge is quoted per year. A trading cost is quoted per trade. Multiply the second by how often you trade and the two numbers are routinely the same size — and it is usually the cost that wins.

One purchase, measured two ways The ruler you pick decides whether your execution looks excellent or expensive time price paid, rising is worse decision target computed here market reopens filled fill quality small staleness — large your own limit price sits here measured against it, every fill looks good the fills were fine · the decision was stale · these need opposite fixes
A single purchase, with the shaded region showing how far the market moved between the price the decision was computed against and the price prevailing when the order could actually be worked. Tightening the limit shrinks the gold bracket and leaves the pink one where it is. The proportions here are schematic; the measured version, which is less tidy, appears further down.

Annualising friction

Consider a strategy that beats its benchmark by 2 points a year, a respectable and believable edge that a great many institutional programmes would be delighted with. Suppose it rebalances often enough to turn its holdings over 40 times a year, which sounds extreme but is what a book that adjusts its positions daily with moderate signal movement actually does, and suppose each unit of turnover costs 15 basis points — 0.15 percent — between the spread paid, the market moving away while you trade, and any commission.

The same strategy, priced two ways

QuantityPer year
Edge over benchmark, before costs+2.0 points
Turnover, times the book40
Cost per unit of turnover0.15 points
Total friction−6.0 points
Net−4.0 points

Illustrative arithmetic on round numbers. No commissions are charged in this example beyond what the 15 basis points already contains, and no borrow, financing or tax is included. Adding any of those makes it worse.

The friction is three times the edge. Nothing about the strategy has changed — the signal is just as good as it was, the backtest that produced the 2 points was honest — and the whole thing is deeply negative. Frequent trading evaluated before costs produces this outcome by default, which describes most of what gets built.

Stated as a budgeting rule rather than as the familiar “costs matter”: friction has to be carried as an annual number and set beside the annual edge before anything is built. A per-trade cost is not comparable to a per-year edge, and the mind does not automatically do the multiplication. Done on paper at the design stage, it kills a large share of otherwise-appealing ideas before they cost anything.

That table is round-number arithmetic; our own research ledger contains the same trade-off measured on real data, three times over, with all three outcomes represented.

The same arithmetic, measured on our own studies

StudyTurnoverVerdict
Minute-scale cross-sectional reversal, crypto1.28× book / daySignal real before costs (t = 5.2); at a 7.5 bp taker fee the net effect drops to t = 2.0 and fails our evidence bar
Funding-rate tilt, crypto0.025× book / dayFirst design in that family to clear the fee floor — by trading 50× slower, not by finding a bigger edge
Option roll-timing variantadds one roll per positionBeats hold-to-expiry when legs are priced free; charging realistic leg costs flips it to −4.4 compound-annual-return points (t = −2.6)

Three of our register-settled studies, 2026. All statistics are autocorrelation-robust. The first and third are edges that exist and are not worth trading; the second is the same asset class made viable purely by cutting turnover. In every case the verdict was decided by the friction line, not the signal line.

Cost as a function of account size

Trading cost as a share of capital is roughly U-shaped in account size, and the two ends hurt for different reasons.

At the large end the problem is impact: your own order moves the price against you, and the effect grows with size, which is why capacity is a real constraint and why the best strategies are often the ones nobody can run at scale. That end is well documented, while the small end gets far less attention and is arguably more brutal.

A strategy is therefore good or bad at a size rather than in general: the same design can be comfortably profitable at one scale and structurally impossible at another with nothing changed but the number in the account.

Implementation shortfall

Checking execution against the limit price you sent is the natural first move — you compare the price you got against your own order: the data is right there and it produces a comforting number. It is also close to meaningless, because your limit price is your own choice — a loose limit fills easily and scores well, and a tight limit fills rarely and scores well on the fills it does get. The measurement is self-referential, evaluating your limit-setting policy against itself and barely registering a cost at all, which is why a lot of shops believe their execution is fine when it is not. We ran this ruler for months and it told us execution was improving.

Measuring against something external fixes that, and the canonical construction has been in the literature since Andre Perold named it in 1988: implementation shortfall — the difference between the return of the portfolio you actually own and the return of the portfolio you would own if every trade had happened instantly and free at the moment the decision was taken. Two reference prices matter, and the distance between them is the whole diagnosis.

Two reference prices

ReferenceWhat a gap against it means
The price prevailing when you tradedFill quality. Spread paid, impact, poor routing, bad timing within the window.
The price your target was computed fromTotal shortfall. Includes everything above plus everything the market did between deciding and arriving.

The difference between the two is decision staleness: a cost, frequently the larger cost, that does not appear anywhere in a conventional fill-quality report.

Measuring both gives a diagnostic. Fills that sit close to the price prevailing when you traded but far from the price your decision was computed against indicate that your execution is fine and your latency is not, since the market has already moved by the time you arrive. Tightening limits will not help and will actively hurt, because tighter limits mean more unfilled orders, more retries, and a longer average delay between deciding and owning — which is the actual problem, made worse.

If the gap to where the market was is small and the gap to where you decided is large, you do not have an execution problem. You have a clock problem.

Our own order flow said roughly this when we finally measured it properly — but only roughly, and the wrinkle is worth more than the headline. The measurement below covers 181 matched equity fills, using daily reference prices.

Our own equity fills against two external references

Measured againstAverageTypical trade
The price prevailing at the session open−15.6 bp−2.6 bp
The prior close the target was computed from−9.9 bp−4.5 bp

Negative means we transacted worse than the reference. A basis point is 0.01 percent. These are our own fills over one measurement window, not a general property of markets. Only trades that could be paired with a clean reference price are included; a minority could not be, and are excluded. Daily reference prices approximate what a proper measurement would use — the price prevailing at each individual moment — so treat the levels as indicative.

Average and typical trade rank the two costs in opposite orders, and both rankings are true. On the typical trade the gap to the decision price is the larger of the two, which is the staleness reading and the one the diagram above draws. On the average the ordering flips, because a minority of unusually bad fills pulls the mean well past the median, and that tail sits in the execution leg rather than the decision leg. We have not isolated where those fills come from, and the one hypothesis worth ruling out is ruled out by the numbers themselves: the fills around the open are the better-executed group, not the worse one. Restricted to that window, where the two references are cleanly separated in time, they sit about 10 basis points from the open print and about 15 from the prior close, reproducing the staleness ordering. Since the full sample averages about 16 basis points from the open, the fills outside that window must average appreciably worse, which is where the tail should be looked for next.

The honest statement is therefore narrower than the tidy one. For the typical trade, and for the open-window population specifically, the decision-to-arrival gap is the more expensive leg, while across the whole sample the average is dominated by a fill-quality tail. Both readings need acting on and they need different actions, which is exactly the point of measuring against two references rather than one. The sign survives every cut: every number in that table is negative. That retracted an earlier internal reading — produced by the self-referential ruler — that execution had been improving. It had not; we had been grading our own limit-setting.

A mean that disagrees with its median carries a lesson worth more than the specific numbers: when an average and a typical value rank two costs in opposite orders, the difference between them is the tail, and it locates where the tail lives. Isolate it before changing any policy, since an average quietly built by a small number of pathological fills will send you to fix the wrong leg.

Unfilled orders

Measurement discipline runs in both directions, as a second episode from the same review shows. That review found that a large share of orders never filled at all, which reads like a serious wound: if the intended trades are not happening, the live book is not the modelled book. A routing change was designed to fix it.

Before shipping it we checked whether the dead orders corresponded to positions that never got established, and largely they did not. Within the biggest group of them — orders carrying no price condition at all, which therefore cannot have failed on price — about 3 in 5 involved an instrument that filled later the same session, because the system recomputes its targets continuously and simply succeeded on a later attempt. Those orders were redundant retries rather than missed trades, and the fix was written and withheld.

An alarming statistic is a hypothesis rather than a finding, and shipping a change that treats noise as a wound carries a real cost: a permanent increase in complexity, a new surface for defects, and a false belief that a problem has been solved.

Three rules we now work by

  1. Friction is budgeted annually and at design time: turnover multiplied by cost per turn, compared against the annual edge, on paper, before any code. If friction is the same order of magnitude as the edge, what has to change is the design rather than the execution.
  2. Execution is never measured against our own order, but against at least one external reference and ideally the two above, since a metric that cannot show a loss is not a metric.
  3. The decision-to-order delay is carried as a cost line. It is usually cheaper to shorten than to out-trade a spread, it is entirely within your control, and almost nobody measures it. We have written separately about how quickly a signal loses its value while a decision is being made.

Scope and limitations

Further reading

The companion piece

Why a strategy that wins the large majority of its trades can still be worth almost nothing.

A High Win Rate Is Not an Edge More Research