Over 887 episodes on real option chains, a stop at three times the credit received fired on 63 to 75 per cent of short-dated index credit spreads at the higher of the two volatility readings used to price the exit, while the underlying touched the short strike on 11.7 per cent. A credit spread loses money only when the underlying settles beyond its short strike, so distance from that strike governs the loss. Premium responds to that distance and to implied volatility together, which is why a stop keyed to a multiple of the credit triggers largely on volatility expansion.
Payoff and premium
A credit spread sells one option and buys a cheaper one further out, capping the worst case in advance. At expiry its payoff depends on one number, where the underlying settles relative to the short strike: on the put side the position keeps the credit above the strike, and below it the loss grows linearly until the long option caps it. Nothing the option market priced along the way enters the payoff.
Before expiry, under Black and Scholes or any standard pricing model, an option's value depends on moneyness and on implied volatility and increases in both. Two positions at the same distance from the short strike can be marked at very different premiums, and one position can double in premium overnight with the underlying unchanged.
A rule that closes the position when premium reaches a multiple of the credit therefore conditions on both variables, only one of which determines the outcome. It must fire on every state in which distance has moved and on other states besides, with an excess set by the size of the volatility term.
Measurement
Ten management rules, including the unmanaged baseline, were swept over short-dated credit spreads on a broad equity index tracker, priced from real option chains and evaluated on the underlying's five-minute intraday path rather than on closing marks. The headline cell is 887 put-side entry days at approximately ten-delta short strikes, entered one session before expiry. Sixteen cells span both sides of the market, two short-strike distances, two spread widths and two holding lengths, and the ranking of rules below reproduces in all of them.
Trigger rates against event rates — 887 episodes
| Quantity | Rate | Note |
|---|---|---|
| Stop at three times the credit received | 63.1% to 74.7% | the two put cells; the one-session call cells run to 78.5% |
| Underlying touches the short strike at any point | 11.7% | 95% interval 9.8% to 14.0% |
| Position finishes in the money | 5.5% | ratio of touch to settlement 2.12 |
| Finished in the money without ever touching | 0 of 887 | coherence check; holds in all sixteen cells |
The stop fires five to six times as often as the underlying reaches the strike and about twelve times as often as the position loses at expiry. Because it depends on volatility, its rate is also the least stable quantity in the table, ranging from about a quarter to three quarters of days depending on the volatility reading used to price the exit.
The touch-to-settlement ratio, a well-known folk factor of about two, measures 2.12 on this cell, 2.54 at the longer holding period and 1.6 to 1.8 on the call side, so a rule acting on the touch alone already acts on an event about twice as frequent as the loss.
Cost of early exits
An early stop pays the toll to exit, and to re-enter where the rule allows it, and it converts positions that would have recovered into realised losses, which is the larger cost. Of the 104 episodes in which the underlying touched the short strike, 55 (52.9 per cent, 95 per cent interval 43.4 to 62.2 per cent) recovered to finish out of the money, so a rule acting on the touch forfeits those recoveries and scratches about three winners for every two losses it caps.
Rules on the headline cell — put side, ten-delta, one session, n = 887
| Rule | Fires on | Win rate | Winners scratched | Losses capped | Value per trade |
|---|---|---|---|---|---|
| Hold to expiry | — | 94.5% | 0 | 0 | +0.09% |
| Premium reaches 3× credit | 63.1% | 36.9–63.4% | 512 | 42 | −2.02% to −6.03% |
| Exit 0.5% before the strike | 28.5% | 72.7–73.6% | 204 | 40 | −1.53% to −2.76% |
| Exit at the strike | 11.7% | 88.3% | 55 | 36 | −0.32% to −0.45% |
| Exit 0.25% past the strike | 7.0% | 92.8% | 15 | 34 | +0.24% to +0.51% |
| Exit when the spread marks at half its width | 7.1% | 92.7–92.9% | 16 | 32 | +0.13% to +0.27% |
| Unconditional flatten in early afternoon | 100% | 58.4–77.8% | 838 | 36 | −2.15% to −3.94% |
Value per trade is a fraction of collateral. Value and win rate are banded between the exit repriced at entry-day implied volatility and at one and a half times it; the 3×-credit rule and the unconditional flatten have the widest bands because they depend most on volatility.
Once the exit sits at or inside the strike, each quarter of a per cent by which it is moved earlier scratches roughly seventy to eighty more winners per 887 episodes to cap one to three more losses. The premium-referenced stop sits at the destructive extreme, capping 42 losses against 512 scratched winners. The three latest rules improve on holding, and all are referenced to price rather than premium: exit a quarter or a half of a per cent past the strike, or exit when the spread marks at half its width, a state a far out-of-the-money spread reaches only once the underlying has arrived at the strike.
The unconditional early-afternoon flatten is the limiting case of a rule with no conditioning variable. It fires on every episode, paying the exit toll and surrendering the remaining time value each time, and costs 2.15 to 3.94 per cent of collateral per trade to cap 36 losses, the same number the exit at the strike caps at a tenth of the trigger rate.
Each quarter of a per cent of earliness scratched about seventy winners to cap one to three losses.
Overnight gap risk
Three conditions were committed before the data was read: a win rate of at least 90 per cent at both endpoints of the volatility band, a worst single loss of no more than four times the credit received, and positive expected value at both endpoints. Four rules meet the first and third; none meets the second, by wide margins. The worst episode costs 73 times the credit unmanaged, 145 and 151 to 161 times under the two late price-referenced stops, and 31 to 36 times under the early stops, which cap the most. The best rule in this cell misses the bar of four by a factor of eight, and the most favourable cell in the sweep misses it by more than a factor of five.
Entered one session before expiry, the position is held overnight, and under every rule the worst outcome is an opening gap through the short strike, where the market reopens beyond the stop and the stop price never trades. The best any rule does is reduce the worst episode from a total loss of collateral to 74 to 88 per cent of it, a reduction set by the size of the gap rather than by the rule. At this holding period the loss tail is gap risk, which no intraday condition on premium or on price reaches.
Management improves expected value modestly, in a direction consistent across all sixteen cells, and cannot make the worst case proportionate to the credit collected. In a compounding simulation at the smallest position size tested, the deepest drawdown falls from 24.2 per cent under the baseline to 14.6 per cent under the best rule at its more favourable volatility reading, while the probability of a drawdown of 15.4 per cent or more falls from 0.90 only to between 0.47 and 0.61, and is essentially 1.00 at every larger position size tested. Position size and the decision to hold overnight, rather than the stop, govern the tail.
Trigger rate against event rate
Before running any backtest, compare the rate at which a proposed rule would have fired historically with the rate at which the event it exists to prevent occurred. A large multiple means the rule conditions on a variable other than the one that governs the event, and the excess is the fraction of its activity that is noise. Here the ratio is five to six against the touch and about twelve against the loss at settlement. A ratio near one means the rule is pointed at the right state variable, though not that acting on it pays.
- A high trigger rate produces high apparent activity and a high count of losses averted, both consistent with severe value destruction. The 3×-credit rule capped more losses than any other rule on the headline cell and was the worst rule by expected value in every cell.
- The premium-referenced trigger rate moved by a factor of three across the volatility band and the price-referenced triggers not at all. A trigger inherits the instability of a second variable it depends on, so a single historical trigger rate cannot characterise a rule whose activity depends on the volatility regime.
- Israelov makes the general point for protective puts: protection carries a negative expected cost outside stress regimes, and its benefit is routinely assessed against the loss avoided rather than the compounded cost of carrying it. A stop is protection bought after entry and paid for in scratched winners, so it should be compared with the unmanaged position rather than with the loss it prevented.
Limitations
- Paths are read at five-minute resolution. Touches are detected exactly from bar extremes, but their timing is known only to the nearest bar and the order of prices inside a bar is unobservable, so stop fills are priced at the worse of the bar's opening price and the trigger level.
- The volatility band brackets an input, so which endpoint is conservative depends on the case: for exits while the position is still out of the money the higher reading is the expensive side, and after a breach higher volatility makes the spread cheaper to buy back, so the entry-day reading is conservative there.
- Differences in expected value between rules are not certified at any evidence bar. The preregistered minimum detectable difference was 2 to 3 per cent per trade against a mechanism of 0.1 to 1 per cent. The paired comparison between the best late stop and holding runs from +0.15 to +0.76 percentage points per trade, with a heteroskedasticity- and autocorrelation-consistent t-statistic of 0.40 to 1.90, so the ordering is consistent across all cells and certified in none. The trigger-rate arithmetic does not depend on an inference bar.
- All measurements come from short holding periods on an index underlying. At longer holding periods the overnight gap is a smaller share of total risk and the tail argument weakens; the conditioning argument does not, since premium depends on both variables at every maturity.
- The first execution pass compared chain strikes against a dividend-adjusted price series, putting strikes and prices on different scales. Three coherence readouts caught the error before any figure was cited, most decisively call-side touch rates below call-side settlement rates, which is impossible because a position cannot finish in the money without touching. After the frames were reconciled, zero of 887 episodes settled in the money without a touch, in every cell.
Further reading
- Fischer Black and Myron Scholes, “The pricing of options and corporate liabilities”, Journal of Political Economy, 1973 — the source of the two-variable dependence the argument rests on.
- Roni Israelov, “Pathetic protection: the elusive benefits of protective puts”, working paper, 2017.
- Roni Israelov and Lars N. Nielsen, “Still not cheap: portfolio protection in calm markets”, Journal of Portfolio Management, 2015.
- Hersh Shefrin and Meir Statman, “The disposition to sell winners too early and ride losers too long”, Journal of Finance, 1985 — on why rules of this kind are adopted in the first place.
What we have tested, and what we rejected
Our validation policy and a plain-language record of the ideas that did not survive it.
Research Integrity More Research