A bar nobody has tried to fail is not a bar.
The gates exist to separate an edge from a coincidence. That claim is testable, so it gets tested: the engine is handed data with no edge in it by construction — returns with the drift removed, or contracts priced at exactly the volatility that generated the path — and the question is how often it promotes something anyway.
Anything that clears the promotion battery on that data is a false positive. There is nothing there to find.
Every run is below, including the one that failed and the re-measurement it forced. Thresholds here are never set by intuition; they are set by what noise scores, which is why a null that comes back badly is worth more than one that comes back clean.
The equity null.
Real price history with the drift removed from log returns, so any apparent trend is an artifact of the sampling rather than a property of the market. Measured 2026-08-06.
Not one of the 65 scored trials reached promotion. That is the result the gates are supposed to produce, and it is the reason the deflated-Sharpe threshold is set where it is rather than where it “felt” right — an earlier placeholder of 0.95 would have rejected one hundred percent of this engine's real output.
The options null, which failed.
A second harness for defined-risk options needed its own thresholds, because an options Sharpe and an equity Sharpe are not comparable numbers. The obvious move — reuse the equity battery — was tested rather than assumed. Driftless random walks, every contract quoted by Black–Scholes at exactly the volatility that generated the path: no volatility risk premium, no drift, no term structure. Measured 2026-08-19.
7.0% of chains with no edge in them cleared a battery built to admit almost nothing. The instinct at that point is to raise the Sharpe floor until the number looks acceptable. That was rejected: raising a global threshold to fix a structural problem is how a gate stops meaning anything. Splitting the result by structure located the fault exactly.
It was one structure, not the gates.
The reading at the time: a covered call is a hundred shares plus a short call, so it carries residual long delta, and on a driftless random walk roughly half of all paths drift up — paths that look exactly like edge. Re-instrumenting the backtester later showed that reading was wrong. The number was real; its cause was a settlement defect, not exposure — the measurement further down this page.
Either way the response was the same and it stands: the structure was removed from the offered set rather than the threshold being moved. It stays implemented and tested. With it excluded, the equity battery transfers essentially unchanged at 0.7%.
A trade-count floor makes this worse rather than better — requiring more trades raises the share of noise that clears the Sharpe bar, because a short-volatility book with a high win rate and a rare large loss simply has not met its tail yet.
Then the cost model changed, so it was run again.
The run above charged a flat spread on every contract. The harness has since moved to a spread surface fitted to the option-chain archive, which is anything but flat — widest at short-dated wings, tightest months out, and wider on the call side than the put side. A null describes the code that was in the box when it ran, and the fill model is part of that code, so the result above expired the day the surface shipped. Re-measured 2026-08-20.
Two things changed with it. A strike-and-expiry bucket the archive never measured is now simply not listed rather than priced off an average — the same refusal the modelled data path makes, because inventing a quote is how a backtest manufactures an edge. And the verdict is no longer a re-implementation of the gates: every scored trial is pushed through the engine's own evaluator, so this rate is literally what fraction of noise would have been promoted, and it cannot drift from the gates it claims to measure.
A clean re-run is not an improvement.
0 of 191 is not evidence the gates got better. With no false positive observed, the most a sample this size rules out is roughly 1.6% — and the earlier result, 0.7% on the structures actually offered, sits comfortably inside that. The two runs agree with each other; that is the finding. Reading the second as progress would be reading noise.
What it does establish is that the thresholds were calibrated against the cost model production actually charges, rather than against a placeholder. The distributions barely moved, mildly compressed if anything, because the fitted surface charges more than the old flat figure at the short-dated end where credit structures trade.
The withdrawn structure is absent here rather than removed after the fact: strategy kinds are drawn from the offered set, so it was never in the sample. That is why this run reports no all-structures figure. The rate above and the rate two sections up answer different questions, and the earlier one is the one that failed.
The battery they were measured against.
Recorded with each result, because a false-positive rate means nothing without the thresholds that produced it. These are read from the engine's live configuration, so if a gate moves this page moves with it.
What these results do not cover.
It establishes that the gates reject noise at a measured rate. It does not establish that anything which survives them is profitable. Surviving a null is necessary, never sufficient.
Dividend-driven early call assignment is not simulated, which biases covered calls optimistic — a second, independent reason that structure is not offered.
Option bars carry no bid or ask. Before the chain archive begins, spreads are extrapolated from later quotes rather than observed, and 0-DTE is excluded entirely.
Each result describes the code that was in the box on the date shown. Changing the fill model, the sizing rule or the structure set invalidates it, and it has to be run again.
The thresholds these runs were testing — every gate, and why each one exists — are set out in full on the method page.