The dealer-positioning corner of the internet is loud on calls and silent on outcomes. Runir keeps the opposite discipline: every level we publish is graded against the next session and kept in public, win or lose. That archive is now 478,548 scored walls deep. Here is what it says.

How often the walls actually hold

Cumulative held rate across every scored wall. The shaded band is the 95% Wilson interval: it starts wide and tightens as the sample grows. We render the uncertainty rather than smooth it away.

Scored in public

84.1% held · 16833 of 20018 tested

Scored through Sep 3

First scored May 27 · early, by design

By wall, and by regime

The same record, split by the level type and by the dealer gamma regime it broke under. Each bar carries its own sample size; the axis starts at the 50% coin-flip so real differences read honestly.

Call wall87.8% n=63,319
Amplifying87.6% n=5,105
Damping88.1% n=46,427
Neutral86.6% n=11,787
Put wall87.2% n=45,098
Amplifying87.8% n=9,260
Damping87.1% n=22,186
Neutral86.9% n=13,652
Gamma flip87.9% n=55,298
Amplifying94.0% n=6,240
Damping93.1% n=11,026
Neutral85.4% n=38,032

50%held rate, scored in public100%

The gamma flip is the tell: it holds 94% of the time when volatility is amplifying versus 85% in a neutral regime (n=6,240). The regime changes the level's meaning, and the archive shows it.

Does the model beat the base rate?

Publishing a hit rate is one thing; a calibrated probability is another. Our breach model is scored against the raw base rate on the same archive, out of sample.

0.700 AUC rank accuracy (0.5 = coin-flip)
−3.8% Brier vs base lower error than the base rate
260,094 predictions over 250 sessions

The lift over the base rate is statistically significant on the pre-registered stratum.

And here is the part a headline number can hide: is a "20% chance it breaks" actually a 20%? Each bin of out-of-sample predictions is plotted against what really happened. Dots on the diagonal mean the probabilities are honest.

Each dot is a bin of out-of-sample predictions (area scales with its n); whiskers are 95% Wilson intervals on the observed rate. 260,094 predictions, expected calibration error 0.7pp. Axes zoomed to 0–35%.

This is the moat: not a prettier chart, a kept record. See it on a live name at $NVDA →, or read how every number here is computed.