← Research

When the model says another pick was better, how much of that shows up?

How it was measured

Only the pick that was made has a result, so “what if you had picked X” is never observed. What can be tested is whether the model's differences between picks show up across many states where different players made different picks.

The share, in three populations

The champions-only figure is two measurements of the same model on different days, pooled. The first, on its own, said 0.675 ± 0.118; the second said 0.969 ± 0.118. They differ by less than twice their combined error, so neither alone is the number.

PopulationDraft statesShare95% range
Champions only90,4080.82 ± 0.080.66–0.98
Picking player known, every candidate13,3120.83 ± 0.150.55–1.12
Picking player known, only champions that player already plays50,5970.56 ± 0.070.42–0.70

Until 5 October 2026 the site used 0.29 for the player-known reading. That number came from a player's own pool, the last row, and was being applied to every candidate. Measured where it is applied, the share is 0.83.

Is it one straight line?

If small gaps were noise and only large ones real, a single share would mislead. Sorting the champions-only states into ten equal groups by the pick's claimed value, what showed up follows the claim across the range, and a curvature term is indistinguishable from zero.

Ten groups of draft statesThe pooled share, 0.82
−4−20+2+4−4−20+2+4Model's claimed value of the pick, against the board (points)
Vertical axis: what showed up, in points. Each point is about 4,500 draft states from the 2–3 October games; whiskers are 95% ranges resampled by game. The dashed line is the claim in full; the solid line is the pooled share.
Claimed (points)Draft statesObserved (points)95% range
−3.164,516−2.96−4.42 – −1.67
−1.434,515−2.14−3.42 – −0.78
−0.714,515−1.18−2.75 – +0.13
−0.184,516−0.49−1.85 – +0.84
+0.264,515−1.72−3.26 – −0.27
+0.674,515+0.55−0.73 – +1.87
+1.094,516+0.63−0.70 – +2.03
+1.584,527+1.320.00 – +2.74
+2.254,503+2.81+1.54 – +4.23
+3.894,516+3.20+1.85 – +4.43

The ranges are wide: each group's result is uncertain by more than a point. The picture supports a line through zero, not any particular bend.

Does scaling help the forecast?

On the same states, three forecasts of who wins: the board alone, the board plus the model's full gap for the pick made, and the board plus the gap at the share the site had been using. Knowing the pick beats the board alone on AUC, log loss and Brier score. Scaling the gap down did not beat the full gap on these games.

ForecastAUCLog lossBrierAccuracyECE
Board alone0.53910.69030.248652.64%0.0074
Board + the model's full gap0.54450.68950.248253.08%0.0062
Board + the gap at the earlier share (0.675)0.54400.68960.248253.12%0.0055

What this changes on the site

Arena reports and game reviews print a cost in observed outcomes: the model's gap times the share, with a 95% range from the share's own uncertainty. An alternative is called likely better only when the low end of that range still clears half a point.

Rankings, grades and ratings read the order of the picks, which a common share cannot change. The draft arena and the game review use the same numbers.

What it does not show

What should we measure next? Ask on the research page. The most-voted questions our data can answer become future studies.