When the model says another pick was better, how much of that shows up?
Published October 5, 2026 · 90,408 draft states from 12,000 ranked games
A draft tool that says “Garen was 3.9 points better” is making a claim that can be checked. Across 90,408 real draft states, picks the model valued higher did win more: by 0.82 of the claimed gap with only champions known, and by 0.83 with the picking player known. KayLoL shows every cost at that share, with its uncertainty, and calls it a calibration: it is not proof that switching champions would have changed your game.
How it was measured
Only the pick that was made has a result, so “what if you had picked X” is never observed. What can be tested is whether the model's differences between picks show up across many states where different players made different picks.
- For each real draft state, every legal candidate is scored by a model trained only on earlier games. The average over the candidates is the board's value; the pick's own value is its score minus that average.
- The game's result is regressed on the board's value and on the pick's own value. The second coefficient is the share: 1 means the model's gaps between picks show up in full, 0 means they carry nothing.
- Every game was played after the model that scored it was built, and the player histories used were frozen before the day began.
- Errors are clustered by game, because the states of one game share one result.
The share, in three populations
The champions-only figure is two measurements of the same model on different days, pooled. The first, on its own, said 0.675 ± 0.118; the second said 0.969 ± 0.118. They differ by less than twice their combined error, so neither alone is the number.
| Population | Draft states | Share | 95% range |
|---|---|---|---|
| Champions only | 90,408 | 0.82 ± 0.08 | 0.66–0.98 |
| Picking player known, every candidate | 13,312 | 0.83 ± 0.15 | 0.55–1.12 |
| Picking player known, only champions that player already plays | 50,597 | 0.56 ± 0.07 | 0.42–0.70 |
Until 5 October 2026 the site used 0.29 for the player-known reading. That number came from a player's own pool, the last row, and was being applied to every candidate. Measured where it is applied, the share is 0.83.
Is it one straight line?
If small gaps were noise and only large ones real, a single share would mislead. Sorting the champions-only states into ten equal groups by the pick's claimed value, what showed up follows the claim across the range, and a curvature term is indistinguishable from zero.
| Claimed (points) | Draft states | Observed (points) | 95% range |
|---|---|---|---|
| −3.16 | 4,516 | −2.96 | −4.42 – −1.67 |
| −1.43 | 4,515 | −2.14 | −3.42 – −0.78 |
| −0.71 | 4,515 | −1.18 | −2.75 – +0.13 |
| −0.18 | 4,516 | −0.49 | −1.85 – +0.84 |
| +0.26 | 4,515 | −1.72 | −3.26 – −0.27 |
| +0.67 | 4,515 | +0.55 | −0.73 – +1.87 |
| +1.09 | 4,516 | +0.63 | −0.70 – +2.03 |
| +1.58 | 4,527 | +1.32 | 0.00 – +2.74 |
| +2.25 | 4,503 | +2.81 | +1.54 – +4.23 |
| +3.89 | 4,516 | +3.20 | +1.85 – +4.43 |
The ranges are wide: each group's result is uncertain by more than a point. The picture supports a line through zero, not any particular bend.
Does scaling help the forecast?
On the same states, three forecasts of who wins: the board alone, the board plus the model's full gap for the pick made, and the board plus the gap at the share the site had been using. Knowing the pick beats the board alone on AUC, log loss and Brier score. Scaling the gap down did not beat the full gap on these games.
| Forecast | AUC | Log loss | Brier | Accuracy | ECE |
|---|---|---|---|---|---|
| Board alone | 0.5391 | 0.6903 | 0.2486 | 52.64% | 0.0074 |
| Board + the model's full gap | 0.5445 | 0.6895 | 0.2482 | 53.08% | 0.0062 |
| Board + the gap at the earlier share (0.675) | 0.5440 | 0.6896 | 0.2482 | 53.12% | 0.0055 |
What this changes on the site
Arena reports and game reviews print a cost in observed outcomes: the model's gap times the share, with a 95% range from the share's own uncertainty. An alternative is called likely better only when the low end of that range still clears half a point.
Rankings, grades and ratings read the order of the picks, which a common share cannot change. The draft arena and the game review use the same numbers.
What it does not show
- It is observational. Players pick champions for reasons the model only partly sees, so a share near 1 is necessary for the advice to be right, not proof that a different pick would have won.
- Four days, one region, one model. The site now serves a later model trained on all three regions, which has not been measured this way.
- The data has no observed pick order, so each side's order is assigned the way training assigns it.
- The player-known figure knows the picking player only; the review's “everyone known” reading knows all ten.
- Inside a player's own pool, picks valued up to about two points above the board showed almost nothing; the effect sat in the tails.
What should we measure next? Ask on the research page. The most-voted questions our data can answer become future studies.