Benchmarks
What every win-prediction model reports about itself
Every paper and website we know of that publishes a number for its own League of Legends (or MOBA) win predictor. The numbers are theirs, written on one scale; a dash means the source does not report that metric, and we never compute one for it. Each result says when the forecast is made, how its test games were kept apart from training, what it was tested on, and how far we would trust it.
Every source was checked on Sep 22, 2026. Our own numbers, with every curve behind them, are on the accuracy page.
0of 57 published reports give all five metrics for their own model (ours not counted).
The most any report gives for one model: 4 of the five: 4 · 3 of the five: 1 · 2 of the five: 10 · 1 of the five: 36 · none of the five, only other measures: 6.
Numbers
All five are decimals. AUC and accuracy run from 0.5, a coin flip, to 1; log loss, Brier and ECE measure error, so lower is better. A dotted underline means we rewrote how the source printed a number, for example 92.2% as 0.922: hover it to see the original. We never change the number itself.
Predicted
When the forecast is made: before the draft, during it, after the last pick but before the game starts, or during the game. Below it, whether it is for one game or map, or for a whole series. Solo queue and pro are grouped by it, so like sits beside like.
Reliability
- High: tested on games played after the ones it learned from, at least 1,000 test games, and no known problem.
- Medium: tested on later games, but fewer of them or an unstated number.
- Unknown: the source does not say how it tested.
- Low: a random split, a tiny test, or a known problem.
Flags
- Red a leak (admitted by the source, or found by us in its code), a score no real game decided, or in-game data.
- Amber a random split, an unstated split, a small test, or a result far above anything tested on later games. Hover a flag for the reason.
After the draft 15
All ten champions are known and the game has not started. The like-for-like set for our two full-draft models.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
KayLoLWebsiteKayLoL /metrics, layer 1: the held-out patchKayLoL · kaylol.gg · 2026-09-09 build (live payload generated 2026-09-23T01:22:01Z) · checked Sep 23, 2026NotesA second layer on the same page scores every day's games with a model frozen before that day (1.9M predictions per mode, na1, 2026-08-27 onward) and agrees with these within 0.004 on every metric. | all_info_puuid — full draft, all ten players identified | After the draftOne game | 0.6554 | 0.6539 | 0.2312 | 0.6105 | 0.0036 | Hightested on later games · 1,164,446 test games | temporalwhole-patch holdout: trained on patches 16.13-16.16, tested on 16.17, which the model never saw | NA ranked solo queue (na1), all tiers; patch 16.17, 2026-08-26 to 09-09; 1,164,446 held-out games, scored on the validation half |
| KayLoL /metrics, layer 1: the held-out patch | all_info_anonymous — full draft, no player identities (champions only) | After the draftOne game | 0.5852 | 0.6813 | 0.2442 | 0.5599 | 0.0032 | Hightested on later games · 1,164,446 test games | temporalwhole-patch holdout: trained on patches 16.13-16.16, tested on 16.17, which the model never saw | NA ranked solo queue (na1), all tiers; patch 16.17, 2026-08-26 to 09-09; 1,164,446 held-out games, scored on the validation half |
WebsiteDraftGap vs LoLDraftAI: A Detailed ComparisonLoLDraftAI (no visible byline; site-wide HTML meta author 'looyyd') · loldraftai.com · 2026-04-19 ('April 19, 2026') · checked Sep 22, 2026NotesFiles published beside the post (exist; not downloaded): https://media.loldraftai.com/blog/draftgap-vs-loldraftai-comparison/predictions.csv.../draftgap-current-patch.json.../draftgap-30-days.json. ECE is described as the 'average gap between stated probability and observed frequency' (bins not stated); no AUC. The post adds: 'The shipped LoLDraftAI product additionally uses side information and fine-grained elo (features DraftGap doesn't have), but they weren't needed to produce this result.' A vendor's comparison against a rival. | LoLDraftAI (side-agnostic: 'Side-agnostic prediction (no blue/red knowledge), matching DraftGap's side-blind nature')Also reports: 95% bootstrap CIs: log loss (0.6811–0.6846), accuracy (55.34–56.42), Brier (0.2441–0.2458) | After the draftOne game | — | 0.6829 | 0.2449 | 0.5588 | 0.0088 | Hightested on later games · 32,750 games in the data, test share not stated | temporal'LoLDraftAI training cutoff at 2026-04-17, 23:25 UTC' | 32,750 matches: ranked solo/duo, Emerald+ ('DraftGap pulls its data from Lolalytics at tier=emerald_plus'), EUW1 and KR; 'Matches where any (champion, role) cell has fewer than 50 games in DraftGap's current-patch dataset are excluded'; exact dates and patches not stated |
PaperDraftRec: Personalized Draft Recommendation for Winning in Multi-Player Online Battle Arena GamesHojoon Lee, Dongyoon Hwang, Hyunseung Kim, Byungkun Lee, Jaegul Choo · WWW 2022 (The ACM Web Conference); arXiv:2204.12750 · 2022-04-27 (arXiv v1, 'Accepted to WWW 2022') · checked Sep 22, 2026NotesThe Dota 2 rows are the same paper's second dataset (Dota 2 is not this file's game). Generic baselines in the same tables, not listed above: MC (majority class, 'Blue for LOL and Radiant for Dota2') 0.5040 LoL and 0.5180 Dota 2; LR 0.5255 during and 0.5323 after the draft (LoL), 0.5750 and 0.6126 (Dota 2); NN 0.5263 and 0.5335 (LoL), 0.5748 and 0.6108 (Dota 2). After the Dota 2 draft, LR (0.6126) prints above DraftRec (0.6110). Outcome prediction is the paper's secondary task; its main task is champion recommendation. Code and data: github.com/dojeon-ai/DraftRec. | DraftRec — LoL, after the draft (all ten champions selected)Also reports: MAE 0.4826 (accuracy and MAE both starred: p<0.01) | After the draftOne game | — | — | — | 0.5618 | — | Hightested on later games · about 27,989 test games | temporal | Last 10% by time of 279,893 LoL matches collected through the Riot Games API: (62,466 players); region and queue not stated; test-set size not stated |
| DraftRec: Personalized Draft Recommendation for Winning in Multi-Player Online Battle Arena Games | DraftRec-no-history — LoL, after the draft (all ten champions selected)Also reports: MAE 0.4893 | After the draftOne game | — | — | — | 0.5432 | — | Hightested on later games · about 27,989 test games | temporal | Last 10% by time of 279,893 LoL matches collected through the Riot Games API: (62,466 players); region and queue not stated; test-set size not stated |
WebsiteLoLTheory vs LoLDraftAI: A Detailed ComparisonLoLDraftAI (no visible byline; site-wide HTML meta author 'looyyd') · loldraftai.com · 2026-04-19 ('April 19, 2026') · checked Sep 22, 2026NotesPer-match predictions published: https://media.loldraftai.com/blog/loltheory-vs-loldraftai-comparison/predictions.csv (exists; not downloaded). 'Parenthetical ranges are 95% bootstrap confidence intervals'. The ECE bin count is not stated; no AUC is reported. The post describes LoLTheory as, 'primarily an Overwolf overlay app (~166,000 downloads)', and notes. Headline: 'On every one of these, LoLDraftAI beats LoLTheory.' A vendor's comparison against a rival, not an independent evaluation. | LoLDraftAI (side-agnostic mode)Also reports: 95% bootstrap CIs: log loss (0.6807–0.6842), accuracy (55.37–56.43), Brier (0.2439–0.2456) | After the draftOne game | — | 0.6824 | 0.2447 | 0.5590 | 0.0100 | Hightested on later games · 32,930 games in the data, test share not stated | temporal'LoLDraftAI training cutoff at 2026-04-17, 23:25 UTC — the latest game timestamp in the model's training data'; 'Eval set = matches played after both cutoffs.' | 32,930 matches; 'Eval set = matches played after both cutoffs'; exact dates and patches of the matches not stated (LoLTheory 'queried live on patch 16.8.1 during April 2026') |
PaperOnline Game Outcome Prediction Model Using Weighted-Based Feature ApproachM. Asyhraf Zamir Zamri, Nurul Aswa Omar, Isredza Rahmi A. Hamid · Fusion: Practice and Applications, Vol. 15, No. 2, pp. 132-144 (2024) · 2024 (received 2023-08-12, revised 2023-12-25, accepted 2024-04-15) · checked Sep 22, 2026NotesSample counts disagree inside the paper (4,552 NA + 12,458 LAN; '12,458 samples... from Gonzalez'; 'a dataset of 28000 samples'). Each feature becomes a 1/0 'dominance' value ('If team A's performance is better than team B's for a feature, assign a value of 1; otherwise, assign 0'), and the text also says 'The feature's value indicator is a binary value of 1 if the blue team won and 0 if the blue team did not win'. The weights are computed from counts of 'data classified as 1' and 'data classified as 0' over the whole dataset, and the paper does not state that the 'class avg' and 'class med' accuracies are scored against the real match outcome; pregame is therefore 'unclear'. The paper frames the model for 'Solo Queue Ranked Match (SQRM)' but does not name its data's queue. F-measure values print between 4.0000 and 5.0160. Table 1's accuracy column is misaligned with the text (the text gives Lee et al. 62.26% at 5 minutes and 73% at 15, Silva et al. 63.91% and 83.54%, Ani et al. over 90%, Do et al. 75.10%, Gonzalez 82-90.48%); the cross-reports follow the text. | weighted-based feature predictor with Naïve Bayes and Support Vector Machine (WEKA BayesNet and SMO)Also reports: Conclusion | After the draftOne game | — | — | — | >0.97 | — | Lownot tested on later games · 4,552 games in the data, test share not statedRandom splitFar above the rest | random(WEKA, Section 3.1) | NA and LAN matches: (Gonzalez); Section 3.1 says 12,458 samples from Gonzalez were used, and Section 3.2 'a dataset of 28000 samples' partitioned into sizes 1000 to 7000 (results also shown at 12458); queue, tier and dates not stated |
PaperPlayer Skill Decomposition in Multiplayer Online Battle ArenasZhengxing Chen, Yizhou Sun, Magy Seif El-Nasr, Truong-Huy D. Nguyen · 2016 Meaningful Play Conference; arXiv:1702.06253 · 2017-02-21 (arXiv v1) · checked Sep 22, 2026NotesThe Dota 2 rows are the same paper's second dataset. Majority-class baseline (BL-MC): 53.22% ± 0.16% LoL, 52.65% ± 0.12% Dota 2.. The weights are learned from match outcomes, so under 10-fold CV a player's weight is fit on that player's other matches, earlier and later. The paper's significance rule: mean1 − mean2 > 2(std1 + std2). | LR-P-C-PC (player + champion + player-champion weights)Also reports: ± 0.16% (std over 10 folds) | After the draftOne game | — | — | — | 0.6024 | — | Lownot tested on later games · 231,212 games in the data, test share not statedRandom split | random | 10-fold CV over 231,212 unique ranked matches played in 2015 by 972 seed players of the top two tiers (Challenger and Master) on the North America server, through the official Riot API (93,098 players, 129 champions); the paper studies '5-vs-5 ranked matches of solo queue' |
PaperUsing Machine Learning to Predict Game Outcomes Based on Player-Champion Experience in League of LegendsTiffany D. Do, Seong Ioi Wang, Dylan S. Yu, Matthew G. McMillian, Ryan P. McMahan · FDG 2021 (16th International Conference on the Foundations of Digital Games); arXiv:2108.02799 · 2021-08-05 (arXiv v1); FDG '21, August 3-6, 2021 · checked Sep 22, 2026NotesFeatures as defined in the paper: champion mastery points ('lifetime experience'), player-champion win rate, season games on the champion, and ranked games on the champion within the last 20. The paper does not say whether the season aggregates were read before each match or at collection time. Other models in Table 1: RF 74.7%, SVC 74.3%, kNN 72.7% (kNN uses only 'the team's average win rate'). The authors prefer the DNN over GBOOST because GBOOST's standard error is much larger. | Deep neural network (5 dense layers)Also reports: ± 1.2% (95% CI); std. dev. 1.9%; std. error 0.60% | After the draftOne game | — | — | — | 0.751 | — | Lownot tested on later games · 5,000 games in the data, test share not statedRandom splitFar above the rest | random | 5,000 unique ranked matches from the North American server, 2020, ranks Iron to Diamond, pulled through the Riot API ; Table 1 gives each model's mean accuracy with a 95% CI, and the paper does not say whether that is the 10-fold mean or the 1,000-match test set |
| Using Machine Learning to Predict Game Outcomes Based on Player-Champion Experience in League of Legends | Gradient boosting (GBOOST)Also reports: ± 1.19% (95% CI); std. dev. 5.25%; std. error 1.66% | After the draftOne game | — | — | — | 0.754 | — | Lownot tested on later games · 5,000 games in the data, test share not statedRandom splitFar above the rest | random | 5,000 unique ranked matches from the North American server, 2020, ranks Iron to Diamond, pulled through the Riot API ; Table 1 gives each model's mean accuracy with a 95% CI, and the paper does not say whether that is the 10-fold mean or the 1,000-match test set |
WebsiteiTero vs LoLDraftAI: A Detailed ComparisonLoLDraftAI (no visible byline; site-wide HTML meta author 'looyyd') · loldraftai.com · 2026-04-19 ('April 19, 2026') · checked Sep 22, 2026NotesFull tournament data published as drafts_scored.jsonl (exists; not downloaded). The post infers iTero's model ('consistent with a gradient-boosted tree ensemble over lolalytics-style aggregates') from 'its API response signature'; that is a rival's inference, not iTero's statement. Its feature table says iTero offers 'per-candidate score only; no team WR', which is why no calibrated metric could be scored for iTero. No real match outcomes are involved, so none of these numbers is a forecast score. | LoLDraftAI (as drafter) — judge: DraftGap (third-party)Also reports: LoLDraftAI win rate 55.0% (45.0–65.0); Mean WR margin (LoLDraftAI) +0.6% (−1.0 to +2.2) | After the draftOne game | — | — | — | — | — | Lownot tested on later games · 100 games in the data, test share not stated · known problemNo game resultsNo held-out testSmall test | no held-out testnone: no games are played; 100 synthetic greedy drafts are scored by a judge model's predicted win rate | 100 head-to-head greedy drafts; patch and date not stated |
| iTero vs LoLDraftAI: A Detailed Comparison | LoLDraftAI (as drafter) — judge: LoLDraftAI's own win-probability head (flagged non-independent)Also reports: LoLDraftAI win rate 93.0% (88.0–98.0); Mean WR margin (LoLDraftAI) +13.2% (+11.6 to +14.9) | After the draftOne game | — | — | — | — | — | Lownot tested on later games · 100 games in the data, test share not stated · known problemNo game resultsNo held-out testSmall test | no held-out testnone: the judge is the same model that drafted ('LoLDraftAI judging drafts it itself picked is textbook circular') | the same 100 drafts |
| iTero vs LoLDraftAI: A Detailed Comparison | LoLDraftAI (predicted gold at 15 min) against the two judges' disagreement — judge disagreement analysisAlso reports: Pearson r = +0.59 (n = 100); Mean |LoLDraftAI − DraftGap| by predicted gold@15 quartile: Q1 32 – 497 g (n 25) 10.4 pp; Q2 497 – 1,108 g (25) 13.3 pp; Q3 1,108 – 2,162 g (25) 12.6 pp; Q4 2,162 – 4,882 g (25) 14.8 pp | After the draftOne game | — | — | — | — | — | Lownot tested on later games · 100 games in the data, test share not stated · known problemNo game resultsNo held-out testSmall test | no held-out testnone | the same 100 drafts |
Website【LoL】AIがドラフト前にプレイヤー情報だけで76.7%の精度で勝敗予測的中 – MMRは本当に機能しているのか? (lolninja.net report of the Reddit post 'AI can predict the winner of 76.7% of your games before draft even begins! Is the MMR system broken?')いちずなイブリン (lolninja.net), reporting a Reddit user's post · lolninja.net · 2025-04-23 ('2025.04.23') · checked Sep 22, 2026NotesSecondary source: it relays the Reddit post https://www.reddit.com/r/leagueoflegends/comments/1k5hmlf/ai_can_predict_the_winner_of_767_of_your_games/ (not read; Reddit is unreachable to the fetcher). pregame is 'unclear' because the rank feature is taken after the games were played. | Reddit poster's neural net, rank + draft | After the draftOne game | — | — | — | 0.747 | — | Lowsplit not stated · 21,000 games in the data, test share not stated · known problemLeak admittedSplit not stated | not stated'80/10/10分割' (an 80/10/10 split); whether random, temporal or by player is not stated | about 21,000 matches after cleaning (15,000 solo queue, 2,000 flex, 4,000 normal draft) from about 30,000 collected; players 'randomly' gathered from Gold/Platinum; NA; patch 15.7; test-set size not stated |
| 【LoL】AIがドラフト前にプレイヤー情報だけで76.7%の精度で勝敗予測的中 – MMRは本当に機能しているのか? (lolninja.net report of the Reddit post 'AI can predict the winner of 76.7% of your games before draft even begins! Is the MMR system broken?') | Reddit poster's neural net, draft only | After the draftOne game | — | — | — | 0.528 | — | Unknownsplit not stated · 21,000 games in the data, test share not statedSplit not stated | not stated'80/10/10分割' (an 80/10/10 split); whether random, temporal or by player is not stated | about 21,000 matches after cleaning (15,000 solo queue, 2,000 flex, 4,000 normal draft) from about 30,000 collected; players 'randomly' gathered from Gold/Platinum; NA; patch 15.7; test-set size not stated |
WebsiteBugfix and Correction of Reddit Post Accuracy Claims (blog index title: 'Correction: Reddit Post Accuracy Claims')LoLDraftAI (first person, no visible byline; site-wide HTML meta author 'looyyd') · loldraftai.com · 2025-04-07 ('April 7, 2025') · checked Sep 22, 2026NotesRetracted Reddit post: https://www.reddit.com/r/leagueoflegends/comments/1joumtm/i_made_an_ai_model_that_predicts_62_of_ranked/ (not read; Reddit is unreachable to the fetcher). The author thanks '/u/Impossible_Concert88 for trying to verify the accuracy claims' and writes: Queue: the post says 'ranked' (from the Reddit title) without naming solo/duo; level set to solo queue on that basis. | LoLDraftAI, fixed model uploaded April 4, 2025Also reports: 'around 55% (as of April 4 2025)' | After the draftOne game | — | — | — | 0.55 | — | Unknownsplit not stated · test size not statedSplit not stated | not statednot stated (validation set after the duplication fix; construction not described) | not stated |
WebsiteHow it Works: LoLDraftAILoLDraftAI (no visible byline; site-wide HTML meta author 'looyyd') · loldraftai.com · 2025-09-04 ('September 4, 2025') · checked Sep 22, 2026NotesTraining logs linked from the post (not opened here): solo queue https://wandb.ai/loyd-team/draftking/runs/jy5hf0bv?panelDisplayName=val_win_prediction_accuracy&panelSectionName=Charts; pro play https://wandb.ai/loyd-team/draftking-pro-finetune/runs/jg3ls0xp?panelDisplayName=val_pro_win_prediction_accuracy&panelSectionName=Charts. 'We mask randomly because the riot data doesn't have information about the original pick order'... 'the model is not aware of "blind pickability".' The page header 'Last model update: September 19 on patch 16.18' is the site-wide banner for the current model, not the one the post describes. | LoLDraftAI solo queue and pro play models (the training-log links point to the May 2025 model) | After the draftOne game | — | — | — | 0.56 | — | Unknownsplit not stated · test size not statedSplit not stated | not stated'on random drafts they have not been trained on' — read from 'the training logs' (Weights & Biases panels val_win_prediction_accuracy / val_pro_win_prediction_accuracy); how the held-out drafts were separated is not stated | not stated (held-out drafts from the training runs; no dates, regions, elo range or n) |
WebsiteLoLDraftAI homepage: headline stat block and FAQ 'How accurate is the model?'LoLDraftAI (site-wide HTML meta author 'looyyd') · loldraftai.com · not stated; page stamped 'Patch 16.18 · Updated September 19' and 'Last model update: September 19 on patch 16.18' · checked Sep 22, 2026NotesFAQ answer in full: No test set, split, n, curve or ECE is given anywhere on the page. The landing page also shows 'Outperforms DraftGap · Outperforms iTero · Outperforms LoLTheory' badges and a 'Side by side' table whose 'Calibrated win %' row gives LoLDraftAI a check and DraftGap, iTero and LoLTheory a '?'; no metric is attached, so no cross-report is recorded. 'Read the full breakdowns' links to the three comparison posts. Other page claims: 'Re-trained weekly', and. | LoLDraftAI (draft only)Also reports: Calibrated · 55% means 55% | After the draftOne game | — | — | — | 0.567 | — | Unknownsplit not stated · test size not statedSplit not stated | not statednot stated | not stated (the page says 'Silver to Challenger' and 'Trained on millions of ranked games' but gives no test set, dates, region or n for the 56.7%) |
| LoLDraftAI homepage: headline stat block and FAQ 'How accurate is the model?' | LoLDraftAI with runes | After the draftOne game | — | — | — | 0.577 | — | Unknownsplit not stated · test size not statedSplit not stated | not statednot stated | not stated |
PaperPredicting League of Legends Match Outcomes Through Machine Learning Models Using Past Match Player PerformanceChaim Joseph A. Cordova, Carl Victor A. Villaceran, Christine F. Peña · 2024 IEEE International Conference on Computing (ICOCO) · 2024-12-12 · checked Sep 22, 2026Notes. The abstract does not say how the data were split or whether the win rates exclude the predicted match. Queue not named ('ranked matches'). | Gradient Boosting, Logistic Regression and Deep Neural Networks (reported together) | After the draftOne game | — | — | — | 0.975 | — | Unknownsplit not stated · 11,000 games in the data, test share not statedSplit not statedFar above the rest | not statednot stated in the abstract | 'over 11,000 ranked matches' (abstract); source, region, tier and dates not stated |
PaperScalable Psychological Momentum Forecasting in EsportsAlfonso White, Daniela M. Romano · SUM '20 (State-based User Modelling) workshop at WSDM 2020; arXiv:2001.11274 · 2020-01-30 (arXiv v1; v2 2020-02-15); SUM '20, February 3-7, 2020 · checked Sep 22, 2026NotesOther Table 3 rows (test %): post-draft AutoLog+Rolling+LR 72.0, AutoLog+LR 71.8, Loginit+Rolling+(momentum)+LR 70.8, Rolling+(momentum)+LR 68.8, LR baseline 68.3; pre-draft teams AutoLog+Rolling+LR 65.6, AutoLog+LR 65.1, LR baseline 62.3, MTL+RNN 61.6; pre-draft solo AutoLog+Rolling+LR 54.03, AutoLog+LR 53.59, LR baseline 53.38, AutoLog+MTL-TL+RNN 52.48, AutoLog+RNN 52.38. Participant profiles were loaded from op.gg 'Once a valid match is found' (season totals and up to the last 20 games); the paper does not say they exclude the target match. Stated limitation. | AutoLog+Rolling+(momentum)+LR (post-draft)Also reports: train accuracy 73.6%; train n 70,194 | After the draftOne game | — | — | — | 0.721 | — | Unknownsplit not stated · 10,000 test gamesSplit not stated | not stated | 10,000 test matches out of 87,743 collected 'from February 5th to September 20th of 2019' (86.4% Solo/Duo queue, 13.6% Flex; 517,269 unique summoners), crawled through the Riot Games API with op.gg profile summaries and champion.gg / op.gg averages; region not stated; how the 10,000 were drawn is not stated |
PaperStrategic Feature Selection for Draft Prediction in League of LegendsManaschai Aonon, Ponguthai Samrankhong, Teerawat Kamnardsiri · 2026 Joint International Conference on Digital Arts, Media and Technology with ECTI Northern Section Conference (ECTI DAMT & NCON) · 2026-02-04 · checked Sep 22, 2026NotesFeature importance per the abstract: meta strength 26.3%, team synergy 24.8%, counter-matchups 20.6%, runes and spells 6.7% combined. | LightGBM (best of 20 algorithms)Also reports: 'minimal generalization gap (0.27%)' | After the draftOne game | 0.922 | — | — | 0.831 | — | Unknownsplit not stated · 86,556 games in the data, test share not statedSplit not statedFar above the rest | not statednot stated in the abstract (it reports 'test accuracy') | '86,556 high-level matches' (abstract); source, region and dates not stated |
During the draft 5
Some picks are still hidden. Where our three draft-time models sit.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
KayLoLWebsiteKayLoL /metrics, layer 1: the held-out patchKayLoL · kaylol.gg · 2026-09-09 build (live payload generated 2026-09-23T01:22:01Z) · checked Sep 23, 2026NotesA second layer on the same page scores every day's games with a model frozen before that day (1.9M predictions per mode, na1, 2026-08-27 onward) and agrees with these within 0.004 on every metric. | draft_all_puuid — mid-draft, all ten players identified (clash / organised draft) | During the draftOne game | 0.6345 | 0.6635 | 0.2357 | 0.5950 | 0.0035 | Hightested on later games · 1,164,446 test games | temporalwhole-patch holdout: trained on patches 16.13-16.16, tested on 16.17, which the model never saw | NA ranked solo queue (na1), all tiers; patch 16.17, 2026-08-26 to 09-09; 1,164,446 held-out games, scored on the validation half |
| KayLoL /metrics, layer 1: the held-out patch | draft_puuid — mid-draft, own identity plus revealed picks (solo-queue champion select) | During the draftOne game | 0.5642 | 0.6864 | 0.2467 | 0.5446 | 0.0041 | Hightested on later games · 1,164,446 test games | temporalwhole-patch holdout: trained on patches 16.13-16.16, tested on 16.17, which the model never saw | NA ranked solo queue (na1), all tiers; patch 16.17, 2026-08-26 to 09-09; 1,164,446 held-out games, scored on the validation half |
| KayLoL /metrics, layer 1: the held-out patch | draft_anonymous — mid-draft, no identities | During the draftOne game | 0.5495 | 0.6888 | 0.2479 | 0.5331 | 0.0042 | Hightested on later games · 1,164,446 test games | temporalwhole-patch holdout: trained on patches 16.13-16.16, tested on 16.17, which the model never saw | NA ranked solo queue (na1), all tiers; patch 16.17, 2026-08-26 to 09-09; 1,164,446 held-out games, scored on the validation half |
PaperDraftRec: Personalized Draft Recommendation for Winning in Multi-Player Online Battle Arena GamesHojoon Lee, Dongyoon Hwang, Hyunseung Kim, Byungkun Lee, Jaegul Choo · WWW 2022 (The ACM Web Conference); arXiv:2204.12750 · 2022-04-27 (arXiv v1, 'Accepted to WWW 2022') · checked Sep 22, 2026NotesThe Dota 2 rows are the same paper's second dataset (Dota 2 is not this file's game). Generic baselines in the same tables, not listed above: MC (majority class, 'Blue for LOL and Radiant for Dota2') 0.5040 LoL and 0.5180 Dota 2; LR 0.5255 during and 0.5323 after the draft (LoL), 0.5750 and 0.6126 (Dota 2); NN 0.5263 and 0.5335 (LoL), 0.5748 and 0.6108 (Dota 2). After the Dota 2 draft, LR (0.6126) prints above DraftRec (0.6110). Outcome prediction is the paper's secondary task; its main task is champion recommendation. Code and data: github.com/dojeon-ai/DraftRec. | DraftRec (LoL, during the draft)Also reports: MAE 0.4842 (accuracy and MAE both starred: p<0.01) | During the draftOne game | — | — | — | 0.5535 | — | Hightested on later games · about 27,989 test games | temporal | Last 10% by time of 279,893 LoL matches collected through the Riot Games API: (62,466 players); region and queue not stated; test-set size not stated |
Website+4% to +7%: How Much the Top Suggestion Helps, by Counterpick SkillLoLDraftAI (no visible byline; site-wide HTML meta author 'looyyd') · loldraftai.com · 2026-04-28 ('April 28, 2026') · checked Sep 22, 2026Notes'This is model-predicted lift, not measured outcomes' — the model scores its own suggestions; no game result enters these numbers. Naive baseline: 'uniformly at random among the revealed slots'; strongest-visible. The model 'implicitly assume[s] the opponent's downstream picks don't change in response to ours'. Only Bottom and Utility (roles) and Silver and Master+ (elo) are broken out in the text; no overall average is printed. This study is the source of the homepage's '+4–7%'. | LoLDraftAI top suggestion (model-predicted lift, self-scored) — pick 1 (Blue, 0 visible picks)Also reports: lift vs naive baseline +3.0%; lift vs strongest-visible baseline +3.0% | During the draftOne game | — | — | — | — | — | Lowsplit not stated · 10,000 games in the data, test share not stated · known problemNo game resultsSplit not stated | not statedhow that test set was held out is not stated | 10,000 ranked games |
| +4% to +7%: How Much the Top Suggestion Helps, by Counterpick Skill | LoLDraftAI top suggestion (model-predicted lift, self-scored) — pick 2 (Red, 1 visible picks)Also reports: lift vs naive baseline +3.8%; lift vs strongest-visible baseline +3.8% | During the draftOne game | — | — | — | — | — | Lowsplit not stated · 10,000 games in the data, test share not stated · known problemNo game resultsSplit not stated | not statedhow that test set was held out is not stated | 10,000 ranked games |
| +4% to +7%: How Much the Top Suggestion Helps, by Counterpick Skill | LoLDraftAI top suggestion (model-predicted lift, self-scored) — pick 3 (Red, 2 visible picks)Also reports: lift vs naive baseline +4.1%; lift vs strongest-visible baseline +3.5% | During the draftOne game | — | — | — | — | — | Lowsplit not stated · 10,000 games in the data, test share not stated · known problemNo game resultsSplit not stated | not statedhow that test set was held out is not stated | 10,000 ranked games |
CodeLeagueOfPredictions: Predictive Analytics for League of Legendsaliciusschroeder (GitHub) · GitHub · README undated; repo created 2023-05-17, last push 2024-10-15 (GitHub API) · checked Sep 22, 2026NotesLevel set to solo queue from 'ranked games in silver elo' (queue not named). 'still in the experimental phase'. Code not audited here; whether any input leaks the outcome is not answerable from the README. | LeagueOfPredictions model | During the draftOne game | — | — | — | 0.879 | — | Unknownsplit not stated · 300,000 games in the data, test share not statedSplit not statedFar above the rest | not statednot stated (README names train.py and validate.py but no protocol) | not stated beyond 'trained the model on 300,000 datasets' and 'ranked games in silver elo'; region, queue, patch and dates not stated |
WebsiteWhat's the Best League of Legends Draft AI?Winrate (winrate.gg; no named author) · winrate.gg · 2026-04-28 ('April 28, 2026') · checked Sep 22, 2026Notes'This is a correlation, not a causal experiment. We didn't tell anyone to follow the recommendation.' 'That's a 5.4-point spread between top and bottom' (Top 10). DraftGap left out ('static aggregate stats... rather than a learned model'); LoLTheory left out ('LoLDraftAI published a head-to-head where they outperformed it'). The article describes iTero's model as 'gradient-boosted trees with linear models'. Only the anonymous winrate.gg model was tested; the personalized model (~40 more features) was not. The realized win rate is not an accuracy, so the five metric fields are null. Our check: The article does not say whether the statistics behind winrate.gg's own suggestions excluded the 12 hours of games being tested. | Winrate anonymous draft model (GBDT, 58 features) — Top 1Also reports: realized win rate when the actual pick was in the tool's Top 1: 58.7% (n 300) | During the draftOne game | — | — | — | — | — | Unknownsplit not stated · 20,000 games in the data, test share not statedSplit not stated | not statedTest games are the ranked games of the 12 hours before the query ('match_timestamp >= toUnixTimestamp64Milli(now64 - INTERVAL 12 HOUR)'); the article does not say whether winrate.gg's model or its champion aggregates had already seen those games | 20,000 'real ranked solo-queue games (queue 420)', 'Gold tier and above', 'Recent live patches', both teams' five roles assigned, from the 12 hours before the query ('ORDER BY sipHash64(match_id) LIMIT 20000'); one champion erased per game, nine visible; region not stated; 'n is the number of games where the actual pick was in the site's Top K' |
| What's the Best League of Legends Draft AI? | Winrate anonymous draft model (GBDT, 58 features) — Top 3Also reports: realized win rate when the actual pick was in the tool's Top 3: 56.0% (n 1,131) | During the draftOne game | — | — | — | — | — | Unknownsplit not stated · 20,000 games in the data, test share not statedSplit not stated | not statedTest games are the ranked games of the 12 hours before the query ('match_timestamp >= toUnixTimestamp64Milli(now64 - INTERVAL 12 HOUR)'); the article does not say whether winrate.gg's model or its champion aggregates had already seen those games | 20,000 'real ranked solo-queue games (queue 420)', 'Gold tier and above', 'Recent live patches', both teams' five roles assigned, from the 12 hours before the query ('ORDER BY sipHash64(match_id) LIMIT 20000'); one champion erased per game, nine visible; region not stated; 'n is the number of games where the actual pick was in the site's Top K' |
| What's the Best League of Legends Draft AI? | Winrate anonymous draft model (GBDT, 58 features) — Top 5Also reports: realized win rate when the actual pick was in the tool's Top 5: 55.2% (n 2,004) | During the draftOne game | — | — | — | — | — | Unknownsplit not stated · 20,000 games in the data, test share not statedSplit not stated | not statedTest games are the ranked games of the 12 hours before the query ('match_timestamp >= toUnixTimestamp64Milli(now64 - INTERVAL 12 HOUR)'); the article does not say whether winrate.gg's model or its champion aggregates had already seen those games | 20,000 'real ranked solo-queue games (queue 420)', 'Gold tier and above', 'Recent live patches', both teams' five roles assigned, from the 12 hours before the query ('ORDER BY sipHash64(match_id) LIMIT 20000'); one champion erased per game, nine visible; region not stated; 'n is the number of games where the actual pick was in the site's Top K' |
Before the draft 3
The players are known, the champions are not.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
Website【LoL】AIがドラフト前にプレイヤー情報だけで76.7%の精度で勝敗予測的中 – MMRは本当に機能しているのか? (lolninja.net report of the Reddit post 'AI can predict the winner of 76.7% of your games before draft even begins! Is the MMR system broken?')いちずなイブリン (lolninja.net), reporting a Reddit user's post · lolninja.net · 2025-04-23 ('2025.04.23') · checked Sep 22, 2026NotesSecondary source: it relays the Reddit post https://www.reddit.com/r/leagueoflegends/comments/1k5hmlf/ai_can_predict_the_winner_of_767_of_your_games/ (not read; Reddit is unreachable to the fetcher). pregame is 'unclear' because the rank feature is taken after the games were played. | Reddit poster's neural net, rank only | Before the draftOne game | — | — | — | 0.767 | — | Lowsplit not stated · 21,000 games in the data, test share not stated · known problemLeak admittedSplit not statedFar above the rest | not stated'80/10/10分割' (an 80/10/10 split); whether random, temporal or by player is not stated | about 21,000 matches after cleaning (15,000 solo queue, 2,000 flex, 4,000 normal draft) from about 30,000 collected; players 'randomly' gathered from Gold/Platinum; NA; patch 15.7; test-set size not stated |
Codeleague_winrate_model ('Source code for Reddit post')ksavino1 (GitHub) · GitHub · 2025-04-22 (repo created; last push 2025-06-07, GitHub API) · checked Sep 22, 2026NotesREADME... Found via the lolninja.net article's code link; one fetch only, not chased further. The two players' profile links in the README are left out here. | rank-based model from the Reddit post — re-check on fresh games with rank at game timeAlso reports: 23/30 right; one tailed p = 0.002611 | Before the draftOne game | — | — | — | — | — | Unknownsplit not stated · test size not statedSplit not stated | not statedfresh games scored with op.gg ranks 'at time the game happened, not current ranks'; relation to the training data not stated | 30 solo/duo games of two players |
PaperScalable Psychological Momentum Forecasting in EsportsAlfonso White, Daniela M. Romano · SUM '20 (State-based User Modelling) workshop at WSDM 2020; arXiv:2001.11274 · 2020-01-30 (arXiv v1; v2 2020-02-15); SUM '20, February 3-7, 2020 · checked Sep 22, 2026NotesOther Table 3 rows (test %): post-draft AutoLog+Rolling+LR 72.0, AutoLog+LR 71.8, Loginit+Rolling+(momentum)+LR 70.8, Rolling+(momentum)+LR 68.8, LR baseline 68.3; pre-draft teams AutoLog+Rolling+LR 65.6, AutoLog+LR 65.1, LR baseline 62.3, MTL+RNN 61.6; pre-draft solo AutoLog+Rolling+LR 54.03, AutoLog+LR 53.59, LR baseline 53.38, AutoLog+MTL-TL+RNN 52.48, AutoLog+RNN 52.38. Participant profiles were loaded from op.gg 'Once a valid match is found' (season totals and up to the last 20 games); the paper does not say they exclude the target match. Stated limitation. | AutoLog+Rolling+(momentum)+LR — pre-draft, teamsAlso reports: train accuracy 66.8%; train n 70,194 | Before the draftOne game | — | — | — | 0.657 | — | Unknownsplit not stated · 10,000 test gamesSplit not stated | not stated | 10,000 test matches out of 87,743 collected 'from February 5th to September 20th of 2019' (86.4% Solo/Duo queue, 13.6% Flex; 517,269 unique summoners), crawled through the Riot Games API with op.gg profile summaries and champion.gg / op.gg averages; region not stated; how the 10,000 were drawn is not stated |
| Scalable Psychological Momentum Forecasting in Esports | AutoLog+MTL+RNN — pre-draft, teamsAlso reports: train accuracy 64.5%; train n 70,194 | Before the draftOne game | — | — | — | 0.644 | — | Unknownsplit not stated · 10,000 test gamesSplit not stated | not stated | 10,000 test matches out of 87,743 collected 'from February 5th to September 20th of 2019' (86.4% Solo/Duo queue, 13.6% Flex; 517,269 unique summoners), crawled through the Riot Games API with op.gg profile summaries and champion.gg / op.gg averages; region not stated; how the 10,000 were drawn is not stated |
| Scalable Psychological Momentum Forecasting in Esports | AutoLog+Rolling+(momentum)+RNN — pre-draft, single player in queueAlso reports: train accuracy 54.36%; train n 701,940 | Before the draftOne game | — | — | — | 0.5430 | — | Unknownsplit not stated · test size not statedSplit not stated | not stated | Single-player samples from the 701,940 individual match histories, 'within the same folds', of the same 2019 collection; test size not stated |
Show a win chance, publish no test
Tools that show a win percentage but, as far as we could find, publish no evaluation of it.
- How the Draft Model Works (winrate.gg)Publishes no accuracy, AUC, log loss, Brier or calibration figure for its own win model (checked 2026-09-22; also the homepage, the articles index and 'Measuring What Matters: From Stats to Win Probability', 2026-02-04, which has none). Training; 'we don't have the actual pick order'. The only evaluation winrate.gg publishes is the 2026-04-28 recommender benchmark (winrate-gg-draft-ai-benchmark-2026-04).
- Global Power Rankings 2026 page and the Worlds 2025 primer's GPR changesNo accuracy, Brier or log loss is published for the revised system (checked 2026-09-22 on the 2026 GPR page and https://lolesports.com/en-US/news/worlds-2025-primer). The primer: 'daily updates rolling out after each match day' and. The 2026 page's 'More Info' link goes to the 2024 dev diary, whose 65% describes the earlier 80/20 system. Agrees with sources.md and benchmarks.md.
- LoL: Can Solo Queue data be used for Professional playNo accuracy, AUC, log loss or Brier for any iTero model: the post says its model became 'significantly more accurate' with solo-queue data and is 'evaluated on how accurately it predicts future games', without a number (checked 2026-09-22).
- iTero homepagePublishes no accuracy or evaluation claim for its own model (checked 2026-09-22). The page shows '+1,000,000 Downloads', '4.4 Ratings' and user testimonials (e.g. 'My winrate has improved to 65% in 105 games'), which are user anecdotes, not an evaluation. iTero is measured only by rivals: loldraftai-vs-itero-2026-04 and winrate-gg-draft-ai-benchmark-2026-04.
- PropsBot.AI: track record and LoL predictions pageNon-MOBA figures on the track-record page, recorded here only as context: NFL '13.7% ROI on 8,210 graded NFL picks' (56.9% win rate) and '73.2% win rate on 20,222 graded NFL picks' (−2.0% ROI); NBA '25.9% ROI on 81,410 graded NBA picks' (47.1% win rate); MLB 'Brier 0.1903 vs Vegas 0.1947'. Logging: Shows no LoL win % record and publishes no LoL evaluation (checked 2026-09-22).
- sport.gg: League of Legends predictionsShows win probabilities but publishes no track record, accuracy, Brier or log loss (checked 2026-09-22). It states a calibration goal and that 'For best-of-five series, probabilities are tighter than best-of-one matches'. Matches comparison-protocol.md §8b.
- DraftGap (GitHub README and draftgap.com)Publishes no evaluation of itself (checked 2026-09-22): the README is seven lines with no metric, the repository holds no other documentation besides AGENTS.md, and draftgap.com returned only a title to the fetcher. Its formula produces a draft win rate, which LoLDraftAI scored in loldraftai-vs-draftgap-2026-04 (the only published evaluation of DraftGap).
- LoLTheory (loltheory.gg)Publishes no evaluation of itself on its site (checked 2026-09-22): the homepage shows '90K+ downloads' and no accuracy, calibration or model-documentation link, and no draft win % on the page read. LoLDraftAI's comparison post says LoLTheory's developer ('Griffin') stated a '54–56% range' on Reddit; recorded as a cross-report in loldraftai-vs-loltheory-2026-04, not verified (Reddit not read).
- LOL Brain (lol-brain.com)Shows a win % but publishes no evaluation (checked 2026-09-22); the draft calculator on the page displays a sample 'Win Probability' of 54% against 46%.
- Smartpick (smartpick.gg)Publishes no evaluation (checked 2026-09-22); the page read showed no win % either, only the slogan 'Smarter picks, stronger drafts, higher win rates'.
- Draft Genius (draftgenius.lol)Shows a win % but publishes no evaluation (checked 2026-09-22); the page displays a draft win-probability bar ('50.0 Ally' / '50.0 Enemy' in the empty state).
- Baron Buff (baronbuff.com)Publishes no evaluation (checked 2026-09-22); the page read does not show a draft win %.
- DraftForge (draftforge.gg)Publishes no evaluation (checked 2026-09-22); the page read does not show a draft win %, and its only quantified claim is the 200,000+ pairs it draws on.
- Drafter.one ('The First League of Legends Drafting Game'; model 'DrafterEye')drafter.one and drafter.one/game return only a title to the fetcher. A search-engine summary (secondary, unverified, apparently from the product's X account @Drafter_GG) says 'Drafter's algorithm achieved a 71% prediction rate during the knock out stage' and '71% prediction on Worlds games (5/7)'; the 2026-09-11 survey logged the same '71% prediction rate' as a search snippet. Not recorded as a metric until read at the source.
After the draft, one map 5
Every pick known, one map at a time. The like-for-like set for our pro model after the draft.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
KayLoLWebsiteKayLoL /metrics: pro tournament mapsKayLoL · kaylol.gg · live payload generated 2026-09-23T01:22:01Z · checked Sep 23, 2026NotesA backfill, as our own scoreboard states: out of sample for each fit, not for the design choices, which were made looking at the same seasons. | pro model — after the draftAlso reports: calibration slope 0.9976 | After the draftOne map | 0.7359 | 0.5904 | 0.2039 | 0.6749 | 0.0254 | Hightested on later games · 1,784 test games | walk-forwardwalk-forward: each map scored by a model fitted only on the seasons before it | 1,784 pro tournament maps, 2026-06-29 onward, 29 leagues |
CodeLoL Match Prediction (Ra1nForest/LOL_Predictions)Ra1nForest (GitHub) · GitHub · 2026-09-12 to 2026-09-23 (repo created, last pushed 2026-09-23 00:06 UTC; README updates dated 2026-08-16, 2026-08-24, 2026-09-13) · checked Sep 22, 2026NotesIn-game Stage 4 results exist but use live game state, so they are left out of self_reported: '73.7% (80.5% at 25 min)' against a 56.3% baseline (28 features); live variant without XP 72.9%; isotonic re-test (2026-08-24) 'ECE 0.0351 → 0.0160', accuracy '77.2% → 77.3%'.. 'There is no server any more (since 2026-09-13)'; the public site ra1nforest.github.io/LOL_Predictions runs the stages in the browser. No forward record is published. | Stage 2 post_draft (112 features)Also reports: baseline 53.3% | After the draftOne map | — | — | — | 0.611 | — | Hightested on later games · 8,453 test games | walk-forwardthe README's evaluation apparatus; comparisons 'are paired', accepted only when |t| > 2.5. The stage table itself does not restate which evaluation produced it | 8,453 games from Oracle's Elixir, LPL / LCK / LEC / LCS, 2022–2026 (fold windows not listed; 'fold 8 now ends 2026-08-10') |
CodeLoL Esports Match Prediction Model (Hget15/lol-esports-predictor)Hget15 (GitHub) · GitHub · 2026-09-14 to 2026-09-16 (repo created and last pushed, GitHub API) · checked Sep 22, 2026NotesV1–V3 log loss is not printed. Top V4 feature: 'patch_wr_delta_diff' at 25.3% importance . The live demo (hget15.github.io/lol-esports-predictor) runs, not the evaluated model. The README credits V4 with 'the single biggest jump — AUC 0.698 → 0.797'. Our check: We read the code (upstream commit e50e9ce, 2026-09-20). The V4 features come from trackers filled with every game of 2021-2026 before any feature is computed, so a team's win rate on the current patch includes the game being predicted and the games after it (src/v4_enhancements.py: build_all_from_raw, line 223; get_patch_wr, line 85). The V1-V3 results do not use those features. | V4 gradient-boosted trees (130 features) — V4 (headline)Also reports: worklog.md: AUC 0.7970, Brier 0.1879; results/model_comparison.csv: AUC 0.7969557771451893, Brier 0.1878867184551877, Log Loss 0.5582568789963737, Accuracy 0.7214140238580157 | After the draftOne map | 0.797 | 0.558 | 0.188 | 0.721 | — | Lowtested on later games · about 7,663 test games · known problemLeak found | temporalworklog.md | held-out test set: last 15% by date of 51,088 Oracle's Elixir games (2021–2026), '~6,874 games' (worklog.md); leagues not listed |
PaperFeature Analysis to League of Legends Victory Prediction on the Picks and Bans PhaseLincoln Magalhães Costa, Rafael Gomes Mantovani, Fracisco [sic] Carlos Monteiro Souza, Geraldo Xexéo · IEEE Conference on Games (CoG) 2021 · 2021 · checked Sep 22, 2026NotesAUC is the only metric. Gini importance from the RF models on the Complete set ranks the blue jungler's win rate on the picked champion first (Figure 4). The player statistics come from. Author name printed 'Fracisco' in the paper. | Random Forest and Logistic Regression | After the draftOne map | 0.97 | — | — | — | — | Lownot tested on later games · 2,840 games in the data, test share not statedRandom splitFar above the rest | random | 2,840 professional LoL matches from Oracle's Elixir, 2021/01/01 to 2021/03/23 (blue side won 46.37%); datasets published on Kaggle (tekpixo/leagueoflegendsprematch2021) |
PaperMachine Learning Methods for Predicting League of Legends Game OutcomeJuan Agustín Hitar-García, Laura Morán-Fernández, Verónica Bolón-Canedo · IEEE Transactions on Games 15(2):171-181 · 2022-02-23 (online, per OpenAlex); journal issue 2023 · checked Sep 22, 2026NotesThe abstract: the accuracy is. | best classifier after feature selection (the abstract names none) | After the draftOne map | — | — | — | >0.70 | — | Unknownsplit not stated · test size not statedSplit not stated | not statednot stated in the abstract | professional matches, pregame, features from 'the game application programming interface (API)'; size, dates and leagues not in the abstract |
Before the draft, one map 6
Ratings and form before a champion is picked. The like-for-like set for our pro model before the draft.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
KayLoLWebsiteKayLoL /metrics: pro tournament mapsKayLoL · kaylol.gg · live payload generated 2026-09-23T01:22:01Z · checked Sep 23, 2026NotesA backfill, as our own scoreboard states: out of sample for each fit, not for the design choices, which were made looking at the same seasons. | pro model — before the draft | Before the draftOne map | 0.7287 | 0.6006 | 0.2081 | 0.6749 | 0.0426 | Hightested on later games · 1,784 test games | walk-forwardwalk-forward: each map scored by a model fitted only on the seasons before it | 1,784 pro tournament maps, 2026-06-29 onward, 29 leagues |
CodeLoL Match Prediction (Ra1nForest/LOL_Predictions)Ra1nForest (GitHub) · GitHub · 2026-09-12 to 2026-09-23 (repo created, last pushed 2026-09-23 00:06 UTC; README updates dated 2026-08-16, 2026-08-24, 2026-09-13) · checked Sep 22, 2026NotesIn-game Stage 4 results exist but use live game state, so they are left out of self_reported: '73.7% (80.5% at 25 min)' against a 56.3% baseline (28 features); live variant without XP 72.9%; isotonic re-test (2026-08-24) 'ECE 0.0351 → 0.0160', accuracy '77.2% → 77.3%'.. 'There is no server any more (since 2026-09-13)'; the public site ra1nforest.github.io/LOL_Predictions runs the stages in the browser. No forward record is published. | Stage 1 pre_draft (79 features)Also reports: baseline 53.3% | Before the draftOne map | — | — | — | 0.599 | — | Hightested on later games · 8,453 test games | walk-forwardthe README's evaluation apparatus; comparisons 'are paired', accepted only when |t| > 2.5. The stage table itself does not restate which evaluation produced it | 8,453 games from Oracle's Elixir, LPL / LCK / LEC / LCS, 2022–2026 (fold windows not listed; 'fold 8 now ends 2026-08-10') |
PaperPandaSkill - Player Performance and Skill Rating in Esports: Application to League of LegendsMaxime De Bois, Flora Parmentier, Raphaël Puget, Matthew Tanti, Jordan Peltier · arXiv · 2025-01-17 (arXiv v1; v2 2025-01-20) · checked Sep 22, 2026NotesNumbers match the repo. In the same Table IV the outcome-only 'PScore + OpenSkill' row prints 65.56% on all games and 65.51% intra-region, above PandaSkill's 64.98% and 64.86% (it trails on inter-region games and on ECE); the paper's own summary: ECE bins are not stated. Generic EWMA rows (all / intra / inter): PScore + EWMA 63.06 / 63.33 / 52.34 (ECE 1.43 / 1.57 / 5.51), PI + EWMA 63.53 / 63.75 / 54.68, PlayeRank + EWMA 62.25 / 62.80 / 51.67. Expert survey (not a win metric): 80.63% majority concordance (SD 6.52), 88.98% unanimity concordance (SD 5.99). The repo's 'chronological replay' is, in the paper's words, a one-month rolling window with a logistic regression trained on the previous year. | PScore + Meta_FFA_OpenSkill (PandaSkill), all games | Before the draftOne map | — | — | — | 0.6498 | 0.0101 | Hightested on later games · 37,388 test games | walk-forward | Professional LoL games from the Leaguepedia API, 2019-09-15 to 2024-09-15 (37,388 games, 4,927 players; LCK, LPL, LEC, LCS, PCS, VCS, CBLOL, LLA, Worlds, MSI and others); each month after the first year is tested; test-set sizes not stated |
CodeLoL Esports Match Prediction Model (Hget15/lol-esports-predictor)Hget15 (GitHub) · GitHub · 2026-09-14 to 2026-09-16 (repo created and last pushed, GitHub API) · checked Sep 22, 2026NotesV1–V3 log loss is not printed. Top V4 feature: 'patch_wr_delta_diff' at 25.3% importance . The live demo (hget15.github.io/lol-esports-predictor) runs, not the evaluated model. The README credits V4 with 'the single biggest jump — AUC 0.698 → 0.797'. Our check: We read the code (upstream commit e50e9ce, 2026-09-20). The V4 features come from trackers filled with every game of 2021-2026 before any feature is computed, so a team's win rate on the current patch includes the game being predicted and the games after it (src/v4_enhancements.py: build_all_from_raw, line 223; get_patch_wr, line 85). The V1-V3 results do not use those features. | V1 best model: logistic regression (63 features) — V1Also reports: worklog.md: AUC 0.6881, Brier 0.2230 | Before the draftOne map | 0.688 | — | 0.223 | 0.635 | — | Hightested on later games · about 7,663 test games | temporalworklog.md | held-out test set: last 15% by date of 51,088 Oracle's Elixir games (2021–2026), '~6,874 games' (worklog.md); leagues not listed |
| LoL Esports Match Prediction Model (Hget15/lol-esports-predictor) | V2 best model: ensemble (79 features) — V2Also reports: worklog.md: AUC 0.6976, Brier 0.2197 | Before the draftOne map | 0.698 | — | 0.220 | 0.641 | — | Hightested on later games · about 7,663 test games | temporalworklog.md | held-out test set: last 15% by date of 51,088 Oracle's Elixir games (2021–2026), '~6,874 games' (worklog.md); leagues not listed |
| LoL Esports Match Prediction Model (Hget15/lol-esports-predictor) | V3 best model: GBDT (106 features) — V3Also reports: worklog.md: AUC 0.6983, Brier 0.2200 | Before the draftOne map | 0.698 | — | 0.220 | 0.643 | — | Hightested on later games · about 7,663 test games | temporalworklog.md | held-out test set: last 15% by date of 51,088 Oracle's Elixir games (2021–2026), '~6,874 games' (worklog.md); leagues not listed |
PaperPre-game paired-comparison modeling of professional League of Legends map outcomesMin-Ren Guan, Shen-Ning Tung · arXiv · 2026-09-07 (arXiv v1) · checked Sep 22, 2026NotesThe paper names its score RPS and says. No accuracy, AUC or ECE is reported. Table 6 gives single-feature ROC AUCs of end-of-game and 15-minute features against the outcome (e.g., final gold 0.997, towers 1.000, gold at 15 minutes 0.784); those are associations, not forecasts. On the market window the two architectures tie (Δ = +0.0003 on 963 walk-forward games). Rank validity (Spearman, 50 teams): 0.438 proposed, 0.436 candidate. | one-stage logistic regression, 'outcome' (d = 1), the proposed modelAlso reports: calibration slope 0.88 (held-out); reported as RPS, which the paper states is numerically identical to the Brier score for a two-outcome game | Before the draftOne map | — | 0.6371 | 0.2230 | — | — | Mediumtested on later games · 912 test gamesSmall test | temporalGlobal time-based holdout: fit on the oldest 80% of games by calendar time, score the newest 20% | Newest 20% by calendar time (n_te = 912) of 5,135 professional maps with a first-pick assignment (LPL 1,822, LCK 1,276, LEC 697, LCP 432, CBLOL 277, LCS 216, Worlds 190, MSI 157, First Stand 68), 2024-2026; lolesports telemetry plus Leaguepedia |
CodeLoL Esports Predictor for Polymarket (farrelo25/lol-esports-predictor)farrelo25 (Hugging Face) · Hugging Face · 2026-05-01 (model repo created and last modified that day, Hugging Face API; the card is undated) · checked Sep 22, 2026Notesno forward record, P&L or ROI is published. Unit: the backtest table counts 'Games'; a 'Bo3/Bo5 series probability extension' sits on top. Author's limitations include 'Champion draft not modeled' and 'Player-level ratings not used'. | XGBoost + isotonic regression calibrationAlso reports: Brier and log loss labelled '(calibrated)'; architecture box elsewhere on the card: '63% accuracy, 0.67 AUC-ROC, 0.225 Brier Score' | Before the draftOne map | 0.672 | 0.639 | 0.224 | 0.630 | — | Unknownsplit not stated · test size not statedSplit not stated | not stated'Validation Results (7,763 train / 1,941 val)'; how the 1,941 validation rows were chosen is not stated (the architecture box says 'Time-series cross-validation (no data leakage)') | 1,941 validation rows; dates, leagues and selection not stated |
Series, or maps and series mixed 3
A best-of-three or best-of-five winner is an easier question than one map, so these sit apart.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
WebsiteRiftcast: Model accuracy (Every model, every market, every league) and Model methodologyRiftcast · riftcast.gg · live page: 'Resolved 2026-09-23 01:54 UTC', 'Models last trained 2026-09-23 01:44 UTC' · checked Sep 22, 2026NotesHit rate and return go in other_metrics and accuracy is left null because the hit rate covers selected picks, not all games. The methodology page's 585/895 equals the sum of the four models' 'All leagues' match-winner cells (176+146+130+133 of 266+229+197+203, my arithmetic; the page names no model). Consensus is 'a weighted blend of the three above, weighted by each model's recent log loss'; log loss itself is not published. Other markets on the page are not win predictions and are not recorded (e.g. Total Maps LightGBM 100% 14/14 +97.9% B 0.154; Series Handicap Winner; Exact Score; total-kills markets). A 'Download CSV' link exists (not downloaded). Every figure moves as matches resolve. | FastTree — match-winner markets, all leagues (series and per-game combined)Also reports: hit rate 66% (176/266 picks correct); return +21.0% on a flat one-unit stake | Before the draftMaps and series | — | — | 0.214 | — | — | Mediumtested on later games · test size not stated | forwardmodels retrain daily at 00:10 UTC; 'Predictions made before match start remain on record permanently' | published value picks only; graded against results scraped from gol.gg; all leagues; latest graded 2026-09-22 |
| Riftcast: Model accuracy (Every model, every market, every league) and Model methodology | LightGBM — match-winner markets, all leagues (series and per-game combined)Also reports: hit rate 64% (146/229 picks correct); return +18.3% on a flat one-unit stake | Before the draftMaps and series | — | — | 0.233 | — | — | Mediumtested on later games · test size not stated | forwardmodels retrain daily at 00:10 UTC; 'Predictions made before match start remain on record permanently' | published value picks only; graded against results scraped from gol.gg; all leagues; latest graded 2026-09-22 |
| Riftcast: Model accuracy (Every model, every market, every league) and Model methodology | PCA Sweep — match-winner markets, all leagues (series and per-game combined)Also reports: hit rate 66% (130/197 picks correct); return +15.9% on a flat one-unit stake | Before the draftMaps and series | — | — | 0.219 | — | — | Mediumtested on later games · test size not stated | forwardmodels retrain daily at 00:10 UTC; 'Predictions made before match start remain on record permanently' | published value picks only; graded against results scraped from gol.gg; all leagues; latest graded 2026-09-22 |
PaperEsportsBench: A Collection of Datasets for Benchmarking Rating Systems in EsportsClayton Thorrez · cthorrez.github.io (preprint); datasets on Hugging Face (EsportsBench/EsportsBench); code on GitHub (cthorrez/esports-bench) · 2024 (preprint, 'Preprint. Under review.', undated; citation year 2024) · checked Sep 22, 2026NotesLog loss counts a draw as 0.5; accuracy excludes draws; LoL draw rate 0.0261 (Table 1). The dataset card (last modified 2026-08-09) lists releases 1.0 to 10.0; its 10.0 line reads (train and test overlap as printed) while its intro says 'Date is complete up to 2026-03-31'. The paper hopes EsportsBench could become 'a dynamic or near-real-time leaderboard' (future work). Rating systems are implemented in the author's riix package. | Elo (riix implementation) — League of Legends | Before the draftMaps and series | — | 0.6371 | — | 0.6340 | — | Hightested on later games · 17,806 test games | walk-forwardhyperparameters chosen by log loss on the train set; ratings updated online one 7-day rating period at a time, predicting every match in a period before updating on it | League of Legends test set: 17,806 matches (train 104,737), Leaguepedia data from 2011 (the dataset card's release 1.0: 'Test: 2023-04-01 to 2024-03-31'); accuracy excludes draws |
| EsportsBench: A Collection of Datasets for Benchmarking Rating Systems in Esports | Glicko (riix implementation) — League of Legends | Before the draftMaps and series | — | 0.6292 | — | 0.6355 | — | Hightested on later games · 17,806 test games | walk-forwardhyperparameters chosen by log loss on the train set; ratings updated online one 7-day rating period at a time, predicting every match in a period before updating on it | League of Legends test set: 17,806 matches (train 104,737), Leaguepedia data from 2011 (the dataset card's release 1.0: 'Test: 2023-04-01 to 2024-03-31'); accuracy excludes draws |
| EsportsBench: A Collection of Datasets for Benchmarking Rating Systems in Esports | Glicko 2 (riix implementation) — League of Legends | Before the draftMaps and series | — | 0.6292 | — | 0.6356 | — | Hightested on later games · 17,806 test games | walk-forwardhyperparameters chosen by log loss on the train set; ratings updated online one 7-day rating period at a time, predicting every match in a period before updating on it | League of Legends test set: 17,806 matches (train 104,737), Leaguepedia data from 2011 (the dataset card's release 1.0: 'Test: 2023-04-01 to 2024-03-31'); accuracy excludes draws |
WebsiteLeague of Elo: All-Time leaderboard entry of the site's own model ('LeagueOfElo')League of Elo · leagueofelo.com · not stated (live board; its season selector lists events from 2015 to 2025) · checked Sep 22, 2026NotesAbout page (https://leagueofelo.com/about): 'AR is derived directly from Brier Score'; 'XP = min(1, 1/3 log10(N))'; 'a 50/50 prediction yielding 75 points'. The highest human account on the All-Time board reads Adjusted AR 79.16, Up/Down 68.4% on 1378 matches (a single user, not a crowd aggregate, so not recorded as a cross-report). | LeagueOfElo (the site's Elo model account) — All-TimeAlso reports: Adjusted AR (All-Time) 79.36; Matches Predicted 9987; the 68.0% is the All-Time 'Up/Down' column; the about page | Before the draftMaps and series | — | — | — | 0.680 | — | Unknownsplit not stated · 9,987 games in the data, test share not statedSplit not stated | not statednot stated: an online Elo updated after each match; whether its board predictions were logged before the matches or filled in afterwards is not stated | 9,987 matches predicted on the All-Time board (events 2015–2025 in the season selector); a 'match' is a single game or a best-of series |
Map or series not stated 5
The source does not say whether it predicts one map or a whole series.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
Codelol_esports-predictions (a-huk)a-huk (GitHub) · GitHub · README undated; results.csv rows dated 11/02/2022 to 01/07/2023; repo created 2021-07-30, last push 2023-07-01 (GitHub API) · checked Sep 22, 2026NotesNo probability score is reported (results.csv holds probabilities; the README reports accuracy only). README: 'We are determining the "stronger" team and saying that it will win' and 'In the result's file, sometimes a > 0.5 probability is listed and the team still wins. Indeed, there was an error in the code, but the prediction output teams were always correct.' Whether a row is a game or a series is not stated. | EnsembleAlso reports: 381/535 | Before the draftMap or series not stated | — | — | — | 0.7121 | — | Mediumtested on later games · test size not stated | forwardforward: a script appends each day's predictions to results.csv | resolved predictions in results.csv for LCS, LEC and LCK (LPL 'left out'); rows dated 11/02/2022 to 01/07/2023 as printed (date format not stated) |
| lol_esports-predictions (a-huk) | tflearn DNNAlso reports: 392/566 | Before the draftMap or series not stated | — | — | — | 0.6926 | — | Mediumtested on later games · test size not stated | forwardforward: a script appends each day's predictions to results.csv | resolved predictions in results.csv for LCS, LEC and LCK (LPL 'left out'); rows dated 11/02/2022 to 01/07/2023 as printed (date format not stated) |
| lol_esports-predictions (a-huk) | AkkioAlso reports: 304/471 | Before the draftMap or series not stated | — | — | — | 0.6454 | — | Mediumtested on later games · test size not stated | forwardforward: a script appends each day's predictions to results.csv | resolved predictions in results.csv for LCS, LEC and LCK (LPL 'left out'); rows dated 11/02/2022 to 01/07/2023 as printed (date format not stated) |
WebsiteModel Accuracy & Track Record (draftlol.ai)draftlol.ai · draftlol.ai · live page, no 'updated' stamp; months Jun 2026 to Sep 2026; read 2026-09-22 · checked Sep 22, 2026NotesIntro: Pre-match caption. Hermes caption. ✓ marks the lower Brier in each row. Two pre-match rows (EMEA Masters Sep 2026, LFL Aug 2026) show '—' for the market, so they have no cross-report. Two Hermes rows are printed truncated as 'North Amer' and 'Prime Leag'. No log loss, AUC or ECE is published, and the page never says whether a prediction is for one game or a series ('match' is not defined). Numbers change as games resolve. | draftlol.ai pre-match model (; 'published uncalibrated')Also reports: N 554 | Before the draftMap or series not stated | — | — | 0.2290 | 0.610 | — | Mediumtested on later games · 554 test gamesSmall test | forward'Every prediction frozen at registration time'; resolved automatically | 554 resolved pro predictions, Jun 2026 to Sep 2026 (the table's months), leagues as listed in the table rows |
| Model Accuracy & Track Record (draftlol.ai) | draftlol.ai Hermes, the at-draft agent (pre-match signals plus 'seven signals including champion win rates and duo synergies'; predicts at draft lock-in)Also reports: N 1855 | After the draftMap or series not stated | — | — | 0.2263 | 0.629 | — | Hightested on later games · 1,855 test games | forward'Every prediction frozen at registration time'; resolved automatically | 1,855 resolved pro predictions at draft lock-in over 20 listed leagues; dates not stated |
WebsiteDev Diary: Unveiling the Global Power RankingsLoL Esports Staff (Riot Games) · lolesports.com · 2024-09-24 ('24 September 2024') · checked Sep 22, 2026NotesNo margin-of-victory term appears in this article (the Stomp Factor arrives with the 2025 revision; see riot-gpr-2025-26-revision). | Global Power Rankings, 80/20 team/league Elo | Before the draftMap or series not stated | — | — | — | 0.65 | — | Unknownsplit not stated · test size not statedSplit not stated | not statednot stated | not stated (no matches, years or held-out set named) |
PaperPredicting the Outcome of MOBA Game: a Case Study on Pregame Prediction for League of LegendsShilong Cong, Hanwei Wu, Shuo Xiong, Yue Wu, Long Zuo · 2024 8th Asian Conference on Artificial Intelligence Technology (ACAIT), pp. 1366-1372 · 2024-11-08 (ACAIT, November 8-10, 2024, Fuzhou) · checked Sep 22, 2026NotesThe level rests on the abstract's wording ('the outcome of the championship'; metrics 'based on the defining characteristics of professional-level matches'); the dataset, unit and split are not described. | modified deep learning model with six metrics | Not statedMap or series not stated | — | — | — | 0.8182 | — | Unknownsplit not stated · test size not statedSplit not stated | not statednot stated in the abstract | not stated in the abstract |
CodeProjektZero-LoL-Model ('Match Modeling')MRittinghouse (GitHub) · GitHub · README metrics dated 'Feb 24, 2021' (experimental model 'Dec 28, 2021'); repo created 2021-06-19, last push 2022-02-28 (GitHub API) · checked Sep 22, 2026NotesThe README says predictions cover 'game or match win probabilities' (unit not fixed). 'this model is intended purely for academic purposes'. | Ensemble (accuracy-weighted average of the component models) | Before the draftMap or series not stated | — | 0.6425 | 0.2255 | 0.6366 | — | Unknownsplit not stated · test size not statedSplit not stated | not statednot stated: each figure is a 'Tested Accuracy... on Feb 24, 2021'; the README says a 'model validator' module holds 'a series of tests to help demonstrate accuracy' but never describes the test set | not stated (Oracle's Elixir pro data; test games, leagues and dates not described) |
When is not stated 2
The source does not say at what point it predicts.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
PaperGCN-WP – Semi-Supervised Graph Convolutional Networks for Win Prediction in EsportsAlexander J. Bisberg, Emilio Ferrara · IEEE, 2022 (Xplore 9893671; the repo lists CoG 2022); arXiv:2207.13191 · 2022-07-26 (arXiv v1) · checked Sep 22, 2026NotesRandom-forest baseline (Table I, lookback of 1, 3, 5 previous games): raw 0.563 ± 0.009, 0.573 ± 0.010, 0.572 ± 0.005; delta 0.560 ± 0.009, 0.569 ± 0.008, 0.578 ± 0.012; Table IV lists 'Random Forest (lookback=5) + delta 0.578'. On timing the paper says both that (Section I) and that a one-convolution GCN's 'fair' prediction is 'predicting two games ahead', with labels attached 'from future game results' (Section IV-C, Fig. 3); hence pregame 'unclear'. | GCN-cheby (1 layer) + delta | Not statedOne map | — | — | — | 0.619 | — | Mediumtested on a different league · test size not stated | cross-leagueTrain on the LPL, validate on the LCK, test on the LCS, 2020 regular season | LCS (North America) 2020 regular-season games from Oracle's Elixir, the test network; international tournaments excluded; game count not stated |
| GCN-WP – Semi-Supervised Graph Convolutional Networks for Win Prediction in Esports | GCN-cheby (2 layer) + delta | Not statedOne map | — | — | — | 0.568 | — | Mediumtested on a different league · test size not stated | cross-leagueTrain on the LPL, validate on the LCK, test on the LCS, 2020 regular season | LCS (North America) 2020 regular-season games from Oracle's Elixir, the test network; international tournaments excluded; game count not stated |
| GCN-WP – Semi-Supervised Graph Convolutional Networks for Win Prediction in Esports | GCN (1 layer) + delta | Not statedOne map | — | — | — | 0.551 | — | Mediumtested on a different league · test size not stated | cross-leagueTrain on the LPL, validate on the LCK, test on the LCS, 2020 regular season | LCS (North America) 2020 regular-season games from Oracle's Elixir, the test network; international tournaments excluded; game count not stated |
PaperStatistical Models for Predicting Results in Professional League of LegendsRobbie Jadowski, Stuart Cunningham · ArtsIT 2021 (10th EAI International Conference), Springer, 2022 · 2022-02-10 (EUDL); conference December 2-3, 2021 · checked Sep 22, 2026NotesThe abstract concludes that more accuracy 'would require substantially more data and contextual information'. MMU e-space lists a camera-ready copy (e-space.mmu.ac.uk/629272/1/EAI_ArtsIT_Robbie___cam_ready.pdf) whose download redirect expired before it could be read. | best performing model (the abstract does not name it) | Not statedOne map | — | — | — | 0.67 | — | Unknownsplit not stated · 306 games in the data, test share not statedSplit not statedSmall test | not statednot stated in the abstract | 306 games; (abstract) |
Dota 2 15
The same question in a different game: context, not competition.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
PaperDraftRec: Personalized Draft Recommendation for Winning in Multi-Player Online Battle Arena GamesHojoon Lee, Dongyoon Hwang, Hyunseung Kim, Byungkun Lee, Jaegul Choo · WWW 2022 (The ACM Web Conference); arXiv:2204.12750 · 2022-04-27 (arXiv v1, 'Accepted to WWW 2022') · checked Sep 22, 2026NotesThe Dota 2 rows are the same paper's second dataset (Dota 2 is not this file's game). Generic baselines in the same tables, not listed above: MC (majority class, 'Blue for LOL and Radiant for Dota2') 0.5040 LoL and 0.5180 Dota 2; LR 0.5255 during and 0.5323 after the draft (LoL), 0.5750 and 0.6126 (Dota 2); NN 0.5263 and 0.5335 (LoL), 0.5748 and 0.6108 (Dota 2). After the Dota 2 draft, LR (0.6126) prints above DraftRec (0.6110). Outcome prediction is the paper's secondary task; its main task is champion recommendation. Code and data: github.com/dojeon-ai/DraftRec. | DraftRec — Dota 2, during the draftAlso reports: MAE 0.4782 | During the draftOne game | — | — | — | 0.5755 | — | Hightested on later games · 50,000 games in the data, test share not stated | temporal | Last 10% by time of a public Kaggle Dota 2 dataset: 50,000 matches of various skill tiers, November 5 to November 18, 2015; test-set size not stated |
| DraftRec: Personalized Draft Recommendation for Winning in Multi-Player Online Battle Arena Games | DraftRec — Dota 2, after the draft (all ten heroes selected)Also reports: MAE 0.4642 | After the draftOne game | — | — | — | 0.6110 | — | Hightested on later games · 50,000 games in the data, test share not stated | temporal | Last 10% by time of a public Kaggle Dota 2 dataset: 50,000 matches of various skill tiers, November 5 to November 18, 2015; test-set size not stated |
| DraftRec: Personalized Draft Recommendation for Winning in Multi-Player Online Battle Arena Games | DraftRec-no-history — Dota 2, during the draftAlso reports: MAE 0.4757 (starred: p<0.01) | During the draftOne game | — | — | — | 0.5745 | — | Hightested on later games · 50,000 games in the data, test share not stated | temporal | Last 10% by time of a public Kaggle Dota 2 dataset: 50,000 matches of various skill tiers, November 5 to November 18, 2015; test-set size not stated |
PaperWin Prediction in Esports: Mixed-Rank Match Prediction in Multi-player Online Battle Arena GamesVictoria Hodge, Sam Devlin, Nick Sephton, Florian Block, Anders Drachen, Peter Cowling · arXiv · 2017-11-17 (arXiv v1, the only version) · checked Sep 22, 2026NotesIn-game results of this version, not recorded as rows because they are not in benchmarks.md §7: Mixed-InGame best 76.17% (RF, CFS selection) and Pro-InGame best 75.22% (LR, single feature Kills R-D), 20-minute mark (Table 3). The authors conclude Hero-model parameters were compared 'on the training data set'; for the in-game models they. | logistic regression with WrapperSubsetEval feature selection — PREGAME hero vectors: mixed test set (Mixed-Hero)Also reports: Table 2, Mixed-Hero (All / Wrapper Select): LR 54.6423 / 58.7519; RF 53.1202 / 58.2953 | After the draftOne game | — | — | — | 0.5875 | — | Mediumtested on later games · about 657 test gamesSmall test | temporal | the latest 34% of 1,933 replays: 270 professional matches (13.97%) and 1,663 public matches with skill score above 6000 ('99.81 percentile'), 27 March to 4 May 2017, Valve replays parsed with Clarity |
| Win Prediction in Esports: Mixed-Rank Match Prediction in Multi-player Online Battle Arena Games | Random Forest with WrapperSubsetEval feature selection — PREGAME hero vectors: professional test set (Pro-Hero, Kiev Major 2017)Also reports: Table 2, Pro-Hero (All / Wrapper Select): LR 47.7876 / 50.4425; RF 50.4425 / 55.7522 | After the draftOne game | — | — | — | 0.5575 | — | Lowtested on a different league · 113 games in the data, test share not statedSmall test | cross-league. The training data run to 4 May 2017, so they include games played after the test tournament | the 113 matches of the Kiev Major 2017 (24-30 April 2017) |
PaperLearning Dota 2 Team CompositionsAtish Agarwala, Michael Pearce · CS229 report (Stanford) · 2014 (Stanford CS229 2014 project archive; undated; data 10/1/14 to 12/3/14) · checked Sep 22, 2026NotesThe Full Heroes Model is Conley & Perry's hero-presence regression, antisymmetrized, re-run on this report's data; it is kept as a self-report because it is the report's headline. The authors note its top five heroes. | Full Heroes Model (logistic regression on hero presence, antisymmetrized; 'previously studied in [1]', Conley & Perry)Also reports: the abstract calls it | After the draftOne game | — | — | — | 0.62 | — | Lownot tested on later games · about 4,000 test gamesRandom split | random | the random 10% of 40,000 public matches (Valve web API), 10/1/14 to 12/3/14, one balance patch, All Pick and Captain's Mode, no leavers, 'a mix of low, middle, and high level play' (the API's skill filter was broken); test count not printed |
| Learning Dota 2 Team Compositions | 2nd Order PCA (d = 7)Also reports: the number of PCA dimensions | After the draftOne game | — | — | — | 0.57 | — | Lownot tested on later games · about 4,000 test gamesRandom split | random | the random 10% of 40,000 public matches, 10/1/14 to 12/3/14, mixed skill |
| Learning Dota 2 Team Compositions | Sorted PCA (d = 7)Also reports: the number of PCA dimensions | After the draftOne game | — | — | — | 0.57 | — | Lownot tested on later games · about 4,000 test gamesRandom split | random | the random 10% of 40,000 public matches, 10/1/14 to 12/3/14, mixed skill |
PaperPlayer Skill Decomposition in Multiplayer Online Battle ArenasZhengxing Chen, Yizhou Sun, Magy Seif El-Nasr, Truong-Huy D. Nguyen · 2016 Meaningful Play Conference; arXiv:1702.06253 · 2017-02-21 (arXiv v1) · checked Sep 22, 2026NotesThe Dota 2 rows are the same paper's second dataset. Majority-class baseline (BL-MC): 53.22% ± 0.16% LoL, 52.65% ± 0.12% Dota 2.. The weights are learned from match outcomes, so under 10-fold CV a player's weight is fit on that player's other matches, earlier and later. The paper's significance rule: mean1 − mean2 > 2(std1 + std2). | LR-P-C-PC — Dota 2Also reports: ± 0.42% | After the draftOne game | — | — | — | 0.5966 | — | Lownot tested on later games · test size not statedRandom split | random | 10-fold CV over 195,538 ranked Dota 2 matches from 2015 (yasp.co), All Pick, Random Draft or Captain modes on North America regions, each with at least one player whose public solo-queue MMR exceeds 4000 (117,874 players plus one universal anonymous player, 111 champions) |
| Player Skill Decomposition in Multiplayer Online Battle Arenas | LR-P-C — Dota 2Also reports: ± 0.39% | After the draftOne game | — | — | — | 0.5953 | — | Lownot tested on later games · test size not statedRandom split | random | 10-fold CV over 195,538 ranked Dota 2 matches from 2015 (yasp.co), All Pick, Random Draft or Captain modes on North America regions, each with at least one player whose public solo-queue MMR exceeds 4000 (117,874 players plus one universal anonymous player, 111 champions) |
| Player Skill Decomposition in Multiplayer Online Battle Arenas | LR-C — Dota 2Also reports: ± 0.41% | After the draftOne game | — | — | — | 0.5916 | — | Lownot tested on later games · test size not statedRandom split | random | 10-fold CV over 195,538 ranked Dota 2 matches from 2015 (yasp.co), All Pick, Random Draft or Captain modes on North America regions, each with at least one player whose public solo-queue MMR exceeds 4000 (117,874 players plus one universal anonymous player, 111 champions) |
PaperPredicting outcomes of professional DotA 2 matchesPetra Grutzik, Joe Higgins, Long Tran · CS229 report (Stanford) · 2017-12-16 · checked Sep 22, 2026NotesThe abstract says the best model's performance, meaning the bettors' 62.8%, which is measured on a different unit. Data were collected with the OpenDota web API (contributions section). | linear-kernel SVM (penalty 2) — dev set: Team and Draft Features (best model)Also reports: training accuracy 72% | After the draftOne map | — | — | — | 0.61 | — | Lownot tested on later games · about 4,258 test gamesRandom split | randomThis number is on the dev set | dev set: random 10% of the 42,578 professional games (Valve-sanctioned tournaments) from November 2011 to before May 1, 2017; data from the OpenDota web API |
| Predicting outcomes of professional DotA 2 matches | linear-kernel SVM — dev set: Draft Features OnlyAlso reports: training accuracy 58% | After the draftOne map | — | — | — | 0.57 | — | Lownot tested on later games · about 4,258 test gamesRandom split | randomrandom 90/10 train/dev split of games before May 1, 2017 | dev set: random 10% of the 42,578 pre-May-2017 professional games |
| Predicting outcomes of professional DotA 2 matches | linear-kernel SVM — dev set: Team Features OnlyAlso reports: training accuracy 72% | Before the draftOne map | — | — | — | 0.59 | — | Lownot tested on later games · about 4,258 test gamesRandom split | randomrandom 90/10 train/dev split of games before May 1, 2017 | dev set: random 10% of the 42,578 pre-May-2017 professional games |
PaperPredicting the winning side of DotA2Kuangyan Song, Tianyi Zhang, Chao Ma · CS229 report (Stanford) · 2015 (Stanford CS229 2015 project archive; undated) · checked Sep 22, 2026NotesUnder the owner's rule no error rate is converted to an accuracy here. The authors warn that the combo improvement 'might because of the increasing number of match data' (6,000 against 3,000 matches). | logistic regression, hero lineup (baseline)Also reports: error rates only. Table 1 test error per fold: 0.47, 0.43, 0.49, 0.42, 0.48, 0.47, 0.42, 0.46, 0.43, 0.46 (training: 0.28, 0.29, 0.30, 0.27, 0.29, 0.25, 0.23, 0.29, 0.33, 0.31); text | After the draftOne game | — | — | — | — | — | Lownot tested on later games · 3,000 test gamesRandom split | random | 300-match held-out folds of 3,000 matches from Valve's Web API (dota2api); modes All Pick, Ranked All Pick, Random Draft; games with absent players or zero kills removed; skill bracket, dates and region not stated |
| Predicting the winning side of DotA2 | logistic regression + 50 hand-picked two-hero combosAlso reports: error rates only. Table 2 (6000 matches, 10 folds) test error per fold: 0.40, 0.41, 0.45, 0.43, 0.41, 0.39, 0.41, 0.43, 0.40, 0.42 (training: 0.25, 0.26, 0.23, 0.28, 0.25, 0.25, 0.24, 0.27, 0.29, 0.28) | After the draftOne game | — | — | — | — | — | Lownot tested on later games · 6,000 test gamesRandom split | random10-fold cross-validation on 6,000 matches | held-out folds of 6,000 matches from the same collection |
| Predicting the winning side of DotA2 | stepwise regression feature selection (heroes and hero combos)Also reports: Figure 1 plots training and testing error against the number of features; No value printed | After the draftOne game | — | — | — | — | — | Unknownsplit not stated · test size not statedSplit not stated | not statednot stated for this experiment | not stated for this experiment |
PaperA Recommender System for Hero Line-Ups in MOBA GamesLucas Hanke, Luiz Chaimowicz · AIIDE-17 (Thirteenth AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment), pp. 43-49 · 2017 (AIIDE-17 proceedings) · checked Sep 22, 2026NotesA baseline of always predicting a radiant win scores 'around 63%' (text; drawn in Figure 7). Recommender success rates, all judged by the neural network: Table 2 (random enemies) 56.9%, 60.5%, 73.4%, 76.4%; Table 3 (enemies among the 10 most picked) 59.9%, 60.2%, 71.7%, 74.9%. | multilayer perceptron on hero line-ups — patch 7.06dAlso reports: training accuracy 89.28% | After the draftOne game | — | — | — | 0.8863 | — | Unknownsplit not stated · 23,000 test gamesSplit not statedFar above the rest | not statedRetrained every hour for 100 epochs; figures after 193 rounds. How the 70/30 partition is drawn, and whether it stays fixed across rounds, is not stated | around 23,000 test matches (with around 54,000 for training) from patch 7.06d, collected from the Steam Web API: very high skill players, no disconnects, pick modes with all heroes available; region not stated |
PaperHow Does He Saw Me? A Recommendation Engine for Picking Heroes in Dota 2Kevin Conley, Daniel Perry · CS229 report (Stanford) · 2013 (Stanford CS229 2013 project archive; the report itself is undated) · checked Sep 22, 2026NotesThe authors note the game version did not change during collection. Their web interface shows 'Chance we win based on picks: 96%' for an example lineup; that is a demo, not a metric. | logistic regression on hero presenceAlso reports: the test accuracy approaches 69.8% at about 18,000 training matches (learning curve) | After the draftOne game | — | — | — | 0.698 | — | Unknownsplit not stated · 56,691 test gamesSplit not stated | not statedHow matches were assigned is not stated | test set of 5,669 of 56,691 matches pulled from Valve's Steam web API between November 5 and December 7 (year not printed; references dated 2013); modes all pick, single draft, all random, random draft, captain's draft, captain's mode, least played; skill 'very-high' ('roughly the top 8% of players'); no leavers; region not stated |
| How Does He Saw Me? A Recommendation Engine for Picking Heroes in Dota 2 | k-nearest neighbors (custom polynomial weights, d = 4) | After the draftOne game | — | — | — | ≈0.70 | — | Unknownsplit not stated · test size not statedSplit not stated | not statedsame 90/10 export (51,022 train / 5,669 test); assignment not stated | the same 5,669-match test set |
PaperOutcome prediction of DOTA2 based on Naïve Bayes classifierKaixiang Wang, Wenqian Shang · 2017 IEEE/ACIS 16th International Conference on Computer and Information Science (ICIS), pp. 591-593 · 2017-05 (conference 24-26 May 2017, Wuhan; per Crossref) · checked Sep 22, 2026NotesWan & Sándor 2023 (§2.2, reference [17]) quote the same two figures. A search-engine summary, not the source, said the classifier was tested on the UCI Machine Learning Repository DOTA2 dataset; unverified. | Naive Bayes on hero line-upsAlso reports: training-set accuracy 85.33% | After the draftOne game | — | — | — | 0.5899 | — | Unknownsplit not stated · test size not statedSplit not stated | not statednot re-verified (benchmarks.md: 'train/test') | not re-verified |
PaperPerformance of Machine Learning Algorithms in Predicting Game Outcome from Drafts in Dota 2Aleksandr Semenov, Peter Romov, Sergey Korolev, Daniil Yashkov, Kirill Neklyudov · AIST 2016 (International Conference on Analysis of Images, Social Networks and Texts), Springer CCIS vol. 661, pp. 26-37 · 2017-02-17 (first online; AIST 2016 conference) · checked Sep 22, 2026NotesSecondary evidence only, not the source itself: Hodge et al. (IEEE ToG) Table I lists this paper as 'FM & XGB', 5,071,858 games, skill 'Normal-V. High', '70%**', footnoted as having 'calculated AUC rather than accuracy'. Akhmedov & Phan 2021 (arXiv:2106.01782) say it was evaluated 'on ten-fold cross-validation' and that; if right, the split is random k-fold and ~0.70 belongs to XGBoost, which conflicts with the abstract's FM-best. A search-engine snippet, not the paper, gave 0.706 / 0.670 / 0.660 for normal / high / very high skill. The authors' preprint (dotascience.com/papers/aist2016, linked from their CEUR-WS Vol-1842 review) no longer resolves, and the Springer full text is behind a subscription. 'public matchmaking' rests on the Normal / High / Very High skill brackets named by Hodge's table and the repo; the abstract says only 'skill level of the players'. | Factorization Machines — normal skill | After the draftOne game | 0.706 | — | — | — | — | Unknownsplit not stated · 5,000,000 games in the data, test share not statedSplit not stated | not statednot stated in the reachable abstract (benchmarks.md: 'unconfirmed') | not re-verified; benchmarks.md: Dota 2 public matchmaking, 5M matches, MMR-bracketed |
| Performance of Machine Learning Algorithms in Predicting Game Outcome from Drafts in Dota 2 | Factorization Machines — very high skill | After the draftOne game | 0.660 | — | — | — | — | Unknownsplit not stated · 5,000,000 games in the data, test share not statedSplit not stated | not statednot stated in the reachable abstract (benchmarks.md: 'unconfirmed') | not re-verified; benchmarks.md: Dota 2 public matchmaking, 5M matches, MMR-bracketed |
PaperPredicting Game Outcome in Dota 2 with NLP and Machine Learning AlgorithmsZiping Wan, Bálint Sándor · thesis, Jönköping University (School of Engineering, MSc AI Engineering; supervisor Tuwe Löfström) · 2023-06 (printed '06. 2023') · checked Sep 22, 2026NotesGeneric baselines omitted from self_reported: Gradient Boost 67.76 %, Logistic Regression 67.21 %, Random Forest 63.36 %, Gaussian Naive Bayes 60.49 %, KNN 54.37 % (Table 3). §4.1 prints confusion counts (991, 574, 500, 1183) beside; the counts alone do not reproduce 73%, and the thesis does not say which split they come from. It does not say whether the Word2Vec vectors and the log5 counter matrix were built from training matches only. Keeping only games of 20 minutes or more selects on game length, known only after the game. The literature review also quotes Hanke & Chaimowicz's recommender 'success rate of up to 74.9%', which is not a prediction metric. | CBOW hero embeddings + attention LSTMAlso reports: Precision.72, Recall.68, F1-score.69 (Table 3). Section 4: at early stopping (epoch 175) 'an accuracy of 73% and a loss of 0.56' (binary cross-entropy), said to hold | After the draftOne game | — | — | — | 0.7313 | — | Unknownsplit not stated · about 4,000 test gamesSplit not stated | not stated(§4.2); §3.3 separately describes '10-fold cross-validation'. Whether the 80/20 assignment is random is not stated | the 20% test share of 20,000 OpenDota matches, 'high-ranking' games with average MMR over 6000 (also called 'professional Dota 2 matches' once), at least 20 minutes long, won and lost matches balanced; dates, patch and region not stated |
PaperReal-time eSports Match Result PredictionYifan Yang, Tian Qin, Yu-Heng Lei · arXiv · 2016-12-10 (arXiv v1) · checked Sep 22, 2026NotesRead from the ar5iv HTML. The player features are the players' MMR scores and percentiles; test matches are excluded from the match-history statistics, but no cut-off at each match's start time is stated. The 93.73% real-time result is in yang-2016-realtime. | logistic regression — Hero + Player + Hero-Player (all prior features) | After the draftOne game | — | — | — | 0.7149 | — | Unknownsplit not stated · 78,362 test gamesSplit not stated | not statedRandom or chronological not stated | the 1-in-10 test share of 78,362 matches 'participated by 19790 players with "very high" skill level', crawled with the Dota 2 API; dates, region and test count not stated |
| Real-time eSports Match Result Prediction | two-layer neural network — Hero + Player + Hero-Player (all prior features) | After the draftOne game | — | — | — | 0.7046 | — | Unknownsplit not stated · 78,362 test gamesSplit not stated | not statedRandom or chronological not stated | the 1-in-10 test share of 78,362 'very high' skill matches (Dota 2 API); dates, region and test count not stated |
| Real-time eSports Match Result Prediction | logistic regression — Hero features only (draft only) | After the draftOne game | — | — | — | 0.6007 | — | Unknownsplit not stated · 78,362 test gamesSplit not stated | not statedRandom or chronological not stated | the 1-in-10 test share of 78,362 'very high' skill matches (Dota 2 API); dates, region and test count not stated |
PaperThe Explainability of Machine Learning Algorithms for Victory Prediction in the Video Game Dota 2Julio Losada-Rodríguez, Pedro A. Castillo, Antonio Mora, Pablo García-Sánchez · Computer Sciences & Mathematics Forum 11(1):26 (ITISE 2025, 11th International Conference on Time Series and Forecasting) · 2025-08-18 · checked Sep 22, 2026NotesThe authors themselves say the results 'should be taken with a grain of salt' because predicting is very difficult, and that the web dataset. The paper counts 138 heroes and uses raw hero IDs per slot as features. | Random Forest (100 estimators, no maximum depth)Also reports: Accuracy (cross-validation) 0.9837; Precision 0.9795; Recall 0.9869; F1-Score 0.9832 | After the draftOne game | 1.00 | — | — | 0.9800 | — | Unknownsplit not stated · about 900 test gamesSplit not statedSmall testFar above the rest | not statedfold count and row assignment not stated | the 20% test share of 4,500 OpenDota matches (API accessed 1 November 2024), hero IDs only, 60% radiant wins; skill, dates and region not stated |
| The Explainability of Machine Learning Algorithms for Victory Prediction in the Video Game Dota 2 | Gradient Boosting (learning rate 0.1, 100 estimators)Also reports: Accuracy (cross-validation) 0.8653; Precision 0.8916; Recall 0.9383; F1-Score 0.9143 | After the draftOne game | 0.95 | — | — | 0.8955 | — | Unknownsplit not stated · about 900 test gamesSplit not statedSmall testFar above the rest | not stated80/20 train/test split plus cross-validation; assignment not stated | the 20% test share of 4,500 OpenDota matches, hero IDs only |
| The Explainability of Machine Learning Algorithms for Victory Prediction in the Video Game Dota 2 | Logistic Regression (maximum 1000 iterations)Also reports: Accuracy (cross-validation) 0.6562; Precision 0.6656; Recall 0.8560; F1-Score 0.7489; on the ROC figure the text says | After the draftOne game | — | — | — | 0.6588 | — | Unknownsplit not stated · about 900 test gamesSplit not statedSmall test | not stated80/20 train/test split plus cross-validation; assignment not stated | the 20% test share of 4,500 OpenDota matches, hero IDs only |
PaperTo win or not to win? A prediction model to determine the outcome of a DotA2 matchKaushik Kalyanaraman · CSE 255 report (University of California, San Diego, Winter 2015) · 2015 (UCSD CSE 255 Winter 2015 report archive; undated; data 01/23/2015 to 02/20/2015) · checked Sep 22, 2026NotesThe printed split counts (17,000 / 1,500) do not match the 30,426 games collected. The author notes that plain logistic regression scores higher on test than on training data and suggests. The genetic algorithm starts from 'an initial population of 220 matches' and the hero communities come from 250 matches; the report does not say these came from the training split only. Hodge et al. (ToG) Table I lists this work at '91%' on '15,146' games, which matches neither figure here. | augmented logistic regression (logistic regression + genetic-algorithm success set)Also reports: test accuracy 'asymptotically approaches a value of 74.1%'; Table I: Precision 0.684, Recall 0.909, F-Score 0.781 | After the draftOne game | — | — | — | 0.741 | — | Unknownsplit not stated · 1,500 test gamesSplit not stated | not statedAssignment method not stated | test set of about 1,500 matches from 30,426 games pulled from Valve's Steam API, 01/23/2015 to 02/20/2015; modes All Pick, Random Draft, Single Draft, All Random, Least Played, Captains Draft; all players above a minimum MMR ('very high skill level'); 10 players to the end; at least 900 seconds; games with Winter Wyvern dropped; region not stated |
| To win or not to win? A prediction model to determine the outcome of a DotA2 match | logistic regression on hero presence (Conley & Perry's features and procedure)Also reports: test accuracy 'asymptotically approaches 69.42%'; Table I: Precision 0.752, Recall 0.787, F-Score 0.769 | After the draftOne game | — | — | — | 0.6942 | — | Unknownsplit not stated · 1,500 test gamesSplit not stated | not statedAssignment method not stated | the same test set of about 1,500 very-high-skill matches (January-February 2015) |
| To win or not to win? A prediction model to determine the outcome of a DotA2 match | genetic algorithm aloneAlso reports: Table I: Precision 0.741, Recall 0.762, F-Score 0.752 | After the draftOne game | — | — | — | — | — | Unknownsplit not stated · 1,500 test gamesSplit not stated | not statedAssignment method not stated | the same test set of about 1,500 very-high-skill matches (January-February 2015) |
WebsiteUsing AI to Predict Dota 2 — The International 8 — Games at 77% AccuracySTRATZ · medium.com/stratz · 2018 (per papers.md; not verified) · checked Sep 22, 2026NotesNeeds a real browser. The 77% is taken from the title, which the search index shows as. | STRATZ AI | Not statedMap or series not stated | — | — | — | 0.77 | — | Unknownsplit not stated · test size not statedSplit not stated | not statednot verified (papers.md: 'none published') | not verified (TI8 games; n not verified) |
Other MOBAs 3
Mobile Legends, Honor of Kings and an unnamed MOBA.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
PaperCHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player GamesKai Wang, Ge Fan, Chaoyun Zhang, Yuyang Jiang, Yuze Liu · CIKM 2026 (Applied Research Track); arXiv · 2026-09-04 (arXiv v1) · checked Sep 22, 2026NotesGeneric baselines in Table 1 (accuracy, Casual / League / Elite): LR 0.5657 / 0.5040 / 0.5912; MLP 0.5753 / 0.5153 / 0.6041; LSTM 0.5931 / 0.5231 / 0.6336; Transformer 0.6012 / 0.5429 / 0.6528. OwO is kept as a self-report because the abstract calls Cupid 'Our prior work'. The online A/B test 'spanned over 30 days'; the novice segment saw 'a 20.73% reduction in 5-minute kill crushing', which is a matchmaking outcome, not a prediction metric. The game is not named. Read from the arXiv HTML. | DAWN (Domain-Aware Win-rate Network) — League modeAlso reports: RMSE 0.4216 | Before the draftOne game | — | — | — | 0.5685 | — | Hightested on later games · 30,000,000 games in the data, test share not stated | temporal | the most recent matches of the League-mode dataset (about 30M samples in total; test size and dates not stated), industrial match logs of 'a popular online MOBA title' |
| CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games | DAWN (Domain-Aware Win-rate Network) — Casual modeAlso reports: RMSE 0.4022 | Before the draftOne game | — | — | — | 0.6612 | — | Hightested on later games · 40,000,000 games in the data, test share not stated | temporalchronological: earliest matches train, most recent test, 5% of training held out for validation | the most recent matches of the Casual-mode dataset (about 40M samples in total; test size and dates not stated) |
| CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games | DAWN (Domain-Aware Win-rate Network) — Elite mode ('the dedicated mode for top-expert players')Also reports: RMSE 0.3987; the abstract gives it as '67.73% win-rate prediction accuracy' | Before the draftOne game | — | — | — | 0.6773 | — | Hightested on later games · 100,000 games in the data, test share not stated | temporalchronological: earliest matches train, most recent test, 5% of training held out for validation | the most recent matches of the Elite-mode dataset (about 0.1M samples in total; test size and dates not stated) |
PaperMatch Outcome Prediction in Draft Pick and In-game Phases of MSC 2025 Mobile Legends using Random Forest and XGBoostDzaky Fadli Firmansyah, Adam Prayogo Kuncoro, Riyanto · Journal of Applied Informatics and Computing (JAIC) 9(6), pp. 3892-3903 · 2025-12-17 (received 2025-10-31, accepted 2025-12-10) · checked Sep 22, 2026NotesThe body is in Indonesian (abstract in English); decimal commas are as printed. Split into two records (this draft phase and firmansyah-2025-ingame) because the paper holds both a pregame and an in-game model. Training data: 553 games from six qualifier tournaments, 54 teams. | Random Forest (n_estimators=100, max_depth=4) — PREGAME: draft phaseAlso reports: macro precision 57%, recall 57%, F1-score 57% | After the draftOne map | 0.56 | — | — | 0.57 | — | Lowtested on later games · 118 games in the data, test share not statedSmall test | temporaltrained on the qualifier games (patch 1.9.68), tested on the MSC 2025 main event (patch 1.9.91); the draft statistics were computed from training data only 'untuk menghindari kebocoran data' (to avoid data leakage) | 118 games of the MSC 2025 main event (patch 1.9.91), 50% Blue and 50% Red wins, 23 teams; drafts from Liquipedia's API checked against match recordings |
| Match Outcome Prediction in Draft Pick and In-game Phases of MSC 2025 Mobile Legends using Random Forest and XGBoost | XGBoost (n_estimators=100, learning_rate=0.05, max_depth=4) — PREGAME: draft phaseAlso reports: macro precision 55%, recall 55%, F1-score 55% | After the draftOne map | 0.58 | — | — | 0.55 | — | Lowtested on later games · 118 games in the data, test share not statedSmall test | temporaltrained on the qualifier games (patch 1.9.68), tested on the MSC 2025 main event (patch 1.9.91) | 118 games of the MSC 2025 main event (patch 1.9.91), 50% Blue and 50% Red wins |
PaperWhich Heroes to Pick? Learning to Draft in MOBA Games with Neural Networks and Tree SearchSheng Chen, Menghui Zhu, Deheng Ye, Weinan Zhang, Qiang Fu, Wei Yang · IEEE Transactions on Games (arXiv:2012.10171) · 2020-12-18 (arXiv v1; revised 2021-08-05) · checked Sep 22, 2026NotesNo absolute metric is printed in the text, so every metric field is null. Read from the ar5iv HTML through targeted extraction. The paper calls the game both 'Honor of Kings' and 'King of Honor'. | neural-network winning-rate predictor — human dataset (top 5% players)Also reports: values shown only as bars (Fig. 7); text | After the draftOne game | — | — | — | — | — | Lownot tested on later games · 30 games in the data, test share not statedRandom splitSmall test | random'averaged over 10-fold cross validation' | 30 million matches of, 95 heroes |
| Which Heroes to Pick? Learning to Draft in MOBA Games with Neural Networks and Tree Search | neural-network winning-rate predictor — AI self-play datasetAlso reports: values shown only as bars (Fig. 7); text | After the draftOne game | — | — | — | — | — | Lownot tested on later games · 30 games in the data, test share not statedRandom splitSmall test | random'averaged over 10-fold cross validation' | 30 million self-play matches of the trained AI, 95 heroes |
Uses in-game data 10
Gold, kills or the clock from inside the game: a different question from a pre-game forecast.
| Model or setting | Tested on | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
PaperWin Prediction in Multi-Player Esports: Live Professional Match PredictionVictoria J. Hodge, Sam Devlin, Nick Sephton, Florian Block, Peter I. Cowling, Anders Drachen · IEEE Transactions on Games, pp. 368-379 (DOI 10.1109/TG.2019.2948469) · 2019-11-19 online; 2021-12-17 in issue (per White Rose Research Online) · checked Sep 22, 2026NotesAll of this paper's own models are in-game; its only pregame figure, 55.8% from [9], is recorded in hodge-2017. Parameters were chosen by 'comparing the results on the training data set'. Table I's caption warns the datasets, and footnotes that. | Random Forest (all features) and logistic regression (CfsSubsetEval features), tied — IN-GAME at the 20-minute mark: professional test set (TI 2017)Also reports: Table III, Pro-InGame (1-Attr / All / CFS Select): LR 71.89 / 72.97 / 74.59; RF 70.81 / 74.59 / 70.81; GBM 72.43 / 73.51 / 71.89; text: 'accuracies ranged from 70.81% to 74.59%' | In gameOne game | — | — | — | 0.7459 | — | Mediumtested on later games · test size not statedIn-game data | temporal | 186 TI 2017 professional matches that had a replay file and lasted 20 minutes or more |
| Win Prediction in Multi-Player Esports: Live Professional Match Prediction | Random Forest (all features) — IN-GAME at the 20-minute mark: mixed test set (5.7K data)Also reports: Table III, Mixed-InGame (1-Attr / All / CFS Select): LR 74.06 / 77.35 / 77.41; RF 76.36 / 77.51 / 77.24; GBM 76.41 / 77.46 / 77.25; text | In gameOne game | — | — | — | 0.7751 | — | Hightested on later games · about 1,953 test gamesIn-game data | temporal'we split the data 66% for training and 34% for testing as per Weka's train/test split ratio with the data sorted in chronological order' | the latest 34% of 5,744 replays: 23.97% professional (1,377) and 4,367 public matches with MMR above 5000 ('96th percentile'), 27 March to 14 July 2017, replays located through the OpenDota API |
| Win Prediction in Multi-Player Esports: Live Professional Match Prediction | live system (Random Forest, one model per minute) — IN-GAME live at ESL One Hamburg 2017: first prediction, 5-minute markAlso reports: The abstract: 'up to 85% accurate after 5 minutes of gameplay' | In gameOne game | — | — | — | 0.85 | — | Lowtested on later games · 28 test gamesIn-game dataSmall test | forwardlive: minute-by-minute predictions generated and saved during each game, the winner added to the log when it ended | the 28 games of ESL One Hamburg 2017 (26-29 October 2017) |
PaperPandaSkill - Player Performance and Skill Rating in Esports: Application to League of LegendsMaxime De Bois, Flora Parmentier, Raphaël Puget, Matthew Tanti, Jordan Peltier · arXiv · 2025-01-17 (arXiv v1; v2 2025-01-20) · checked Sep 22, 2026NotesNumbers match the repo. In the same Table IV the outcome-only 'PScore + OpenSkill' row prints 65.56% on all games and 65.51% intra-region, above PandaSkill's 64.98% and 64.86% (it trails on inter-region games and on ECE); the paper's own summary: ECE bins are not stated. Generic EWMA rows (all / intra / inter): PScore + EWMA 63.06 / 63.33 / 52.34 (ECE 1.43 / 1.57 / 5.51), PI + EWMA 63.53 / 63.75 / 54.68, PlayeRank + EWMA 62.25 / 62.80 / 51.67. Expert survey (not a win metric): 80.63% majority concordance (SD 6.52), 88.98% unanimity concordance (SD 5.99). The repo's 'chronological replay' is, in the paper's words, a one-month rolling window with a logistic regression trained on the previous year. | role-based XGBoost performance model (the source of PScore) — NOT a pre-match forecast: the winner of a game from that game's own end-game statisticsAlso reports: SD 0.60 (accuracy), SD 0.03 (ECE); 91.79% (SD 0.56) without monotonicity constraints | In gameOne map | — | — | — | 0.9074 | 0.0093 | Lownot tested on later games · test size not statedIn-game dataRandom split | random | the same Leaguepedia corpus, per-player end-of-game statistics |
PaperMatch Outcome Prediction in Draft Pick and In-game Phases of MSC 2025 Mobile Legends using Random Forest and XGBoost (in-game phase)Dzaky Fadli Firmansyah, Adam Prayogo Kuncoro, Riyanto · Journal of Applied Informatics and Computing (JAIC) 9(6), pp. 3892-3903 · 2025-12-17 (received 2025-10-31, accepted 2025-12-10) · checked Sep 22, 2026NotesFor test games the 'last snapshot' is picked by converting each game's final duration to the nearest snapshot time ('mengonversi durasi pertandingan ke interval terdekat'), which uses information known only after the game. Snapshot values were copied by hand from screenshots of match recordings. The authors state the in-game models are a snapshot evaluation, not a real-time system. Split off from firmansyah-2025, which holds the draft phase; cross-reports are there. Snapshot totals over training and test: 671 at minute 9, 582 at 12, 342 at 15, 57 at 18. | Random Forest — IN-GAME: last snapshotAlso reports: macro precision 89%, recall 88%, F1-score 88% | In gameOne map | 0.94 | — | — | 0.88 | — | Lowtested on later games · test size not stated · known problemLeakIn-game data | temporaltrained on qualifier snapshots (patch 1.9.68), tested on MSC 2025 main-event snapshots (patch 1.9.91) | main-event test snapshots (304 across all snapshot times) from MSC 2025 match recordings |
| Match Outcome Prediction in Draft Pick and In-game Phases of MSC 2025 Mobile Legends using Random Forest and XGBoost (in-game phase) | XGBoost — IN-GAME: last snapshotAlso reports: macro precision 86%, recall 84%, F1-score 84% | In gameOne map | 0.94 | — | — | 0.84 | — | Lowtested on later games · test size not stated · known problemLeakIn-game data | temporaltrained on qualifier snapshots (patch 1.9.68), tested on MSC 2025 main-event snapshots (patch 1.9.91) | main-event test snapshots from MSC 2025 match recordings |
| Match Outcome Prediction in Draft Pick and In-game Phases of MSC 2025 Mobile Legends using Random Forest and XGBoost (in-game phase) | Random Forest — IN-GAME: minute 9Also reports: macro precision 76%, recall 76%, F1-score 76% | In gameOne map | 0.86 | — | — | 0.76 | — | Mediumtested on later games · test size not statedIn-game data | temporaltrained on qualifier snapshots (patch 1.9.68), tested on MSC 2025 main-event snapshots (patch 1.9.91) | main-event minute-9 test snapshots |
PaperMachine Learning Models on MOBA Gaming: League of Legends Winner PredictionKaan Arık · Acta Infologica 7(1):139-151 (2023) · 2023-06-15 (published online; submitted 2022-09-26, accepted 2023-05-09) · checked Sep 22, 2026NotesThe confusion-matrix discussion counts 1,295 misclassified samples (591 winners predicted as losers, 704 losers as winners). Nine further classifiers are in Table 2 (RF 0.9481, LDA 0.9449, Ridge 0.9416, Extra Trees 0.9393, AdaBoost 0.9353, DT 0.9263, KNN 0.8879, QDA 0.7951, NB 0.6481 accuracy). Table 4 lists 'Costa and et al. (Costa et al., 2021)' with 'TSSTN' and no number, and 'Mora-Cantallops and Sicilia' with 'k-means'. DergiPark's landing page shows 'Published January 2, 2024' while the PDF says 'Published Online: 15.06.2023'. | LightGBMAlso reports: recall 0.971; precision 0.966; F1 0.969; the abstract rounds the accuracy to 'LightGBM (0.97)' | In gameOne game | 0.996 | — | — | 0.968 | — | Lownot tested on later games · test size not statedIn-game dataRandom split | randomTable 1 'Data partition for train and testing': '%70 - %30', train 92979, test 39849, total 132828; Section 4: (pyCaret). The paper does not say whether Table 2 is scored on the 30% test set or in cross-validation, nor how the 70/30 partition was drawn. | Riot API data; rows are players ; game sessions under 600 seconds removed; region, tier and queue not stated |
| Machine Learning Models on MOBA Gaming: League of Legends Winner Prediction | Logistic RegressionAlso reports: recall 0.956; precision 0.954; F1 0.955; abstract 'Logistic Regression (0.96)' | In gameOne game | 0.988 | — | — | 0.955 | — | Lownot tested on later games · test size not statedIn-game dataRandom split | randomTable 1 'Data partition for train and testing': '%70 - %30', train 92979, test 39849, total 132828; Section 4: (pyCaret). The paper does not say whether Table 2 is scored on the 30% test set or in cross-validation, nor how the 70/30 partition was drawn. | Riot API data; rows are players ; game sessions under 600 seconds removed; region, tier and queue not stated |
| Machine Learning Models on MOBA Gaming: League of Legends Winner Prediction | SVMAlso reports: AUC printed as 0.000 in Table 2; recall 0.954; precision 0.946; F1 0.950; abstract 'SVM and GBC... (0.95)' | In gameOne game | 0.000 | — | — | 0.950 | — | Lownot tested on later games · test size not statedIn-game dataRandom split | randomTable 1 'Data partition for train and testing': '%70 - %30', train 92979, test 39849, total 132828; Section 4: (pyCaret). The paper does not say whether Table 2 is scored on the 30% test set or in cross-validation, nor how the 70/30 partition was drawn. | Riot API data; rows are players ; game sessions under 600 seconds removed; region, tier and queue not stated |
PaperE-Sports Player Performance Metrics for Predicting the Outcome of League of Legends Matches Considering Player RolesFarnod Bahrololloomi, Fabio Klonowski, Sebastian Sauer, Robin Horst, Ralf Dörner · SN Computer Science 4 (2023) · 2023-03-02 (received 2022-06-21, accepted 2022-12-30) · checked Sep 22, 2026NotesThe per-player model is a regression whose target is xpPerMin ; its 15 features after reduction include 'win'. That model's own metrics (R², MAE, RMSD, MAD and an 'accuracy' over 100 shuffled runs) describe the regression, not a win prediction. To predict a match; the paper does not say whether the evaluated match is excluded. Per-role result. | heuristic: each team's sum of player scores from the per-player ML model | In gameOne game | — | — | — | 0.86 | — | Lownot tested on later games · 2,901 games in the data, test share not statedIn-game dataNo held-out test | no held-out testNo held-out test for the 86% | 2,901 games ('Only 5v5 ranked solo games in the standard map Summoner's Rift are considered here'; 29,010 player stats), crawled from the Riot API by querying each player's last 20 games, 2 iterations; region and dates not stated |
PaperApplications of Linear and Ensemble-Based Machine Learning for Predicting Winning Teams in League of LegendsSupratik Chowdhury, Mominul Ahsan, Phoebe Barraclough · Applied Sciences 15(10):5241 (MDPI) · 2025-05-08 · checked Sep 22, 2026NotesThe abstract says a population-representative dataset and that the combined models. Gold open access (CC BY), but not readable here. | best-performing model combining pre-game and in-game features (logistic regression, random forest, C5.0, Gradient Boost and XGBoost were built) | In gameOne game | — | — | — | 0.768 | — | Unknownsplit not stated · test size not statedIn-game dataSplit not stated | not statednot stated in the abstract | 'a dataset derived from the Riot API' that 'more accurately reflects the general player population' (abstract); size, dates and split not stated |
WebsiteBugfix and Correction of Reddit Post Accuracy Claims (blog index title: 'Correction: Reddit Post Accuracy Claims')LoLDraftAI (first person, no visible byline; site-wide HTML meta author 'looyyd') · loldraftai.com · 2025-04-07 ('April 7, 2025') · checked Sep 22, 2026NotesRetracted Reddit post: https://www.reddit.com/r/leagueoflegends/comments/1joumtm/i_made_an_ai_model_that_predicts_62_of_ranked/ (not read; Reddit is unreachable to the fetcher). The author thanks '/u/Impossible_Concert88 for trying to verify the accuracy claims' and writes: Queue: the post says 'ranked' (from the Reddit title) without naming solo/duo; level set to solo queue on that basis. | LoLDraftAI, leaky model (its true accuracy) — the model live | In gameOne game | — | — | — | 0.52 | — | Lowsplit not stated · test size not stated · known problemLeak admittedIn-game dataSplit not stated | not statednot stated | not stated |
PaperPredicting the probability of winning in League of LegendsAdam Czak (supervisor: Petr Novák) · thesis, Czech Technical University in Prague (bachelor's, Faculty of Information Technology) · 2025-05-16 (thesis date); defended 2025-06-17 · checked Sep 22, 2026NotesSection 4.2. The conclusion says 'Over 150,000 matches'. Validation-set tuning results (Tables 5.1-5.3) are not listed. The declaration says AI tools were used in writing the thesis. Queue not named. | logistic regressionAlso reports: F1 0.7395; the abstract prints 'an accuracy of 72.49% and an ECE of 0.0062'; ECE 'with 20 bins' | In gameOne game | — | — | — | 0.7249 | 0.0062 | Unknownsplit not stated · about 24,753 test gamesIn-game dataSplit not stated | not statedWhether matches were assigned at random or by time is not stated. | 15% test split of 165,017 matches downloaded 'from 02 May 2025 through 11 May 2025' (4,405,382 minute snapshots) from the match histories of Challenger, Grandmaster and Master players 'from the major LoL regions', via the Riot API; the logistic regression used a dataset cut by 25%, the RNN one shrunk by a further 75% |
| Predicting the probability of winning in League of Legends | feedforward neural network with data-uncertainty lossAlso reports: F1 0.7352 | In gameOne game | — | — | — | 0.7204 | 0.0068 | Unknownsplit not stated · about 24,753 test gamesIn-game dataSplit not stated | not statedWhether matches were assigned at random or by time is not stated. | 15% test split of 165,017 matches downloaded 'from 02 May 2025 through 11 May 2025' (4,405,382 minute snapshots) from the match histories of Challenger, Grandmaster and Master players 'from the major LoL regions', via the Riot API; the logistic regression used a dataset cut by 25%, the RNN one shrunk by a further 75% |
PaperReal-time eSports Match Result Prediction (real-time part)Yifan Yang, Tian Qin, Yu-Heng Lei · arXiv · 2016-12-10 (arXiv v1) · checked Sep 22, 2026NotesSplit off from yang-2016 so each record's pregame flag is true for every row it holds. benchmarks.md §7 lists this figure as out of scope. The paper | logistic regression, the Attribute Sequence Model and their combinations (the abstract does not say which reaches the figure) — IN-GAME: predicting at the 40th minute | In gameOne game | — | — | — | 0.9373 | — | Unknownsplit not stated · test size not statedIn-game dataSplit not stated | not statedRandom or chronological not stated | test share of the 20,631 (of 78,362) 'very high' skill matches with replay data from the OpenDota API; dates not stated |
PaperVictory prediction in League of Legends using Feature Selection and Ensemble methodsR. Ani, Vishnu Harikumar, Arjun K. Devan, O. S. Deepa · 2019 International Conference on Intelligent Computing and Control Systems (ICCS), pp. 74-77 · 2019-05-01 (OpenAlex / Semantic Scholar date) · checked Sep 22, 2026NotesCosta et al. 2021 re-implemented this paper's banned-champion feature set on 2,840 professional games and got AUC 'lower than 0.6' (see costa-2021). pregame is false because the carried headline is the combined pre-game plus in-game set; per Bahrololloomi et al.'s citation, the pre-match set alone gave 95.52%. | ensemble classifier after feature selection (Random Forest according to citing papers) — combined pre-game and in-game feature set (per the repo) | In gameMap or series not stated | — | — | — | 0.9975 | — | Unknownsplit not stated · test size not statedIn-game dataSplit not stated | not statednot stated in the reachable abstract | not re-verified; the abstract names no dataset (citing papers describe professional games) |
When one source reports another's numbers
Comparison posts and papers that measure someone else's model. Kept apart from the self-reports, and dated like them.
| How they got it | Tested on | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| DraftRec: Personalized Draft Recommendation for Winning in Multi-Player Online Battle Arena GamesHojoon Lee, Dongyoon Hwang, Hyunseung Kim, Byungkun Lee, Jaegul Choo | OptMatch (Gong et al. 2020, KDD): graph neural network hero embeddings — LoL, during the draftAlso reports: MAE 0.4944 | — | — | — | 0.5411 | — | re-implemented on the same data | temporal | Last 10% by time of 279,893 LoL matches collected through the Riot Games API: (62,466 players); region and queue not stated; test-set size not stated |
| DraftRec: Personalized Draft Recommendation for Winning in Multi-Player Online Battle Arena Games | OptMatch (Gong et al. 2020, KDD): graph neural network hero embeddings — LoL, after the draft (all ten champions selected)Also reports: MAE 0.4928 | — | — | — | 0.5449 | — | re-implemented on the same data | temporal | Last 10% by time of 279,893 LoL matches collected through the Riot Games API: (62,466 players); region and queue not stated; test-set size not stated |
| DraftRec: Personalized Draft Recommendation for Winning in Multi-Player Online Battle Arena Games | OptMatch (Gong et al. 2020, KDD): graph neural network hero embeddings — Dota 2, during the draftAlso reports: MAE 0.4842 | — | — | — | 0.5751 | — | re-implemented on the same data | temporal | Last 10% by time of a public Kaggle Dota 2 dataset: 50,000 matches of various skill tiers, November 5 to November 18, 2015; test-set size not stated |
| Using Machine Learning to Predict Game Outcomes Based on Player-Champion Experience in League of LegendsTiffany D. Do, Seong Ioi Wang, Dylan S. Yu, Matthew G. McMillian, Ryan P. McMahan | Chen et al. 2017, Player Skill Decomposition in MOBAs (arXiv:1702.06253) | — | — | — | 0.6024 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Scalable Psychological Momentum Forecasting in EsportsAlfonso White, Daniela M. Romano | Ong, Deolalikar and Peng 2015, Player Behavior and Optimal Team Composition (arXiv:1503.02230) — post-draft LoLAlso reports: train accuracy 74.8%; train n 117,000 | — | — | — | 0.704 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Scalable Psychological Momentum Forecasting in Esports | Ong, Deolalikar and Peng 2015 (arXiv:1503.02230) — post-draft LoLAlso reports: train accuracy 74.8%; train n 117,000 | — | — | — | 0.688 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Scalable Psychological Momentum Forecasting in Esports | Chen et al. 2017 (arXiv:1702.06253) — post-draft LoLAlso reports: train n 208,091 | — | — | — | 0.6024 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Player Skill Decomposition in Multiplayer Online Battle ArenasZhengxing Chen, Yizhou Sun, Magy Seif El-Nasr, Truong-Huy D. Nguyen | TrueSkill (Herbrich, Minka and Graepel 2006) — LoLAlso reports: ± 0.16% | — | — | — | 0.5524 | — | re-implemented on the same data | temporal | The latest 10% by match creation time of the same LoL dataset |
| Player Skill Decomposition in Multiplayer Online Battle Arenas | TrueSkill (Herbrich, Minka and Graepel 2006) — Dota 2Also reports: ± 0.12% | — | — | — | 0.5205 | — | re-implemented on the same data | temporal | The latest 10% by match creation time of the same Dota 2 dataset |
| Online Game Outcome Prediction Model Using Weighted-Based Feature ApproachM. Asyhraf Zamir Zamri, Nurul Aswa Omar, Isredza Rahmi A. Hamid | Gonzalez R., 'ML-Prediction-LoL' (GitHub, 2022)Also reports: Table 1 prints '>90' | — | — | — | 0.8200–0.9048 | — | copied from their work | not statednot stated in this source | not stated in this source |
| Online Game Outcome Prediction Model Using Weighted-Based Feature Approach | Chen et al. 2017 (arXiv:1702.06253) | — | — | — | 0.6024 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Online Game Outcome Prediction Model Using Weighted-Based Feature Approach | Do et al. 2021 (cited as 'Tiffany et al.') | — | — | — | 0.7510 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Machine Learning Models on MOBA Gaming: League of Legends Winner PredictionKaan Arık | Almeida et al. 2017 (CISTI) — 'MOBA' | — | — | — | 0.77 | — | copied from their paper (Table 4 'Related works', column 'Best Accuracy') | not statednot stated in this source | not stated in this source |
| Machine Learning Models on MOBA Gaming: League of Legends Winner Prediction | Ani et al. 2019 (ICCS) — 'MOBA' | — | — | — | 0.99 | — | copied from their paper (Table 4 'Related works', column 'Best Accuracy') | not statednot stated in this source | not stated in this source |
| Machine Learning Models on MOBA Gaming: League of Legends Winner Prediction | Porokhnenko et al. 2019 (FRUCT) — 'MOBA' | — | — | — | 0.70 | — | copied from their paper (Table 4 'Related works', column 'Best Accuracy') | not statednot stated in this source | not stated in this source |
| E-Sports Player Performance Metrics for Predicting the Outcome of League of Legends Matches Considering Player RolesFarnod Bahrololloomi, Fabio Klonowski, Sebastian Sauer, Robin Horst, Ralf Dörner | Ani et al. 2019 (cited as 'Harikumar et al.') — pre-match features | — | — | — | 0.9552 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| E-Sports Player Performance Metrics for Predicting the Outcome of League of Legends Matches Considering Player Roles | Ani et al. 2019 (cited as 'Harikumar et al.') — within-match features | — | — | — | 0.9818 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| E-Sports Player Performance Metrics for Predicting the Outcome of League of Legends Matches Considering Player Roles | Ani et al. 2019 (cited as 'Harikumar et al.') — pre-match and within-match combined | — | — | — | 0.9975 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Feature Analysis to League of Legends Victory Prediction on the Picks and Bans PhaseLincoln Magalhães Costa, Rafael Gomes Mantovani, Fracisco [sic] Carlos Monteiro Souza, Geraldo Xexéo | Ani et al. 2019 (ICCS): the Banned Champions feature set — Banned Champions feature set | <0.6 | — | — | — | — | re-implemented on the same data | random | 2,840 professional LoL matches from Oracle's Elixir, 2021/01/01 to 2021/03/23 (blue side won 46.37%); datasets published on Kaggle (tekpixo/leagueoflegendsprematch2021) |
| Feature Analysis to League of Legends Victory Prediction on the Picks and Bans Phase | Hodge et al. 2019 (IEEE Transactions on Games) — professional Dota 2 championship, after 5 minutes of gameplay | — | — | — | 0.85 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Feature Analysis to League of Legends Victory Prediction on the Picks and Bans Phase | Silva, Pappa and Chaimowicz 2018 — 20 to 25 minutes interval | — | — | — | 0.835 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| GCN-WP – Semi-Supervised Graph Convolutional Networks for Win Prediction in EsportsAlexander J. Bisberg, Emilio Ferrara | SCOPE (Bisberg and Cardona-Rivera 2019, AIIDE): Elo with cross-validated parameters; the first author co-wrote it — as listed in Table IV | — | — | — | 0.597 | — | re-implemented on the same data (the SCOPE procedure, with kill difference as margin of victory) | temporal | 2020 regular season; Table IV does not name the league |
| GCN-WP – Semi-Supervised Graph Convolutional Networks for Win Prediction in Esports | SCOPE (Bisberg and Cardona-Rivera 2019, AIIDE) — LCS | — | — | — | 0.591 | — | re-implemented on the same data | temporal2018 initialization, 2019 parameter validation, 2020 test | LCS 2020 regular season |
| GCN-WP – Semi-Supervised Graph Convolutional Networks for Win Prediction in Esports | SCOPE (Bisberg and Cardona-Rivera 2019, AIIDE) — LPL | — | — | — | 0.597 | — | re-implemented on the same data | temporal2018 initialization, 2019 parameter validation, 2020 test | LPL 2020 regular season |
| PandaSkill - Player Performance and Skill Rating in Esports: Application to League of LegendsMaxime De Bois, Flora Parmentier, Raphaël Puget, Matthew Tanti, Jordan Peltier | OpenSkill (Joshy 2024), updates from the team outcome only — all games | — | — | — | 0.6556 | 0.0153 | ran the rating system on the same data (the 'PScore + OpenSkill' row: 'The skill rating updates only depend on the team outcome of the game and not the player's in-game performance') | walk-forward | Professional LoL games from the Leaguepedia API, 2019-09-15 to 2024-09-15 (37,388 games, 4,927 players; LCK, LPL, LEC, LCS, PCS, VCS, CBLOL, LLA, Worlds, MSI and others); each month after the first year is tested; test-set sizes not stated |
| PandaSkill - Player Performance and Skill Rating in Esports: Application to League of Legends | OpenSkill (Joshy 2024), updates from the team outcome only — intra-region games | — | — | — | 0.6551 | 0.0139 | ran the rating system on the same data (the 'PScore + OpenSkill' row: 'The skill rating updates only depend on the team outcome of the game and not the player's in-game performance') | walk-forward | Professional LoL games from the Leaguepedia API, 2019-09-15 to 2024-09-15 (37,388 games, 4,927 players; LCK, LPL, LEC, LCS, PCS, VCS, CBLOL, LLA, Worlds, MSI and others); each month after the first year is tested; test-set sizes not stated |
| PandaSkill - Player Performance and Skill Rating in Esports: Application to League of Legends | OpenSkill (Joshy 2024), updates from the team outcome only — inter-region games | — | — | — | 0.6739 | 0.0759 | ran the rating system on the same data (the 'PScore + OpenSkill' row: 'The skill rating updates only depend on the team outcome of the game and not the player's in-game performance') | walk-forward | Professional LoL games from the Leaguepedia API, 2019-09-15 to 2024-09-15 (37,388 games, 4,927 players; LCK, LPL, LEC, LCS, PCS, VCS, CBLOL, LLA, Worlds, MSI and others); each month after the first year is tested; test-set sizes not stated |
| Pre-game paired-comparison modeling of professional League of Legends map outcomesMin-Ren Guan, Shen-Ning Tung | Cattelan, Varin and Firth (2013), dynamic Bradley–Terry — global holdoutAlso reports: reported as RPS, which the paper states is numerically identical to the Brier score for a two-outcome game | — | — | 0.2351 | — | — | re-implemented on the same data ('Three deliberate departures from Cattelan, Varin, and Firth's (2013) own specification only strengthen this baseline'; λ swept and the best configuration reported, 'an oracle selection made on the evaluation set itself') | temporalGlobal time-based holdout: fit on the oldest 80% of games by calendar time, score the newest 20% | Newest 20% by calendar time (n_te = 912) of 5,135 professional maps with a first-pick assignment (LPL 1,822, LCK 1,276, LEC 697, LCP 432, CBLOL 277, LCS 216, Worlds 190, MSI 157, First Stand 68), 2024-2026; lolesports telemetry plus Leaguepedia |
| Pre-game paired-comparison modeling of professional League of Legends map outcomes | Cattelan, Varin and Firth (2013), dynamic Bradley–Terry — global holdoutAlso reports: reported as RPS, which the paper states is numerically identical to the Brier score for a two-outcome game | — | — | 0.2391 | — | — | re-implemented on the same data ('Three deliberate departures from Cattelan, Varin, and Firth's (2013) own specification only strengthen this baseline'; λ swept and the best configuration reported, 'an oracle selection made on the evaluation set itself') | temporalGlobal time-based holdout: fit on the oldest 80% of games by calendar time, score the newest 20% | Newest 20% by calendar time (n_te = 912) of 5,135 professional maps with a first-pick assignment (LPL 1,822, LCK 1,276, LEC 697, LCP 432, CBLOL 277, LCS 216, Worlds 190, MSI 157, First Stand 68), 2024-2026; lolesports telemetry plus Leaguepedia |
| Pre-game paired-comparison modeling of professional League of Legends map outcomes | Cattelan, Varin and Firth (2013), dynamic Bradley–Terry — global holdoutAlso reports: reported as RPS, which the paper states is numerically identical to the Brier score for a two-outcome game | — | — | 0.2462 | — | — | re-implemented on the same data ('Three deliberate departures from Cattelan, Varin, and Firth's (2013) own specification only strengthen this baseline'; λ swept and the best configuration reported, 'an oracle selection made on the evaluation set itself') | temporalGlobal time-based holdout: fit on the oldest 80% of games by calendar time, score the newest 20% | Newest 20% by calendar time (n_te = 912) of 5,135 professional maps with a first-pick assignment (LPL 1,822, LCK 1,276, LEC 697, LCP 432, CBLOL 277, LCS 216, Worlds 190, MSI 157, First Stand 68), 2024-2026; lolesports telemetry plus Leaguepedia |
| Predicting the probability of winning in League of LegendsAdam Czak (supervisor: Petr Novák) | Jalovaara 2024 (Aalto University master's thesis) | — | — | — | ≈0.76 | 0.0090 | copied from their work | not statednot stated in this source | not stated in this source |
| Predicting the probability of winning in League of Legends | LoLDraftAI (loldraftai.com, accessed 2025-04-23)Also reports: 'exact details aren't public as of writing this' | — | — | — | ≈0.55 | — | copied from their site | not statednot stated in this source | not stated in this source |
| Real-time eSports Match Result PredictionYifan Yang, Tian Qin, Yu-Heng Lei | Conley et al. [4] (Conley & Perry 2013, Stanford CS229), hero-selection features — Hero features onlyAlso reports: the paper adds that both previous works 'claimed to have achieved over 70% prediction accuracy' on data with | — | — | — | 0.5879 | — | re-implemented on the same data ('We train their features using LR on our dataset.') | not statedthis paper's 9:1 split | the same test share of 78,362 'very high' skill matches |
| Real-time eSports Match Result Prediction | Kinkade et al. [8] (Kinkade & Lim 2015, UC San Diego), hero features plus hero-against-hero winning rate — Hero features only | — | — | — | 0.5869 | — | re-implemented on the same data ('We train their features using LR on our dataset.') | not statedthis paper's 9:1 split | the same test share of 78,362 'very high' skill matches |
| How Does He Saw Me? A Recommendation Engine for Picking Heroes in Dota 2Kevin Conley, Daniel Perry | Dota2cp (dota2cp.com), a hero-recommendation web application | — | — | — | 0.63 | — | copied from the Dota2cp author's own claim ('the author reports a 63% accuracy') | not statednot stated | not stated |
| To win or not to win? A prediction model to determine the outcome of a DotA2 matchKaushik Kalyanaraman | Dota2cp (dota2cp.com), a hero-recommendation engine | — | — | — | 0.63 | — | copied from the Dota2cp author's claim | not statednot stated | not stated |
| To win or not to win? A prediction model to determine the outcome of a DotA2 match | Conley & Perry 2013 [4] | — | — | — | 0.698 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Predicting Game Outcome in Dota 2 with NLP and Machine Learning AlgorithmsZiping Wan, Bálint Sándor | Conley & Perry 2013 [14] | — | — | — | 0.698 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Predicting Game Outcome in Dota 2 with NLP and Machine Learning Algorithms | Conley & Perry 2013 [14] | — | — | — | 0.70 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Predicting Game Outcome in Dota 2 with NLP and Machine Learning Algorithms | Song, Zhang, Ma 2015 [15]Also reports: 'Their best achieved accuracy was around 53%.' The Song report itself prints only error rates | — | — | — | ≈0.53 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Predicting outcomes of professional DotA 2 matchesPetra Grutzik, Joe Higgins, Long Tran | GosuGamers.net bettors: for each match, the outcome with the largest betting amount — per 'match' (a set of games whose format varies; draws possible) | — | — | — | 0.628 | — | bettors' consensus | forwardbets are placed before the match ('bettors only have information before a match') | 16,800+ matches in GosuGamers.net betting records, May 2013 to November 2017 |
| Win Prediction in Multi-Player Esports: Live Professional Match PredictionVictoria J. Hodge, Sam Devlin, Nick Sephton, Florian Block, Peter I. Cowling, Anders Drachen | Agarwala & Pearce 2014 [18] — Pre-Game | — | — | — | 0.62 | — | copied from their paper (literature table) | not statednot stated in this source's table | Table I lists 40,000 games, skill 'Low-High' (dataset size, not test size) |
| Win Prediction in Multi-Player Esports: Live Professional Match Prediction | Conley & Perry 2013 [11] — Pre-Game | — | — | — | 0.70 | — | copied from their paper (literature table) | not statednot stated in this source's table | Table I lists 18,000 games, skill 'Top 8%' |
| Win Prediction in Multi-Player Esports: Live Professional Match Prediction | Johansson & Wikström 2015 [21] (Master's thesis, Blekinge Institute of Technology) — IN-GAME: 5 minsAlso reports: text: after 20 minutes '99%' | — | — | — | 0.82 | — | copied from their paper (literature table) | not statednot stated in this source's table | Table I lists 59,208 games, skill 'V. High' |
| Match Outcome Prediction in Draft Pick and In-game Phases of MSC 2025 Mobile Legends using Random Forest and XGBoostDzaky Fadli Firmansyah, Adam Prayogo Kuncoro, Riyanto | Hamir & Andono 2025 [7] (JPTI), Mobile LegendsAlso reports: this source says the feature construction 'tidak dijelaskan secara jelas' (is not clearly explained) | — | — | — | 0.85 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Match Outcome Prediction in Draft Pick and In-game Phases of MSC 2025 Mobile Legends using Random Forest and XGBoost | Chowdhury, Ahsan, Barraclough 2025 [8] (Applied Sciences), League of Legends | — | — | — | 0.768 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| Match Outcome Prediction in Draft Pick and In-game Phases of MSC 2025 Mobile Legends using Random Forest and XGBoost | Stanlly, Putra, Qomariyah 2022 [10] (IEEE IAICT), Dota 2 item and hero dataAlso reports: 'sekitar' means 'about' | — | — | — | ≈0.91 | — | copied from their paper | not statednot stated in this source | not stated in this source |
| LoLTheory vs LoLDraftAI: A Detailed ComparisonLoLDraftAI (no visible byline; site-wide HTML meta author 'looyyd') | LoLTheoryAlso reports: 95% bootstrap CIs: log loss (0.6885–0.6928), accuracy (53.87–54.93), Brier (0.2476–0.2496) | — | 0.6907 | 0.2486 | 0.5441 | 0.0326 | queried their live tool: 'For each match we call LoLTheory's /team-comp/team-performance/0 endpoint' with 'the correct rank-range for its elo bucket' | temporal'Eval set = matches played after both cutoffs'; LoLTheory was 'queried live on patch 16.8.1 during April 2026' and (the post does not say whether that window could contain the evaluated matches) | 32,930 matches; 'Eval set = matches played after both cutoffs'; exact dates and patches of the matches not stated (LoLTheory 'queried live on patch 16.8.1 during April 2026') |
| DraftGap vs LoLDraftAI: A Detailed ComparisonLoLDraftAI (no visible byline; site-wide HTML meta author 'looyyd') | DraftGapAlso reports: 95% bootstrap CIs: log loss (0.6851–0.6887), accuracy (54.12–55.20), Brier (0.2460–0.2478) | — | 0.6869 | 0.2469 | 0.5466 | 0.0199 | reproduced their model offline: 'DraftGap's predictions above were produced by a Python port of their open-source formula, run against the exact JSON datasets their website was serving on 2026-04-18 at 12:57 UTC' | temporal'DraftGap snapshot frozen at 2026-04-18, 12:57 UTC' | 32,750 matches: ranked solo/duo, Emerald+ ('DraftGap pulls its data from Lolalytics at tier=emerald_plus'), EUW1 and KR; 'Matches where any (champion, role) cell has fewer than 50 games in DraftGap's current-patch dataset are excluded'; exact dates and patches not stated |
| iTero vs LoLDraftAI: A Detailed ComparisonLoLDraftAI (no visible byline; site-wide HTML meta author 'looyyd') | iTero — judge: DraftGap and judge: LoLDraftAIAlso reports: No iTero-only figure is printed; the head-to-head result is LoLDraftAI's win rate over iTero's drafts: 55.0% (45.0–65.0) judged by DraftGap, 93.0% (88.0–98.0) judged by LoLDraftAI | — | — | — | — | — | queried their live tool: iTero's picks came from 'its API' (endpoint, settings and date not stated) | no held-out testnone: synthetic drafts scored by judge models | the same 100 drafts |
| What's the Best League of Legends Draft AI?Winrate (winrate.gg; no named author) | iTero — Top 1Also reports: realized win rate when the actual pick was in iTero's Top 1: 52.7% (n 332) | — | — | — | — | — | queried their live tool: (iTero via HTTP) | not statednot stated for iTero | 20,000 'real ranked solo-queue games (queue 420)', 'Gold tier and above', 'Recent live patches', both teams' five roles assigned, from the 12 hours before the query ('ORDER BY sipHash64(match_id) LIMIT 20000'); one champion erased per game, nine visible; region not stated; 'n is the number of games where the actual pick was in the site's Top K' |
| What's the Best League of Legends Draft AI? | LoLDraftAI — Top 1Also reports: realized win rate when the actual pick was in LoLDraftAI's Top 1: 48.5% (n 472) | — | — | — | — | — | queried their live tool: (LoLDraftAI via HTTP) | not statednot stated for LoLDraftAI | 20,000 'real ranked solo-queue games (queue 420)', 'Gold tier and above', 'Recent live patches', both teams' five roles assigned, from the 12 hours before the query ('ORDER BY sipHash64(match_id) LIMIT 20000'); one champion erased per game, nine visible; region not stated; 'n is the number of games where the actual pick was in the site's Top K' |
| What's the Best League of Legends Draft AI? | iTero — Top 3Also reports: realized win rate when the actual pick was in iTero's Top 3: 51.9% (n 962) | — | — | — | — | — | queried their live tool: (iTero via HTTP) | not statednot stated for iTero | 20,000 'real ranked solo-queue games (queue 420)', 'Gold tier and above', 'Recent live patches', both teams' five roles assigned, from the 12 hours before the query ('ORDER BY sipHash64(match_id) LIMIT 20000'); one champion erased per game, nine visible; region not stated; 'n is the number of games where the actual pick was in the site's Top K' |
| Model Accuracy & Track Record (draftlol.ai)draftlol.ai | market (pre-match table's 'Mkt Brier'; the page does not name the market or when its price was taken)Also reports: N 554 (same rows) | — | — | 0.2001 | — | — | market prices | forward'Every prediction frozen at registration time'; resolved automatically | the same 554 pre-match rows |
| Model Accuracy & Track Record (draftlol.ai) | Polymarket ('compared against Polymarket odds captured at the same moment', i.e. at draft lock-in)Also reports: N 1855 (same rows) | — | — | 0.2181 | — | — | market prices | forward'Every prediction frozen at registration time'; resolved automatically | the same 1,855 at-draft rows |
| LoL Esports Predictor for Polymarket (farrelo25/lol-esports-predictor)farrelo25 (Hugging Face) | PandaSkill (De Bois et al., 2025) | — | — | — | 0.907 | — | copied from their paper (listed under the card's 'Methodology & References') | not statednot stated in the card | not stated in the card |
Ask us to audit a source
Point us at a paper or site we do not list, or tell us a row is wrong. The button on every row fills in its link.
Requests
Open checks 9
Things we could not finish ourselves. Each stays here until it is done.
- winrate.gg draft recommendation benchmarkNeeds a browserOur notes quote confidence intervals the page text does not contain; they may exist only inside its charts.
- Drafter.oneNeeds a browserA '71% on 5 of 7 Worlds games' claim was seen only in a search summary; the page does not render for our fetcher.
- DraftGapNeeds a browserThe site does not render for our fetcher; its GitHub README publishes no test of the model.
- STRATZ, The International 8 post (2018)Not reachableMedium refuses our fetcher, so the 77% we list is carried from our notes and not re-checked.
- Wang & Shang 2017 (ICIS)Not reachableThe publisher page did not load; 58.99% is carried from a thesis that quotes it.
- Semenov et al. 2016 (AIST)Abstract onlyBehind a paywall; the abstract prints no numbers, so the AUCs are carried from our notes.
- Ani et al. 2019 (ICCS)Abstract onlyThe abstract has no numbers; 99.75% is carried from a paper that cites it.
- Jadowski & Cunningham 2021Abstract onlyThe repository's PDF link expires before it can be read.
- Kellich (Charles University thesis)Not read yetThe full text is linked now; nobody has read it.
Found but not yet audited (24)
- Ong, Deolalikar & Peng 2015 (arXiv 1503.02230), a post-draft LoL model
- Kim & Lee 2019 (ISSAT) and Kim & Lee 2020 (KIPS), LoL pre-game models
- Lin 2016, a LoL pre-game model
- Lee, Hong & Yang 2020 (ICTC), LoL in-game
- Jalovaara 2024 (Aalto thesis) and Jafari & Olbe 2021 (KTH thesis), LoL in-game
- Kim et al. 2017 (CSCW), 'What makes a strong team?'
- Galán-Sales et al. 2024 (CAEPIA)
- Arizona State University repository: 'Developing a Model to Predict Match Outcomes in League of Legends'
- Gonzalez, ML-Prediction-LoL (GitHub)
- Nanzhi Wang et al. 2018 (ICMAI), Dota 2
- Semenov et al. 2017, a Dota 2 literature review (CEUR-WS 1842)
- Kinkade & Lim 2015 (UC San Diego report), Dota 2
- Yu et al. 2018, MOBA-Slice (arXiv 1807.08360)
- W. Wang 2016 (National College of Ireland thesis), Dota 2
- Stanlly et al. 2022 (IAICT)
- Putro 2024; Hamir & Andono 2025; Sena & Emanuel 2023, Mobile Legends
- Dota2cp, a website claiming 63%
- 'Team Radiant or Dire?' (ACM, 10.1145/3635059.3635067)
- arXiv 1803.10402, Dota 2
- 'Trouncing in Dota 2' (AIIDE 2020)
- arXiv 2008.06313, Honor of Kings, in-game
- Almeida et al. 2017 and Porokhnenko et al. 2019, Dota 2
- TechLabs Aachen, predicting pro win chances from the draft (Medium)
- cthorrez/esports-bench (GitHub), the code behind EsportsBench