We built real win-probability models for two sports from public box-score data, tested them strictly out-of-sample, and held each to the only bar that counts: the market's price. The two sports answer in opposite shapes — and neither model beats the market.
| ▲Skill — passing / QB rating | AUC 0.658 |
| Line — rush eff, sacks, TFL/QB-hits | AUC 0.618 |
Out-of-sample AUC of each pre-game rating (0.50 = coin flip). Skill clearly leads.
A logistic win-probability model on those ratings, tested out-of-sample: Brier 0.221 vs a 0.249 baseline — 11.4% skill, 64.0% accuracy over 447 games. Genuinely predictive.
Against Polymarket's moneyline, out-of-sample over 443 games: model Brier 0.220 (64.1%) vs market 0.200 (69.8%). The market wins — it prices in injuries, lines and sharp money the model never sees.
| ▲Pitching — starter ERA/WHIP + staff | AUC 0.576 |
| Offense — rolling runs scored | AUC 0.522 |
Out-of-sample AUC of each pre-game rating (0.50 = coin flip). Pitching clearly leads.
A logistic win-probability model on those ratings, tested out-of-sample: Brier 0.246 vs a 0.250 baseline — 1.5% skill, 55.6% accuracy over 1195 games. Barely better than a coin flip — the sport is near-random game to game.
Against Polymarket's moneyline, out-of-sample over 640 games: model Brier 0.246 (56.1%) vs market 0.247 (55.8%). A statistical tie (paired t = -0.2) — and pre-vig, so a real bet still loses the spread. No edge.
In the NFL, passing quality — not the line of scrimmage — drives winning, and a public model has real skill (~11.4%). But the market still clearly beats it: there's signal to capture, and the market captures more.
In MLB, pitching and offense matter about equally and both barely beat a coin flip. The model ties the market — not because it's good, but because baseball is nearly random game to game and there's little for anyone to know. The sport itself is the ceiling.
Different reasons, same bottom line — a public-data model doesn't out-price these markets — reached from a completely fresh direction, exactly like the stock funnel and the arena.
Every model is opened up on its own page — the data it eats, the machinery step by step, its live weights where they exist, one fully worked prediction, and exactly which file each knob lives in. Written so someone who has never seen a model can follow, and so tweaking one starts from understanding instead of archaeology.
Not a backtest — this logs each upcoming game's model probability and the market's price before first pitch, then scores it when the game ends. A record hindsight can't fake, accruing through the season.
Over 329 scored games so far: model Brier 0.246 (57%) vs market-at-prediction 0.246 (54%). Dead even so far — as the backtest predicted.
Against the closing line — the harder bar, captured in the last ~90 minutes before first pitch — over 283 games: model Brier 0.247 (55%) vs close 0.247 (54%).
| date | matchup (starters) | model | market | close | result |
|---|---|---|---|---|---|
| 09-15 | SEA @ LAAKade Anderson vs Reid Detmers | 56% | — | — | — |
| 09-15 | SD @ COLCasey Mize vs Tomoyuki Sugano | 45% | — | — | — |
| 09-14 | SF @ STLLanden Roupp vs Quinn Mathews | 54% | 55% | — | — |
| 09-14 | NYY @ MINWill Warren vs Bailey Ober | 47% | 46% | — | — |
| 09-14 | ATL @ CHCAJ Smith-Shawver vs David Peterson | 55% | 55% | — | — |
| 09-14 | DET @ TORTroy Melton vs Jose Soriano | 48% | 54% | — | — |
| 09-14 | LAD @ CINTarik Skubal vs Nick Lodolo | 38% | 41% | — | — |
| 09-14 | CHW @ CLEDavis Martin vs Gavin Williams | 54% | 55% | — | — |
| 09-13 | LAD @ MIAJustin Wrobleski vs Eury Perez | 47% | 44% | — | — |
| 09-13 | SD @ SFNick Pivetta vs Logan Webb | 49% | — | — | — |
| 09-13 | TEX @ ARICal Quantrill vs Eduardo Rodriguez | 53% | — | — | — |
| 09-13 | SEA @ ATHBryce Miller vs Jacob Lopez | 47% | — | — | — |
| 09-13 | PIT @ CHCBubba Chandler vs Matthew Boyd | 57% | — | — | — |
| 09-13 | CHW @ STLSean Burke vs Michael McGreevy | 52% | — | — | — |
| 09-13 | CLE @ MINJoey Cantillo vs Joe Ryan | 54% | — | — | voided |
| 09-13 | CIN @ MILChase Burns vs Shane Drohan | 60% | — | — | — |
| 09-13 | HOU @ TBHayden Wesneski vs Freddy Peralta | 54% | — | — | — |
| 09-13 | BAL @ TORTrevor Rogers vs Dylan Cease | 55% | 54% | — | — |
| 09-13 | PHI @ ATLAndrew Painter vs Grant Holmes | 57% | 53% | — | — |
| 09-13 | NYM @ NYYChristian Scott vs Cam Schlittler | 62% | 60% | — | — |
| 09-13 | KC @ BOSNoah Cameron vs Payton Tolle | 60% | 56% | — | — |
| 09-13 | LAA @ WSHGrayson Rodriguez vs Jake Irvin | 57% | 51% | — | — |
| 09-13 | COL @ DETGabriel Hughes vs Jackson Jobe | 61% | 59% | — | — |
| 09-13 | SEA @ ATHBryan Woo vs Gage Jump | 46% | 43% | — | SEA won |
Home-team win probability. market is the price when the prediction was logged — often days early, often not yet listed; close is the price captured in the last ~90 minutes before first pitch, the harder bar. Rows are never edited or deleted — a game that never resolves as predicted (postponed, canceled), or a row whose recorded market price can't be trusted, is voided: kept in the table, counted toward nothing.
Correction, 2026-07-21: the first days of ledger rows were logged by a market matcher that could pair a game with a different game's price — MLB series share a team pair, and the matcher keyed on the pair alone (it also read "Red Sox" as the White Sox). Every pre-fix row that carried a market price is voided above rather than silently corrected; the model probabilities were unaffected. The market comparison restarts from rows logged after the fix, which matches by scheduled start time.