MLB model vs market · single · 2026-07-25
model 33% vs ask 27c → +5.5c net of costs (shadow: the model has not earned belief)
The bar is the edge: the gap between what the model thought this outcome was worth and what the exchange was asking for it. A ticket is only issued when that gap clears 5¢ after costs — the exchange spread, and a 1¢ haircut on Polymarket quotes because roughly 30% of them are already stale by the time we could act on one.
Model
33.5%
Paid
27¢
Edge
+5.5¢
Stake
18.9u
Full Kelly is growth-optimal only if the edge estimate is exactly right; overestimating it by 2× makes full Kelly growth-negative, which is easily done. Quarter Kelly keeps roughly 55% of the maximum log-growth and never lost more than half the bankroll across 2,000 simulated paths.
| Leg | Model | Price | Edge | Result |
|---|---|---|---|---|
| COL moneyline @ 0.27COL @ MIL | 33% | 27¢ | +6.5¢ | lost |
Not the same question as “did it win”. A bet at 38% loses nearly two times in three when it is priced perfectly, so a loss is not evidence of a mistake and a win is not evidence of skill. The closing price is the better judge — it is the market’s final answer, and it is calibrated almost perfectly.
Wrong bet, wrong result.
The market closed at 24¢ against the 27¢ we paid — it moved −3.5¢ away from us. That is evidence the price we took was not cheap.
Lost — staked 18.9u, returned −18.9u
This number cannot promote or kill a strategy on its own. A genuine 5–8% edge could not be proven at p<0.05 with 672 bets; it takes well over a thousand. P&L is reported here because hiding it would be its own kind of dishonesty, not because it decides anything yet.
Four features, each a difference between the two sides, each oriented so a bigger number favours the home team. They are standardised, weighted, summed, and squashed into a probability.
away starter ERA − home starter ERA
away WHIP − home WHIP
home runs scored/g − away runs scored/g
away runs allowed/g − home runs allowed/g
Published 2026-07-25. Each feature is standardised to the training mean and spread, multiplied by its weight, summed with an intercept of 0.120, and squashed through a logistic into a probability. Note what the fit actually learned: the starter ERA edge — the number a human would lead with — carries a weight near zero, and team offense carries the most. The model is not the story a fan would tell about a pitching matchup.
These are the live weights, not the ones frozen at this ticket’s issue time — the ratings behind them advance with every game, so the exact feature values this ticket saw are not recoverable. The stored 33.5% above point-in-time; the weight shape below is what the model looks like today.
The research verdict is that this model does NOT beat the market (backtest Brier parity at best). Shadow tier: if the model is as beatable as measured, these tickets will lose and the record will say so — that is the point.
tier: shadow
One ticket proves nothing in either direction. The record that matters is the mean closing-line value across every ticket a strategy ever issued, with the failures still in the denominator — which is what the desk reports, and why its promotion gate is still two unfilled bars.