MLB model vs market · single · 2026-07-29
model 52% vs ask 40c → +10.6c net of costs (shadow: the model has not earned belief)
The bar is the edge: the gap between what the model thought this outcome was worth and what the exchange was asking for it. A ticket is only issued when that gap clears 5¢ after costs — the exchange spread, and a 1¢ haircut on Polymarket quotes because roughly 30% of them are already stale by the time we could act on one.
Model
51.6%
Paid
40¢
Edge
+10.6¢
Stake
25.0u
Full Kelly is growth-optimal only if the edge estimate is exactly right; overestimating it by 2× makes full Kelly growth-negative, which is easily done. Quarter Kelly keeps roughly 55% of the maximum log-growth and never lost more than half the bankroll across 2,000 simulated paths. This ticket hit the 25-unit cap (2.5% of bankroll), so it was cut again.
| Leg | Model | Price | Edge | Result |
|---|---|---|---|---|
| MIA moneyline @ 0.40PHI @ MIA | 52% | 40¢ | +11.6¢ | pending |
Not the same question as “did it win”. A bet at 38% loses nearly two times in three when it is priced perfectly, so a loss is not evidence of a mistake and a win is not evidence of skill. The closing price is the better judge — it is the market’s final answer, and it is calibrated almost perfectly.
Wrong bet, wrong result.
The market closed at 40¢ against the 40¢ we paid — it moved −0.5¢ away from us. That is evidence the price we took was not cheap.
Still open. Settles after the game resolves.
Four features, each a difference between the two sides, each oriented so a bigger number favours the home team. They are standardised, weighted, summed, and squashed into a probability.
away starter ERA − home starter ERA
away WHIP − home WHIP
home runs scored/g − away runs scored/g
away runs allowed/g − home runs allowed/g
Published 2026-07-25. Each feature is standardised to the training mean and spread, multiplied by its weight, summed with an intercept of 0.120, and squashed through a logistic into a probability. Note what the fit actually learned: the starter ERA edge — the number a human would lead with — carries a weight near zero, and team offense carries the most. The model is not the story a fan would tell about a pitching matchup.
These are the live weights, not the ones frozen at this ticket’s issue time — the ratings behind them advance with every game, so the exact feature values this ticket saw are not recoverable. The stored 51.6% above point-in-time; the weight shape below is what the model looks like today.
The research verdict is that this model does NOT beat the market (backtest Brier parity at best). Shadow tier: if the model is as beatable as measured, these tickets will lose and the record will say so — that is the point.
tier: shadow