MLB model vs market · single · 2026-07-27
model 57% vs ask 50c → +5.6c net of costs (shadow: the model has not earned belief)
The bar is the edge: the gap between what the model thought this outcome was worth and what the exchange was asking for it. A ticket is only issued when that gap clears 5¢ after costs — the exchange spread, and a 1¢ haircut on Polymarket quotes because roughly 30% of them are already stale by the time we could act on one.
Model
56.6%
Paid
50¢
Edge
+5.6¢
Stake
25.0u
Full Kelly is growth-optimal only if the edge estimate is exactly right; overestimating it by 2× makes full Kelly growth-negative, which is easily done. Quarter Kelly keeps roughly 55% of the maximum log-growth and never lost more than half the bankroll across 2,000 simulated paths. This ticket hit the 25-unit cap (2.5% of bankroll), so it was cut again.
| Leg | Model | Price | Edge | Result |
|---|---|---|---|---|
| MIA moneyline @ 0.50PHI @ MIA | 57% | 50¢ | +6.6¢ | won |
Not the same question as “did it win”. A bet at 38% loses nearly two times in three when it is priced perfectly, so a loss is not evidence of a mistake and a win is not evidence of skill. The closing price is the better judge — it is the market’s final answer, and it is calibrated almost perfectly.
Right bet, right result.
The market closed at 56¢ against the 50¢ we paid — it moved +5.5¢ toward us. Closing prices are the best-calibrated number available, so a move in our direction is evidence the price we took was genuinely cheap
Won — staked 25.0u, returned +25.0u
This number cannot promote or kill a strategy on its own. A genuine 5–8% edge could not be proven at p<0.05 with 672 bets; it takes well over a thousand. P&L is reported here because hiding it would be its own kind of dishonesty, not because it decides anything yet.
Four features, each a difference between the two sides, each oriented so a bigger number favours the home team. They are standardised, weighted, summed, and squashed into a probability.
away starter ERA − home starter ERA
away WHIP − home WHIP
home runs scored/g − away runs scored/g
away runs allowed/g − home runs allowed/g
Published 2026-07-25. Each feature is standardised to the training mean and spread, multiplied by its weight, summed with an intercept of 0.120, and squashed through a logistic into a probability. Note what the fit actually learned: the starter ERA edge — the number a human would lead with — carries a weight near zero, and team offense carries the most. The model is not the story a fan would tell about a pitching matchup.
These are the live weights, not the ones frozen at this ticket’s issue time — the ratings behind them advance with every game, so the exact feature values this ticket saw are not recoverable. The stored 56.6% above point-in-time; the weight shape below is what the model looks like today.
The research verdict is that this model does NOT beat the market (backtest Brier parity at best). Shadow tier: if the model is as beatable as measured, these tickets will lose and the record will say so — that is the point.
tier: shadow
One ticket proves nothing in either direction. The record that matters is the mean closing-line value across every ticket a strategy ever issued, with the failures still in the denominator — which is what the desk reports, and why its promotion gate is still two unfilled bars.