MLB model vs market · single · 2026-07-24
model 38% vs ask 28c → +9.1c net of costs (shadow: the model has not earned belief)
The bar is the edge: the gap between what the model thought this outcome was worth and what the exchange was asking for it. A ticket is only issued when that gap clears 5¢ after costs — the exchange spread, and a 1¢ haircut on Polymarket quotes because roughly 30% of them are already stale by the time we could act on one.
Model
38.1%
Paid
28¢
Edge
+9.1¢
Stake
25.0u
Full Kelly is growth-optimal only if the edge estimate is exactly right; overestimating it by 2× makes full Kelly growth-negative, which is easily done. Quarter Kelly keeps roughly 55% of the maximum log-growth and never lost more than half the bankroll across 2,000 simulated paths. This ticket hit the 25-unit cap (2.5% of bankroll), so it was cut again.
| Leg | Model | Price | Edge | Result |
|---|---|---|---|---|
| KC moneyline @ 0.28KC @ DET | 38% | 28¢ | +10.1¢ | pending |
Not the same question as “did it win”. A bet at 38% loses nearly two times in three when it is priced perfectly, so a loss is not evidence of a mistake and a win is not evidence of skill. The closing price is the better judge — it is the market’s final answer, and it is calibrated almost perfectly.
Right price, no result.
The market closed at 30¢ against the 28¢ we paid — it moved +1.5¢ toward us. Closing prices are the best-calibrated number available, so a move in our direction is evidence the price we took was genuinely cheap The ticket was voided before it could resolve, so this judges the price and nothing else.
Voided — stake returned, no result recorded.
ledger: starter scratch: Noah Cameron → Beck Way A void is not a loss and not a win; it is removed from the record rather than counted as either, because a ticket whose premise disappeared never tested anything. The row stays visible so the void cannot quietly improve the record.
Four features, each a difference between the two sides, each oriented so a bigger number favours the home team. They are standardised, weighted, summed, and squashed into a probability.
away starter ERA − home starter ERA
away WHIP − home WHIP
home runs scored/g − away runs scored/g
away runs allowed/g − home runs allowed/g
Published 2026-07-25. Each feature is standardised to the training mean and spread, multiplied by its weight, summed with an intercept of 0.120, and squashed through a logistic into a probability. Note what the fit actually learned: the starter ERA edge — the number a human would lead with — carries a weight near zero, and team offense carries the most. The model is not the story a fan would tell about a pitching matchup.
These are the live weights, not the ones frozen at this ticket’s issue time — the ratings behind them advance with every game, so the exact feature values this ticket saw are not recoverable. The stored 38.1% above point-in-time; the weight shape below is what the model looks like today.
The research verdict is that this model does NOT beat the market (backtest Brier parity at best). Shadow tier: if the model is as beatable as measured, these tickets will lose and the record will say so — that is the point.
tier: shadow