The MLB shadow test
Not deployedThe model forecast well and bet badly. Where it disagreed most with the market, it predicted 64.7% and delivered 54.0%.
A Monte Carlo simulator that prices MLB hitter props from pitch-level Statcast data. I wrote a pass/fail gate into the methodology document two days before the first prediction, then ran it against live sportsbook prices for 87 days without betting a dollar.
It failed. A head-to-head regression against the market price showed why: the model carries genuine information the closing line doesn't (β = 0.231, z = 2.20) but earns only about a quarter of the weight the edge formula was giving it. The decision rule was wrong, not the model — and neither was good enough to ship.