NBA Game Prediction: Can Features Beat Elo?
2026-07-24
Predicting NBA winners over 68 years of games, benchmarked honestly against FiveThirtyEight's Elo. A lesson in how hard a good baseline is to beat.
classification temporal validation calibration sports analytics
Problem
A well-built baseline is often the hardest thing to beat. Elo ratings already encode team strength and home advantage, so the interesting question is whether engineered features like rest days and recent form can actually improve on them. I set out to answer that honestly rather than to manufacture a win.
Data
FiveThirtyEight NBA Elo (Public, every NBA/BAA game 1947 to 2015 (63,000 modeled games))
Approach
A strict temporal split (train on seasons through 2010, test on 2011 to 2015, never a random split), leakage-safe rest and rolling-form features built from each team's prior games only, and three baselines to clear: always-home, higher-Elo, and FiveThirtyEight's own forecast. Models were judged on probabilities, not just picks: accuracy, log-loss, Brier score, and a calibration curve.
Result
My logistic model matched FiveThirtyEight's Elo at about 67% accuracy but did not beat it. Home court, rest, and form are already baked into Elo. That is the right result to report: knowing why a strong baseline holds is more valuable than a fabricated improvement. Along the way I found and documented a framing quirk (the recorded team is almost always home, so that feature is near-constant).

demo video
a short recorded walkthrough goes here once the project is complete
What I'd do differently
Beating Elo needs signal Elo cannot see: injuries, load management, travel distance, lineup changes, or betting-market movement. Those are the features I would add next.