← Back to blog

Published on Thu Sep 17 2026 15:08:00 GMT+0000 (Coordinated Universal Time) by cresencio

NFL Week 1: A 9–7 Record Hid the Cost of Confidence

My Week 1 ensemble finished 9–7. The less-confident version made exactly the same picks and earned better probability scores.

That is the most useful result from the first full week. Before kickoff, the Week 1 preview promised to freeze the forecasts and grade them honestly. The Seattle recap showed the confidence adjustment helping on a win. The Rams recap showed it making a miss expensive.

With all 16 games complete, the accounting is less flattering than the winning record suggests. The sharp final probabilities scored worse than the raw blend on both Brier score and log loss. They also scored worse than assigning every game a 50% home-win probability.

One week is a small sample. It is still a complete week that belongs in the record.

The full board, with nothing removed

These are the original post-audit picks published on September 2. The percentages are the final forecast’s probability for its selected winner. Each final score links to an official team source; all 16 results also match the saved evaluation data.

MatchupFrozen pickFinal scorePick result
NE @ SEASEA 70.9%SEA 13–10Correct
SF vs. LALA 93.6%SF 27–7Wrong
CHI @ CARCHI 72.9%CHI 59–37Correct
TB @ CINCIN 61.7%CIN 33–27Correct
NO @ DETDET 91.7%DET 31–30, OTCorrect
BUF @ HOUHOU 62.3%BUF 36–31Wrong
BAL @ INDIND 70.5%BAL 41–23Wrong
CLE @ JAXJAX 96.2%JAX 34–10Correct
ATL @ PITPIT 79.4%PIT 20–13Correct
NYJ @ TENTEN 61.5%NYJ 23–10Wrong
ARI @ LACLAC 89.2%ARI 26–14Wrong
MIA @ LVMIA 74.3%LV 27–13Wrong
GB @ MINGB 56.4%MIN 39–22Wrong
WAS @ PHIPHI 95.2%PHI 24–22Correct
DAL @ NYGNYG 64.8%NYG 28–20Correct
DEN @ KCKC 50.8%KC 31–10Correct

Los Angeles was the designated home team for the Melbourne game. That designation also determines the probability orientation in the saved data.

The ten games with unanimous agreement among the three active models finished 6–4. The six disagreement games finished 3–3 for the ensemble. Agreement helped identify a group of favorites; it did not make the group safe.

The scores behind the record

Week 1 used Elo, Logistic regression, and Bayesian, with effective weights of approximately 21.1%, 40.6%, and 38.3%. XGBoost and Random Forest were inactive and receive no grade. The raw and final ensembles below are two stages of that same three-model blend.

ForecastCorrect winnersMean Brier scoreMean log loss
Elo9 of 160.22980.6509
Logistic regression8 of 160.26470.7362
Bayesian8 of 160.24660.6855
Raw weighted ensemble9 of 160.24500.6823
Final ensemble, T = 0.49 of 160.28310.8106
Constant 50% reference0.25000.6931

Lower is better. Every score uses the original full-precision home-win probability, p, and an outcome, y, of 1 for a home win or 0 for an away win. Binary Brier score is (p - y)^2. Log loss is -[y × ln(p) + (1 - y) × ln(1 - p)], using the natural logarithm. The table averages each measure across all 16 games.

Elo earned the best scores this week. Bayesian picked fewer winners than the final ensemble and still scored better on both probability measures. That is why I want more than a record: the cost of a mistake depends on how much probability the forecast left for the outcome that happened.

The 50% row is a simple probability reference, not a betting strategy or a claim that coin flips are the best forecasting system. The raw blend narrowly beat it this week. The final blend did not.

The confidence layer added a measurable cost

Temperature scaling at T = 0.4 moved every raw probability farther from 50% without changing a single selected winner. Across the whole slate, that raised mean Brier score by 0.0381 and mean log loss by 0.1283.

The five loudest picks from the preview—Jacksonville, Philadelphia, the Rams, Detroit, and the Chargers—went 3–2. Their average final confidence was 93.2%. The Rams and Chargers losses imposed large penalties because those forecasts left the eventual winners little room.

Here is the entire confidence breakdown, including the quieter picks:

Final favorite probabilityGamesCorrectAverage forecast confidence
50% to under 60%2153.6%
60% to under 70%4262.6%
70% to under 80%5373.6%
80% and above5393.2%

These groups contain only two to five games each. They describe this slate; they cannot establish reliable long-run win rates. The narrower conclusion is already enough: sharpening made this week’s saved probabilities worse on both scoring rules.

That does not identify an optimal replacement temperature. Choosing a setting because it would have made these 16 completed games look better would require a separate evaluation on future games. The inherited calibration evidence was limited before Week 1, and this result adds an unfavorable observation rather than settling the whole question.

Revisit the disagreements, too

The preseason watchlist should not disappear behind the Rams loss.

Elo alone picked Baltimore among the active models, and Baltimore won. Elo also picked Minnesota while the other two models and the ensemble favored Green Bay. Both went into Elo’s correct column.

The Giants game went the other way: Logistic favored New York against Elo and Bayesian, pulled the ensemble toward the Giants, and got the winner right. In Kansas City, Bayesian was the only active model favoring the Chiefs, and the weighted blend just crossed 50% in their direction.

Those cases are useful to preserve because they show where the component forecasts differed. They do not establish that I should follow whichever model was right most recently. Even this small set has different models supplying the correct dissent.

A close win and a big win obey the same rule

Detroit’s overtime win, Philadelphia’s two-point win, and Kansas City’s twenty-one-point win all count as correct winner forecasts. Their margins do not validate or invalidate the probabilities.

Kansas City’s 50.8% was nearly a coin flip before the game. Winning 31–10 does not retrospectively turn it into a confident forecast. Detroit winning by one point does not make 91.7% a failed margin prediction. No margin was predicted in either case.

That is the lesson from Seattle applied consistently across a whole week. These saved outputs forecast winners. I cannot borrow the final scores to claim success on totals, spreads, or the shape of a game.

The Week 2 preview starts a new forecast record, and Detroit at Buffalo gets its own closer look. This first record stays intact: 9–7 on winners, with extra confidence that cost more than it earned.

Source note: Forecasts come from the September 2 post-audit comparison used in the original Week 1 preview. Final scores were checked against the official team sources linked in the table and the local Week 1 evaluation export dated September 16. All grades were recalculated from the original probability values, rather than rounded report percentages. No forecasts were regenerated, models retrained, or production records changed for this review. This is a 16-game winner-probability evaluation, not a betting-return analysis.

Written by cresencio

← Back to blog