Published on Thu Sep 17 2026 15:08:00 GMT+0000 (Coordinated Universal Time) by cresencio
NFL Week 1: A 9–7 Record Hid the Cost of Confidence
My Week 1 ensemble finished 9–7. The less-confident version made exactly the same picks and earned better probability scores.
That is the most useful result from the first full week. Before kickoff, the Week 1 preview promised to freeze the forecasts and grade them honestly. The Seattle recap showed the confidence adjustment helping on a win. The Rams recap showed it making a miss expensive.
With all 16 games complete, the accounting is less flattering than the winning record suggests. The sharp final probabilities scored worse than the raw blend on both Brier score and log loss. They also scored worse than assigning every game a 50% home-win probability.
One week is a small sample. It is still a complete week that belongs in the record.
The full board, with nothing removed
These are the original post-audit picks published on September 2. The percentages are the final forecast’s probability for its selected winner. Each final score links to an official team source; all 16 results also match the saved evaluation data.
| Matchup | Frozen pick | Final score | Pick result |
|---|---|---|---|
| NE @ SEA | SEA 70.9% | SEA 13–10 | Correct |
| SF vs. LA | LA 93.6% | SF 27–7 | Wrong |
| CHI @ CAR | CHI 72.9% | CHI 59–37 | Correct |
| TB @ CIN | CIN 61.7% | CIN 33–27 | Correct |
| NO @ DET | DET 91.7% | DET 31–30, OT | Correct |
| BUF @ HOU | HOU 62.3% | BUF 36–31 | Wrong |
| BAL @ IND | IND 70.5% | BAL 41–23 | Wrong |
| CLE @ JAX | JAX 96.2% | JAX 34–10 | Correct |
| ATL @ PIT | PIT 79.4% | PIT 20–13 | Correct |
| NYJ @ TEN | TEN 61.5% | NYJ 23–10 | Wrong |
| ARI @ LAC | LAC 89.2% | ARI 26–14 | Wrong |
| MIA @ LV | MIA 74.3% | LV 27–13 | Wrong |
| GB @ MIN | GB 56.4% | MIN 39–22 | Wrong |
| WAS @ PHI | PHI 95.2% | PHI 24–22 | Correct |
| DAL @ NYG | NYG 64.8% | NYG 28–20 | Correct |
| DEN @ KC | KC 50.8% | KC 31–10 | Correct |
Los Angeles was the designated home team for the Melbourne game. That designation also determines the probability orientation in the saved data.
The ten games with unanimous agreement among the three active models finished 6–4. The six disagreement games finished 3–3 for the ensemble. Agreement helped identify a group of favorites; it did not make the group safe.
The scores behind the record
Week 1 used Elo, Logistic regression, and Bayesian, with effective weights of approximately 21.1%, 40.6%, and 38.3%. XGBoost and Random Forest were inactive and receive no grade. The raw and final ensembles below are two stages of that same three-model blend.
| Forecast | Correct winners | Mean Brier score | Mean log loss |
|---|---|---|---|
| Elo | 9 of 16 | 0.2298 | 0.6509 |
| Logistic regression | 8 of 16 | 0.2647 | 0.7362 |
| Bayesian | 8 of 16 | 0.2466 | 0.6855 |
| Raw weighted ensemble | 9 of 16 | 0.2450 | 0.6823 |
| Final ensemble, T = 0.4 | 9 of 16 | 0.2831 | 0.8106 |
| Constant 50% reference | — | 0.2500 | 0.6931 |
Lower is better. Every score uses the original full-precision home-win probability, p, and an outcome, y, of 1 for a home win or 0 for an away win. Binary Brier score is (p - y)^2. Log loss is -[y × ln(p) + (1 - y) × ln(1 - p)], using the natural logarithm. The table averages each measure across all 16 games.
Elo earned the best scores this week. Bayesian picked fewer winners than the final ensemble and still scored better on both probability measures. That is why I want more than a record: the cost of a mistake depends on how much probability the forecast left for the outcome that happened.
The 50% row is a simple probability reference, not a betting strategy or a claim that coin flips are the best forecasting system. The raw blend narrowly beat it this week. The final blend did not.
The confidence layer added a measurable cost
Temperature scaling at T = 0.4 moved every raw probability farther from 50% without changing a single selected winner. Across the whole slate, that raised mean Brier score by 0.0381 and mean log loss by 0.1283.
The five loudest picks from the preview—Jacksonville, Philadelphia, the Rams, Detroit, and the Chargers—went 3–2. Their average final confidence was 93.2%. The Rams and Chargers losses imposed large penalties because those forecasts left the eventual winners little room.
Here is the entire confidence breakdown, including the quieter picks:
| Final favorite probability | Games | Correct | Average forecast confidence |
|---|---|---|---|
| 50% to under 60% | 2 | 1 | 53.6% |
| 60% to under 70% | 4 | 2 | 62.6% |
| 70% to under 80% | 5 | 3 | 73.6% |
| 80% and above | 5 | 3 | 93.2% |
These groups contain only two to five games each. They describe this slate; they cannot establish reliable long-run win rates. The narrower conclusion is already enough: sharpening made this week’s saved probabilities worse on both scoring rules.
That does not identify an optimal replacement temperature. Choosing a setting because it would have made these 16 completed games look better would require a separate evaluation on future games. The inherited calibration evidence was limited before Week 1, and this result adds an unfavorable observation rather than settling the whole question.
Revisit the disagreements, too
The preseason watchlist should not disappear behind the Rams loss.
Elo alone picked Baltimore among the active models, and Baltimore won. Elo also picked Minnesota while the other two models and the ensemble favored Green Bay. Both went into Elo’s correct column.
The Giants game went the other way: Logistic favored New York against Elo and Bayesian, pulled the ensemble toward the Giants, and got the winner right. In Kansas City, Bayesian was the only active model favoring the Chiefs, and the weighted blend just crossed 50% in their direction.
Those cases are useful to preserve because they show where the component forecasts differed. They do not establish that I should follow whichever model was right most recently. Even this small set has different models supplying the correct dissent.
A close win and a big win obey the same rule
Detroit’s overtime win, Philadelphia’s two-point win, and Kansas City’s twenty-one-point win all count as correct winner forecasts. Their margins do not validate or invalidate the probabilities.
Kansas City’s 50.8% was nearly a coin flip before the game. Winning 31–10 does not retrospectively turn it into a confident forecast. Detroit winning by one point does not make 91.7% a failed margin prediction. No margin was predicted in either case.
That is the lesson from Seattle applied consistently across a whole week. These saved outputs forecast winners. I cannot borrow the final scores to claim success on totals, spreads, or the shape of a game.
The Week 2 preview starts a new forecast record, and Detroit at Buffalo gets its own closer look. This first record stays intact: 9–7 on winners, with extra confidence that cost more than it earned.
Source note: Forecasts come from the September 2 post-audit comparison used in the original Week 1 preview. Final scores were checked against the official team sources linked in the table and the local Week 1 evaluation export dated September 16. All grades were recalculated from the original probability values, rather than rounded report percentages. No forecasts were regenerated, models retrained, or production records changed for this review. This is a 16-game winner-probability evaluation, not a betting-return analysis.
Written by cresencio
← Back to blog