Published on Tue Sep 22 2026 16:54:59 GMT+0000 (Coordinated Universal Time) by cresencio
NFL Week 2: Eleven Picks Above 90%. Five Winners
The original Week 2 ensemble went 8–8. Eleven of its favorites carried probabilities above 90%. Only five won.
The preview put that confidence on the page before kickoff. It also recorded the reason to inspect it carefully: Bayesian had been excluded from the blend, leaving Logistic regression with almost twice Elo’s weight. Temperature scaling then pushed eleven favorites above 90%, although none reached that level in the raw blend.
Now the results are in. The extra confidence made both probability scores worse, by considerably more than it did in Week 1. Elo picked ten winners. The ensemble picked eight and left very little probability for several outcomes that happened.
There is also a correction to account for. Bayesian’s exclusion came from a feature-timing bug that was fixed during the week. That fix belongs in the explanation. It does not replace the forecasts already published.
The original board, graded in full
These are the sixteen picks from the original weekly preview. Final scores match the saved evaluation export and the NFL’s Week 2 results. Percentages refer to the selected team’s chance of winning.
| Matchup | Original final pick | Final score | Pick result |
|---|---|---|---|
| DET @ BUF | BUF 79.4% | BUF 41–31 | Correct |
| CAR @ ATL | ATL 98.3% | CAR 34–3 | Wrong |
| NO @ BAL | BAL 99.5% | NO 24–17 | Wrong |
| MIN @ CHI | CHI 51.5% | MIN 9–3 | Wrong |
| CIN @ HOU | CIN 93.4% | CIN 20–6 | Correct |
| PIT @ NE | PIT 91.7% | NE 20–3 | Wrong |
| GB @ NYJ | NYJ 94.0% | GB 20–17 | Wrong |
| CLE @ TB | TB 94.5% | CLE 23–19 | Wrong |
| PHI @ TEN | PHI 96.6% | PHI 24–20 | Correct |
| JAX @ DEN | JAX 95.6% | DEN 20–13 | Wrong |
| LV @ LAC | LV 88.8% | LV 26–14 | Correct |
| SEA @ ARI | SEA 87.4% | SEA 31–7 | Correct |
| WAS @ DAL | DAL 92.4% | DAL 37–20 | Correct |
| MIA @ SF | SF 99.3% | SF 35–13 | Correct |
| IND @ KC | KC 99.4% | KC 33–30 | Correct |
| NYG @ LA | NYG 89.8% | LA 28–6 | Wrong |
LA denotes the Rams; LAC denotes the Chargers.
The eleven favorites above 90% averaged 95.9% confidence and finished 5–6. The three above 99%—Baltimore, San Francisco, and Kansas City—went 2–1. Baltimore’s loss is enough to make that group’s confidence costly even though most of its picks won.
An 8–8 record does not describe the whole problem
The raw blend used 34.2% Elo and 65.8% Logistic. Bayesian’s original probabilities remained available for separate evaluation but had zero ensemble weight. XGBoost and Random Forest were unavailable and receive no grade.
| Forecast | Correct winners | Mean Brier score | Mean log loss |
|---|---|---|---|
| Elo | 10 of 16 | 0.2470 | 0.6913 |
| Logistic regression | 8 of 16 | 0.4102 | 1.5279 |
| Original Bayesian, excluded from blend | 7 of 16 | 0.2488 | 0.6907 |
| Raw weighted ensemble | 8 of 16 | 0.3148 | 0.8628 |
| Final ensemble, T = 0.4 | 8 of 16 | 0.4151 | 1.5203 |
| Constant 50% reference | — | 0.2500 | 0.6931 |
Lower is better. Binary Brier score is (p - y)^2; log loss is the negative natural logarithm of the probability assigned to the winner. Calculations use full-precision home-win probabilities and a home-win outcome of 1 or 0, then average over all sixteen games.
Temperature scaling changed no winners. It increased mean Brier score by 0.1002 and mean log loss by 0.6575. The raw blend already scored worse than the 50% reference. Sharpening made the same set of picks substantially more expensive.
Atlanta and Baltimore show why. The final forecast left Carolina about 1.7% and New Orleans about 0.5%. Those losses produced log losses of 4.0551 and 5.2573, respectively. A correct 99% pick earns a small loss close to zero; a wrong one can add a large penalty. That asymmetry is the point of scoring the probabilities.
Bayesian’s original forecast had the lowest log loss in this table despite the fewest correct picks. Its estimates stayed near 50%, so its mistakes were less expensive. That does not vindicate the bug that produced those estimates. It shows why a winner count and a probability score answer different questions.
The disagreements went 2–4 for the ensemble
Elo and Logistic disagreed in six games. The ensemble followed Logistic in all six, winning Cincinnati and Las Vegas while missing Minnesota, Green Bay, Denver, and the Rams. Elo got the other four right.
The ten games where the two contributors agreed finished 6–4. Agreement did not prevent the Atlanta, Baltimore, Pittsburgh, or Tampa Bay misses.
This closes the watchlist from the preview without selecting only the convenient examples. Cincinnati’s 93.4% pick won. Jacksonville’s 95.6% and the Jets’ 94.0% picks lost. Chicago’s 51.5% lean also lost, but it had left Minnesota nearly half the probability. Those are all one loss or one win in the record; their scoring costs differ sharply.
Giants–Rams has two forecasts to preserve
The Monday preview recorded a corrected forecast before that game’s kickoff. Bayesian’s inputs had been repaired to use completed Week 1 games instead of pregame snapshots with missing cumulative values. Bayesian returned to the blend and switched from the Rams to the Giants.
The corrected ensemble still picked New York, now at 92.3% rather than the original 89.8%. The Rams won 28–6. Both versions missed, and the update was more confident in the losing side.
Here is the promised accounting for that game. Elo and Logistic were unchanged between versions.
| Forecast | Pregame pick | Brier score | Log loss |
|---|---|---|---|
| Elo, both versions | LA 80.4% | 0.0384 | 0.2182 |
| Logistic, both versions | NYG 97.0% | 0.9400 | 3.4917 |
| Original Bayesian, zero blend weight | LA 56.1% | 0.1924 | 0.5774 |
| Corrected Bayesian | NYG 77.0% | 0.5926 | 1.4690 |
| Original raw ensemble | NYG 70.5% | 0.4971 | 1.2210 |
| Original final ensemble | NYG 89.8% | 0.8070 | 2.2860 |
| Corrected raw ensemble | NYG 73.0% | 0.5327 | 1.3089 |
| Corrected final ensemble | NYG 92.3% | 0.8521 | 2.5651 |
The weekly totals above use the original board throughout. The corrected Monday forecast stays here as a separately identified pregame update. Mixing versions after seeing the results would make the weekly comparison less meaningful.
Fixing an input bug can be necessary even when the first corrected pick loses. Likewise, Elo winning this disagreement does not prove it should receive all the weight next week.
Two weeks into the record
Through 32 games, Elo is 19–13, the ensemble is 17–15, Logistic is 16–16, and the original Bayesian forecasts are 15–17. The raw and final ensembles have the same 17 winners. Their pooled mean Brier scores are 0.2799 raw versus 0.3491 final; pooled log loss is 0.7726 versus 1.1654.
That combines the actual weekly blends: three contributors in Week 1, two in the original Week 2 board. It is not a test of one unchanged model combination.
The narrower finding is consistent across both weeks: the retained T = 0.4 adjustment worsened the saved probabilities on both scoring rules. Thirty-two games do not establish the best replacement temperature. They do give me a concrete result to carry into the Week 3 preview, where Bayesian returns but the confidence question remains.
Source note: The weekly review uses the versioned original pregame forecasts and the September 22 grading export. The corrected Giants–Rams row comes from the evidence retained with its September 21 preview. All probabilities and scores were recalculated before rounding; source hashes and full-precision values are preserved in the companion editorial evidence. This article evaluates winner forecasts, not spreads, totals, or betting returns. Drafting it did not regenerate forecasts, alter model settings, or settle the paper ledger.
Written by cresencio
← Back to blog