← Back to blog

Published on Tue Sep 22 2026 16:54:59 GMT+0000 (Coordinated Universal Time) by cresencio

NFL Week 2: Eleven Picks Above 90%. Five Winners

The original Week 2 ensemble went 8–8. Eleven of its favorites carried probabilities above 90%. Only five won.

The preview put that confidence on the page before kickoff. It also recorded the reason to inspect it carefully: Bayesian had been excluded from the blend, leaving Logistic regression with almost twice Elo’s weight. Temperature scaling then pushed eleven favorites above 90%, although none reached that level in the raw blend.

Now the results are in. The extra confidence made both probability scores worse, by considerably more than it did in Week 1. Elo picked ten winners. The ensemble picked eight and left very little probability for several outcomes that happened.

There is also a correction to account for. Bayesian’s exclusion came from a feature-timing bug that was fixed during the week. That fix belongs in the explanation. It does not replace the forecasts already published.

The original board, graded in full

These are the sixteen picks from the original weekly preview. Final scores match the saved evaluation export and the NFL’s Week 2 results. Percentages refer to the selected team’s chance of winning.

MatchupOriginal final pickFinal scorePick result
DET @ BUFBUF 79.4%BUF 41–31Correct
CAR @ ATLATL 98.3%CAR 34–3Wrong
NO @ BALBAL 99.5%NO 24–17Wrong
MIN @ CHICHI 51.5%MIN 9–3Wrong
CIN @ HOUCIN 93.4%CIN 20–6Correct
PIT @ NEPIT 91.7%NE 20–3Wrong
GB @ NYJNYJ 94.0%GB 20–17Wrong
CLE @ TBTB 94.5%CLE 23–19Wrong
PHI @ TENPHI 96.6%PHI 24–20Correct
JAX @ DENJAX 95.6%DEN 20–13Wrong
LV @ LACLV 88.8%LV 26–14Correct
SEA @ ARISEA 87.4%SEA 31–7Correct
WAS @ DALDAL 92.4%DAL 37–20Correct
MIA @ SFSF 99.3%SF 35–13Correct
IND @ KCKC 99.4%KC 33–30Correct
NYG @ LANYG 89.8%LA 28–6Wrong

LA denotes the Rams; LAC denotes the Chargers.

The eleven favorites above 90% averaged 95.9% confidence and finished 5–6. The three above 99%—Baltimore, San Francisco, and Kansas City—went 2–1. Baltimore’s loss is enough to make that group’s confidence costly even though most of its picks won.

An 8–8 record does not describe the whole problem

The raw blend used 34.2% Elo and 65.8% Logistic. Bayesian’s original probabilities remained available for separate evaluation but had zero ensemble weight. XGBoost and Random Forest were unavailable and receive no grade.

ForecastCorrect winnersMean Brier scoreMean log loss
Elo10 of 160.24700.6913
Logistic regression8 of 160.41021.5279
Original Bayesian, excluded from blend7 of 160.24880.6907
Raw weighted ensemble8 of 160.31480.8628
Final ensemble, T = 0.48 of 160.41511.5203
Constant 50% reference0.25000.6931

Lower is better. Binary Brier score is (p - y)^2; log loss is the negative natural logarithm of the probability assigned to the winner. Calculations use full-precision home-win probabilities and a home-win outcome of 1 or 0, then average over all sixteen games.

Temperature scaling changed no winners. It increased mean Brier score by 0.1002 and mean log loss by 0.6575. The raw blend already scored worse than the 50% reference. Sharpening made the same set of picks substantially more expensive.

Atlanta and Baltimore show why. The final forecast left Carolina about 1.7% and New Orleans about 0.5%. Those losses produced log losses of 4.0551 and 5.2573, respectively. A correct 99% pick earns a small loss close to zero; a wrong one can add a large penalty. That asymmetry is the point of scoring the probabilities.

Bayesian’s original forecast had the lowest log loss in this table despite the fewest correct picks. Its estimates stayed near 50%, so its mistakes were less expensive. That does not vindicate the bug that produced those estimates. It shows why a winner count and a probability score answer different questions.

The disagreements went 2–4 for the ensemble

Elo and Logistic disagreed in six games. The ensemble followed Logistic in all six, winning Cincinnati and Las Vegas while missing Minnesota, Green Bay, Denver, and the Rams. Elo got the other four right.

The ten games where the two contributors agreed finished 6–4. Agreement did not prevent the Atlanta, Baltimore, Pittsburgh, or Tampa Bay misses.

This closes the watchlist from the preview without selecting only the convenient examples. Cincinnati’s 93.4% pick won. Jacksonville’s 95.6% and the Jets’ 94.0% picks lost. Chicago’s 51.5% lean also lost, but it had left Minnesota nearly half the probability. Those are all one loss or one win in the record; their scoring costs differ sharply.

Giants–Rams has two forecasts to preserve

The Monday preview recorded a corrected forecast before that game’s kickoff. Bayesian’s inputs had been repaired to use completed Week 1 games instead of pregame snapshots with missing cumulative values. Bayesian returned to the blend and switched from the Rams to the Giants.

The corrected ensemble still picked New York, now at 92.3% rather than the original 89.8%. The Rams won 28–6. Both versions missed, and the update was more confident in the losing side.

Here is the promised accounting for that game. Elo and Logistic were unchanged between versions.

ForecastPregame pickBrier scoreLog loss
Elo, both versionsLA 80.4%0.03840.2182
Logistic, both versionsNYG 97.0%0.94003.4917
Original Bayesian, zero blend weightLA 56.1%0.19240.5774
Corrected BayesianNYG 77.0%0.59261.4690
Original raw ensembleNYG 70.5%0.49711.2210
Original final ensembleNYG 89.8%0.80702.2860
Corrected raw ensembleNYG 73.0%0.53271.3089
Corrected final ensembleNYG 92.3%0.85212.5651

The weekly totals above use the original board throughout. The corrected Monday forecast stays here as a separately identified pregame update. Mixing versions after seeing the results would make the weekly comparison less meaningful.

Fixing an input bug can be necessary even when the first corrected pick loses. Likewise, Elo winning this disagreement does not prove it should receive all the weight next week.

Two weeks into the record

Through 32 games, Elo is 19–13, the ensemble is 17–15, Logistic is 16–16, and the original Bayesian forecasts are 15–17. The raw and final ensembles have the same 17 winners. Their pooled mean Brier scores are 0.2799 raw versus 0.3491 final; pooled log loss is 0.7726 versus 1.1654.

That combines the actual weekly blends: three contributors in Week 1, two in the original Week 2 board. It is not a test of one unchanged model combination.

The narrower finding is consistent across both weeks: the retained T = 0.4 adjustment worsened the saved probabilities on both scoring rules. Thirty-two games do not establish the best replacement temperature. They do give me a concrete result to carry into the Week 3 preview, where Bayesian returns but the confidence question remains.

Source note: The weekly review uses the versioned original pregame forecasts and the September 22 grading export. The corrected Giants–Rams row comes from the evidence retained with its September 21 preview. All probabilities and scores were recalculated before rounding; source hashes and full-precision values are preserved in the companion editorial evidence. This article evaluates winner forecasts, not spreads, totals, or betting returns. Drafting it did not regenerate forecasts, alter model settings, or settle the paper ledger.

Written by cresencio

← Back to blog