Published on Thu Sep 17 2026 15:08:01 GMT+0000 (Coordinated Universal Time) by cresencio
NFL Week 2 2026: The Forecast Got Louder. The Ensemble Got Smaller
The saved Week 2 forecast has eleven favorites above 90%. That is the headline I trust least without reading the numbers underneath it.
The Week 1 recap records a 9–7 start for both the raw and final ensembles. The same winners produced different probability scores: mean Brier score was 0.2450 raw and 0.2831 final, with lower being better. Sharpening hurt that sixteen-game average. It is a small sample, but it belongs in the record before the next set of confident picks.
The Rams follow-up showed the cost on one game: the wrong pick at 93.6% received a much larger penalty than the same wrong pick at 74.5%. Week 2 gives us another slate to evaluate, not a reason to forget that result.
This week adds another question. The saved file contains Elo, Logistic, and Bayesian forecasts, but only Elo and Logistic contribute to the final ensemble. Before treating a 99% pick as three models reaching the same conclusion, we need to explain why.
Three forecasts on the page, two in the blend
The pipeline checks how much each model’s home-win probabilities vary across the slate. If their standard deviation falls below 0.05, that model receives zero weight in the blend.
Bayesian crosses that threshold this week. Its home-win probabilities run from 51.1% to 57.7%, with a standard deviation of about 0.0218. The probabilities still appear in the export, but they do not influence the final number.
That rule is a mechanical screening choice. A narrow range is not, by itself, evidence that Bayesian is inaccurate or that excluding it will improve the forecast. Those are questions for evaluation.
| Model | Available in the saved file? | Effective Week 2 weight |
|---|---|---|
| Elo | Yes | 34.2% |
| Logistic regression | Yes | 65.8% |
| Bayesian | Yes, excluded by the slate-variance check | 0% |
| XGBoost | No | 0% |
| Random Forest | No | 0% |
XGBoost and Random Forest remain unavailable under the early-season rolling-data requirements. The Bayesian exclusion is a separate event. It leaves Logistic with almost twice Elo’s weight.
This is not a weight change I made for the article. Reconstructing the September 15 export with those two contributing models and the retained T = 0.4 temperature adjustment reproduces every saved final probability to numerical precision.
The raw blend and final ensemble select the same winner. Temperature scaling changes how strongly that winner is favored.
The complete Week 2 board
The table preserves the saved forecast, including Bayesian’s separate opinion. The Bayesian column has zero ensemble weight this week. Raw percentages are reconstructed before temperature scaling; final percentages are the exported predictions.
| Matchup | Elo pick | Logistic pick | Bayesian pick, zero weight | Raw blend | Final ensemble |
|---|---|---|---|---|---|
| DET @ BUF | BUF | BUF | BUF | BUF 63.1% | BUF 79.4% |
| CAR @ ATL | ATL | ATL | ATL | ATL 83.4% | ATL 98.3% |
| NO @ BAL | BAL | BAL | BAL | BAL 89.1% | BAL 99.5% |
| MIN @ CHI | MIN | CHI | CHI | CHI 50.6% | CHI 51.5% |
| CIN @ HOU | HOU | CIN | HOU | CIN 74.3% | CIN 93.4% |
| PIT @ NE | PIT | PIT | NE | PIT 72.3% | PIT 91.7% |
| GB @ NYJ | GB | NYJ | NYJ | NYJ 75.1% | NYJ 94.0% |
| CLE @ TB | TB | TB | TB | TB 75.7% | TB 94.5% |
| PHI @ TEN | PHI | PHI | TEN | PHI 79.3% | PHI 96.6% |
| JAX @ DEN | DEN | JAX | DEN | JAX 77.4% | JAX 95.6% |
| LV @ LAC | LAC | LV | LAC | LV 69.6% | LV 88.8% |
| SEA @ ARI | SEA | SEA | ARI | SEA 68.5% | SEA 87.4% |
| WAS @ DAL | DAL | DAL | DAL | DAL 73.1% | DAL 92.4% |
| MIA @ SF | SF | SF | SF | SF 88.1% | SF 99.3% |
| IND @ KC | KC | KC | KC | KC 88.7% | KC 99.4% |
| NYG @ LA | LA | NYG | LA | NYG 70.5% | NYG 89.8% |
Schedule checked against the NFL: Detroit–Buffalo opens September 17; fourteen games follow September 20; Giants–Rams closes September 21.
LA denotes the Rams; LAC denotes the Chargers.
None of the raw favorites reaches 90%. After the temperature adjustment, eleven do. Baltimore, Kansas City, and San Francisco each exceed 99% in the final forecast.
That is a substantial increase in certainty. It should be visible beside the picks, not hidden behind them.
The games that will tell us the most
Jacksonville at Denver: two very different forecasts
Elo gives Denver 65.8%. Logistic gives Jacksonville 99.8%. Bayesian also picks Denver at 57.7%, but contributes no weight.
The raw blend favors Jacksonville at 77.4%, and temperature scaling takes it to 95.6%. This is one of the clearest tests of Logistic’s extreme probabilities and the weight they carry. Calling it broad agreement would misrepresent the forecast.
Cincinnati at Houston: the same disagreement, another road favorite
Elo favors Houston at 69.4%; Logistic favors Cincinnati at 97.0%. Bayesian’s separate estimate also favors Houston.
The final Cincinnati pick reaches 93.4%, up from 74.3% before sharpening. The result will matter, but the postgame review also needs to retain how much probability each model assigned to the winner. Counting this as one correct or incorrect pick will not settle which estimate was better.
Green Bay at the Jets: 94% without agreement between the contributors
Elo takes Green Bay at 65.9%. Logistic gives the Jets 96.4%. The weighted blend lands on New York at 75.1%, then rises to 94.0%.
Elo and Logistic disagree in six games this week, and the final ensemble takes Logistic’s side in all six. The other three are Minnesota–Chicago, Las Vegas–Chargers, and Giants–Rams. That concentration of influence deserves a separate look after the games.
Minnesota at Chicago: a small lean that stays small
This is the exception to the loud board. Elo gives Minnesota 51.5%; Logistic gives Chicago 51.7%. Their weighted average gives Chicago 50.6%, becoming 51.5% after the adjustment.
The saved pick is Chicago. The useful interpretation is that the two contributors leave this game very close to even. A winner here will not turn a marginal forecast into a strong one retrospectively.
New Orleans at Baltimore: agreement still leaves a confidence question
All three available forecasts favor Baltimore, but their probabilities are far apart: 77.6% Elo, 95.1% Logistic, and 53.4% Bayesian. The two-model raw blend gives Baltimore 89.1%. The final output gives it 99.5%.
The direction is shared. The degree of confidence is not. Even a Baltimore win would be one observation, not evidence that this class of forecast deserves 99.5% confidence.
Tonight starts the next record
Buffalo is the saved pick against Detroit: 63.1% raw, 79.4% final. All three available models choose the Bills, while Elo and Logistic supply the actual blend. The Lions–Bills preview examines the matchup and puts the possible grading outcomes on the page before kickoff.
There is no score prediction attached to that 79.4%. The same is true of every number on this board. Margin, totals, and against-the-spread outputs are unavailable in the saved Week 2 forecast, so these probabilities cannot support a claimed spread or totals edge.
I am also not using a single game’s result to explain the extreme Logistic estimates as proven changes in team quality. The retained export lets us inspect the probabilities and reproduce the blend; it does not establish a verified current-form explanation for every input.
What gets graded next
The Week 2 review will grade the full slate with the same rules used in the earlier recaps: correct winners, binary Brier score, and natural-log loss, calculated from full-precision saved probabilities. Raw and final ensemble scores belong beside each other. Bayesian can be evaluated as an available forecast even though it has zero blend weight.
The retained T = 0.4 setting still carries the earlier calibration limitation: it was fitted using a different model combination and included fitted-sample inputs. Removing one contributor does not supply new evidence that sharpening improves the remaining pair.
The predictions are recorded. The confidence still has work to do.
Source note: Forecasts come from the regular-season Week 2 full-precision export generated September 15, 2026, and retained in the prediction repository on September 16. The code and configuration at that retained revision reproduce all sixteen final probabilities from Elo and Logistic, with a maximum absolute difference below 3 × 10⁻¹⁶. Percentages are rounded only for display. This editorial review did not regenerate predictions, retrain models, or change the production ledger.
Written by cresencio
← Back to blog