Published on Fri Sep 18 2026 18:48:02 GMT+0000 (Coordinated Universal Time) by cresencio
Buffalo Won 41–31. This Time, the Extra Confidence Helped
I picked Buffalo. The Bills won 41–31. The extra confidence earned its credit on this one.
The pregame preview gave Buffalo a 63.1% chance in the reconstructed raw blend and 79.4% after temperature scaling. It also put the scoring consequences of either winner on the page before kickoff.
Thursday delivered the Buffalo-win scenario. Every available model picked the correct winner, and the final ensemble earned the best probability scores of the five forecasts shown. That belongs in the record just as plainly as the expensive Rams miss.
Buffalo took control early
The Bills opened their new Highmark Stadium with a 41–31 win and moved to 2–0. Josh Allen accounted for five touchdowns, three passing and two rushing. Buffalo led 27–10 at halftime. The Bills’ official recap records the result; the official game page confirms the scoring by quarter.
Detroit fell behind 21–0 early in the second quarter and spent the rest of the night chasing. One important sequence came with Buffalo ahead 14–0: a defensive holding penalty erased a Lions interception, and the Bills scored three plays later. Detroit kept scoring but never completed the comeback, as described in the Lions’ recap.
Those details explain the game. The saved forecast made a narrower claim: Buffalo was more likely to win. It did not predict Allen’s touchdowns, the penalty, or the ten-point margin.
Grade the numbers that were on the page
These are the original probabilities, preserved from the pregame preview. The scores use full precision rather than the rounded percentages in the table.
| Forecast | Pregame Bills win probability | Brier score | Log loss |
|---|---|---|---|
| Elo | 58.84% | 0.1694 | 0.5304 |
| Logistic regression | 65.39% | 0.1198 | 0.4249 |
| Bayesian | 56.80% | 0.1866 | 0.5656 |
| Raw weighted ensemble, reconstructed | 63.15% | 0.1358 | 0.4597 |
| Final ensemble, T = 0.4 | 79.35% | 0.0426 | 0.2313 |
Lower is better. With Buffalo winning, binary Brier score is (1 − p)² and natural-log loss is −ln(p), where p is Buffalo’s pregame win probability.
The final forecast assigned the most probability to the outcome that happened, so it receives the smallest penalty. Logistic scored best among the three individual models. Bayesian also picked Buffalo correctly, even though it did not contribute to this week’s ensemble.
That last distinction matters. Bayesian’s probabilities varied too little across the Week 2 slate to pass the pipeline’s existing inclusion rule. Its forecast remained available to inspect, but the effective blend was 34.2% Elo and 65.8% Logistic. XGBoost and Random Forest were unavailable. The raw and final rows are two stages of that blend, not additional independent models.
A correct Bayesian pick does not prove the exclusion rule was wrong. A correct ensemble pick does not prove it was right. This game gives us one scored observation for each available forecast.
The confidence adjustment helped on this result
Temperature scaling raised Buffalo’s probability by 16.21 percentage points, from 63.15% to 79.35%, without changing the selected winner.
Because Buffalo won, that reduced Brier score by 0.0932 and log loss by 0.2284. The resulting 0.0426 Brier score and 0.2313 log loss match the Buffalo-win scenario published before kickoff.
The Rams recap showed the larger penalty when sharpening backed the loser. Buffalo shows the reward when it backs the winner. Both outcomes count under the same rule.
The ten-point margin adds no extra credit. A one-point Buffalo win would produce exactly these scores. This was a winner-probability forecast; the saved row had no predicted score, margin, or total.
The full record still favors the raw blend
The Week 1 recap covered all sixteen opening games. Adding this Thursday result takes both ensemble stages to 10 correct winners in 17 games.
| Forecast | Correct winners | Mean Brier score | Mean log loss |
|---|---|---|---|
| Raw ensemble | 10 of 17 | 0.2385 | 0.6692 |
| Final ensemble, T = 0.4 | 10 of 17 | 0.2689 | 0.7765 |
This includes every game from the frozen Week 1 board plus Detroit at Buffalo, not just the matchups that received individual articles. It is a running editorial record through Thursday, not a completed Week 2 grade.
Buffalo narrows the gap. It does not erase it. Across these seventeen forecasts, the raw blend still has the lower average penalty on both measures, despite making exactly the same winner selections.
There is also a change in the system being counted: Week 1’s blend included Bayesian; this Week 2 blend did not. The table tracks the forecasts actually issued across those two model combinations. It is not a controlled comparison of an unchanged model.
What this win earns
A 79.4% forecast cannot be certified by one win, any more than it can be disproved by one loss. To assess that confidence, I need more forecasts made before their outcomes are known and a record that keeps the misses.
The retained temperature setting was fitted on a different model combination using evidence that included fitted-sample predictions. Thursday’s result does not resolve that limitation. It does supply a favorable observation that deserves to stay alongside the unfavorable ones.
Buffalo was the right pick. Sharpening improved this game’s scores. The accumulated record still favors the raw probabilities.
The remaining fifteen games in the Week 2 preview are the next entries to grade. The original Buffalo–Detroit forecast stays as written.
Source note: The final score and game sequence were checked against the official Bills and Lions sources linked above. Probabilities and grading scenarios come from the editorial evidence saved September 17 before kickoff for game 2026_02_DET_BUF. The raw blend is the same pregame reconstruction, not a new forecast. Running averages use the full-precision sixteen-game evidence behind the Week 1 recap plus this result. No forecast was regenerated, model retrained, or production evaluation or paper-betting ledger changed for this article.
Written by cresencio
← Back to blog