← Back to blog

Published on Thu Sep 10 2026 17:39:32 GMT+0000 (Coordinated Universal Time) by cresencio

Seattle won 13–10. My ensemble picked Seattle. So did all three active models.

That is a correct call to put in the record. It is not a reason to pretend the system has answered every question I brought into the season.

In the pregame preview, I wrote that Seattle was the pick while the strength of the system’s confidence remained an open question. The result gives me the first observation to grade. It does not close that question.

The pick was right. The score was not a prediction.

Seattle entered the fourth quarter trailing 10–3, then scored ten unanswered points to win 13–10.

The result makes it tempting to say the less-confident models understood the game better. Logistic gave Seattle 57.2%. Bayesian was barely past a coin flip at 53.8%. The final ensemble was much louder at 70.9%.

But none of those numbers predicted a margin.

I predicted Seattle to win. I did not predict Seattle to win comfortably, and I did not predict 13–10.

That is a distinction I want to get better at making. In the preview, I described watching whether the game would look more like Elo’s substantial favorite or stay closer to the Logistic and Bayesian view. I would phrase that differently now: those models gave different probabilities of a Seattle win, not different forecasts of how close the score would be.

A narrow Seattle win does not make 53.8% more correct than 70.9%. A comfortable win would not prove 70.9% either. Both are probability estimates for the same event.

The same boundary applies to the 23 combined points. This Week 1 forecast did not include a total-points prediction. I cannot claim the system anticipated a low-scoring game after seeing one.

Grading what was actually saved

The cleanest way to evaluate this result is to leave the pregame numbers alone and score them against what happened.

These are the corrected, post-audit probabilities retained before the game. I calculated the scores from the full-precision values, not the rounded percentages displayed here.

ForecastSeattle win probabilityBrier scoreLog loss
Elo70.91%0.08460.3437
Logistic regression57.23%0.18290.5580
Bayesian53.81%0.21330.6196
Raw weighted ensemble58.81%0.16970.5309
Final ensemble, T = 0.470.89%0.08470.3440

Lower is better for both scores. Every row picked the correct winner, but the probability scores distinguish how much confidence each forecast assigned to that outcome.

For this result, the binary Brier score is (1 - p)^2, where p is Seattle’s pregame win probability. Log loss is -ln(p), using the natural logarithm. Because Seattle won, forecasts that assigned Seattle more probability earned lower scores.

Elo finished just ahead of the final ensemble, although the difference is tiny. This is a ranking for one observation, not a declaration that Elo is the best model.

Under this winner-only task, a 13–10 Seattle win receives the same score as a hypothetical 35–10 Seattle win. Evaluating the margin would require a different prediction.

The confidence adjustment helped—on this game

The active ensemble combined Elo, Logistic, and Bayesian into a raw Seattle probability of 58.81%. Temperature scaling at T = 0.4 raised the published forecast to 70.89%. That transformation added confidence, not another observation about either team.

With a Seattle win, the adjustment improved the Brier score from 0.1697 to 0.0847 and log loss from 0.5309 to 0.3440.

That is a favorable first result for the adjustment. It is not validation of the temperature setting.

The other possible outcome shows the tradeoff. Had New England won, the raw ensemble’s Brier score would have been 0.3458, while the final ensemble’s would have been 0.5026. The sharper forecast gets more credit when its side wins and takes a larger penalty when it loses.

Whether that tradeoff works over time is the question. One game cannot tell me whether forecasts around 71% win at anything like that rate across comparable predictions.

The calibration evidence was already a limitation before kickoff. The retained temperature setting came from a different model combination and included fitted-sample predictions rather than a clean set of forecasts on unseen games. Seattle winning does not repair that evidence.

What I am taking into the rest of Week 1

The opener gives the system one correct winner. More importantly, it gives me a concrete example of why keeping only wins and losses would throw away useful information.

The raw and final ensembles had the same pick. A simple accuracy count treats them identically. Their probability scores tell a different, more useful story: the confidence adjustment helped on this observation, and I can measure that without declaring it settled.

I want to preserve that distinction across the season. The record needs the original model probabilities, the raw blend, the adjusted forecast, and the outcome—not a rewritten explanation of what the model supposedly knew.

I am also leaving the pregame article intact. This follow-up is where the result and the lesson belong, including the correction to how I described probability versus margin.

There is no reason in this result alone to change the models or the temperature setting. The next step is to grade the remaining saved forecasts with the same rules, not to make the system sound smarter after a win.

Seattle was the right pick. How confidently the system should make that kind of pick is still a question for more evidence.

The full Week 1 preview has the rest of the board. This is the first entry in the follow-up record, not a verdict on the season.

Source note: Pregame probabilities come from the frozen post-audit Week 1 comparison used for the preview. The final score was verified against the official Patriots box score, and quarter totals against the official scoring summary. The full-precision probabilities were checked against the original audit CSV. Brier scores and log losses were calculated from those saved probabilities; no model was retrained or forecast refreshed for this article. This is a winner-probability review, not a spread, total, or betting-return analysis.

Written by cresencio

← Back to blog