Published on Tue Oct 06 2026 19:06:53 GMT+0000 (Coordinated Universal Time) by cresencio
NFL Week 4: Eleven Winners, but Confidence Still Matters
My original Week 4 ensemble picked 11 of 16 winners. The research version picked 10. Both finished the week watching Atlanta beat the Saints they had selected, 45-24.
That Monday result belongs beside every correct pick. New Orleans had a 67.6% win probability in the original forecast and 53.0% in the revised one. Neither called the winner. Both did expect the game to go over the saved 48.5-point total, which it cleared with 69 points. Those are separate predictions, and they deserve separate grades. Official Saints recap
Week 4 leaves a mixed comparison. The original forecast won on the number of correct winners. The research version improved the spread results and took a smaller log-loss penalty on its winner probabilities, but did worse on totals. And the original blend before its confidence adjustment scored better than either final version on both probability measures.
The Week 2 recap asked whether sharpening the probabilities was helping. A better win-loss record makes that question easier to overlook. This week gives another reason to keep asking it.
All sixteen games
These are the original issued forecast and a preserved research comparison. The research winner probabilities came from a reconstruction and were then frozen unchanged before kickoff. That separate record does not turn the broader research comparison into a second published forecast record. Its spread and total revisions are graded against their own saved lines.
None of the saved values has been changed to fit the results. Each percentage is the named team’s probability of winning; each final score links to an official team recap.
| Matchup | Original pick | Research pick | Final score |
|---|---|---|---|
| PIT @ CLE | CLE 70.4% | CLE 51.9% | CLE 27-24 |
| ARI @ NYG | NYG 56.9% | ARI 53.4% | NYG 36-24 |
| DAL @ HOU | DAL 50.7% | HOU 52.4% | DAL 34-30 |
| DEN @ SF | SF 97.1% | SF 72.4% | SF 24-14 |
| DET @ CAR | DET 57.6% | DET 53.2% | CAR 32-26 |
| GB @ TB | TB 54.9% | GB 52.9% | GB 17-14 |
| IND vs. WAS | WAS 77.3% | WAS 54.4% | IND 30-13 |
| JAX @ CIN | JAX 68.6% | JAX 52.4% | JAX 22-17 |
| KC @ LV | KC 57.5% | KC 54.1% | KC 30-27 |
| LAC @ SEA | SEA 97.0% | SEA 72.2% | SEA 30-23 |
| LA @ PHI | LA 77.8% | LA 52.9% | LA 24-20 |
| MIA @ MIN | MIN 90.8% | MIN 63.7% | MIN 15-10 |
| NE @ BUF | BUF 96.2% | BUF 73.2% | NE 29-26 |
| NYJ @ CHI | CHI 88.9% | CHI 64.9% | CHI 23-12 |
| TEN @ BAL | BAL 94.5% | BAL 75.8% | BAL 24-18 |
| ATL @ NO | NO 67.6% | NO 53.0% | ATL 45-24 |
LA denotes the Rams; LAC denotes the Chargers. Washington was the designated home team against Indianapolis in London.
Only three winner picks changed. Moving from the Giants to Arizona and from Dallas to Houston cost the research version two correct picks. Moving from Tampa Bay to Green Bay recovered one. That accounts for the entire difference between 11-5, or 68.75%, and 10-6, or 62.50%.
The thirteen shared picks went 9-4. Agreement still left both versions wrong on Detroit, Washington, Buffalo and New Orleans.
Buffalo shows why confidence needs its own score
New England’s 29-26 win in Buffalo counted as one miss for each version. The original forecast, however, had left the Patriots only 3.81% of the win probability. The research version had left them 26.78%.
That difference matters even though neither picked New England. The original forecast’s log-loss penalty for this game was 3.2683, against 1.3175 for the research forecast. A probability score asks how much room the forecast left for what happened, not just which side was above 50%.
Here is the complete winner-probability comparison, including the original raw blend so that the confidence adjustment gets a fair test.
| Forecast | Correct winners | Mean Brier score | Mean log loss |
|---|---|---|---|
| Original raw blend | 11 of 16 | 0.2109 | 0.6113 |
| Original final, T = 0.4 | 11 of 16 | 0.2177 | 0.6650 |
| Research final, raw T = 1 | 10 of 16 | 0.2210 | 0.6327 |
| Constant 50% reference | — | 0.2500 | 0.6931 |
Lower is better. Binary Brier score is the average squared distance between a probability and the outcome, written (p - y)^2. Log loss is the negative natural logarithm of the probability assigned to the winner, so a confident miss can be especially costly. These means use all sixteen games and the full-precision saved probabilities.
The research version had lower log loss than the original final forecast, but slightly higher Brier score and one fewer winner. The scoring rules can disagree about which set of probabilities did better because they penalize errors differently.
The cleaner comparison is within the original forecast. Sharpening it at T = 0.4 changed no picks and raised mean Brier score by 0.0068 and mean log loss by 0.0537. It helped the confident wins, including San Francisco and Seattle, but those gains did not outweigh the penalties elsewhere. The original raw blend also beat the research version on both probability scores this week.
That is evidence about these sixteen forecasts. It does not establish that one model will be better next week.
Better against the spread, worse on totals
Picking the winner and covering the spread are different jobs. Kansas City won by three but did not cover the saved -4.5. Minnesota won by five but did not cover -11.5. The research version picked the Chiefs and Vikings to win while preferring Las Vegas +4.5 and Miami +11.5 against the spread. Both combinations were correct.
The original spread forecast took the favorites in those two games and lost both. Across the full week, the research version improved from five correct spread picks to seven. It still had more losses than wins.
| Market and version | Win-loss-push | Accuracy, excluding pushes | Mean Brier score | Mean log loss |
|---|---|---|---|---|
| Original spread | 5-10-1 | 33.33% | 0.2866 | 0.7670 |
| Research spread | 7-8-1 | 46.67% | 0.2548 | 0.7028 |
| Original total | 9-7-0 | 56.25% | 0.2491 | 0.6922 |
| Research total | 8-8-0 | 50.00% | 0.2538 | 0.7008 |
Seattle’s seven-point win over the Chargers pushed the saved seven-point spread. That game is excluded from the spread accuracy and probability-score denominators, leaving fifteen decisions. All sixteen totals settled without a push. Every grade uses the line saved with its forecast, not a later closing line.
The research spread probabilities improved on both scoring measures, but remained slightly worse than a constant 50% reference. Totals moved in the other direction: fewer correct picks and higher losses on both measures.
Monday illustrates why a correct total should not be oversold. The research probability of over 48.5 was only 50.05%. Atlanta and New Orleans producing 69 points made that a winning over pick; it did not turn the forecast into a confident call after the fact.
These are grades for the full set of directional forecasts. The original list of twelve moneyline recommendations finished 8-4, a different group from the sixteen winner picks above. There was no corresponding final research recommendation portfolio. Neither the win counts nor these probability scores establish a betting profit or an advantage over the market.
What this week earns, and what it does not
The research probabilities use T = 1, which leaves the raw blend unchanged. A proposed T = 0.60 adjustment failed the earlier held-out validation checks. Using the raw fallback is reasonable; calling it successfully calibrated would not be. Its calibration status remains unvalidated.
The two raw blends are also different forecasts. Comparing the original final version with the research version changes more than temperature, so their difference cannot all be credited to removing sharpening. The original raw-versus-final rows isolate that adjustment more directly.
Week 4 gives me an 11-5 original record to keep, a 10-6 research comparison to examine, and another instance where extra confidence worsened the original blend’s probability scores. The spread improvement deserves continued attention; the weaker totals deserve the same scrutiny. Neither gets to stand in for a longer test of the whole model.
Source note: All sixteen finals were checked against the official team recaps linked above. Winner, spread and total scores use unchanged saved forecasts, full-precision probabilities and each forecast’s saved line. Displayed figures are rounded only after calculation. This is a complete Week 4 comparison, not a cumulative season record or a claim about future returns.
Written by cresencio
← Back to blog