Published on Wed Sep 30 2026 16:20:00 GMT+0000 (Coordinated Universal Time) by cresencio
NFL Week 4: Two Forecasts, Three Changed Picks
Week 4 brings all five models into the blend. A rebuilt research forecast changes three winners and sharply reduces the confidence attached to most of the board. Both versions stay on the record.
The Week 3 preview tracked three contributors and six favorites above 90%. Atlanta’s 35–14 win at Green Bay then supplied another example of sharpening increasing the penalty for a wrong pick. Week 4 adds XGBoost and Random Forest, now that three completed weeks supply their rolling inputs.
But availability exposed a deeper issue: several models were receiving current statistics without learning from current-season outcomes. The rebuild addresses that distinction. The useful preview is one shared board showing where the two versions agree, where they switch sides, and where the same pick carries a very different probability.
What changed in the calculation
The original workflow updated Elo and refitted Logistic, but XGBoost, Random Forest and Bayesian retained their 2023–2025 fits. Logistic also used pooled season aggregates for historical training rows, rather than features restricted to earlier games. Ensemble weights remained fixed from August.
The rebuilt package brings all five families through 2026 Week 3, including all 48 completed games. Historical features now use only earlier-game data. Models are updated or refitted on the expanded history, and weights are recomputed from chronological out-of-sample forecasts.
| Model | Original Week 4 weight | Rebuilt weight |
|---|---|---|
| Elo | 14.57% | 21.51% |
| Logistic regression | 28.04% | 19.53% |
| XGBoost | 4.88% | 20.37% |
| Bayesian | 26.48% | 19.00% |
| Random Forest | 26.02% | 19.58% |
XGBoost’s influence rises substantially; Logistic and Bayesian lose weight. These are fitted weights, not a decision to give every model an equal vote. The historical tests were reconstructed retrospectively, so they are not recovered forecasts actually issued at the time.
Calibration remains unresolved. The original uses T = 0.4 to sharpen its raw blend. A replacement T = 0.60, selected on 2024 data, failed independent 2025 holdout checks. The rebuild therefore uses T = 1: the raw, explicitly uncalibrated blend. Less extreme probabilities are not proof of better calibration.
Training, features, weights and the probability adjustment all changed. Their effects should not be attributed to retraining alone.
The complete Week 4 board
Each percentage is the named team’s estimated win probability. The original’s raw and final columns are two stages of one forecast. The rebuilt version needs one column because its raw and final values are identical. Bold rebuilt picks change the selected winner.
| Matchup | Original raw | Original final | Rebuilt research |
|---|---|---|---|
| PIT @ CLE | CLE 58.6% | CLE 70.4% | CLE 51.9% |
| IND @ WAS | WAS 62.0% | WAS 77.3% | WAS 54.4% |
| ARI @ NYG | NYG 52.8% | NYG 56.9% | ARI 53.4% |
| DAL @ HOU | DAL 50.3% | DAL 50.7% | HOU 52.4% |
| GB @ TB | TB 52.0% | TB 54.9% | GB 52.9% |
| JAX @ CIN | JAX 57.8% | JAX 68.6% | JAX 52.4% |
| LAR @ PHI | LAR 62.3% | LAR 77.8% | LAR 52.9% |
| NE @ BUF | BUF 78.4% | BUF 96.2% | BUF 73.2% |
| NYJ @ CHI | CHI 69.6% | CHI 88.9% | CHI 64.9% |
| TEN @ BAL | BAL 75.7% | BAL 94.5% | BAL 75.8% |
| MIA @ MIN | MIN 71.4% | MIN 90.8% | MIN 63.7% |
| DEN @ SF | SF 80.3% | SF 97.1% | SF 72.4% |
| KC @ LV | KC 53.0% | KC 57.5% | KC 54.1% |
| LAC @ SEA | SEA 80.0% | SEA 97.0% | SEA 72.2% |
| DET @ CAR | DET 53.1% | DET 57.6% | DET 53.2% |
| ATL @ NO | NO 57.3% | NO 67.6% | NO 53.0% |
LAR means the Rams; LAC means the Chargers. Washington is the designated home team against Indianapolis in London on October 4, not at its usual home stadium.
Thirteen winners stay the same. Three change. Both versions have six unanimous five-model selections. The ensemble combines weighted probabilities, so agreement counts alone do not explain the picks.
The three switches are not the same kind of change
Arizona at the Giants: the same votes, different strengths
The original favors the Giants at 56.9%; the rebuild takes Arizona at 53.4%. None of the five component models changes its selected winner.
Elo and Logistic still pick New York; XGBoost, Bayesian and Random Forest still pick Arizona. But Logistic’s Giants estimate falls from 75.2% to 52.9%, while Random Forest’s Arizona estimate rises from 66.1% to 72.0%. Alongside the new weights, those shifts move the blend from the two-model minority to the three-model majority.
Dallas at Houston: a near-even pick crosses the line
Dallas goes from 50.7% original to a 52.4% Houston research pick. Again, every component keeps the same side: Elo, XGBoost and Random Forest favor Houston; Logistic and Bayesian favor Dallas.
Elo still gives Houston 67.6%, but now carries 21.51% of the blend instead of 14.57%. Other probabilities and weights move too, so Elo’s larger role is part of the explanation, not an isolated cause. Neither version makes this a strong selection.
Green Bay at Tampa Bay: two models actually change sides
Here, the components do move. Bayesian and Random Forest switch from Tampa Bay to Green Bay, taking the Packers from two votes to four. Random Forest now gives Green Bay 60.4%; XGBoost’s lone Tampa Bay preference is only 50.4%.
The ensemble changes from 54.9% Tampa Bay to 52.9% Green Bay. That is a modest Packers lean after their Atlanta loss, not evidence that a rebound is assured.
Removing temperature sharpening alone cannot reverse a winner. These switches reflect changes to the underlying blend.
The biggest differences are often in unchanged picks
Minnesota stays the pick against Miami, but falls from 90.8% to 63.7%, the largest change in final win probability on the board: 27.1 percentage points. Four models favor Minnesota in both versions; Random Forest still takes Miami. The original raw blend was already much lower, at 71.4%.
The original has five favorites above 90%; the rebuild has none. Baltimore leads the research board at 75.8%, followed by Buffalo at 73.2%, San Francisco at 72.4% and Seattle at 72.2%. Baltimore’s original raw estimate was 75.7%, showing that the rebuild does not simply reduce every underlying estimate.
Washington offers the clearest distinction between agreement and confidence. All five rebuilt models favor it against Indianapolis, but their estimates occupy a narrow 53.3%–55.8% range. The ensemble lands at 54.4%, down from 77.3% original. Unanimous picks can still describe an uncertain game.
Thursday and Monday frame the test
Pittsburgh at Cleveland opens Thursday, October 1, at 7:15 p.m. Central. Cleveland keeps four model votes, yet falls from 70.4% to 51.9%. Elo still favors Pittsburgh at 60.7%. The opener tests a much smaller Cleveland advantage than the original final number suggested.
Atlanta visits New Orleans on Monday, October 5, at 7:15 p.m. Central. Both versions favor the Saints, 67.6% original and 53.0% rebuilt, with four model votes each. Atlanta’s win at Green Bay is included in the expanded history, but this comparison cannot isolate that game’s contribution to the change.
What gets graded next
The original was published September 29. The research forecast was generated that evening at 8:28 p.m. Central and separately verified and frozen, unchanged, at 10:20 p.m., before the slate began. It is a later pregame forecast and has not replaced the original.
I’ll grade all sixteen games using correct winners, binary Brier score and natural-log loss from full-precision probabilities. The original raw blend, original final output and rebuilt forecast will be reported separately, alongside the component models. The three switches and the unchanged high-confidence picks remain on the watchlist whichever way they finish.
These are outright-winner forecasts. A win probability supplies neither a predicted margin nor evidence of a betting-price advantage. One week’s results will add evidence for the rebuild; they will not settle its long-term quality.
Source note: Both saved Week 4 versions, component probabilities, weights, timestamps and source hashes are retained in the companion editorial evidence. The research column uses the September 29 pregame moneyline freeze; a separate market-interpretation correction left those values unchanged. Writing this preview did not regenerate forecasts, refit models or alter grading records.
Written by cresencio
← Back to blog