← Back to blog

Published on Tue Sep 22 2026 16:55:00 GMT+0000 (Coordinated Universal Time) by cresencio

NFL Week 3: Bayesian Is Back. The Confidence Test Continues

Week 3 restores a third model to the ensemble. It does not erase the first two weeks of results.

The Week 2 recap records an 8–8 ensemble, with eleven favorites above 90% going 5–6. Through two weeks, sharpening the raw blend at T = 0.4 has worsened both Brier score and log loss. That evidence belongs beside the next board.

This September 22 forecast has six favorites above 90%, down from eleven in the original Week 2 preview. Bayesian contributes again after the feature-timing correction explained in the Giants–Rams preview. There are more voices in the blend, but the final probabilities still deserve scrutiny.

What changed in the calculation

Bayesian now uses completed games strictly before the forecast week. The checked Week 3 state includes two completed games for every team. Its home-win probabilities have a standard deviation of about 0.178, comfortably above the pipeline’s 0.05 exclusion threshold.

The original Week 2 Bayesian output had failed that check because a feature-timing bug left the forecasts unusually flat. Correcting the inputs restores Bayesian’s contribution without changing the base weights or temperature.

ModelEffective Week 3 weight
Elo21.09%
Logistic regression40.58%
Bayesian38.33%
XGBoostUnavailable
Random ForestUnavailable

Passing the variation check establishes that Bayesian can participate under the existing rule. It does not establish that the new blend is calibrated. Week 2 and Week 3 also have different opponents and new results available, so the change in confidence cannot be attributed entirely to Bayesian returning.

The complete Week 3 board

Each percentage below is the named team’s win probability. The raw column is the weighted blend before temperature scaling; the final column is the saved output after it. All three component columns contribute this week.

MatchupEloLogisticBayesianRaw blendFinal ensemble
ATL @ GBGB 70.5%GB 56.6%GB 66.3%GB 63.3%GB 79.5%
LAC @ BUFBUF 73.1%BUF 95.8%BUF 82.3%BUF 85.9%BUF 98.9%
CAR @ CLECLE 58.4%CAR 82.4%CAR 67.9%CAR 68.2%CAR 87.1%
NYJ @ DETDET 83.7%NYJ 52.4%NYJ 55.4%DET 54.0%DET 60.0%
HOU @ INDHOU 57.8%IND 76.6%HOU 51.0%IND 58.8%IND 70.8%
NE @ JAXJAX 58.8%JAX 94.3%JAX 67.9%JAX 76.7%JAX 95.1%
KC @ MIAKC 55.6%KC 98.1%KC 79.9%KC 82.2%KC 97.9%
TEN @ NYGNYG 64.5%TEN 75.1%TEN 59.6%TEN 60.8%TEN 75.0%
CIN @ PITPIT 64.0%CIN 73.9%CIN 60.5%CIN 60.8%CIN 74.9%
SEA @ WASSEA 68.0%SEA 97.3%SEA 67.1%SEA 79.5%SEA 96.7%
ARI @ SFSF 79.0%SF 97.4%SF 82.6%SF 87.9%SF 99.3%
MIN @ TBMIN 59.1%MIN 97.5%MIN 57.2%MIN 74.0%MIN 93.2%
BAL @ DALBAL 54.0%DAL 69.6%BAL 59.0%DAL 53.7%DAL 59.1%
LV @ NONO 63.1%LV 78.7%LV 53.2%LV 60.1%LV 73.6%
LA @ DENDEN 60.6%LA 81.4%LA 58.4%LA 63.7%LA 80.4%
PHI @ CHIPHI 59.1%CHI 62.4%CHI 60.3%CHI 57.0%CHI 67.0%

LA means the Rams; LAC means the Chargers. Dallas is the designated home team for the Baltimore game in Rio de Janeiro. The NFL schedule runs from September 24 through September 28.

Seven games have unanimous picks among the three contributors. Nine have disagreement. In three of those nine, the ensemble chooses the side supported by only one model.

That is possible because the ensemble combines probabilities and weights, rather than taking a majority vote.

Three games where one model carries the pick

Jets at Detroit: Elo supplies the strongest opinion

Logistic and Bayesian lean toward the Jets at 52.4% and 55.4%. Elo gives Detroit 83.7%. That stronger estimate is enough to pull the raw blend to 54.0% Detroit, even though Elo has the smallest weight.

The final Detroit probability rounds to 60.0%. The useful distinction is between two mild Jets preferences and one much stronger Lions preference. Counting the votes would miss it.

Houston at Indianapolis: Logistic pulls toward the Colts

Elo gives Houston 57.8% and Bayesian gives Houston 51.0%. Logistic gives Indianapolis 76.6%. The raw blend favors the Colts at 58.8%, becoming 70.8% after sharpening.

This is worth tracking alongside Detroit because the same minority-pick pattern has a different model driving it. The postgame review needs each model’s probability, not just the observation that two models disagreed with the ensemble.

Baltimore versus Dallas: a small raw lean in Rio

Elo and Bayesian take Baltimore at 54.0% and 59.0%. Logistic takes Dallas at 69.6%. The ensemble reaches 53.7% Dallas raw and 59.1% final.

Dallas’s home designation sets the orientation of the saved row. It should not be read as a claim that this is an ordinary game in Arlington or that the model has separately validated the effect of the international venue.

The loudest picks still need a probability test

San Francisco is the strongest final pick at 99.3% against Arizona, followed by Buffalo at 98.9% against the Chargers. Kansas City, Seattle, Jacksonville, and Minnesota also exceed 90%.

All six have unanimous component picks. Their raw probabilities range from 74.0% to 87.9%, however, and the components are not equally confident. Minnesota is a useful example: Elo and Bayesian put the Vikings near 59% and 57%, while Logistic gives them 97.5%. The final blend reaches 93.2%.

None of the raw favorites reaches 90%. In seven games, the final home-win probability lies outside the range of the three component estimates. Temperature scaling can do that mathematically; it is not additional agreement among the models.

The retained calibration setting was fitted using a different model combination and included fitted-sample predictions. Its name is not evidence that it improves the current three-model blend. The first two weekly reviews have measured the opposite effect on their saved forecasts.

Thursday opens the next test

Green Bay is the pick against Atlanta: 63.3% raw, 79.5% final. All three contributors favor the Packers. This is a less extreme forecast than the six above 90%, but the scoring rule is the same.

After the week, I will grade all sixteen picks, each available component, and both ensemble stages. The seven unanimous games, the nine disagreement games, and the three minority picks stay on the watchlist whether they win or lose.

No score, margin, spread-cover, or totals forecast is available in this saved export. These winner probabilities do not establish a betting edge or tell us how close a game should be.

The input correction is in place. Week 3 now gets its own record of whether the resulting probabilities earn their confidence.

Source note: This is the September 22, 2026 forecast snapshot, before Week 3 begins. The full-precision export and ensemble diagnostic cover all sixteen games; independent reconstruction reproduces every saved final probability within numerical precision. The companion editorial evidence retains the input hashes, component probabilities, raw blend, and final output. Later forecast refreshes should be identified separately. Drafting this preview did not retrain models, refresh odds, or change the prediction or paper-betting records.

Written by cresencio

← Back to blog