Published on Sat Sep 12 2026 17:08:41 GMT+0000 (Coordinated Universal Time) by cresencio
The 49ers Won 27–7. My 93.6% Rams Pick Was Wrong
I picked the Rams. San Francisco won 27–7. Every active model picked the wrong winner.
The uncomfortable part is how little room the final forecast left for the result that happened. In the pregame preview, the raw ensemble gave Los Angeles 74.5%. Temperature scaling pushed it to 93.6%, leaving San Francisco just 6.4%.
I put the grading rules on the page before kickoff. Now the hypothetical 49ers-win row is the actual result. There is no reason to soften it because the number looks bad.
How the game got away from Los Angeles
The Rams’ official recap records a 27–7 loss at the Melbourne Cricket Ground. Los Angeles led 7–3 in the second quarter, but San Francisco went into halftime ahead 10–7 and extended the lead to 17–7 on its opening possession of the second half.
The Rams then reached the 49ers’ one-yard line with a chance to cut the deficit. Matthew Stafford was stopped on fourth-and-goal. San Francisco answered with an 11-play, 99-yard touchdown drive to make it 24–7 early in the fourth quarter.
That sequence helps explain the game. It does not explain the forecast by itself. I did not predict the goal-line stop, the long response, or the twenty-point margin. The 49ers’ official account confirms the final result and their 1–0 start.
The numbers stay exactly where they were
These are the frozen post-audit probabilities used in the preview. Scores below were calculated from their full-precision values, not the displayed percentages.
| Forecast | Pregame Rams win probability | Brier score | Log loss |
|---|---|---|---|
| Elo | 62.72% | 0.3933 | 0.9866 |
| Logistic regression | 87.78% | 0.7706 | 2.1024 |
| Bayesian | 66.89% | 0.4474 | 1.1052 |
| Raw weighted ensemble | 74.49% | 0.5548 | 1.3660 |
| Final ensemble, T = 0.4 | 93.58% | 0.8756 | 2.7451 |
Lower is better. With p defined as the Rams’ win probability and the Rams losing, binary Brier score is p². Natural-log loss is -ln(1 - p).
Elo took the smallest penalty because it assigned San Francisco the most probability: 37.3%. That is a better score on this observation, not a correct pick. Elo still favored the Rams. So did Bayesian and Logistic. The raw and final ensembles are two stages of the same blend, not additional independent models.
The final ensemble performed worst on both measures. Its 2.7451 log loss is the direct cost of giving the eventual winner only about 6.4%.
The confidence adjustment made the miss more expensive
The blend was already wrong before temperature scaling. Removing that step would not have produced a 49ers pick: the raw forecast still favored Los Angeles at 74.49%.
What the adjustment changed was the size of the penalty. It added 19.09 percentage points to the Rams’ probability. With a San Francisco win, Brier score rose from 0.5548 to 0.8756, and log loss rose from 1.3660 to 2.7451.
Those are exactly the values in the pregame scenario table. The downside was not discovered after the loss. This game realized it.
Saying a 6.4% event can happen is true, but it is not an evaluation. The forecast assigned a small probability to the winner and receives a large penalty. We need to retain that penalty in the record, just as we retained the sharper forecast’s benefit when Seattle won.
One correct pick, one wrong pick, different costs
The Seattle follow-up showed temperature scaling improving the scores on a correct winner. This game shows the other side of the same tradeoff.
Across only these two reviewed games—New England at Seattle and San Francisco against Los Angeles—the raw and final ensembles both went 1–1. Their average probability scores are different:
| Forecast | Correct winners | Mean Brier score | Mean log loss |
|---|---|---|---|
| Raw ensemble | 1 of 2 | 0.3623 | 0.9484 |
| Final ensemble | 1 of 2 | 0.4802 | 1.5445 |
The extra penalty in Melbourne outweighed the improvement in Seattle on both measures. A win-loss count cannot show that.
This is a two-game accounting, not a full Week 1 report or a calibration study. It does not establish the best model or the right temperature. It does establish that sharpening has hurt the average scores in this small record so far.
The margin does not change the scoring rule
The twenty-point loss makes the pick look worse as a football story. It does not add a separate penalty to a binary winner forecast. A one-point San Francisco win would receive exactly the same Brier score and log loss.
That is the same boundary I drew after Seattle’s narrow win. I cannot say a close win failed to validate confidence, then use a blowout loss as though the model had issued a margin prediction. The winner forecast missed; the margin is a different outcome that this saved file did not predict.
Melbourne also remains a question for investigation, not a ready-made excuse. Before kickoff, I noted that the comparison CSV did not establish a venue-specific travel adjustment. The final score cannot tell us how much travel, the designated home-team treatment, or any individual feature contributed to the error. That requires a separate audit of inputs and model behavior.
What changes after this result
The record changes. The original probabilities do not.
I am preserving the pregame article and adding this loss beside the Seattle win. There is no retrospective forecast refresh, no claim that a less-confident Rams pick secretly anticipated San Francisco, and no new spread or total analysis attached to a winner-only model.
The calibration evidence was limited before this game: the retained temperature setting came from a different model combination and included fitted-sample inputs. That remains a reason to evaluate the current lineup on unseen games. One bad result is not enough to choose a replacement setting, but it belongs in that evaluation without qualification or omission.
The Rams were the wrong pick. The extra confidence made the miss more costly. Both facts stay in the record.
The full Week 1 preview preserves the rest of the saved board. This recap grades two specific reviewed games; it does not stand in for grading the entire slate.
Source note: Final score and game sequence were checked against the official Rams and 49ers recaps linked above. Forecast inputs come from the retained pregame evidence for game 2026_01_SF_LA; the two-game averages also use the saved Seattle grading evidence. Calculations use single-event binary Brier score and natural-log loss. No model was retrained and no production prediction ledger was updated for this editorial review.
Written by cresencio
← Back to blog