What the AI forecasters say, how much of it is their own, and whether any of it goes with being right. No AI judge is involved: every figure is a comparison against the market price or a plain count over the rationale text.
| Forecaster | Prompt | Games | Avg gap to market | Within 3 pts | Brier vs market |
|---|---|---|---|---|---|
| Claudeprovisional | No line in prompt | 25 | 9.1 pts | 16% | +0.019 ± 0.041 |
| Claude | Line in prompt | 52 | 6.5 pts | 69% | −0.006 ± 0.035 |
| Claude | Line + “don’t copy it” | 42 | 11.6 pts | 31% | +0.027 ± 0.054 |
| GLMprovisional | No line in prompt | 13 | 9.3 pts | 15% | −0.031 ± 0.051 |
| GLM | Line + “don’t copy it” | 57 | 4.8 pts | 53% | +0.015 ± 0.020 |
| GPTprovisional | No line in prompt | 25 | 8.6 pts | 24% | +0.009 ± 0.039 |
| GPT | Line in prompt | 54 | 6.3 pts | 48% | +0.013 ± 0.030 |
| GPT | Line + “don’t copy it” | 41 | 10.2 pts | 46% | +0.027 ± 0.051 |
The gap is between the model’s forecast and the Polymarket price captured at the same moment. “Within 3 pts” is a forecast that is, for practical purposes, the market’s number. “Brier vs market” is the model’s Brier minus the market’s on those exact games (positive = worse than simply using the price; ± is a 95% margin). The prompt column is what the model could see: “No line in prompt” — question text and any team/fighter facts only; “Line in prompt” — sportsbook consensus shown, no guidance; “Line + “don’t copy it”” — line shown with an explicit instruction to reason independently. The instruction was added on 14 Sep 2026. Effect on copying (share within 3 pts): Claude 69% → 31%; GPT 48% → 46%. Before 19 Sep the line in the prompt was keyed to the home team without saying which side the question asked about, and for many US-sports games it described the other side.
| AI forecast was… | Games | Model Brier | Market Brier | Difference | Model closer |
|---|---|---|---|---|---|
| Within 3 points of the market | 136 | 0.240 | 0.240 | +0.000 ± 0.003 | 70 of 136 |
| 3 to 8 points away | 80 | 0.226 | 0.231 | −0.006 ± 0.010 | 37 of 80 |
| 8 or more points away | 93 | 0.264 | 0.218 | +0.046 ± 0.045 | 31 of 93 |
Every AI forecast pooled. “Model closer” counts games where the model’s probability was nearer the outcome than the market’s.
| Share of rationales that mention… | GPT | Claude | GLM | DeepSeek |
|---|---|---|---|---|
| Average length | 80 words | 71 words | 93 words | 53 words |
| Market / odds | 69% | 74% | 67% | 0% |
| Record / results | 88% | 99% | 99% | 60% |
| Recent form | 48% | 61% | 59% | 20% |
| Home / venue | 74% | 76% | 73% | 80% |
| Matchup / style | 34% | 21% | 64% | 20% |
| Head-to-head | 3% | 10% | 9% | 0% |
| Says what it doesn’t know | 75% | 45% | 51% | 20% |
| Games | 120 | 119 | 70 | 5 |
Keyword counts over each model’s rationale on resolved games. They record what a rationale mentions, not whether it is right or whether the claim is true — a model can “mention” an injury to say it has no injury information, and the models draw on their own memory of teams and fighters that this site cannot check.
| Rationale mentions… | Brier if it does | Brier if not | Difference | |
|---|---|---|---|---|
| Market / odds | 0.244 (218) | 0.239 (96) | +0.006 ± 0.033 | no detectable difference |
| Record / results | 0.241 (296) | 0.279 (18) | −0.038 ± 0.079 | no detectable difference |
| Recent form | 0.247 (172) | 0.238 (142) | +0.009 ± 0.032 | no detectable difference |
| Home / venue | 0.245 (234) | 0.237 (80) | +0.008 ± 0.038 | no detectable difference |
| Matchup / style | 0.247 (112) | 0.240 (202) | +0.007 ± 0.033 | no detectable difference |
| Head-to-head | 0.246 (22) | 0.243 (292) | +0.003 ± 0.055 | no detectable difference |
| Says what it doesn’t know | 0.245 (180) | 0.239 (134) | +0.006 ± 0.032 | no detectable difference |
Every AI forecast pooled; games in brackets. Difference is Brier when the rationale mentions it minus Brier when it doesn’t, so negative would mean better forecasts. None of 7 clear the noise, and with seven comparisons one or two would by chance alone.
“Steveson is an Olympic gold medalist wrestler with freak athleticism for heavyweight, now 4-0 with four first-round KO/TKO finishes, showing rapidly developing striking to complement elite grappling. Sharaf is 4-2, on a two-fight skid, KO'd in round two by both Asplund and Tafa, indicating a vulnerable chin and a level well below UFC-caliber heavyweights. The main caveat is heavyweight variance—one punch can end any fight—and Steveson's limited pro experience (only four bouts). Still, the skill gap, finishing momentum, and Sharaf's durability concerns point strongly to an early Steveson finish, tempered slightly by division volatility.”
“Sportsbook consensus implies Shelton around a 19.7% chance, and that is strong base-rate evidence. Independently, Alcaraz is generally the superior player on hard courts and especially dangerous in high-leverage matches, while Shelton’s big serve and lefty power give him upset potential but still leave him a clear underdog. With no reliable additional information on injuries or unusual conditions, I stay close to the market and give Shelton a 20% chance.”
“Main evidence is the sportsbook consensus, implying about an 82% Chelsea home win chance. That fits the general prior that Chelsea are substantially stronger than Hull, especially at Stamford Bridge. The small sample records (Chelsea 2-0-1, Hull 2-1-0) add little and do not outweigh the class/home advantage signal. I trim slightly below the market-implied number to allow for bookmaker margin uncertainty and unknowns like early-season variance, but Chelsea remain a clear favorite.”
“Sportsbook consensus gives Chelsea ~82% win probability. Chelsea are playing at home (2-0-1 record) against Hull City (2-1-0 away). Hull City's away record is actually decent at 2-1-0, suggesting they're competitive. Chelsea's home record is solid but not perfect. The sportsbook margin typically inflates favorites slightly, so I'll shade slightly below the 82% consensus. Chelsea are likely a Premier League side while Hull may be newly promoted, making Chelsea strong favorites. Settling at ~79%.”
“The sportsbook consensus gives Shelton ~80% probability of winning (home team). Wait - the home team probability is 0.8027, and the question asks if Shelton beats Alcaraz. If Shelton is the home team, that implies ~80% chance for Shelton. However, Alcaraz is typically ranked higher and has historically dominated Shelton. The sportsbook data suggests Shelton is favored (~80%), possibly due to home crowd advantage at US Open. I'll weight the sportsbook data significantly but apply some skepticism given Alcaraz's superior ranking. Estimate: ~75% for Shelton winning, so ~0.75 probability YES.”
“The sportsbook consensus gives Zverev only ~21.8% probability of winning, suggesting Khachanov is heavily favored in this match. This is unusual given Zverev's typically higher ranking, but I should trust the sportsbook data as meaningful signal. Without additional context about injuries or form, and given three sportsbooks agree on this pricing, I'll align closely with the market consensus. Zverev winning at ~22% probability reflects the available evidence.”
Forecasts of 70% or more (or 30% or less) that went the other way, ranked by the Brier cost of the miss.