🤖 AI benchmark: hit-rate of 7 models

Prematch En directo (in-play)

Seven external AI models (Hermes contour) independently analyze the same partidos — predicting the resultado (1X2), total (Más de/Menos de), both equipos to marcador (BTTS) and the exact marcador. Here we honestly compare their pronósticos against the real result after the final whistle and combine everything into a single accuracy rating. An informational and analytical snapshot, not betting advice.

⚠️ Data is still accumulating — counting starts from 09.07.2026, so all models are compared on the same events (early test pronósticos are excluded). The sample is still small and not representative. Right now the snapshot holds 387 partido(es), 1470 settled AI pronósticos (Basketball). The figures below are N, not «a percentage you can trust»: the more partidos are played out, the more reliable the snapshot becomes. We show it transparently from day one, not only once the sample becomes «convenient».

Leaderboard · Basketball

Model N (settled) 1X2 Total puntos Exact marcador Composite accuracy
Claude
325 63.7%(207/325) 50.8%(120/236) 0.0%(0/240) 40.8%(327/801)
Kimi
195 63.6%(124/195) 51.5%(84/163) 0.0%(0/163) 39.9%(208/521)
Google AI
416 64.3%(267/415) 47.4%(184/388) 0.0%(0/391) 37.8%(451/1194)
ChatGPT
76 60.5%(46/76) 42.4%(25/59) 0.0%(0/59) 36.6%(71/194)
GLM 5.2
224 63.8%(143/224) 42.5%(94/221) 0.9%(2/222) 35.8%(239/667)
Qwen
197 64.0%(126/197) 35.8%(63/176) 0.6%(1/176) 34.6%(190/549)
DeepSeek
37 51.4%(19/37) 38.2%(13/34) 0.0%(0/37) 29.6%(32/108)

grey — sample <5, not representative; «—» — the model has not made a settled pronóstico yet.

Composite accuracy — the share of correct pronósticos en all mostrado mercados together: (sum of correct pronósticos) ÷ (sum of all settled pronósticos) en the mercados 1X2 + Total puntos + Exact marcador. Each mercado-pronóstico weighs equally. This is hit-rate, not profitability — for money/ROI by model see /ai-agent. Total: a push (marcador exactly en la línea) is excluded from the denominator. «Exact marcador» — the full final marcador was guessed correctly (H and A matched); pronósticos with no recognized marcador do not count toward the denominator.

Composite model rating · all mercados · Basketball

Bar height = the model's composite accuracy en all applicable mercados on the current sample. Sorted from best to worst.

40.8% (327/801)
Opus 4.8
39.9% (208/521)
Kimi 2.6
37.8% (451/1194)
Gemini 3.5 Flash
36.6% (71/194)
GPT 5.5
35.8% (239/667)
GLM 5.2
34.6% (190/549)
Qwen 3.7 Plus
29.6% (32/108)
DeepSeek V4 Pro

Bars are AI models by version; grey/dimmed — sample <5, not representative. The snapshot is informational, not betting advice.

Accuracy by mercado · Basketball

Where each model is strong: one mini-bar per applicable mercado, with the percentage and (hits/sample).

Claude Composite 40.8%
1X2
63.7% (207/325)
Total puntos
50.8% (120/236)
Exact marcador
0.0% (0/240)
Kimi Composite 39.9%
1X2
63.6% (124/195)
Total puntos
51.5% (84/163)
Exact marcador
0.0% (0/163)
Google AI Composite 37.8%
1X2
64.3% (267/415)
Total puntos
47.4% (184/388)
Exact marcador
0.0% (0/391)
ChatGPT Composite 36.6%
1X2
60.5% (46/76)
Total puntos
42.4% (25/59)
Exact marcador
0.0% (0/59)
GLM 5.2 Composite 35.8%
1X2
63.8% (143/224)
Total puntos
42.5% (94/221)
Exact marcador
0.9% (2/222)
Qwen Composite 34.6%
1X2
64.0% (126/197)
Total puntos
35.8% (63/176)
Exact marcador
0.6% (1/176)
DeepSeek Composite 29.6%
1X2
51.4% (19/37)
Total puntos
38.2% (13/34)
Exact marcador
0.0% (0/37)

The model's favorite by 1X2 = the max of P1/X/P2 in its probabilities; for sports without a empate (tennis, volleyball, etc.) the «X» option doesn't participate. grey — sample <5, not representative. Not betting advice.