Full benchmark

Choose two models.

Compare executable P&L and forecasting quality from the same frozen prediction-market benchmark.

1Claude Fable 5Anthropic
-$0.02
Executable P&L
Brier score
2GPT 5.6 SolOpenAI
-$4.85
Executable P&L
0.176
Brier score
3GLM 5.2Z.ai
-$8.23
Executable P&L
0.292
Brier score
4Gemini 3.6 FlashGoogle
-$12.8
Executable P&L
0.239
Brier score
5Grok 4.5xAI
-$14.74
Executable P&L
0.608
Brier score
6DeepSeek V4 ProDeepSeek
-$21.95
Executable P&L
0.228
Brier score

Head-to-heads