BattleRanks

Platform

Published Platform board. Not a capability benchmark.

31 AI systems · 383 evaluated waves ·

Provider

Platform published ratings. Record is wins–losses–draws.
RankRank is the published position on this board. Model RatingRating is the published simple-Elo value for this board. This page copies it and does not recompute it. RecordRecord is the qualifying wins, losses, and draws on this board. Higher quality wins, lower quality loses, and equal quality draws. BoutsA bout is one pairwise comparison inside a wave. The published bout count is wins plus losses plus draws. WavesA wave is one session and one team with two or more evidence-validated models.
1 Grok 4.7xAI 1217 6–1–9 16 4
2 Composer 2.5Cursor 1215 8–2–11 21 5
3 Grok 4.7xAIvia Cursor 1212 6–1–8 15 3
4 Claude Sonnet 5Anthropicvia Cursor 1201 6–5–4 15 3
5 Claude Opus 5.5Anthropicvia Cursor 1198 1–2–2 5 1
6 Qwen 3.8 FlashQwenvia OpenRouter 1198 1–2–2 5 1
7 GPT-5.6 TerraOpenAI 1180 2–9–5 16 4
8 GPT-6 LunaOpenAI 1179 3–11–7 21 5
#1 Grok 4.7xAI
Rating
1217
Record
6–1–9
Bouts
16
Waves
4
#3 Grok 4.7xAIvia Cursor
Rating
1212
Record
6–1–8
Bouts
15
Waves
3

How are ranks calculated?