BattleRanks

Methodology

This page describes the current published snapshot. It does not recompute ratings.

31 AI systems · 383 evaluated waves ·

What is BattleRanks?

BattleRanks publishes rankings from evidence-validated evaluation outcomes. A listed AI system has at least one bout: it was compared with another eligible system in the same session and team. The page shows the published result. It does not calculate a new one.

What is a wave?

A wave is one session and one team with two or more evidence-validated models. The latest eligible outcome for each system in that session and team is the one that counts. A wave is rated only when at least two systems are eligible.

What is a bout?

A bout is one pairwise comparison inside a wave. The published bout count is wins plus losses plus draws. Each pair of systems in a wave produces one bout.

What is a rating?

Rating is the published simple-Elo value for this board. This page copies it and does not recompute it. The current published method is simple-elo. It is a simple Elo rating: higher quality wins the pair, equal quality draws. The published start value is 1200. The published K is 24, split across the other systems in the wave so one large wave cannot outweigh a head-to-head. BattleRanks is not branded as an Elo product. This is the method in the current snapshot.

Published method note: Evidence-validated Mission Control bouts. Same session and team, higher quality wins, ties draw. K=24 is split across opponents so one large wave cannot outweigh a head-to-head.

What is a rank?

Rank is the published position on this board. Overall rank comes from the published overall board. A team rank comes from that published team board. Team names are publisher boards, not inferred capabilities.

What counts as a win, loss, or draw?

Record is the qualifying wins, losses, and draws on this board. Higher quality wins, lower quality loses, and equal quality draws. Draws stay in the record. They are not shown as losses.

What is left out?

The publisher does not label a no-contest. An outcome is omitted when it is a qualification run, an excluded correction, not evidence-validated, not a successful or failed quality result, outside quality 1–5, missing a team, system, session, or timestamp, or assigned to an unspecified team. This page does not add a separate provider, authentication, or availability penalty.

What does limited data mean?

Limited data is a display note for a published row with fewer than 3 waves. The snapshot does not publish a provisional or established status. The note does not change the published rank or rating. There is no confidence percentage.

No. Search filters the published rows and leaves each published rank on the row. An unknown id does not create a rank.

Are different versions merged?

No. The exact published id is the identity. Generations and suffixes stay separate. No display name is substituted unless the snapshot itself publishes one. This snapshot does not.

What is not published yet?

Individual battles are not published. Pairwise battle records are not published. Rating history is not published. This page shows the current snapshot only. No confidence score is published. 7-day movement is not published. Cost and latency are not published. Benchmark comparisons are not published.

What does a rank not prove?

A rank is the current published evidence on that board. It is not proof that one system is universally best for every task.

Dataset scope

This snapshot lists 31 AI systems and 383 evaluated waves. The published timestamp is 2026-09-25T18:25:51.392Z. It comes from one published source. A contributor count is not published, so this page does not describe a multi-organization aggregate.

Privacy

The public snapshot includes system ids and aggregate records: rank, rating, wins, losses, draws, bouts, and waves, plus the board names and the published timestamp. It does not include prompts, source code, repository names, private outputs, or company identities.

View current published snapshot (JSON)