Benchmarks measure exam technique. This measures what happened when real people tried to get work done — arena.ai’s data, re-served here and refreshed weekly.
Not ours. The data is arena.ai’s agent leaderboard, built from over a million real agent sessions, and it is linked on every view. We add no scoring of our own — we re-present it in plain English, and say plainly when our copy has gone stale rather than quietly serving old numbers.
Scores are relative to the field average, so a negative figure means below average rather than bad. Every score carries a confidence interval, and models sitting close together are not meaningfully apart. Two models a fraction of a point from each other have not been shown to differ.
A ranking built from a million sessions moves slowly, so polling hourly would spend courtesy on nothing. The data refreshes once a week, a scheduled check re-tests the connection every Monday, and this page states its own age so a broken refresh cannot go unnoticed.
A live ranking of AI coding agents by what real users experienced — task completion, steerability, error recovery and tool hallucination, from over a million sessions. Free, updated weekly. Runs live in your browser; nothing is stored. Part of checkmysite.pro — a free suite of website checks with no account and no paid tier.