Elo scores from blind pairwise votes: people compare two outputs without knowing which model made each. Gaps under about 20 points are noise — read the margins, not the ranks.
The Artificial Analysis image, music, video and speech arenas, and Arena’s code and vision boards. All use the same method: two outputs from the same prompt, shown without labels, and a human picks one. Elo updates from the votes. We link every board and add no scoring of our own.
A 1340 in the image arena and a 1548 in the code arena are not the same measurement. Each board has its own population of models and its own voters, so an Elo is only meaningful against the other numbers on that same board. The bars here are scaled inside each category for exactly that reason.
There is no credible leaderboard for finance, law or medicine. Domain expertise is not something the public arenas measure separately, and a proxy would be worse than an honest gap. Any tool showing you a "best AI for finance" ranking has invented the number.
Which AI is best at making images, music, video, voices and code — from the public blind-vote arenas, with the Elo gaps that tell you when two models are actually tied. Free, no signup. Runs live in your browser; nothing is stored. Part of checkmysite.pro — a free suite of website checks with no account and no paid tier.