LM Arena
LLM evaluation platform via blind pairwise battles: submit a prompt, two anonymous models respond, vote the better one. Chess-style Elo (Bradley-Terry). Includes text, vision, image, video arenas. In May 2026, Claude Opus 4.6 holds the number 1 text spot (Elo 1418), followed by Gemini 3.1 Pro (1406) and GPT-5.2 (1402).
Strengths
- Transparent blind pairwise methodology (Elo, not vibes)
- Tens of millions of cumulative votes, largest public blind-preference dataset
- Extended surface: text, vision, WebDev, Copilot, Search, Image, Video arenas
- Lets you battle your own persona prompts before prod
Limitations
- Elo score masks use-case gaps (a model may be number 1 in text but number 5 in coding)
- No prompt-by-prompt reproducibility guarantee
- Voter bias toward long well-formatted outputs
Best for
- PMs and Tech Leads picking an LLM for product integration
- Curious people wanting to compare 2 models on their own prompt without signup