LM Arena
LLM evaluation platform via blind pairwise battles: submit a prompt, two anonymous models respond, vote the better one. Chess-style Elo (Bradley-Terry). Includes text, vision, image, video arenas. The platform now goes by Arena and lives at arena.ai. It publishes a ranking per task type (text, web development, vision, documents, search, image, video, agents), and that split is what matters: the overall leader is not necessarily the leader on your task, and the podium shifts from month to month.
Strengths
- Transparent blind pairwise methodology (Elo, not vibes)
- Tens of millions of cumulative votes, largest public blind-preference dataset
- Extended surface: text, vision, WebDev, Copilot, Search, Image, Video arenas
- Lets you battle your own persona prompts before prod
Limitations
- Elo score masks use-case gaps (a model may be number 1 in text but number 5 in coding)
- No prompt-by-prompt reproducibility guarantee
- Voter bias toward long well-formatted outputs
Best for
- PMs and Tech Leads picking an LLM for product integration
- Curious people wanting to compare 2 models on their own prompt without signup