

- Published on 30 Oct 2024
- Last updated on 2 Aug 2025
- Reading Time: 18 minutes
Chatbot Arena Categories
By grouping tasks into categories, we can assess models’ strengths and weaknesses in a more granular way.
Definitions, Methods, and Insights
Introduction
While the overall Chatbot Arena leaderboard provides a simple score for each model, people use LLMs for diverse purposes. This raises the question: which model is best for a specific use-case? To offer deeper insights into this question, we have been adding various categories to our leaderboard. Over time, we’ve introduced a wide array of categories, including hard prompts, instruction-following, math prompts, coding prompts, refusal handling, longer queries, multi-turn conversations, various languages, and style control. Today, we’re excited to announce the release of a new category: creative writing!
In this blog post, we’ll explain:
- Key insights from our categorical analysis, including how the topic distribution varies with time
- The deployment process for new categories
- The definition of each category currently in Chatbot Arena
- How the community can contribute to improve and add categories
Why Categorize?
Language models don’t shine equally in different areas. Some tasks may require the precise execution of instructions, while others push the model’s ability to reason through complex math problems or handle long, multi-turn conversations. By grouping tasks into categories, we can assess models’ strengths and weaknesses in a more granular way. We acknowledge that a high-ranking on the overall leaderboard doesn’t imply the model will excel across the board in every situation. Categories help elucidate these nuances, allowing our users to identify which models are best suited for their specific needs.



















