LLM Leaderboards

All worthwhile Large Language Model (LLM) Leaderboards on a single page. Every entry is checked for freshness: retired or broken leaderboards are removed, and new noteworthy ones are added as they gain traction.

The Language Model (LLM) Leaderboard, a critical ranking system in the natural language processing (NLP) field, plays a crucial role in evaluating and comparing the performance of diverse language models. It boosts competition, aids in model development, and sets a standard for measuring model effectiveness across tasks such as text generation, translation, sentiment analysis, and question answering. Despite their importance in promoting innovation and highlighting leading models, LLM leaderboards face scrutiny over their actual impact on NLP advancement. Recent studies have pointed out significant issues, including biases in human judgment, data contamination risks, and the inadequacy of evaluation methods, particularly those based on multiple-choice tests. These challenges have spurred demands for task-specific benchmarks that offer a finer, more accurate evaluation of language models for particular use cases. Additionally, there are growing concerns about leaderboard result manipulations, which could deter real progress, diminish trust within the NLP community, and slow down innovation. Addressing these problems by implementing strict measures to ensure the integrity and fairness of LLM leaderboards is essential for fostering genuine advancement and maintaining the competitive spirit of NLP development.

  • Arena (formerly LMArena / LMSYS Chatbot Arena)

    Arena (formerly LMArena, originally LMSYS Chatbot Arena) is a crowdsourced platform that ranks LLMs and other models using Elo-style ratings derived from millions of blind, head-to-head human votes across text, coding, image, and video categories. As of January 2026 the project rebranded from LMArena to Arena and now operates as an independent company; the old lmsys HuggingFace-space leaderboard is a stale legacy mirror and no longer the canonical source.

  • Artificial Analysis

    Artificial Analysis runs independent, standardized evaluations of frontier AI models across intelligence, price, and speed, aggregating results from benchmarks like MMLU-Pro and Humanity's Last Exam into a composite Intelligence Index. It covers hundreds of models from all major providers and is widely cited for its cost/performance comparisons alongside raw capability scores.

  • LiveBench

    LiveBench is a contamination-limited benchmark that releases new questions monthly, sourced from recent papers, datasets, and news, with objective ground-truth answers scored without an LLM judge. It spans 21 tasks across reasoning, coding, agentic coding, math, data analysis, language, and instruction-following, and was developed by researchers from Abacus.AI, NYU, and other institutions.

  • OpenCompass: CompassRank

    CompassRank is dedicated to exploring the most advanced language and visual models, offering a comprehensive, objective, and neutral evaluation reference for the industry and research community.

  • HELM (Stanford CRFM)

    HELM (Holistic Evaluation of Language Models), maintained by Stanford's Center for Research on Foundation Models, evaluates models across dozens of scenarios and metrics on a standardized, reproducible pipeline rather than a single headline score. Beyond its core leaderboards it hosts domain-specific spinoffs such as MedHELM for medicine and VHELM for vision-language models, including an Arabic-language leaderboard launched with partners in January 2026.

  • Scale SEAL Leaderboards

    Scale AI's SEAL (Safety, Evaluations, and Alignment Lab) leaderboards rank frontier models using private, curated datasets assessed by vetted domain experts, aiming to resist gaming and contamination. Leaderboards cover coding, reasoning, instruction-following, and other capabilities and are refreshed regularly as new models are released.

  • Vellum LLM Leaderboard

    Vellum's LLM Leaderboard compares recent frontier models on non-saturated benchmarks (e.g. Humanity's Last Exam, GPQA Diamond, SWE-Bench), deliberately excluding benchmarks like MMLU that no longer separate top models. Companion pages track open-source models and coding-specific performance separately.

  • OpenRouter Rankings

    OpenRouter Rankings track real-world LLM usage across the OpenRouter API marketplace, ranking models by token volume and app adoption rather than benchmark accuracy. It offers a practical, usage-based view of which models developers actually choose in production, broken down by category and time period.

  • SWE-bench

    SWE-bench, created by Princeton NLP, tests whether LLM agents can resolve real GitHub issues drawn from popular Python repositories like Django and scikit-learn, measuring end-to-end software-engineering ability rather than isolated code snippets. The official leaderboard tracks agents and models on the original benchmark plus its verified subset; related spin-offs (SWE-bench Pro, SWE-bench-Live, SWE-rebench) are run independently by other teams.

  • Aider LLM Leaderboards

    Aider's leaderboards measure how well LLMs perform real code edits inside Aider's agentic edit loop, with the flagship Polyglot benchmark testing 225 difficult Exercism exercises across six programming languages. Models are scored on pass rate after two attempts and on how reliably they follow Aider's structured diff format, with cost per run tracked alongside accuracy.

  • LiveCodeBench

    LiveCodeBench is a contamination-free coding benchmark that continuously collects new problems from recent programming contests, evaluating LLMs on code generation, self-repair, execution, and test-output prediction. Because problems are timestamped, it can measure performance specifically on problems released after a model's training cutoff.

  • BigCodeBench Leaderboard

    BigCodeBench evaluates LLMs on 1,140 practical programming tasks that require invoking multiple function calls from 139 libraries across 7 domains, going beyond simpler single-function code-generation tests. The leaderboard, maintained by the BigCode project, is still receiving new model submissions as of mid-2026.

  • Code Arena (formerly WebDev Arena)

    Code Arena (formerly WebDev Arena) is Arena's coding-focused leaderboard, where two models build a web app from the same prompt and community votes determine a Bradley-Terry ranking. It focuses specifically on front-end web development and agentic build tasks rather than general chat quality.

  • Berkeley Function-Calling Leaderboard

    The Berkeley Function-Calling Leaderboard (BFCL), now on version 4, evaluates how accurately LLMs call functions and tools across simple, multiple, and parallel function-call scenarios spanning Python, Java, JavaScript, REST API, and SQL, plus holistic agentic evaluation. It uses real-world data, is updated periodically, and ships an installable evaluation package (bfcl-eval) for reproducing results.

  • EQ-Bench: Emotional Intelligence

    EQ-Bench has grown into a family of benchmarks hosted at the same site: EQ-Bench 3 measures emotional intelligence in challenging roleplay scenarios via an Elo-based scoring system, alongside companion leaderboards for creative writing (Creative Writing v3), judgment ability (Judgemark v4), longform writing, and other model behaviors. All are maintained by the same independent creator and updated as new models are released.

  • Massive Text Embedding Benchmark (MTEB) Leaderboard

    The Massive Text Embedding Benchmark (MTEB) has expanded well beyond its original 2022 scope of 8 tasks and 58 datasets: the multilingual MMTEB now covers 131 tasks across 250+ languages, and companion benchmarks (MIEB for images, MAEB for audio) extend evaluation to other modalities. Scores are self-reported by model providers using the open-source evaluation code, and the leaderboard is continuously updated with new submissions.

  • Open ASR Leaderboard

    The Open ASR Leaderboard ranks speech-to-text models by word error rate and real-time factor across multiple datasets, languages, and licenses. Hosted on Hugging Face, it lets users filter by model, dataset, or language and is actively maintained as new speech models are released.

  • AlpacaEval Leaderboard

    AlpacaEval is a fast, cheap, automatic evaluation framework that uses an LLM judge to estimate a model's win rate against a reference model on instruction-following tasks. Its current AlpacaEval 2.0 iteration uses length-controlled (LC) win rates to reduce the tendency of LLM judges to favor longer responses.

  • Design Arena

    Design Arena is a crowdsourced leaderboard for AI-generated design output, where models respond to the same creative prompt (e.g. building a website) and users vote head-to-head on the results. It has collected millions of votes from users in 190+ countries and includes category-specific boards such as website generation.

  • TIMETOACT LLM Benchmarks (formerly Trustbit)

    Following TIMETOACT GROUP's acquisition of Austrian consulting firm Trustbit, this monthly benchmark series continues under the TIMETOACT brand, evaluating LLMs on real enterprise workloads such as document processing, CRM integration, external API integration, marketing support, and code generation. Each edition tracks a composite performance score alongside cost and latency as separate decision factors.

  • Oobabooga benchmark

    A small, hand-written 48-question multiple-choice test of academic knowledge and reasoning, designed to avoid overlap with training data; still receiving new model entries as of 2026. Its creator has since shifted primary benchmarking effort to a separate project, LocalBench, which measures GGUF quantization quality via KL divergence rather than general model capability.

  • Uncensored General Intelligence Leaderboard

    A measurement of the amount of uncensored/controversial information an LLM knows. It is calculated from the average score of 5 subjects LLMs commonly refuse to talk about. The leaderboard is made of roughly 60 questions/tasks, measuring both "willingness to answer" and "accuracy" in controversial fact-based questions. I'm choosing to keep the questions private so people can't train on them and devalue the leaderboard.

Was this helpful?
Our Social Media →  
Original data from HuggingFace, Arena and various public git repos.
Check out Ag3ntum — our secure, self-hosted AI agent for server management.
Release v20260328a