# Benchmarks & Leaderboards

> The shared scoring systems the community uses to compare models: LLM leaderboards and human-preference arenas like LMSYS Chatbot Arena, software engineering benchmarks like the Aider Leaderboard, and agent evaluations with live model rankings. Internal eval tooling for your own apps lives in Observability & Evals.

27 tools in 3 subcategories. Web page: https://ailandscape.org/category/benchmarks-leaderboards · Each tool also has a markdown version at /tool/{slug}.md

## Software Engineering Benchmarks

- [Aider Leaderboard](https://ailandscape.org/tool/aider-leaderboard): Aider's polyglot coding leaderboard measuring LLM performance across languages
- [BigCodeArena](https://ailandscape.org/tool/bigcodearena): A human-in-the-loop platform for evaluating code generation through execution
- [CORE-Bench Hard](https://ailandscape.org/tool/core-bench-hard): Agent installs deps, runs published scientific code, and answers questions
- [Live-SWE-agent](https://ailandscape.org/tool/live-swe-agent): Can software engineering agents self-evolve on the fly? A living benchmark
- [LiveCodeBench](https://ailandscape.org/tool/livecodebench): Holistic and contamination-free evaluation of large language models for code
- [LiveCodeBench Pro](https://ailandscape.org/tool/livecodebench-pro): Codeforces, ICPC, and IOI problems; regularly updated to prevent contamination
- [Multi-SWE-bench](https://ailandscape.org/tool/multi-swe-bench): Multilingual benchmark for issue resolving across multiple programming languages
- [SWE-bench](https://ailandscape.org/tool/swe-bench): Evaluates LLM performance on real-world software issues collected from GitHub
- [SWE-bench Multilingual](https://ailandscape.org/tool/swe-bench-multilingual): 300 curated SWE-bench tasks from 42 repositories across 9 programming languages
- [SWE-Bench Pro Commercial](https://ailandscape.org/tool/swe-bench-pro-commercial): Scale AI's real-world software engineering benchmark using a commercial dataset
- [SWE-Bench Pro Public](https://ailandscape.org/tool/swe-bench-pro-public): Public version of Scale AI's SWE-Bench Pro for real-world software engineering
- [SWE-DEV](https://ailandscape.org/tool/swe-dev): Evaluating and training autonomous feature-driven software development agents
- [SWE-rebench](https://ailandscape.org/tool/swe-rebench): Continuously evolving and decontaminated benchmark for software engineering LLMs

## Agent & General Benchmarks

- [APEX-Agents](https://ailandscape.org/tool/apex-agents): Measures whether frontier AI agents can execute real long-horizon tasks
- [ARC-AGI-2](https://ailandscape.org/tool/arc-agi-2): Stress-tests efficiency and capability of state-of-the-art AI reasoning systems
- [Context-Bench](https://ailandscape.org/tool/context-bench): A benchmark for agentic context engineering
- [MCP Atlas](https://ailandscape.org/tool/mcp-atlas): Evaluates how well language models handle real-world tool use through MCP
- [Modu Merge Rate Leaderboard](https://ailandscape.org/tool/modu-merge-rate-leaderboard): Ranking top coding agents by real-world pull request merge success rates
- [OSWorld](https://ailandscape.org/tool/osworld): Benchmarks multimodal agents on open-ended tasks in real computer environments
- [PR Arena](https://ailandscape.org/tool/pr-arena): Software engineering agents head-to-head on real pull request tasks
- [Repo Bench](https://ailandscape.org/tool/repo-bench): Measures large-context reasoning, edit precision, and instruction adherence
- [Terminal-Bench](https://ailandscape.org/tool/terminal-bench): Benchmark measuring AI agent capabilities in a terminal environment (v1)
- [Terminal-Bench 2.0](https://ailandscape.org/tool/terminal-bench-2-0): Benchmark measuring AI agent capabilities in a terminal environment (v2)
- [Vending-Bench 2](https://ailandscape.org/tool/vending-bench-2): Measuring AI model performance on running a business over long time horizons
- [τ-bench](https://ailandscape.org/tool/bench): Benchmarking AI agents in collaborative real-world scenarios

## Model Rankings

- [LMSYS Chatbot Arena](https://ailandscape.org/tool/lmsys-chatbot-arena): Human-preference LLM ranking via pairwise comparisons
- [OpenRouter Rankings](https://ailandscape.org/tool/openrouter-rankings): Model market share, use-case categories, and app rankings on OpenRouter

---

Part of [AI Landscape](https://ailandscape.org), an open map of the AI ecosystem. Index for AI assistants: https://ailandscape.org/llms.txt
