47 benchmark routes · checked 15 August 2026

How Should You Compare AI Benchmarks?

Use human votes for preference, task benchmarks for capability, agent benchmarks for tool work, and systems tests for speed or energy. Never turn one leaderboard into a universal model ranking.

The benchmark map

What should you measure before choosing an AI model?

The smallest credible evaluation stack usually combines one public benchmark, one production-like harness and one local gold set.

01PreferenceWhich response do people prefer?
02ReasoningCan it solve held-out questions?
03CodingCan it repair or write working code?
04AgentsCan it use tools across multiple steps?
05MultimodalCan it reason over image, audio or video?
06SafetyHow and where does it fail?
07SystemsWhat speed, cost and energy deliver it?
Interactive benchmark atlas

From one answer to a long-running world.

A score becomes more production-like as the test adds state, tools and time—but usually becomes harder and costlier to reproduce. Explore the shape before choosing the scoreboard.

All lenses are visible. Click any benchmark to jump to its caveats and official route.

Embedded benchmark demo

Vote first. Reveal the rubric second.

These are editorial sample outputs—not live model results. The point is to feel how a clear rubric changes a preference vote.

Scenario 1 of 3

A policy answer with a missing exception

A retailer allows returns within 30 days only when goods are unopened. A customer reports an opened product is faulty on day 25. What should support say?

Hidden scoring rubricCorrectness · uncertainty · next action · unsupported claims

For real anonymous head-to-head voting with model identities revealed after the vote, use Arena. Its public leaderboard aggregates human preferences; it does not replace a task-specific acceptance test.

Searchable database

Find the scoreboard that matches the job.

Showing all 47 benchmark routes. Scores change; the durable value here is knowing what each route measures and misses.

Human preferenceLive voting

Arena

People submit a prompt, compare two anonymous outputs and vote; model identities are revealed afterwards and votes feed human-preference leaderboards for text, image generation and editing, and video.

Use for
Broad product preference and style across text and creative media
Watch
Prompt mix, voter mix and presentation bias mean preference does not equal task accuracy
Image preferenceOpen voting data

ImgSys

A human pairwise text-to-image preference arena focused on open-source image generators, with open preference data.

Use for
Comparing how voters prefer open image-generator outputs
Watch
Prompt and voter mix, model versions, presentation and sparse matchups can all move the ranking
Preference proxyOpen evaluator

AlpacaEval

An automated, reproducible instruction-following evaluation that compares model outputs against references with a judge model.

Use for
Fast iteration on general chat behavior
Watch
Judge-model, verbosity and style bias; not independent human preference
Open the project →
Holistic evaluationLiving suite

Stanford HELM

A transparent evaluation framework covering scenarios and multiple metrics rather than collapsing model behavior into one score.

Use for
Multi-metric, reproducible model audits
Watch
Scenario coverage still differs from your deployment distribution
Explore HELM →
Open modelsReproducible runs

Open LLM Leaderboard

Hugging Face’s route for comparing open models on standardized academic evaluations with result artifacts and reproducible tooling.

Use for
Shortlisting downloadable base and instruct models
Watch
Quantization, prompt templates and contamination can move results
See leaderboard guidance →
Composite model indexUsed on this site

Artificial Analysis Intelligence Index

A versioned composite of reasoning, knowledge, coding and agentic evaluations, paired with provider-observed price, token use, latency and output-speed measurements.

Use for
A broad quality/cost/speed shortlist under one published methodology
Watch
Index version, reasoning effort, provider route and benchmark weights can change the rank; finish with your workload
Word-sense disambiguationUsed on this site

SenseBench

Models choose the correct WordNet sense for an English word in sentence context. The leaderboard recomputes scores from verified run artifacts and shows confidence intervals and cost.

Use for
A narrow, auditable check of lexical disambiguation and cost
Watch
Near-equal scores here say little about coding, tool use, long-horizon reliability or multimodal work
Social + emotional behaviorOpen transcripts

EQ‑Bench

Multi-turn roleplays and analysis tasks probe empathy, emotional reasoning, social dexterity and response tailoring, with per-model transcripts and an open runner.

Use for
A structured social-behavior signal that academic accuracy suites omit
Watch
A small, subjective set scored by an LLM judge; judge choice, roleplay style and the benchmark’s definition of “EQ” shape the result
Read the method and inspect transcripts →
Knowledge + reasoningUsed on this site

MMLU / MMLU‑Pro

MMLU is the familiar broad academic test; MMLU‑Pro raises difficulty, expands to ten answer choices and requires more reasoning.

Use for
Broad within-table knowledge comparison
Watch
Public static questions invite saturation and training contamination
Expert science448 questions

GPQA

Graduate-level biology, physics and chemistry questions designed to be difficult even with unrestricted web access.

Use for
Hard scientific QA and oversight research
Watch
Multiple choice and a narrow expert-domain slice
Read the benchmark paper →
Novel adaptationCost-aware

ARC-AGI

Abstract tasks designed around generalizing to novel problems; ARC-AGI-3 adds interactive environments and reports cost alongside performance.

Use for
Novel-task adaptation and reasoning efficiency
Watch
Purpose-built solvers may not transfer to language workflows
Open the leaderboard →
Frontier knowledgeUsed on this site

Humanity’s Last Exam

HLE uses difficult, broad expert-written questions, including multimodal items, to keep a closed-ended academic evaluation useful beyond saturated older tests.

Use for
Frontier expert knowledge and calibration
Watch
Tool access, answer revisions and dataset version change comparability
Competition mathsUsed on this site

AIME

Model cards commonly reuse American Invitational Mathematics Examination problems as a short-answer mathematical-reasoning evaluation.

Use for
Exact-answer competition maths within the same year and protocol
Watch
Only 15 problems per exam; sampling, consensus and public solutions can dominate
Fresh competition mathsRolling contests

MathArena

Evaluates models on newly released mathematics competitions and publishes problem-level outputs and leaderboards, reducing the chance that the exact questions appeared in pretraining.

Use for
Current mathematical reasoning and, where offered, expert-graded proof work
Watch
Small contests create wide uncertainty; tool access, sampling, answer extraction and grader protocol must match before comparing scores
Inspect competitions, outputs and method →
Instruction followingUsed on this site

IFEval + Estonian suites

IFEval checks verifiable instruction constraints. TartuNLP and EKI extend the local evidence with IFEval‑et, Grammar‑et, Word‑Meanings‑et and Estonian benchmark tasks.

Use for
Precise format compliance and language-specific shortlisting
Watch
Constraint compliance is not factuality; translations and adaptations need separate validation
Long contextUsed on this site

RULER

RULER generates configurable synthetic tasks across retrieval and multi-hop categories to test how much of an advertised context window remains effective.

Use for
Comparing effective context as sequence length grows
Watch
Synthetic retrieval is easier to grade than messy long-document work
Fresh code generationRolling set

LiveCodeBench

A continuously updated coding evaluation built to reduce contamination and test generation, execution, repair and self-test behavior.

Use for
Current code reasoning on executable tasks
Watch
Competitive-programming tasks are not repository maintenance
Open LiveCodeBench →
Code editingMulti-language

Aider Polyglot

Exercises code editing across multiple programming languages in an open-source coding-assistant harness.

Use for
Editing quality in an actual assistant workflow
Watch
Aider’s prompts, edit format and tooling are part of the result
See Aider leaderboards →
Interactive web developmentBlind human votes

WebDev Arena

Models build rendered web applications from real user prompts, including vision inputs, and people compare the anonymous results head to head.

Use for
One-shot frontend usefulness, visual result and prompt adherence
Watch
Human taste, prompt mix and presentation dominate; a preferred render does not prove accessibility, maintainability, security or multi-turn repository work
Read the method and open the leaderboard →
Function generationUsed on this site

HumanEval

HumanEval grades generated Python functions against tests and remains a common compact model-card signal for code generation.

Use for
Same-harness function synthesis comparisons
Watch
Small public Python tasks are saturated and unlike repository engineering
Inspect the original harness →
Tool useV4 · 2026

Berkeley Function Calling

BFCL evaluates selecting and calling functions across single-turn, multi-turn and agentic tasks, including hallucination and format sensitivity.

Use for
Tool routing and structured calls
Watch
Correct syntax does not prove the tool result or workflow is correct
Open BFCL →
General assistantsHeld-out answers

GAIA

Realistic questions that require reasoning, web browsing, multimodal inputs and tool use, with private answers retained for leaderboard evaluation.

Use for
Research assistants and multi-step tool work
Watch
Agent scaffold and search access can matter as much as the model
Explore GAIA →
Customer-service agentsPolicy + tools

τ-bench

Tests agents in tool-using conversations with simulated users and domain policies such as retail and airline service.

Use for
Multi-turn policy compliance and tool execution
Watch
Simulated users and domains remain a proxy for production
Inspect τ-bench →
MCP tool useUsed on this site

MCP Atlas

Agents discover and orchestrate tools from noisy menus across real MCP servers, with multi-step calls, parameter typing, error recovery and answer synthesis.

Use for
Tool discovery and end-to-end MCP workflows
Watch
Judge version, retries, tool-call budget and server state affect results
Open the MCP Atlas leaderboard →
Task horizonUsed on this site

METR Time Horizons

METR fits success against the time human experts need for multi-step software and reasoning tasks, producing an interpretable duration at a chosen reliability level.

Use for
Tracking reliable task length rather than isolated skill
Watch
Task mix, human baselines and the 50% or 80% threshold change the horizon
Explore the current time-horizon chart →
Terminal agents2.1 + 3.0 · used here

Terminal‑Bench

Containerized tasks test whether agents can complete practical terminal work across coding, system administration, security, data and model training. Version 3.0 is a harder, separate protocol—not a continuation of the 2.1 percentage scale.

Use for
Hands-on terminal autonomy in reproducible environments
Watch
Agent scaffold, task version, timeout, rollout count, context and exploit resistance are part of the score
Recorded terminal work2026 · live suite

TerminalWorld

TerminalWorld turns real terminal recordings into reproducible tasks, tests and a human-verified subset spanning everyday developer and infrastructure work.

Use for
Broad terminal workflows with task and cost views
Watch
Synthesized instructions and tests can preserve artifacts of the source recordings
Browse TerminalWorld tasks →
Computer useLong-horizon GUI

OSWorld 2.0

Agents operate real desktop applications through screenshots, mouse and keyboard across long workflows with dynamic state and cross-application dependencies.

Use for
Whole-computer work beyond browser or API-only agents
Watch
Step budget, environment image and partial-credit rules can shift the result
Explore OSWorld 2.0 →
Web researchHard retrieval

BrowseComp

BrowseComp asks agents to locate obscure, entangled facts that can require long search paths, while keeping final answers short enough to grade.

Use for
Persistent web search and strategic source discovery
Watch
Short factual answers do not measure a complete cited research report
Read the benchmark and caveats →
Professional workflowsCLI · used on this site

Agents’ Last Exam

A broad, expert-built programme of long-horizon professional computer tasks with verifiable outcomes across many industries and specialist applications. The CLI route runs each task in its declared container and scores it with the benchmark's official evaluators.

Use for
Economically meaningful end-to-end agent work
Watch
Task version, harness, context, effort, output budget and task-specific timeout must match before comparing results
Live economic outcomeReal-money tracker

The GPT Investor Portfolio

A public ledger tracks time-stamped stock selections attributed to autonomous agents, their dollar and percentage returns and the return of SPY over the same stated interval.

Use for
Inspecting a longitudinal, real-money outcome record instead of a simulated finance quiz
Watch
Models, prompts, dates, holding periods and portfolio construction differ; the publisher controls the record, so this is not a controlled model comparison or financial advice
Inspect the live portfolio and holdings →
Vision + knowledgeExpert tasks

MMMU / MMMU-Pro

College-level multimodal questions across disciplines using charts, diagrams, images and text; Pro tightens robustness against shortcutting.

Use for
Visual reasoning with domain knowledge
Watch
Exam-style accuracy does not measure document-workflow reliability
Open MMMU →
Video generationOpen evaluation suite

VBench / VBench 2.0

An open video-generation evaluation suite spanning technical quality and intrinsic faithfulness, with prompt suites, metrics, code and a leaderboard.

Use for
Structured comparison of video-generation systems
Watch
Automated dimensions and standardized prompts are proxies; sampling, settings and model version matter, and the local custom-input subset is narrower than the full standard suite
Creative preferenceCrowdsourced pairs

Design Arena

A crowdsourced pairwise benchmark across image, video, editing, web and design work and other creative outputs.

Use for
Exploring current human preference across creative-output routes
Watch
Subjective preference, a live model pool, stochastic routing, low-vote entries and prompt enhancement mean this is not correctness
Visual languageHolistic suite

VHELM

Stanford’s living vision-language evaluation extends HELM’s transparent, multi-metric approach to models that reason over images and text.

Use for
Broad visual-language comparison
Watch
Aggregate coverage still needs a local image and document set
Explore VHELM →
Text → 3D worldLive human arena

MineBench

Models read a natural-language build prompt and emit raw voxel-block coordinates. The site renders both worlds and humans vote blind to produce an Elo ranking.

Use for
Spatial composition, instruction following and inspectable creative output
Watch
Human aesthetic preference, prompt mix and generation budget are not geometric correctness
Vote, explore builds or use the sandbox →
Voxel generationCommunity leaderboard

VoxelBench

A related benchmark evaluates language models on making voxel builds from text prompts and publishes a live leaderboard, keeping the generated world as the inspectable artifact.

Use for
Text-to-voxel build comparison
Watch
Leaderboard details alone do not expose a fully reproducible evaluation harness
Inspect the live VoxelBench leaderboard →
Safety evaluationMultiple risks

HELM Safety

A transparent safety-evaluation route spanning multiple risk categories, models and scenarios rather than one refusal rate.

Use for
Structured safety and risk comparison
Watch
Public prompts can be trained against; deployment permissions still matter
Open HELM Safety →
Jailbreak robustnessUsed on this site

JailbreakBench + HarmBench

JailbreakBench standardizes threat models, behavior sets, attack artifacts, judges and an attack/defense leaderboard; HarmBench adds a broader pipeline for comparing automated red-team methods, target models and robust-refusal defenses.

Use for
Reproducible attack success and defense comparisons under a named protocol
Watch
Results are dual-use and judge-dependent; public attacks invite overfitting, while low attack success can also mean unhelpful over-refusal rather than safe behavior
Hardware + servingAudited submissions

MLPerf

Industry-standard training and inference suites compare systems under defined scenarios, including throughput, latency and power submissions.

Use for
Architecture-neutral hardware and systems procurement
Watch
Submitted configurations may be heavily optimized and unlike your stack
Model + VRAM fitPractitioner route

GPU Battle: Can You Run It?

A third-party practitioner route for model-to-VRAM fit plus recorded throughput and efficiency evidence across LLM, image, video, embedding and related AI workload families—not a universal hardware ranking or endorsement.

Use for
Shortlisting a model and GPU configuration before a hands-on run
Watch
Fit and headline values depend on the exact model, quantization, runtime, context, settings and test conditions
Check model-to-VRAM fit →
AI hardware comparisonPractitioner route

GPU Battle: AI GPU Benchmarks

A third-party cross-card AI hardware comparison table for LLM tokens per second, image-generation performance and related measures—not a universal ranking or endorsement.

Use for
Scanning published cross-GPU evidence while building a hardware shortlist
Watch
Retain the table’s measured-versus-estimated labels; estimated and measured entries are not equivalent procurement evidence
Inspect AI GPU benchmarks →
Fit + ownership guidesPractitioner route

GPU Battle: Buyer’s Guides

A third-party route for VRAM-specific model-fit guides, buy-versus-rent break-even paths and deeper hardware explainers—not a universal purchase recommendation or endorsement.

Use for
Framing a buy-versus-rent decision after a workload has passed acceptance
Watch
Prices, rental rates, availability, utilization and regional electricity and tax assumptions change; recompute with your workload and quote
Read the buyer’s guides →
Inference energyMeasured systems

ML.ENERGY

A benchmark and leaderboard for measuring inference energy under realistic service environments across models, tasks and system choices.

Use for
Energy-aware serving and optimization
Watch
Grid carbon, utilization and workload mix remain site-specific
Open ML.ENERGY →
Coding-harness efficiencyUsed on this site

Nawk Harness Efficiency

Holds one locally served DeepSeek V4 Flash configuration, eight repository bug fixes and one grading method constant while comparing Pi, OpenCode, Claude Code and Nanocoder on quality, generated tokens and wall-clock time.

Use for
Seeing scaffold cost, work style and run-to-run noise when the model stays fixed
Watch
One practitioner’s codebase, model and eight-task distribution; run counts differ and the study is not peer reviewed
Compression effectsUsed on this site

ACBench

The Agent Compression Benchmark tests how quantization and pruning change workflow generation, tool use, long-context understanding and real-world application behavior.

Use for
Choosing a smaller or quantized build without assuming task parity
Watch
Compression effects vary by model, method and task; rerun your exact package
Simple agent harness demo

The model is only one layer of the test.

Turn on the controls and watch an ungrounded answer become a reviewable workflow. This deterministic demo runs entirely in your browser.

TaskFind the renewal date in a contract and calculate the last day to give 60 days’ notice.

TraceNot run
  1. Choose controls, then run the task.
No reviewable answer yet.
Benchmark traps

A leaderboard can be correct and still mislead you.

Treat the score as a measurement produced by a dataset, prompt, harness, budget, judge and date—not as a property floating inside the model.

01 · Contamination

The test leaked into training.

Public static questions can become training data. Prefer held-out, rolling or newly collected tasks and record the cutoff.

02 · Saturation

Everyone clusters near the ceiling.

A benchmark above roughly 95% no longer separates frontier systems well. Retire it or add harder, fresher cases.

03 · Harness lift

The scaffold won the benchmark.

Search, retries, tools, context construction and verification can dominate agent scores. Name and version the complete system.

04 · Judge bias

The evaluator likes a style.

Automated judges can reward length, confidence or familiar phrasing. Use human checks and position-swapped comparisons.

05 · Hidden cost

A high score used far more inference.

Report tokens, reasoning level, attempts, latency and dollars per task. Capability without efficiency is an incomplete result.

06 · Distribution shift

The benchmark is not your workflow.

A coding or exam score may not survive your documents, languages, tools, error costs or user population. End with local acceptance tests.

Measurement guidance: Stanford AI Measurement Science. A 2025 audit of SWE-bench scoring reported corrections affecting 24.4% of Verified leaderboard entries and changing 11 rankings—useful evidence that benchmark infrastructure also needs verification: UTBoost paper.

The final benchmark

A model earns deployment by passing your representative work at an acceptable cost and failure rate.