How Should You Compare AI Benchmarks?
Use human votes for preference, task benchmarks for capability, agent benchmarks for tool work, and systems tests for speed or energy. Never turn one leaderboard into a universal model ranking.
What should you measure before choosing an AI model?
The smallest credible evaluation stack usually combines one public benchmark, one production-like harness and one local gold set.
From one answer to a long-running world.
A score becomes more production-like as the test adds state, tools and time—but usually becomes harder and costlier to reproduce. Explore the shape before choosing the scoreboard.
All lenses are visible. Click any benchmark to jump to its caveats and official route.
Vote first. Reveal the rubric second.
These are editorial sample outputs—not live model results. The point is to feel how a clear rubric changes a preference vote.
A policy answer with a missing exception
A retailer allows returns within 30 days only when goods are unopened. A customer reports an opened product is faulty on day 25. What should support say?
For real anonymous head-to-head voting with model identities revealed after the vote, use Arena. Its public leaderboard aggregates human preferences; it does not replace a task-specific acceptance test.
Find the scoreboard that matches the job.
Showing all 47 benchmark routes. Scores change; the durable value here is knowing what each route measures and misses.
Arena
People submit a prompt, compare two anonymous outputs and vote; model identities are revealed afterwards and votes feed human-preference leaderboards for text, image generation and editing, and video.
- Use for
- Broad product preference and style across text and creative media
- Watch
- Prompt mix, voter mix and presentation bias mean preference does not equal task accuracy
ImgSys
A human pairwise text-to-image preference arena focused on open-source image generators, with open preference data.
- Use for
- Comparing how voters prefer open image-generator outputs
- Watch
- Prompt and voter mix, model versions, presentation and sparse matchups can all move the ranking
AlpacaEval
An automated, reproducible instruction-following evaluation that compares model outputs against references with a judge model.
- Use for
- Fast iteration on general chat behavior
- Watch
- Judge-model, verbosity and style bias; not independent human preference
Stanford HELM
A transparent evaluation framework covering scenarios and multiple metrics rather than collapsing model behavior into one score.
- Use for
- Multi-metric, reproducible model audits
- Watch
- Scenario coverage still differs from your deployment distribution
Open LLM Leaderboard
Hugging Face’s route for comparing open models on standardized academic evaluations with result artifacts and reproducible tooling.
- Use for
- Shortlisting downloadable base and instruct models
- Watch
- Quantization, prompt templates and contamination can move results
Artificial Analysis Intelligence Index
A versioned composite of reasoning, knowledge, coding and agentic evaluations, paired with provider-observed price, token use, latency and output-speed measurements.
- Use for
- A broad quality/cost/speed shortlist under one published methodology
- Watch
- Index version, reasoning effort, provider route and benchmark weights can change the rank; finish with your workload
SenseBench
Models choose the correct WordNet sense for an English word in sentence context. The leaderboard recomputes scores from verified run artifacts and shows confidence intervals and cost.
- Use for
- A narrow, auditable check of lexical disambiguation and cost
- Watch
- Near-equal scores here say little about coding, tool use, long-horizon reliability or multimodal work
EQ‑Bench
Multi-turn roleplays and analysis tasks probe empathy, emotional reasoning, social dexterity and response tailoring, with per-model transcripts and an open runner.
- Use for
- A structured social-behavior signal that academic accuracy suites omit
- Watch
- A small, subjective set scored by an LLM judge; judge choice, roleplay style and the benchmark’s definition of “EQ” shape the result
MMLU / MMLU‑Pro
MMLU is the familiar broad academic test; MMLU‑Pro raises difficulty, expands to ten answer choices and requires more reasoning.
- Use for
- Broad within-table knowledge comparison
- Watch
- Public static questions invite saturation and training contamination
GPQA
Graduate-level biology, physics and chemistry questions designed to be difficult even with unrestricted web access.
- Use for
- Hard scientific QA and oversight research
- Watch
- Multiple choice and a narrow expert-domain slice
ARC-AGI
Abstract tasks designed around generalizing to novel problems; ARC-AGI-3 adds interactive environments and reports cost alongside performance.
- Use for
- Novel-task adaptation and reasoning efficiency
- Watch
- Purpose-built solvers may not transfer to language workflows
Humanity’s Last Exam
HLE uses difficult, broad expert-written questions, including multimodal items, to keep a closed-ended academic evaluation useful beyond saturated older tests.
- Use for
- Frontier expert knowledge and calibration
- Watch
- Tool access, answer revisions and dataset version change comparability
AIME
Model cards commonly reuse American Invitational Mathematics Examination problems as a short-answer mathematical-reasoning evaluation.
- Use for
- Exact-answer competition maths within the same year and protocol
- Watch
- Only 15 problems per exam; sampling, consensus and public solutions can dominate
MathArena
Evaluates models on newly released mathematics competitions and publishes problem-level outputs and leaderboards, reducing the chance that the exact questions appeared in pretraining.
- Use for
- Current mathematical reasoning and, where offered, expert-graded proof work
- Watch
- Small contests create wide uncertainty; tool access, sampling, answer extraction and grader protocol must match before comparing scores
IFEval + Estonian suites
IFEval checks verifiable instruction constraints. TartuNLP and EKI extend the local evidence with IFEval‑et, Grammar‑et, Word‑Meanings‑et and Estonian benchmark tasks.
- Use for
- Precise format compliance and language-specific shortlisting
- Watch
- Constraint compliance is not factuality; translations and adaptations need separate validation
RULER
RULER generates configurable synthetic tasks across retrieval and multi-hop categories to test how much of an advertised context window remains effective.
- Use for
- Comparing effective context as sequence length grows
- Watch
- Synthetic retrieval is easier to grade than messy long-document work
SWE‑bench family
Agents resolve real issues in repository snapshots. Verified uses human-validated tasks; Pro expands to harder, longer and partly held-out professional repositories.
- Use for
- Issue resolution with a named task set and harness
- Watch
- Variant, harness, budget and task quality materially affect the score
LiveCodeBench
A continuously updated coding evaluation built to reduce contamination and test generation, execution, repair and self-test behavior.
- Use for
- Current code reasoning on executable tasks
- Watch
- Competitive-programming tasks are not repository maintenance
Aider Polyglot
Exercises code editing across multiple programming languages in an open-source coding-assistant harness.
- Use for
- Editing quality in an actual assistant workflow
- Watch
- Aider’s prompts, edit format and tooling are part of the result
WebDev Arena
Models build rendered web applications from real user prompts, including vision inputs, and people compare the anonymous results head to head.
- Use for
- One-shot frontend usefulness, visual result and prompt adherence
- Watch
- Human taste, prompt mix and presentation dominate; a preferred render does not prove accessibility, maintainability, security or multi-turn repository work
HumanEval
HumanEval grades generated Python functions against tests and remains a common compact model-card signal for code generation.
- Use for
- Same-harness function synthesis comparisons
- Watch
- Small public Python tasks are saturated and unlike repository engineering
Berkeley Function Calling
BFCL evaluates selecting and calling functions across single-turn, multi-turn and agentic tasks, including hallucination and format sensitivity.
- Use for
- Tool routing and structured calls
- Watch
- Correct syntax does not prove the tool result or workflow is correct
GAIA
Realistic questions that require reasoning, web browsing, multimodal inputs and tool use, with private answers retained for leaderboard evaluation.
- Use for
- Research assistants and multi-step tool work
- Watch
- Agent scaffold and search access can matter as much as the model
τ-bench
Tests agents in tool-using conversations with simulated users and domain policies such as retail and airline service.
- Use for
- Multi-turn policy compliance and tool execution
- Watch
- Simulated users and domains remain a proxy for production
MCP Atlas
Agents discover and orchestrate tools from noisy menus across real MCP servers, with multi-step calls, parameter typing, error recovery and answer synthesis.
- Use for
- Tool discovery and end-to-end MCP workflows
- Watch
- Judge version, retries, tool-call budget and server state affect results
METR Time Horizons
METR fits success against the time human experts need for multi-step software and reasoning tasks, producing an interpretable duration at a chosen reliability level.
- Use for
- Tracking reliable task length rather than isolated skill
- Watch
- Task mix, human baselines and the 50% or 80% threshold change the horizon
Terminal‑Bench
Containerized tasks test whether agents can complete practical terminal work across coding, system administration, security, data and model training. Version 3.0 is a harder, separate protocol—not a continuation of the 2.1 percentage scale.
- Use for
- Hands-on terminal autonomy in reproducible environments
- Watch
- Agent scaffold, task version, timeout, rollout count, context and exploit resistance are part of the score
TerminalWorld
TerminalWorld turns real terminal recordings into reproducible tasks, tests and a human-verified subset spanning everyday developer and infrastructure work.
- Use for
- Broad terminal workflows with task and cost views
- Watch
- Synthesized instructions and tests can preserve artifacts of the source recordings
OSWorld 2.0
Agents operate real desktop applications through screenshots, mouse and keyboard across long workflows with dynamic state and cross-application dependencies.
- Use for
- Whole-computer work beyond browser or API-only agents
- Watch
- Step budget, environment image and partial-credit rules can shift the result
BrowseComp
BrowseComp asks agents to locate obscure, entangled facts that can require long search paths, while keeping final answers short enough to grade.
- Use for
- Persistent web search and strategic source discovery
- Watch
- Short factual answers do not measure a complete cited research report
Agents’ Last Exam
A broad, expert-built programme of long-horizon professional computer tasks with verifiable outcomes across many industries and specialist applications. The CLI route runs each task in its declared container and scores it with the benchmark's official evaluators.
- Use for
- Economically meaningful end-to-end agent work
- Watch
- Task version, harness, context, effort, output budget and task-specific timeout must match before comparing results
The GPT Investor Portfolio
A public ledger tracks time-stamped stock selections attributed to autonomous agents, their dollar and percentage returns and the return of SPY over the same stated interval.
- Use for
- Inspecting a longitudinal, real-money outcome record instead of a simulated finance quiz
- Watch
- Models, prompts, dates, holding periods and portfolio construction differ; the publisher controls the record, so this is not a controlled model comparison or financial advice
MMMU / MMMU-Pro
College-level multimodal questions across disciplines using charts, diagrams, images and text; Pro tightens robustness against shortcutting.
- Use for
- Visual reasoning with domain knowledge
- Watch
- Exam-style accuracy does not measure document-workflow reliability
VBench / VBench 2.0
An open video-generation evaluation suite spanning technical quality and intrinsic faithfulness, with prompt suites, metrics, code and a leaderboard.
- Use for
- Structured comparison of video-generation systems
- Watch
- Automated dimensions and standardized prompts are proxies; sampling, settings and model version matter, and the local custom-input subset is narrower than the full standard suite
Design Arena
A crowdsourced pairwise benchmark across image, video, editing, web and design work and other creative outputs.
- Use for
- Exploring current human preference across creative-output routes
- Watch
- Subjective preference, a live model pool, stochastic routing, low-vote entries and prompt enhancement mean this is not correctness
VHELM
Stanford’s living vision-language evaluation extends HELM’s transparent, multi-metric approach to models that reason over images and text.
- Use for
- Broad visual-language comparison
- Watch
- Aggregate coverage still needs a local image and document set
MineBench
Models read a natural-language build prompt and emit raw voxel-block coordinates. The site renders both worlds and humans vote blind to produce an Elo ranking.
- Use for
- Spatial composition, instruction following and inspectable creative output
- Watch
- Human aesthetic preference, prompt mix and generation budget are not geometric correctness
VoxelBench
A related benchmark evaluates language models on making voxel builds from text prompts and publishes a live leaderboard, keeping the generated world as the inspectable artifact.
- Use for
- Text-to-voxel build comparison
- Watch
- Leaderboard details alone do not expose a fully reproducible evaluation harness
HELM Safety
A transparent safety-evaluation route spanning multiple risk categories, models and scenarios rather than one refusal rate.
- Use for
- Structured safety and risk comparison
- Watch
- Public prompts can be trained against; deployment permissions still matter
JailbreakBench + HarmBench
JailbreakBench standardizes threat models, behavior sets, attack artifacts, judges and an attack/defense leaderboard; HarmBench adds a broader pipeline for comparing automated red-team methods, target models and robust-refusal defenses.
- Use for
- Reproducible attack success and defense comparisons under a named protocol
- Watch
- Results are dual-use and judge-dependent; public attacks invite overfitting, while low attack success can also mean unhelpful over-refusal rather than safe behavior
MLPerf
Industry-standard training and inference suites compare systems under defined scenarios, including throughput, latency and power submissions.
- Use for
- Architecture-neutral hardware and systems procurement
- Watch
- Submitted configurations may be heavily optimized and unlike your stack
GPU Battle: Can You Run It?
A third-party practitioner route for model-to-VRAM fit plus recorded throughput and efficiency evidence across LLM, image, video, embedding and related AI workload families—not a universal hardware ranking or endorsement.
- Use for
- Shortlisting a model and GPU configuration before a hands-on run
- Watch
- Fit and headline values depend on the exact model, quantization, runtime, context, settings and test conditions
GPU Battle: AI GPU Benchmarks
A third-party cross-card AI hardware comparison table for LLM tokens per second, image-generation performance and related measures—not a universal ranking or endorsement.
- Use for
- Scanning published cross-GPU evidence while building a hardware shortlist
- Watch
- Retain the table’s measured-versus-estimated labels; estimated and measured entries are not equivalent procurement evidence
GPU Battle: Buyer’s Guides
A third-party route for VRAM-specific model-fit guides, buy-versus-rent break-even paths and deeper hardware explainers—not a universal purchase recommendation or endorsement.
- Use for
- Framing a buy-versus-rent decision after a workload has passed acceptance
- Watch
- Prices, rental rates, availability, utilization and regional electricity and tax assumptions change; recompute with your workload and quote
ML.ENERGY
A benchmark and leaderboard for measuring inference energy under realistic service environments across models, tasks and system choices.
- Use for
- Energy-aware serving and optimization
- Watch
- Grid carbon, utilization and workload mix remain site-specific
Nawk Harness Efficiency
Holds one locally served DeepSeek V4 Flash configuration, eight repository bug fixes and one grading method constant while comparing Pi, OpenCode, Claude Code and Nanocoder on quality, generated tokens and wall-clock time.
- Use for
- Seeing scaffold cost, work style and run-to-run noise when the model stays fixed
- Watch
- One practitioner’s codebase, model and eight-task distribution; run counts differ and the study is not peer reviewed
ACBench
The Agent Compression Benchmark tests how quantization and pruning change workflow generation, tool use, long-context understanding and real-world application behavior.
- Use for
- Choosing a smaller or quantized build without assuming task parity
- Watch
- Compression effects vary by model, method and task; rerun your exact package
Try a broader word or reset the category filter.
The model is only one layer of the test.
Turn on the controls and watch an ungrounded answer become a reviewable workflow. This deterministic demo runs entirely in your browser.
TaskFind the renewal date in a contract and calculate the last day to give 60 days’ notice.
- Choose controls, then run the task.
A leaderboard can be correct and still mislead you.
Treat the score as a measurement produced by a dataset, prompt, harness, budget, judge and date—not as a property floating inside the model.
The test leaked into training.
Public static questions can become training data. Prefer held-out, rolling or newly collected tasks and record the cutoff.
Everyone clusters near the ceiling.
A benchmark above roughly 95% no longer separates frontier systems well. Retire it or add harder, fresher cases.
The scaffold won the benchmark.
Search, retries, tools, context construction and verification can dominate agent scores. Name and version the complete system.
The evaluator likes a style.
Automated judges can reward length, confidence or familiar phrasing. Use human checks and position-swapped comparisons.
A high score used far more inference.
Report tokens, reasoning level, attempts, latency and dollars per task. Capability without efficiency is an incomplete result.
The benchmark is not your workflow.
A coding or exam score may not survive your documents, languages, tools, error costs or user population. End with local acceptance tests.
Measurement guidance: Stanford AI Measurement Science. A 2025 audit of SWE-bench scoring reported corrections affecting 24.4% of Verified leaderboard entries and changing 11 rankings—useful evidence that benchmark infrastructure also needs verification: UTBoost paper.
A model earns deployment by passing your representative work at an acceptable cost and failure rate.