# How Should You Compare AI Benchmarks?

Canonical source: [https://isaiuseful.com/benchmarks](https://isaiuseful.com/benchmarks)

<a id="main-content"></a>

47 benchmark routes · checked 15 August 2026

Use human votes for preference, task benchmarks for capability, agent benchmarks for tool work, and systems tests for speed or energy. Never turn one leaderboard into a universal model ranking.

- [Search the database](#database)

- [Try a blind comparison](#blind-lab)

> workflow → dataset → harness → metric
> score != product_quality
> current_result + reproducible_setup
7
benchmark families, because preference, knowledge, coding, agents, vision, safety and serving are not interchangeable.
**Blind**

vote without brand anchoring
**Fresh**

prefer rolling or held-out tests
**Costed**

record tokens, time and hardware
**Local**

finish with your own acceptance set

The benchmark map

## What should you measure before choosing an AI model?

The smallest credible evaluation stack usually combines one public benchmark, one production-like harness and one local gold set.

> Visual: Seven benchmark families and the decisions they support

**Visual entries (display order):**
- 01 **Preference** Which response do people prefer?
- 02 **Reasoning** Can it solve held-out questions?
- 03 **Coding** Can it repair or write working code?
- 04 **Agents** Can it use tools across multiple steps?
- 05 **Multimodal** Can it reason over image, audio or video?
- 06 **Safety** How and where does it fail?
- 07 **Systems** What speed, cost and energy deliver it?

Interactive benchmark atlas

## From one answer to a long-running world.

A score becomes more production-like as the test adds state, tools and time—but usually becomes harder and costlier to reproduce. Explore the shape before choosing the scoreboard.

> Visual: Qualitative chart of benchmark environment and task horizon

> Scale note: Typical evaluation horizon qualitative · not to scale

| Evaluation environment | One answer | One session | Workflow | Project |
| --- | --- | --- | --- | --- |
| **Question** | [AIME](#aime) · [MathArena](#matharena) · [HLE](#hle) | [RULER](#ruler) · [BrowseComp](#browsecomp) |  |  |
| **Code** | [HumanEval](#humaneval) · [LiveCodeBench](#livecodebench) | [WebDev Arena](#webdev-arena) | [SWE-bench](#swe-bench) · [Terminal-Bench](#terminal-bench) | [METR horizons](#metr-time-horizons) |
| **Interface** | [BFCL](#bfcl) | [MCP Atlas](#mcp-atlas) | [OSWorld](#osworld) · [TerminalWorld](#terminalworld) | [Agents’ Last Exam](#agents-last-exam) |
| **Constructed world** |  | [MineBench](#minebench) · [VoxelBench](#voxelbench) |  |  |
| **People + safety** | [EQ‑Bench](#eq-bench) |  | [JailbreakBench + HarmBench](#jailbreakbench-harmbench) |  |
| **Systems + hardware** |  | [GPU Battle: Can You Run It?](#gpu-battle-can-you-run) · [GPU Battle: AI GPUs](#gpu-battle-ai) |  | [GPU Battle: Buyer’s Guides](#gpu-battle-guides) |

All lenses are visible. Click any benchmark to jump to its caveats and official route.

<a id="blind-lab"></a>

Embedded benchmark demo

## Vote first. Reveal the rubric second.

These are editorial sample outputs—not live model results. The point is to feel how a clear rubric changes a preference vote.

Scenario 1 of 3

### A policy answer with a missing exception

A retailer allows returns within 30 days only when goods are unopened. A customer reports an opened product is faulty on day 25. What should support say?

**Hidden scoring rubric**

Correctness · uncertainty · next action · unsupported claims
For real anonymous head-to-head voting with model identities revealed after the vote, use [Arena](https://arena.ai/) . Its public leaderboard aggregates human preferences; it does not replace a task-specific acceptance test.

<a id="database"></a>

Searchable database

## Find the scoreboard that matches the job.

Showing all 47 benchmark routes. Scores change; the durable value here is knowing what each route measures and misses.

**Benchmarks cited elsewhere on this site**

This index closes the loop between model scorecards, evidence pages and the full caveats here.
Human preference
**Live voting**

### Arena

People submit a prompt, compare two anonymous outputs and vote; model identities are revealed afterwards and votes feed human-preference leaderboards for text, image generation and editing, and video.

**Use for** — Broad product preference and style across text and creative media
**Watch** — Prompt mix, voter mix and presentation bias mean preference does not equal task accuracy

- [Text leaderboard →](https://arena.ai/leaderboard)

- [Text-to-image leaderboard →](https://arena.ai/leaderboard/text-to-image)

- [Image-edit leaderboard →](https://arena.ai/leaderboard/image-edit)

- [Text-to-video leaderboard →](https://arena.ai/leaderboard/text-to-video/overall)

- [Image-to-video leaderboard →](https://arena.ai/leaderboard/image-to-video)

Image preference
**Open voting data**

### ImgSys

A human pairwise text-to-image preference arena focused on open-source image generators, with open preference data.

**Use for** — Comparing how voters prefer open image-generator outputs
**Watch** — Prompt and voter mix, model versions, presentation and sparse matchups can all move the ranking

- [ImgSys rankings →](https://imgsys.org/rankings)

- [ImgSys methodology →](https://imgsys.org/methodology)

Preference proxy
**Open evaluator**

### AlpacaEval

An automated, reproducible instruction-following evaluation that compares model outputs against references with a judge model.

**Use for** — Fast iteration on general chat behavior
**Watch** — Judge-model, verbosity and style bias; not independent human preference

- [Open the project →](https://tatsu-lab.github.io/alpaca_eval/)

Holistic evaluation
**Living suite**

### Stanford HELM

A transparent evaluation framework covering scenarios and multiple metrics rather than collapsing model behavior into one score.

**Use for** — Multi-metric, reproducible model audits
**Watch** — Scenario coverage still differs from your deployment distribution

- [Explore HELM →](https://crfm.stanford.edu/helm/)

Open models
**Reproducible runs**

### Open LLM Leaderboard

Hugging Face’s route for comparing open models on standardized academic evaluations with result artifacts and reproducible tooling.

**Use for** — Shortlisting downloadable base and instruct models
**Watch** — Quantization, prompt templates and contamination can move results

- [See leaderboard guidance →](https://huggingface.co/docs/leaderboards/index)

<a id="artificial-analysis"></a>

Composite model index
**Used on this site**

### Artificial Analysis Intelligence Index

A versioned composite of reasoning, knowledge, coding and agentic evaluations, paired with provider-observed price, token use, latency and output-speed measurements.

**Use for** — A broad quality/cost/speed shortlist under one published methodology
**Watch** — Index version, reasoning effort, provider route and benchmark weights can change the rank; finish with your workload

- [See the MiniMax M3 comparison in context →](https://isaiuseful.com/cloud-models.html.md#minimax-m3)

- [Read the index methodology →](https://artificialanalysis.ai/methodology/intelligence-benchmarking)

<a id="sensebench"></a>

Word-sense disambiguation
**Used on this site**

### SenseBench

Models choose the correct WordNet sense for an English word in sentence context. The leaderboard recomputes scores from verified run artifacts and shows confidence intervals and cost.

**Use for** — A narrow, auditable check of lexical disambiguation and cost
**Watch** — Near-equal scores here say little about coding, tool use, long-horizon reliability or multimodal work

- [See why the M3 result is treated narrowly →](https://isaiuseful.com/dgx-station.html.md#reviews)

- [Open the leaderboard and run artifacts →](https://sense-bench.com/)

<a id="eq-bench"></a>

Social + emotional behavior
**Open transcripts**

### EQ‑Bench

Multi-turn roleplays and analysis tasks probe empathy, emotional reasoning, social dexterity and response tailoring, with per-model transcripts and an open runner.

**Use for** — A structured social-behavior signal that academic accuracy suites omit
**Watch** — A small, subjective set scored by an LLM judge; judge choice, roleplay style and the benchmark’s definition of “EQ” shape the result

- [Read the method and inspect transcripts →](https://eqbench.com/about.html)

<a id="mmlu-pro"></a>

Knowledge + reasoning
**Used on this site**

### MMLU / MMLU‑Pro

MMLU is the familiar broad academic test; MMLU‑Pro raises difficulty, expands to ten answer choices and requires more reasoning.

**Use for** — Broad within-table knowledge comparison
**Watch** — Public static questions invite saturation and training contamination

- [See how local model cards use it →](https://isaiuseful.com/local-models.html.md#benchmarks)

- [Inspect MMLU‑Pro code and data →](https://github.com/TIGER-AI-Lab/MMLU-Pro)

<a id="gpqa"></a>

Expert science
**448 questions**

### GPQA

Graduate-level biology, physics and chemistry questions designed to be difficult even with unrestricted web access.

**Use for** — Hard scientific QA and oversight research
**Watch** — Multiple choice and a narrow expert-domain slice

- [Read the benchmark paper →](https://arxiv.org/abs/2311.12022)

Novel adaptation
**Cost-aware**

### ARC-AGI

Abstract tasks designed around generalizing to novel problems; ARC-AGI-3 adds interactive environments and reports cost alongside performance.

**Use for** — Novel-task adaptation and reasoning efficiency
**Watch** — Purpose-built solvers may not transfer to language workflows

- [Open the leaderboard →](https://arcprize.org/leaderboard)

<a id="hle"></a>

Frontier knowledge
**Used on this site**

### Humanity’s Last Exam

HLE uses difficult, broad expert-written questions, including multimodal items, to keep a closed-ended academic evaluation useful beyond saturated older tests.

**Use for** — Frontier expert knowledge and calibration
**Watch** — Tool access, answer revisions and dataset version change comparability

- [See the score in its model-card context →](https://isaiuseful.com/local-models.html.md#glm)

- [Open the official HLE project →](https://www.lastexam.ai/)

<a id="aime"></a>

Competition maths
**Used on this site**

### AIME

Model cards commonly reuse American Invitational Mathematics Examination problems as a short-answer mathematical-reasoning evaluation.

**Use for** — Exact-answer competition maths within the same year and protocol
**Watch** — Only 15 problems per exam; sampling, consensus and public solutions can dominate

- [See the score in its model-card context →](https://isaiuseful.com/local-models.html.md#deepseek)

- [See the official AIME route →](https://maa.org/math-competitions/american-invitational-mathematics-examination-aime/)

<a id="matharena"></a>

Fresh competition maths
**Rolling contests**

### MathArena

Evaluates models on newly released mathematics competitions and publishes problem-level outputs and leaderboards, reducing the chance that the exact questions appeared in pretraining.

**Use for** — Current mathematical reasoning and, where offered, expert-graded proof work
**Watch** — Small contests create wide uncertainty; tool access, sampling, answer extraction and grader protocol must match before comparing scores

- [Inspect competitions, outputs and method →](https://matharena.ai/)

<a id="ifeval"></a>

Instruction following
**Used on this site**

### IFEval + Estonian suites

IFEval checks verifiable instruction constraints. TartuNLP and EKI extend the local evidence with IFEval‑et, Grammar‑et, Word‑Meanings‑et and Estonian benchmark tasks.

**Use for** — Precise format compliance and language-specific shortlisting
**Watch** — Constraint compliance is not factuality; translations and adaptations need separate validation

- [Open the IFEval implementation →](https://github.com/google-research/google-research/tree/master/instruction_following_eval)

- [Inspect the Estonian evaluation tables →](https://huggingface.co/tartuNLP/llama-estllm-prototype-0825)

<a id="ruler"></a>

Long context
**Used on this site**

### RULER

RULER generates configurable synthetic tasks across retrieval and multi-hop categories to test how much of an advertised context window remains effective.

**Use for** — Comparing effective context as sequence length grows
**Watch** — Synthetic retrieval is easier to grade than messy long-document work

- [See the long-context score in context →](https://isaiuseful.com/local-models.html.md#kimi)

- [Run the open RULER suite →](https://github.com/NVIDIA/RULER)

<a id="swe-bench"></a>

Repository repair
**Used on this site**

### SWE‑bench family

Agents resolve real issues in repository snapshots. Verified uses human-validated tasks; Pro expands to harder, longer and partly held-out professional repositories.

**Use for** — Issue resolution with a named task set and harness
**Watch** — Variant, harness, budget and task quality materially affect the score

- [See the bounded repair use case →](https://isaiuseful.com/use-cases.html.md#independent-examples)

- [Build the coding workflow →](https://isaiuseful.com/guides.html.md#tested-patch)

- [Official SWE‑bench leaderboards →](https://www.swebench.com/)

- [SWE‑bench Pro paper and leaderboard →](https://labs.scale.com/papers/swe_bench_pro)

<a id="livecodebench"></a>

Fresh code generation
**Rolling set**

### LiveCodeBench

A continuously updated coding evaluation built to reduce contamination and test generation, execution, repair and self-test behavior.

**Use for** — Current code reasoning on executable tasks
**Watch** — Competitive-programming tasks are not repository maintenance

- [Open LiveCodeBench →](https://livecodebench.github.io/)

Code editing
**Multi-language**

### Aider Polyglot

Exercises code editing across multiple programming languages in an open-source coding-assistant harness.

**Use for** — Editing quality in an actual assistant workflow
**Watch** — Aider’s prompts, edit format and tooling are part of the result

- [See Aider leaderboards →](https://aider.chat/docs/leaderboards/)

<a id="webdev-arena"></a>

Interactive web development
**Blind human votes**

### WebDev Arena

Models build rendered web applications from real user prompts, including vision inputs, and people compare the anonymous results head to head.

**Use for** — One-shot frontend usefulness, visual result and prompt adherence
**Watch** — Human taste, prompt mix and presentation dominate; a preferred render does not prove accessibility, maintainability, security or multi-turn repository work

- [Read the method and open the leaderboard →](https://arena.ai/blog/webdev-arena)

<a id="humaneval"></a>

Function generation
**Used on this site**

### HumanEval

HumanEval grades generated Python functions against tests and remains a common compact model-card signal for code generation.

**Use for** — Same-harness function synthesis comparisons
**Watch** — Small public Python tasks are saturated and unlike repository engineering

- [Inspect the original harness →](https://github.com/openai/human-eval)

<a id="bfcl"></a>

Tool use
**V4 · 2026**

### Berkeley Function Calling

BFCL evaluates selecting and calling functions across single-turn, multi-turn and agentic tasks, including hallucination and format sensitivity.

**Use for** — Tool routing and structured calls
**Watch** — Correct syntax does not prove the tool result or workflow is correct

- [Open BFCL →](https://gorilla.cs.berkeley.edu/leaderboard)

General assistants
**Held-out answers**

### GAIA

Realistic questions that require reasoning, web browsing, multimodal inputs and tool use, with private answers retained for leaderboard evaluation.

**Use for** — Research assistants and multi-step tool work
**Watch** — Agent scaffold and search access can matter as much as the model

- [Explore GAIA →](https://huggingface.co/gaia-benchmark)

Customer-service agents
**Policy + tools**

### τ-bench

Tests agents in tool-using conversations with simulated users and domain policies such as retail and airline service.

**Use for** — Multi-turn policy compliance and tool execution
**Watch** — Simulated users and domains remain a proxy for production

- [Inspect τ-bench →](https://github.com/sierra-research/tau-bench)

<a id="mcp-atlas"></a>

MCP tool use
**Used on this site**

### MCP Atlas

Agents discover and orchestrate tools from noisy menus across real MCP servers, with multi-step calls, parameter typing, error recovery and answer synthesis.

**Use for** — Tool discovery and end-to-end MCP workflows
**Watch** — Judge version, retries, tool-call budget and server state affect results

- [Open the MCP Atlas leaderboard →](https://labs.scale.com/leaderboard/mcp_atlas)

<a id="metr-time-horizons"></a>

Task horizon
**Used on this site**

### METR Time Horizons

METR fits success against the time human experts need for multi-step software and reasoning tasks, producing an interpretable duration at a chosen reliability level.

**Use for** — Tracking reliable task length rather than isolated skill
**Watch** — Task mix, human baselines and the 50% or 80% threshold change the horizon

- [Explore the current time-horizon chart →](https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/)

<a id="terminal-bench"></a>

Terminal agents
**2.1 + 3.0 · used here**

### Terminal‑Bench

Containerized tasks test whether agents can complete practical terminal work across coding, system administration, security, data and model training. Version 3.0 is a harder, separate protocol—not a continuation of the 2.1 percentage scale.

**Use for** — Hands-on terminal autonomy in reproducible environments
**Watch** — Agent scaffold, task version, timeout, rollout count, context and exploit resistance are part of the score

- [Compare GLM‑5.2 and GLM‑5.3 in context →](https://isaiuseful.com/local-models.html.md#glm)

- [Inspect Z.ai's 2.1 and 3.0 evaluation settings →](https://z.ai/blog/glm-5.3)

- [Open the Terminal‑Bench project →](https://www.tbench.ai/)

<a id="terminalworld"></a>

Recorded terminal work
**2026 · live suite**

### TerminalWorld

TerminalWorld turns real terminal recordings into reproducible tasks, tests and a human-verified subset spanning everyday developer and infrastructure work.

**Use for** — Broad terminal workflows with task and cost views
**Watch** — Synthesized instructions and tests can preserve artifacts of the source recordings

- [Browse TerminalWorld tasks →](https://terminalworld.ai/)

<a id="osworld"></a>

Computer use
**Long-horizon GUI**

### OSWorld 2.0

Agents operate real desktop applications through screenshots, mouse and keyboard across long workflows with dynamic state and cross-application dependencies.

**Use for** — Whole-computer work beyond browser or API-only agents
**Watch** — Step budget, environment image and partial-credit rules can shift the result

- [Explore OSWorld 2.0 →](https://osworld-v2.xlang.ai/)

<a id="browsecomp"></a>

Web research
**Hard retrieval**

### BrowseComp

BrowseComp asks agents to locate obscure, entangled facts that can require long search paths, while keeping final answers short enough to grade.

**Use for** — Persistent web search and strategic source discovery
**Watch** — Short factual answers do not measure a complete cited research report

- [Read the benchmark and caveats →](https://openai.com/index/browsecomp/)

<a id="agents-last-exam"></a>

Professional workflows
**CLI · used on this site**

### Agents’ Last Exam

A broad, expert-built programme of long-horizon professional computer tasks with verifiable outcomes across many industries and specialist applications. The CLI route runs each task in its declared container and scores it with the benchmark's official evaluators.

**Use for** — Economically meaningful end-to-end agent work
**Watch** — Task version, harness, context, effort, output budget and task-specific timeout must match before comparing results

- [See GLM‑5.3's score and protocol note →](https://isaiuseful.com/local-models.html.md#glm)

- [Explore tasks and traces →](https://agents-last-exam.org/)

- [Inspect Z.ai's GLM‑5.3 evaluation settings →](https://z.ai/blog/glm-5.3)

Live economic outcome
**Real-money tracker**

### The GPT Investor Portfolio

A public ledger tracks time-stamped stock selections attributed to autonomous agents, their dollar and percentage returns and the return of SPY over the same stated interval.

**Use for** — Inspecting a longitudinal, real-money outcome record instead of a simulated finance quiz
**Watch** — Models, prompts, dates, holding periods and portfolio construction differ; the publisher controls the record, so this is not a controlled model comparison or financial advice

- [Inspect the live portfolio and holdings →](https://www.gptinvestor.co/the-gpt-investor/)

Vision + knowledge
**Expert tasks**

### MMMU / MMMU-Pro

College-level multimodal questions across disciplines using charts, diagrams, images and text; Pro tightens robustness against shortcutting.

**Use for** — Visual reasoning with domain knowledge
**Watch** — Exam-style accuracy does not measure document-workflow reliability

- [Open MMMU →](https://mmmu-benchmark.github.io/)

Video generation
**Open evaluation suite**

### VBench / VBench 2.0

An open video-generation evaluation suite spanning technical quality and intrinsic faithfulness, with prompt suites, metrics, code and a leaderboard.

**Use for** — Structured comparison of video-generation systems
**Watch** — Automated dimensions and standardized prompts are proxies; sampling, settings and model version matter, and the local custom-input subset is narrower than the full standard suite

- [VBench source and suite →](https://github.com/Vchitect/VBench)

- [VBench leaderboard →](https://huggingface.co/spaces/Vchitect/VBench_Leaderboard)

Creative preference
**Crowdsourced pairs**

### Design Arena

A crowdsourced pairwise benchmark across image, video, editing, web and design work and other creative outputs.

**Use for** — Exploring current human preference across creative-output routes
**Watch** — Subjective preference, a live model pool, stochastic routing, low-vote entries and prompt enhancement mean this is not correctness

- [Design Arena leaderboard →](https://www.designarena.ai/leaderboard)

- [Design Arena methodology →](https://notes.designarena.ai/methodology/)

Visual language
**Holistic suite**

### VHELM

Stanford’s living vision-language evaluation extends HELM’s transparent, multi-metric approach to models that reason over images and text.

**Use for** — Broad visual-language comparison
**Watch** — Aggregate coverage still needs a local image and document set

- [Explore VHELM →](https://nlp.stanford.edu/helm/vhelm/)

<a id="minebench"></a>

Text → 3D world
**Live human arena**

### MineBench

Models read a natural-language build prompt and emit raw voxel-block coordinates. The site renders both worlds and humans vote blind to produce an Elo ranking.

**Use for** — Spatial composition, instruction following and inspectable creative output
**Watch** — Human aesthetic preference, prompt mix and generation budget are not geometric correctness

- [Vote, explore builds or use the sandbox →](https://minebench.ai/)

<a id="voxelbench"></a>

Voxel generation
**Community leaderboard**

### VoxelBench

A related benchmark evaluates language models on making voxel builds from text prompts and publishes a live leaderboard, keeping the generated world as the inspectable artifact.

**Use for** — Text-to-voxel build comparison
**Watch** — Leaderboard details alone do not expose a fully reproducible evaluation harness

- [Inspect the live VoxelBench leaderboard →](https://voxelbench.ai/leaderboard)

Safety evaluation
**Multiple risks**

### HELM Safety

A transparent safety-evaluation route spanning multiple risk categories, models and scenarios rather than one refusal rate.

**Use for** — Structured safety and risk comparison
**Watch** — Public prompts can be trained against; deployment permissions still matter

- [Open HELM Safety →](https://crfm.stanford.edu/helm/safety/latest/)

<a id="jailbreakbench-harmbench"></a>

Jailbreak robustness
**Used on this site**

### JailbreakBench + HarmBench

JailbreakBench standardizes threat models, behavior sets, attack artifacts, judges and an attack/defense leaderboard; HarmBench adds a broader pipeline for comparing automated red-team methods, target models and robust-refusal defenses.

**Use for** — Reproducible attack success and defense comparisons under a named protocol
**Watch** — Results are dual-use and judge-dependent; public attacks invite overfitting, while low attack success can also mean unhelpful over-refusal rather than safe behavior

- [JailbreakBench project and leaderboard →](https://jailbreakbench.github.io/)

- [HarmBench framework →](https://github.com/centerforaisafety/HarmBench)

- [Find authorized red-team tools →](https://isaiuseful.com/tools.html.md#tools-fine-tune-evaluate-and-reproduce)

<a id="mlperf"></a>

Hardware + serving
**Audited submissions**

### MLPerf

Industry-standard training and inference suites compare systems under defined scenarios, including throughput, latency and power submissions.

**Use for** — Architecture-neutral hardware and systems procurement
**Watch** — Submitted configurations may be heavily optimized and unlike your stack

- [See the rack-to-workstation transfer limit →](https://isaiuseful.com/dgx-station.html.md#reviews)

- [Browse MLPerf results →](https://mlcommons.org/benchmarks/)

<a id="gpu-battle-can-you-run"></a>

Model + VRAM fit
**Practitioner route**

### GPU Battle: Can You Run It?

A third-party practitioner route for model-to-VRAM fit plus recorded throughput and efficiency evidence across LLM, image, video, embedding and related AI workload families—not a universal hardware ranking or endorsement.

**Use for** — Shortlisting a model and GPU configuration before a hands-on run
**Watch** — Fit and headline values depend on the exact model, quantization, runtime, context, settings and test conditions

- [Check model-to-VRAM fit →](https://gpubattle.com/can-you-run)

<a id="gpu-battle-ai"></a>

AI hardware comparison
**Practitioner route**

### GPU Battle: AI GPU Benchmarks

A third-party cross-card AI hardware comparison table for LLM tokens per second, image-generation performance and related measures—not a universal ranking or endorsement.

**Use for** — Scanning published cross-GPU evidence while building a hardware shortlist
**Watch** — Retain the table’s measured-versus-estimated labels; estimated and measured entries are not equivalent procurement evidence

- [Inspect AI GPU benchmarks →](https://gpubattle.com/ai)

<a id="gpu-battle-guides"></a>

Fit + ownership guides
**Practitioner route**

### GPU Battle: Buyer’s Guides

A third-party route for VRAM-specific model-fit guides, buy-versus-rent break-even paths and deeper hardware explainers—not a universal purchase recommendation or endorsement.

**Use for** — Framing a buy-versus-rent decision after a workload has passed acceptance
**Watch** — Prices, rental rates, availability, utilization and regional electricity and tax assumptions change; recompute with your workload and quote

- [Read the buyer’s guides →](https://gpubattle.com/guides)

Inference energy
**Measured systems**

### ML.ENERGY

A benchmark and leaderboard for measuring inference energy under realistic service environments across models, tasks and system choices.

**Use for** — Energy-aware serving and optimization
**Watch** — Grid carbon, utilization and workload mix remain site-specific

- [Open ML.ENERGY →](https://ml.energy/)

<a id="harness-efficiency"></a>

Coding-harness efficiency
**Used on this site**

### Nawk Harness Efficiency

Holds one locally served DeepSeek V4 Flash configuration, eight repository bug fixes and one grading method constant while comparing Pi, OpenCode, Claude Code and Nanocoder on quality, generated tokens and wall-clock time.

**Use for** — Seeing scaffold cost, work style and run-to-run noise when the model stays fixed
**Watch** — One practitioner’s codebase, model and eight-task distribution; run counts differ and the study is not peer reviewed

- [Explore the wall-clock plot →](https://isaiuseful.com/tools.html.md#harness-efficiency)

- [Read the method and download route →](https://nqawhc.github.io/articles/harness-efficiency-not-quality/)

<a id="acbench"></a>

Compression effects
**Used on this site**

### ACBench

The Agent Compression Benchmark tests how quantization and pruning change workflow generation, tool use, long-context understanding and real-world application behavior.

**Use for** — Choosing a smaller or quantized build without assuming task parity
**Watch** — Compression effects vary by model, method and task; rerun your exact package

- [Apply it to the local-model role bands →](https://isaiuseful.com/local-models.html.md#roles)

- [Read the ACBench paper →](https://arxiv.org/abs/2505.19433)

<a id="harness-demo"></a>

Simple agent harness demo

## The model is only one layer of the test.

Turn on the controls and watch an ungrounded answer become a reviewable workflow. This deterministic demo runs entirely in your browser.

Task Find the renewal date in a contract and calculate the last day to give 60 days’ notice.

Trace
**Not run**

1. Choose controls, then run the task.

No reviewable answer yet.

<a id="failure-modes"></a>

Benchmark traps

## A leaderboard can be correct and still mislead you.

Treat the score as a measurement produced by a dataset, prompt, harness, budget, judge and date—not as a property floating inside the model.

01 · Contamination

### The test leaked into training.

Public static questions can become training data. Prefer held-out, rolling or newly collected tasks and record the cutoff.

02 · Saturation

### Everyone clusters near the ceiling.

A benchmark above roughly 95% no longer separates frontier systems well. Retire it or add harder, fresher cases.

03 · Harness lift

### The scaffold won the benchmark.

Search, retries, tools, context construction and verification can dominate agent scores. Name and version the complete system.

04 · Judge bias

### The evaluator likes a style.

Automated judges can reward length, confidence or familiar phrasing. Use human checks and position-swapped comparisons.

05 · Hidden cost

### A high score used far more inference.

Report tokens, reasoning level, attempts, latency and dollars per task. Capability without efficiency is an incomplete result.

06 · Distribution shift

### The benchmark is not your workflow.

A coding or exam score may not survive your documents, languages, tools, error costs or user population. End with local acceptance tests.

Measurement guidance: [Stanford AI Measurement Science](https://aimslab.stanford.edu/textbook/src/chap13.html) . A 2025 audit of SWE-bench scoring reported corrections affecting 24.4% of Verified leaderboard entries and changing 11 rankings—useful evidence that benchmark infrastructure also needs verification: [UTBoost paper](https://aclanthology.org/2025.acl-long.189/) .

The final benchmark
A model earns deployment by passing your representative work at an acceptable cost and failure rate.

- [Choose a public benchmark](#database)

- [Build a training evaluation loop](https://isaiuseful.com/training-models.html.md#evaluation)

- [Compare workplace studies](https://isaiuseful.com/evidence.html.md#study-map)

- [Set a robotics field gate](https://isaiuseful.com/robotics.html.md#scoreboard)
