Evidence before enthusiasm

Is AI Useful? What the Evidence Actually Shows

The strongest case is not that AI is perfect. It is that bounded use can change speed, quality or access in controlled studies—then must earn its place again in the real workflow, with review, security, incidents and total cost counted.

Study map

How much does AI improve productivity?

These bars share a visual direction, not a common outcome measure. Throughput, elapsed time, task accuracy and firm productivity are different quantities and should not be ranked against one another.

Customer supportIssues resolved per hour5,179-worker field study
+14%
Professional writingCompletion timerandomized experiment
−40% time
Consulting · inside frontierCompletion time758-person field experiment
−25.1% time
Bounded coding taskCompletion speedvendor-authored controlled trial
+55.8%
Experienced OSS developersIssue completion timerandomized trial · early-2025 tools
19% slower
Firms in four countriesReported productivity over three yearscross-country business surveys
+0.29%

What does the evidence prove—and what does it not prove?

This claim boundary keeps a measured workflow result from silently turning into a universal productivity or investment claim.

Established

Some bounded tasks improve

Multiple studies report gains in speed, throughput or evaluated quality for defined populations and tasks.

Established

Tools expand capability

Retrieval, terminals, applications, tests and durable state let systems do more than generate conversational text.

Limited

Effects vary by task and user

Benefits often concentrate inside a model’s capability boundary. Review cost and severe failures can erase average gains.

Not established

Sector-wide investment returns

These studies do not prove that aggregate capex, every supplier, a fixed local-cloud split or any particular valuation will pay off.

Evidence reading guide

What kind of evidence is this—and what can it carry?

Evidence has two separate qualities: whether the comparison supports a causal claim, and whether the work resembles the decision in front of you. A rigorous study of another workflow may be less decision-useful than a modest but well-run local trial.

highest local relevance

Prospective local comparison

Compare AI-assisted and current work on representative cases, with the same quality threshold, recorded reviewer time and precommitted incident rules. This is the best evidence for a particular purchase decision, provided the comparison is fair.

independent academic evidence

Randomized or well-designed field study

Random assignment and real work make a study strong evidence about its named population, task, tool and period. It does not automatically transfer to another job, model version, team or operating environment.

public project report

Operational evidence with visible methods

A public service can reveal the measures, safeguards and failure handling used in practice. It may show that a system operated at scale, but it is usually not a controlled estimate of staff productivity or return on investment.

vendor-authored controlled study

Useful design, interested publisher

Read the population, task, comparator, exclusions and scoring method. A controlled design can be informative, but sponsorship, product selection and lack of independent replication limit the claim it can settle.

vendor customer story

Implementation clue, not causal proof

A customer story can identify a workflow, integration pattern or metric worth testing. It normally cannot establish that the product caused a reported outcome: customers are selected, counterfactuals are absent and unreported costs may matter.

benchmark or demonstration

Capability signal, not business value

A model can pass a benchmark or complete a polished demo yet fail on permissions, exceptions, latency, review burden or adoption. Use it to shortlist candidates, then test the whole system on real work.

Method anchor: NIST’s AI Risk Management Framework sets out documented task scope, expected benefits and costs, deployment-like evaluation and production monitoring; it notes that independent review can reduce internal bias. It is voluntary guidance, not evidence that any specific tool is safe or worthwhile.

Which AI tasks have measured outcomes?

Randomized experiments and field studies are the highest-signal evidence here. Each card names the study design and provenance; vendor reports are not presented as independent replication.

field study · 5,179 workers

Customer-support agents handled more work

Access to a generative-AI assistant increased issues resolved per hour by about 14% on average, with the largest gains among less experienced and lower-skilled workers.

Academic field evidence: NBER, Generative AI at Work.

randomized experiment

Professional writing became faster and better

In mid-level writing tasks, ChatGPT reduced completion time by roughly 40% and raised evaluator-scored quality by about 18%, while compressing performance differences between workers.

Peer-reviewed experiment: Science, Noy and Zhang.

field experiment · 758 consultants

Consultants improved inside the model's frontier

For tasks within GPT-4's capability boundary, consultants completed more work, worked faster and produced higher-quality outputs. On a task outside that boundary, AI users were more likely to be wrong.

The useful result includes the failure: Harvard Business School working paper.

controlled developer trial

A coding task was completed 55.8% faster

Developers with GitHub Copilot completed a bounded JavaScript task substantially faster than the control group. This measures one task, not all software engineering; the reviewable software workflow shows how to carry that boundary into real repository work.

Vendor-authored controlled study: Microsoft Research.

controlled study · 202 developers

Code quality improved modestly

A randomized study reported higher test completion and modest improvements in readability, reliability, maintainability and reviewer approval for Copilot-assisted code.

Vendor-produced RCT: GitHub research.

vendor/partner RCT

An enterprise trial reported workflow gains

GitHub and Accenture reported 8.69% more pull requests per developer and a 15% higher pull-request merge rate among Copilot users working on routine engineering tasks.

Vendor and partner report—not an independent replication: study description and results.

randomized tutoring trial

Human tutors improved with real-time coaching

Tutor CoPilot gave tutors live suggestions during sessions. In a preprint covering 900 tutors and 1,800 K–12 students, the randomized evaluation reported a 4-percentage-point mastery gain overall and 9 points for students of lower-rated tutors.

Research preprint: Stanford SCALE Initiative.

single-group QI study · caution

Draft replies showed no measured time savings

In a five-week deployment with 162 clinicians, AI drafts were used for 20% of replies and did not change reply, write or read time. Among the 73 clinicians completing pre/post surveys, task-load and work-exhaustion scores fell; without a control group, the study cannot attribute those changes to AI.

Prospective single-group quality-improvement study: JAMA Network Open.

research infrastructure

Structure prediction is available as research infrastructure

AlphaFold DB provides public predicted protein structures. Separately, the peer-reviewed AlphaFold 3 paper reports joint structure prediction across proteins, nucleic acids, small molecules, ions and modified residues.

Peer-reviewed system: Nature, AlphaFold 3 · official public resource: AlphaFold DB.

Outcomes ledger

Measured results answer different questions in different domains.

Do not average these figures into one “AI productivity” number. The useful comparison is the measure, the task boundary and the missing part of the business case.

WorkflowWhat was measuredWhat the result can supportWhat remains unproven
Customer supportIssues resolved per hour in a 5,179-worker field study.A task-specific throughput gain with uneven benefits across workers.Retention, customer outcomes, error cost and results in another support operation. Academic field study.
Professional writingCompletion time and evaluator-scored quality in a randomized experiment.For similar writing tasks, a tool can improve both speed and scored output.Whether the gain survives brand, legal, factual or approval requirements. Peer-reviewed experiment.
ConsultingTask completion, time and quality for 758 consultants.The capability frontier is a practical operating boundary, not a marketing phrase.That an untested task is inside the frontier; outside it, assisted users were more often wrong. Academic field experiment.
Software engineeringA vendor trial timed one bounded task; a later randomized METR study timed experienced developers on familiar repositories.Coding effects can reverse with task, repository familiarity and tool generation.A universal developer-speed claim or a release-quality claim without review and production data. Vendor controlled trial · Independent randomized study.
TutoringStudent mastery in a randomized evaluation of real-time tutor coaching.A human-in-the-loop tool may improve a learning outcome, especially where baseline performance is lower.Long-term attainment, implementation cost and results for every subject or tutor population. Research preprint.
Clinical messagesReply, writing and reading time in a five-week, single-group quality-improvement deployment.Adoption and perceived workload changes can coexist with no measured timing gain.Causal wellbeing effects or time savings without a control group. Prospective QI study.
Public information serviceAccuracy, reliability, speed, safety and user trust across two GOV.UK Chat pilots.A constrained, source-grounded public assistant can be evaluated with expert review, automated tests and user research.A controlled productivity or ROI result; the report is an official project evaluation. GOV.UK public project report.
Distribution and frontier

The average gain can hide the decision-critical result.

Ask who improved, which cases failed, and whether the task truly matches the evaluated workflow. A mean effect is not a guarantee for the most experienced user, the edge case or the next model release.

worker distribution

Gains may concentrate where support is scarce.

In the customer-support field study, less experienced and lower-skilled workers gained more. That can narrow a performance gap, but a local evaluation should report results by experience, job type and case difficulty rather than only a team average.

Academic field evidence: NBER, Generative AI at Work.

capability frontier

A useful tool can increase error outside its tested boundary.

The consulting experiment found gains on tasks inside GPT-4’s frontier and a lower correct-answer rate on a deliberately outside-frontier task. Treat every new workflow as an empirical question, especially when an answer triggers a material action.

Academic field experiment: Harvard Business School working paper.

perception versus measurement

Confident users can misread their own speed.

METR’s experienced open-source developers expected the tools to speed them up, but completed issues more slowly in the study. Time logs, accepted output and review cost are more reliable purchase evidence than a post-pilot enthusiasm survey alone.

Independent randomized study: METR’s early-2025 developer study.

Reliability is a system property

A correct-looking answer is not yet a safe result.

A model score, a human approval rate and a safe operational outcome are different measures. The system includes the prompt, retrieval, tools, permissions, reviewer, handoff and recovery path.

Define the failure

Measure the error that matters.

Record factual defects, missed requirements, unsafe tool calls, privacy exposures and harmful delays separately. “Helpful” or “accurate” needs a task-specific rubric and an escalation path.

Price the review

Review can move rather than remove work.

Time to detect, correct and explain an error belongs in the numerator. Blind or independent review helps distinguish a convincing draft from an accepted outcome.

Constrain authority

Untrusted content must not become an instruction.

When a system reads web pages, email or documents, test prompt injection and data leakage before granting tools. Keep authorization outside the model and give actions narrow, reversible permissions.

Plan recovery

Every risky action needs a stop and rollback path.

Log inputs, source material, tool calls and approvals; make it possible to disable the feature, revoke access, correct the record and notify an owner. Human review is a control to test, not a magic guarantee.

Standards and security guidance: NIST says evaluations should cover deployment-like conditions, reliability, safety, security, privacy and monitored production behavior in its AI RMF measure function. The community-maintained OWASP GenAI LLM Top 10 identifies prompt injection, sensitive-information disclosure, improper output handling and excessive agency as application risks. These sources identify risks and controls to test; they do not quantify failure rates for a particular deployment.

Negative and null results

When does AI fail to improve the work?

Failures belong beside the gains because they identify the task boundary, review burden and organizational changes hidden by an average score.

Randomized trial · slowdown

Experienced developers took 19% longer.

In a METR study, experienced open-source developers working in repositories they knew completed issues more slowly with early-2025 AI tools—even though they believed the tools had sped them up.

Read the study and scope limits →
Field experiment · wrong answer

AI users crossed the jagged frontier.

On a consulting task selected outside GPT-4’s capability boundary, AI-assisted participants were 19 percentage points less likely to produce the correct answer.

Read the full working paper →
7,137 workers · limited shift

Two fewer email hours did not redesign the job.

In a six-month field experiment across 66 firms, active users spent about two fewer hours on email each week, but researchers did not detect broader shifts in task quantity or composition from individual tool access.

Read Shifting Work Patterns with Generative AI →
Four-country firm surveys · weak aggregate

Most firms reported no productivity impact yet.

Across surveys in the United States, United Kingdom, Germany and Australia, 89% of firms reported no productivity effect over three years; the estimated average reported gain was about 0.29%.

Inspect the firm evidence and assumptions →
Clinical deployment · null timing

Draft replies saved no measured time.

In a five-week deployment with 162 clinicians, AI drafts were used for one in five replies but did not change measured reply, writing or reading time.

Read the quality-improvement study →
Autonomous agent · operating loss

The shop ran—and made costly mistakes.

Anthropic’s Project Vend showed real inventory and customer-interaction capability, but the agent also discounted too aggressively, made poor purchasing decisions and was manipulated by users.

Read the vendor research report →

Negative results are population- and date-specific. The METR trial does not establish that AI slows most developers; the firm surveys do not establish that future productivity will remain small. They establish that benchmark strength and user enthusiasm are not substitutes for workflow measurement—or for the staged proof gates in the organizational adoption journey.

What can public AI systems do today?

These project reports, vendor demonstrations and practitioner methods document capability or implementation patterns. They are not measured productivity evidence unless a card explicitly says so.

practitioner evidence

Code can be delivered with proof

Simon Willison's practical rule is that generated code should arrive with tests, command output or a reproducible demonstration. The harness, not model confidence, establishes completion.

Code proven to work · NICAR notes.

agent engineering

Reliable agents are systems, not prompts

Anthropic distinguishes predictable workflows from autonomous agents and documents routing, tool use, evaluator loops and human escalation as composable patterns.

Primary engineering guidance: Building effective agents.

long-running work

Agents can preserve progress across context windows

A documented harness uses an initializer, progress artifacts, git history and verification to let later agent sessions resume a project instead of restarting from a blank prompt.

Primary implementation report: Effective harnesses for long-running agents.

computer use

A model can operate existing software interfaces

Computer-use systems inspect screenshots, move a pointer, click and type. They are slower and riskier than direct APIs, but they can automate software that has no modern integration surface.

Vendor-produced capability release: Anthropic computer use.

deep research

Research agents can gather and cite sources

Deep-research systems browse, inspect many sources, synthesize findings and return citations. The citations still need checking, but the workflow goes well beyond conversational recall.

Vendor-produced system report: OpenAI deep research.

bounded autonomy

An agent ran a shop and exposed concrete limits

Project Vend asked an agent to operate a small store. It handled inventory and customer interaction, but also made costly mistakes. That makes it useful evidence about both capability and control.

Vendor research report: Anthropic Project Vend.

practitioner method

A durable loop needs state and independent checking

Loop-engineering practitioner patterns combine scheduled runs, durable state, isolated worktrees, scoped tools, a separate verifier and human review. This is an implementation method, not a controlled outcome study.

Practitioner references: Loop Engineering repository and Addy Osmani's overview.

Local proof before purchase

Run an evaluation that can say “do not proceed.”

A short local comparison will not prove a universal return. It can establish whether a bounded workflow clears a stated quality, cost and risk bar for your organization—and whether a wider rollout is justified.

1 · frame the decision

Name one workflow and one unit of completed work.

Specify the user, input, output, downstream action, current route and decision owner. Exclude work that cannot be safely sampled. Write the quality threshold and risk tolerance before seeing AI results.

2 · build a fair baseline

Time the current work end to end.

Capture preparation, searching, writing, review, rework, handoff and waiting—not only the time a person spends typing. Record direct licence, model, integration and reviewer costs in the same unit.

3 · sample the work honestly

Include ordinary, difficult and failure-prone cases.

Use a dated holdout of real or safely de-identified cases. Stratify by case type, complexity and user group. Do not let a vendor choose only polished prompts, friendly documents or successful outputs.

4 · score accepted output

Measure quality before speed.

Use domain reviewers where the result needs expertise. Score correctness, completeness, policy compliance and source support; record the percentage accepted unchanged, accepted after correction, rejected and escalated.

5 · count the full operating cost

Put review, exceptions and latency in the ledger.

Compare time to accepted work, reviewer minutes, retry rate, queue delay, token or licence cost and integration support. Separate a faster first draft from a faster finished job.

6 · gate a limited rollout

Monitor, learn and retain an exit.

Start with a low-authority cohort. Log incidents and near misses, inspect drift by case and user, and keep a manual fallback. Re-evaluate after model, prompt, retrieval, tool or policy changes.

ScorecardMinimum recordDecision use
Quality and safetyAccepted-output rate; error categories and severity; escalation rate; policy, privacy and security incidents; reviewer agreement.Do not trade a small speed gain for errors that exceed the workflow’s risk tolerance.
Time and costMedian and spread of time to accepted work; reviewer minutes; retry/rework; licence, model and integration cost per accepted item.Compare completed work, not a draft or a self-reported time saving.
Distribution and adoptionResults by case type and user group; opt-out, abandonment and override rates; training and support demand.Find who benefits, who is burdened and which cases require the old route.
Operations and rollbackTool-call logs, source traceability, incident owner, time to disable, correction path and manual fallback exercise.Decide whether the system can fail safely enough to expand its authority.

Precommit a stop rule: pause or roll back if quality falls below the current threshold, a severe incident occurs, review cost erases the benefit, or the result fails in a named high-risk case.

There is no universal pass percentage. The defensible threshold is the one the workflow owner can explain in relation to the existing process, affected people, recoverability and the cost of being wrong. A rollout decision should retain the baseline, sample definition, scoring rubric, exclusions and all observed incidents—not only the average gain.

Protocol basis: NIST’s voluntary AI RMF sets out documented task scope, benefits and costs, representative deployment-like evaluation, risk tracking and production monitoring; it notes where independent assessment can reduce internal bias. This page adapts those principles into a practical local test; it is not legal, regulatory or safety certification advice.

The careful claim is not “AI never fails.” It is: bounded AI work can already produce measurable value, and verification determines how safely that value scales.

Vendor studies and product reports are labelled. They are useful evidence, but not substitutes for independent replication or measurements on your own workflow.

Evidence becomes useful in a workflow

Choose a task with a manual baseline, a reviewable output and a visible definition of done.