More issues resolved per hour
AI assistance for 5,172 support agents. Largest gains among less experienced workers.
QJE · workplace field study ↗For some tasks, the value is already measurable. Faster writing. More support issues resolved. Better tutoring. Whether those gains become profit depends on the full cost of getting the work right.
The evidence, the limits, and the economics beyond the chatbot.
By isaiuseful.com · Updated
Measured outcomes
Six concrete results, from everyday work to scientific infrastructure.
AI assistance for 5,172 support agents. Largest gains among less experienced workers.
QJE · workplace field study ↗Professional writing finished faster, with higher evaluator-scored quality.
Science · randomized experiment ↗A bounded JavaScript task with Copilot. This is a task result, not all engineering.
Microsoft Research · controlled trial ↗Human tutors received real-time AI coaching. Gain measured in percentage points.
Tutor CoPilot · randomized trial ↗Open research infrastructure from AlphaFold DB. Predictions still need validation.
EMBL-EBI + DeepMind · public database ↗AlphaEvolve improved data-centre scheduling, according to Google's production report.
DeepMind · vendor-reported deployment ↗Real usefulness. Different kinds of evidence. These measures cannot be added or averaged into an AI ROI. They include independent research, vendor studies and public infrastructure.
Sources and limits: isaiuseful.com/evidence.html · Reviewed 14 Sep 2026Time on task
How much time did the AI-assisted group take? Each study's unaided group is indexed to 100. Lower means less time.
Read the task boundary, not just the average. Separate experiments, tools and populations; point estimates, not forecasts. The bars share a time index, not a common experiment. Figma's estimate averages three tasks among successful completions; benefits differed by role and task. METR's 2026 follow-up could not reliably quantify the current effect.
Sources: Microsoft Research; Science (Noy + Zhang); HBS (Dell'Acqua et al.); Figma (Stewart et al., September 2026); METR.From usefulness to profitability
It becomes profit when the value of accepted work exceeds model, integration, review and failure costs. These studies establish useful capabilities and task-specific gains. They do not establish sector-wide profits or investment returns.
The question is bigger than whether people will pay for a chatbot. It is how much useful work AI can deliver, what that work is worth, and who captures the value.
AI can be profitable for a business using it. That does not mean every deployment pays off, every model developer makes money, or every data centre is a good investment. The useful test is acceptable work at a lower full cost than the alternative.
Tool-using agents expand what can be evaluated: they can inspect repositories, edit files, run commands and tests, and prepare changes for review. Codex is one documented example. Those capabilities make the business case broader, while keeping verification central.12
isaiuseful.com · Editorial analysis · Sources checked 14 September 2026
The question “Is AI profitable?” bundles together several different businesses.
For the organisation using AI, the question is whether accepted output, genuinely avoided expenditure or additional profitable business outweighs implementation and operating costs.
For the company serving a model, the question is whether customer revenue covers delivery: hardware depreciation, power, networking, licensing and operations at realistic utilisation. Positive margins at this layer do not automatically cover the cost of developing the model.
For the frontier-model developer, training, research, staff and the rest of the business also need funding. Those costs are real. Describing inference as profitable does not make them disappear.
For the investor, even a useful technology and a profitable business can be a poor purchase at the wrong price.
There is no contradiction in an adopter saving money while its supplier loses money overall. There is also no guarantee that today's favourable customer economics survive future price changes. Both statements belong in the analysis.
But one company's losses do not tell you whether another company's workflow is worthwhile. An enterprise running an available open-weight model does not have to repeat the original research programme; its own deployment still needs to cover licensing where applicable, hardware or hosting, integration and operation.
Economic value, supplier profitability and investment returns are connected. They are not interchangeable.
A job is a bundle of tasks. Replacing every responsibility attached to a job title is not a prerequisite for changing how many people are needed to deliver the work.
The International Labour Organization's 2025 occupational-exposure research explicitly examines tasks and distinguishes job transformation from complete automation. Its finding that human involvement remains necessary is not a finding that staffing levels must remain unchanged.3
Consider a deliberately simplified example. A ten-person team performs 1,600 hours of work each month. Automation removes 800 hours, while checking results, managing exceptions and correcting mistakes adds back 160. The remaining workload is 960 hours: six people's worth rather than ten.
That is an illustration, not a forecast. It assumes the remaining work can be redistributed, the necessary skills are available and output requirements stay constant. Under those conditions, four positions could disappear without the occupation disappearing. Alternatively, the business could retain the team and produce more.
This also explains why “AI only assists the worker” is not an economic rebuttal. Assistance can reduce labour requirements, avoid future hiring, increase capacity or improve the service delivered.
None of that establishes that half of all jobs will disappear. Nor does half the task count necessarily mean half the working time. It establishes something more useful: task-level automation can have large economic effects long before whole-job replacement becomes possible.
A study published in The Quarterly Journal of Economics examined 5,172 customer-support agents and found a 15% average increase in issues resolved per hour with AI assistance. Effects varied across workers. This is evidence of workplace productivity, not an audited calculation of the employer's net AI profit - but it measures delivered work rather than enthusiasm.4
There are negative results, too. METR's early-2025 experiment found that experienced open-source developers took 19% longer on the studied tasks when allowed to use AI. That result matters: review, correction and workflow friction can outweigh the assistance.5
The update matters as well. In February 2026, METR reported that its subsequent experiment faced substantial selection effects, including developers declining work without AI, and difficulties measuring concurrent agent use. The researchers believed speedups had probably improved but explicitly said their data could not reliably establish the magnitude. That does not invalidate the earlier finding or prove a universal new speedup.6
There is also evidence of agents producing operational improvements. Google DeepMind reported that a scheduling heuristic discovered by AlphaEvolve, and deployed in production, recovered 0.7% of Google's worldwide compute resources on average. This is a vendor-reported result, not independently audited profitability. Nevertheless, the claimed output is additional usable infrastructure capacity - not a better-sounding paragraph.7
The conclusion is neither “AI always saves money” nor “AI never saves money.” Useful work is measurable. So are the costs that can erase its value.
The economic question is not merely:
How many people will pay for a chatbot subscription?
It is:
How much useful work can these systems deliver, and what is that work worth after accounting for compute, integration, supervision and failures?
A practical calculation is:
Net AI value = realised benefits − model and infrastructure costs − integration and maintenance − review and rework − expected failure costs.
Benefits must be counted without double-counting. An hour cannot simultaneously be claimed as eliminated payroll and as an hour an existing employee uses to generate additional revenue. Saved time is not automatically cash saved. Additional revenue is not all profit. Severe risks also need acceptance limits, not just an optimistic average cost.
Suppose, as an illustration, an AI-assisted process genuinely avoids €2,000 of external spending each month. Its model, review and maintenance costs total €700, and implementation costs €3,900. It produces €1,300 of monthly net savings and recovers the initial cost in three months, assuming the savings persist and no material costs have been omitted.
The supplier's company-wide accounts cannot answer that customer-level calculation. Equally, the customer's savings cannot prove the supplier's business is sustainable.
This is why AI should be compared with the cost of the work it delivers - not confined to the price someone will pay for a conversational interface.
A tool-using agent can select an action, invoke a tool, examine the result and choose its next step. In software, a compiler or test runner supplies feedback from outside the language model. That does not guarantee correctness, but it is materially different from accepting an unverified answer from a chat window.8
Now compare two hypothetical instructions:
“Help me fix this bug.”
“Within this budget, investigate incoming failures, propose fixes, run the required checks and send me the changes that need approval.”
The second can generate work without requiring a person to initiate every investigation. One bounded objective becomes a queue of tasks. The owner still sets priorities and authority; the system performs more of the intermediate execution.
Anthropic's June 2025 engineering report provides a concrete indication of the demand difference: in its data, agents typically used about four times as many tokens as chat interactions, and multi-agent systems about fifteen times as many. These are deployment-specific observations, not universal multipliers or direct measures of GPU consumption. Anthropic also emphasised that the task's value must justify the extra cost.9
The point is not that burning more tokens creates value. It is that counting people and subscriptions can miss how much computation a useful objective generates.
An organisation does not necessarily need more employees - or more people staring at chat windows - to create more demand for machine-executed work.
Imagine an operations agent identifying a problem, asking a specialist system to investigate, commissioning a simulation and submitting a verified recommendation for approval. With suitable permissions and budgets, some services could be purchased rather than manually integrated and ordered each time.
This is a scenario, not a claim that fully autonomous commerce is already mature. But its technical ingredients exist. The A2A project released version 1.0 of its agent-communication protocol in March 2026. Google Cloud's April 2026 platform announcement also describes Agent Payments Protocol integration and payment mandates.1011
Interoperability and payment infrastructure do not establish the size of the resulting economy. They make a particular kind of economy easier to build: one in which software can delegate work to other software across organisational boundaries.
The important expansion is not merely doing today's work more cheaply. It is doing work that was previously too expensive to commission.
Consider a small business that cannot justify a bespoke internal application, daily operational analysis or extensive product experimentation. If reliable delivery becomes sufficiently inexpensive, some of those activities become viable. Those are illustrative possibilities - not additional spending that can be assumed before customers obtain results.
There is a stopping rule: the next unit of computation must be worth more than it costs. Agents repeatedly buying services from one another do not create net value merely by generating transactions. But successful applications can generate income that funds further computation; today's software budgets are not necessarily a permanent ceiling.
A common objection to AI infrastructure investment is that models will become cheaper to run. That is a reason to examine the demand response, not assume it disappears.
Stanford's 2025 AI Index documented a more than 280-fold decline in query prices for models reaching approximately GPT-3.5-level performance on MMLU between November 2022 and October 2024. This was a comparison of API prices at a benchmark threshold - not audited production costs or equivalence across every real-world capability.12
The economic mechanism is straightforward. A task that is not worth doing at €100 might be worth doing at €5. Lower prices can attract additional users, increase the frequency of a workflow and make new categories of work viable.
That does not make demand literally infinite. For comparable tasks:
Total compute consumption = completed task volume × compute required per completed task.
If task volume grows twentyfold while compute per task falls fivefold, consumption grows fourfold. Reverse those changes and consumption falls. Efficiency alone does not determine the outcome.
Nor must efficiency savings be spent on doing precisely the same task. They could fund broader searches, more candidate solutions, additional tests or work on harder problems. Whether those are worthwhile has to be measured.
The strongest demand argument is not “models will remain inefficient.” It is “greater efficiency and capability can make more useful work economical.”
Local and cloud AI need not be competing end states. A local model can perform routine work while a remote system handles selected difficult parts.
Research presented at ICML 2025 demonstrated one version of that arrangement. On its evaluated long-document tasks, MINIONS reduced cloud inference costs by 5.7 times on average while retaining 97.9% of remote-only benchmark performance. Those figures describe that experiment's cost and performance trade-off - not total ownership costs or a universal promise for every application.13
Now consider the wider market:
Suppose local agents become 100 times more numerous, while each requires one-tenth as much cloud assistance. With comparable cloud tasks, total cloud demand still increases tenfold.
This is arithmetic, not a forecast. It shows why “each user needs less cloud” and “the market needs more cloud” can both be true.
A cheap, private, always-available local agent could become a way to initiate work and route difficult parts to more capable remote systems. It could replace individual cloud calls while expanding the number of people and processes using cloud intelligence.
And the frontier does not have to stand still. Even suppose a device in 2030 can perform relevant tasks at the level of a 2025 frontier model. That does not establish parity with the strongest cloud system available in 2030.
Cloud capability can also come from spending more computation on search, verification and revision - not just using larger model weights. DeepMind's February 2026 Aletheia research describes such a loop and improvements from inference-time scaling on its mathematical evaluations.14
Our expectation is that the strongest cloud systems to retain an advantage on demanding work. The size and commercial value of that advantage are uncertain. A local model can be excellent at ordinary tasks without eliminating the market for more capable systems.
If a company can profitably execute additional work with agents, insufficient affordable compute can become a production bottleneck rather than merely an inconvenient subscription limit.
There is evidence of physical demand expanding despite efficiency improvements. The International Energy Agency reported that electricity consumption by AI-focused data centres increased 50% in 2025. Its April 2026 central projection has that consumption tripling between 2025 and 2030, while identifying constraints across electricity infrastructure and other supply chains. Electricity consumption is not the same as paid, useful compute, and a projection is not a guarantee.15
The implication for Europe is not “buy any GPU at any price.” It is that access to suitable compute deserves the same operational scrutiny as other important production inputs.
A generic availability figure is not enough. A customer may need a particular region, interconnect, delivery date, capacity commitment, data controls and price. Capacity that fails those requirements is not automatically a substitute.
This is also where the investment argument needs discipline. A growing market can still contain overbuilt locations, poorly matched hardware, weak contracts and operators whose financing costs exceed their returns. Useful infrastructure is not necessarily a good investment at every valuation.
Compute could become strategically essential while some compute investments fail. There is no contradiction.
AI can pay its way without eliminating entire jobs, without every provider becoming profitable simultaneously, and without a person manually requesting every piece of work.
The evidence supports measurable benefits in particular workflows, meaningful failure cases and architectures that go beyond one-shot conversation. The larger economic thesis is an inference from those building blocks - not a claim that future demand or returns have already been proved.
Here is the editorial proposition:
If increasingly capable agents continue finding valuable uses for additional computation, efficiency gains can expand demand faster than they relieve capacity constraints. Powerful on-device models may accelerate that process rather than end it.
That proposition can be tested. Do customers renew after experimentation? Does the all-in cost of accepted work fall? Can organisations convert the benefit into avoided spending, better service or profitable output? Does additional computation improve results enough to justify its cost?
Those are stronger questions than whether a chatbot can replace a whole person.
A demonstration still needs a business case. Take a representative workflow. Give a capable system the context, tools and limited authority it needs. Measure the complete result - including review, mistakes and operating costs - against the current alternative.
The market is not simply people paying to chat. It is the work that becomes worth doing when useful intelligence becomes cheaper to deploy.
OpenAI. Codex - current product overview. Accessed 14 September 2026. Vendor product documentation.↩
OpenAI. Introducing Codex. 16 May 2025; subsequently updated. Historical vendor capability documentation.↩
International Labour Organization. One in four jobs at risk of being transformed by GenAI, new ILO-NASK Global Index shows. 20 May 2025. Institutional occupational-exposure research summary.↩
The Quarterly Journal of Economics. Generative AI at Work. 2025, volume 140, issue 2, pages 889-942. Peer-reviewed workplace field study.↩
METR. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. 10 July 2025. Randomized developer productivity experiment.↩
METR. We are Changing our Developer Productivity Experiment Design. 24 February 2026. Research-methodology update.↩
Google DeepMind. AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. 14 May 2025. Vendor-reported production deployment.↩
Anthropic. Building effective agents. 19 December 2024. Vendor engineering guidance.↩
Anthropic. How we built our multi-agent research system. 13 June 2025. Vendor engineering observations.↩
A2A Protocol / Linux Foundation. A2A Protocol Ships v1.0: Production-Ready Standard for Agent-to-Agent Communication. 12 March 2026. Primary interoperability specification announcement.↩
Google Cloud. Introducing Gemini Enterprise Agent Platform, powering the next wave of agents. 22 April 2026. Vendor platform and integration announcement.↩
Stanford HAI. 2025 AI Index Report - Research and Development. April 2025. Research synthesis.↩
Proceedings of the 42nd International Conference on Machine Learning / PMLR. Cost-efficient Collaboration between On-device and Cloud Language Models. July 2025. Peer-reviewed systems research.↩
Google DeepMind. Accelerating mathematical and scientific discovery with Gemini Deep Think. February 2026. Vendor research report.↩
International Energy Agency. Key Questions on Energy and AI - Executive summary. April 2026. Institutional energy analysis and scenario projection.↩
Study populations, methods and limits. Follow each source to inspect the result in context.
Access to a generative-AI assistant increased issues resolved per hour by 15% on average, with the largest gains among less experienced and lower-skilled workers.
Academic field evidence: QJE, Generative AI at Work.
In mid-level writing tasks, ChatGPT reduced completion time by roughly 40% and raised evaluator-scored quality by about 18%, while compressing performance differences between workers.
Peer-reviewed experiment: Science, Noy and Zhang.
Within GPT-4's tested capability boundary, consultants took 25.1% less time and produced higher-quality work. Outside that boundary, AI users were 19 percentage points less likely to answer correctly.
The useful result includes the failure: Harvard Business School working paper.
David Autor and coauthors studied 133 patent lawyers at 11 U.S. firms in Google's three-month trial; 91 completed the full protocol. Independent patent attorneys graded work blind to assignment. AI access raised drafting quality by 0.34 standard deviations (SD) at 10 days and 0.38 SD at 90 days.
On a separate patent-review task without AI at day 90, senior lawyers with prior AI access scored 0.45 SD above controls. Juniors showed no detectable average gain, with more poor and more good scores. This separates assisted output from unaided judgment; it does not establish universal skill erosion.
Limits: Google conducted and funded the experiment using its own tool, and participating firms do patent work for Google. The small sample and three-month window leave longer-term skill retention unresolved. Compare the apprentices' immediate comprehension test.
Working paper, 7 October 2026: Patent-drafting experiment (PDF) · Google Research summary.
Researchers from FAU, ifo, Stanford, TUM and BIBB randomized ChatGPT Pro access among 673 final-year IT apprentices at 12 German vocational schools. Scores rose by 26-44 percentage points of the maximum available, roughly doubling performance. On two tasks beyond the curriculum, gains were 26 and 31 percentage points.
Four unaided comprehension questions after each task showed no evidence of reduced immediate understanding. Unlike the three-month patent trial, this measures immediate comprehension, not durable skill acquisition or retention.
Limits: controls had no internet, matching exam conditions. Treat the gains as a potential upper bound relative to ordinary work with online resources. Open-ended answers were AI-graded using teacher-validated rubrics, blind to assignment, with checks across grading models. Task scores do not establish whole-job productivity or net ROI.
Research preprint, preregistered as AEARCTR-16936: Task Expansion with Generative AI (PDF), sections 3.3, 3.5 and 5.
Figma researchers randomized 50 product designers and 50 product managers to three standardized UI tasks with or without Figma Make. Among successfully completed tasks, estimated completion time was about 20% lower when averaging the three tasks equally; the cumulative-time estimate was 27%.
Product managers' cumulative gains were 33-35%, depending on the model. Designers' 17% estimate was marginally significant (p=0.049), driven only by the hardest task, and was not significant in the role-interaction model.
Limits: all four authors are affiliated with Figma. Timing excludes failed tasks, and Make changed completion probability, so the groups of successful completers can differ. These bounded tasks do not establish a whole-job productivity gain or net ROI.
Research preprint, submitted 22 September 2026: Does AI Save Time on Product Design?
Developers with GitHub Copilot completed a bounded JavaScript task substantially faster than the control group. This measures one JavaScript task, not all software engineering.
Vendor-authored controlled study: Microsoft Research.
Using activity data and AI-usage telemetry for more than 500,000 GitHub developers, Demirer, Musolff and Yang estimate cumulative commit gains of 30% for autocomplete, 180% for interactive coding agents and 240% for autonomous coding agents. At the autonomous-agent stage, the cumulative gain falls to 80% for projects and 30% for actual releases.
The authors interpret this as consistent with a weak-link effect: human bottlenecks limit how much faster coding becomes shipped output. This observational matched event study is a working paper, not a randomized trial or a measure of net profit. Count accepted and released work alongside activity.
NBER w35275, revised September 2026: Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools. Disclosed interests: Mert Demirer and Leon Musolff are paid Microsoft research consultants and previously held Microsoft postdoctoral positions.
A randomized study reported higher test completion and modest improvements in readability, reliability, maintainability and reviewer approval for Copilot-assisted code.
Vendor-produced RCT: GitHub research.
GitHub and Accenture reported 8.69% more pull requests per developer and a 15% higher pull-request merge rate among Copilot users working on routine engineering tasks.
Vendor and partner report, not an independent replication: study description and results.
Tutor CoPilot gave tutors live suggestions during sessions. In a preprint covering 900 tutors and 1,800 K-12 students, the randomized evaluation reported a 4-percentage-point mastery gain overall and 9 points for students of lower-rated tutors.
Research preprint: Stanford SCALE Initiative.
UMD evaluated a GPT-4o RAG study assistant with 2,379 students and 30 instructors, randomized at instructor level. In the same-course subset (1,353 students, 22 instructors), access reduced final grades by 0.37 standard deviations, about 4 points out of 100. The adjusted full-sample grade effect was not statistically detectable; recorded learning-platform participation fell about 0.9 SD in both samples.
Estimated grade losses for first-generation students were more than twice those of their peers. Only about 15% of students offered access used it; 73.8% of requests sought direct answers, including information, explanations or solutions. The analysis plan was registered after data collection, before analysis. This evaluates access as implemented, not the effect on users alone or all AI tutoring.
EdWorkingPaper 26-1598, Tables 3, 4 and A4. Compare the human-coaching intervention above; grounding and instructional design are separate choices.
AlphaFold DB provides open access to more than 200 million predicted protein structures. These are predictions, not experimentally confirmed structures or approved drugs. Separately, the peer-reviewed AlphaFold 3 paper reports joint structure prediction across proteins, nucleic acids, small molecules, ions and modified residues.
Peer-reviewed system: Nature, AlphaFold 3 · official public resource: AlphaFold DB.
Google DeepMind reported that an AlphaEvolve-discovered heuristic recovered an average 0.7% of Google's worldwide compute resources. It described more than a year in production; the figure is vendor-reported and is not a net-profit measure.
The creator of Linux and Git described AI as useful in kernel development, including code review, while acknowledging the burden on maintainers. This is a practitioner's judgment, not a productivity trial or a profitability study.
Original mailing-list statement · Reported quotation and context. Portrait: Krd, LinuxCon Europe 2014, Wikimedia Commons, CC BY-SA 4.0. Resized and displayed with a crop.
Slowdowns, null results and operating losses belong in the same picture.
In a METR study, experienced open-source developers working in repositories they knew completed issues more slowly with early-2025 AI tools, even though they believed the tools had sped them up. METR’s February 2026 follow-up reported selection and timing problems, so it could not reliably quantify the current speedup.
In a six-month field experiment across 66 firms, active users spent about two fewer hours on email each week, but researchers did not detect broader shifts in task quantity or composition from individual tool access.
Across surveys in the United States, United Kingdom, Germany and Australia, 89% of firms reported no productivity effect over three years; the estimated average reported gain was about 0.29%.
In a five-week deployment with 162 clinicians, AI drafts were used for one in five replies but did not change measured reply, writing or reading time.
Anthropic’s Project Vend showed real inventory and customer-interaction capability, but the agent also discounted too aggressively, made poor purchasing decisions and was manipulated by users.
Negative results are population- and date-specific. The METR trial does not establish that AI slows most developers; the firm surveys do not establish that future productivity will remain small.