# isaiuseful.com full Markdown corpus
> Generated from the 18 canonical editorial HTML pages in sitemap order.
The linked HTML pages remain the source of truth. This file is a deterministic convenience export; regenerate it with `python3 scripts/generate-agent-markdown.py`.
## https://isaiuseful.com/ (`/index.html.md`)
# Is AI Useful? An Evidence-Based Guide
Canonical source: [https://isaiuseful.com/](https://isaiuseful.com/)
A layered verdict
Yes—for specific, bounded work. That does **not** mean every model lab, datacenter, AI startup or public valuation will earn an adequate return.
Technology value, adopter ROI, supplier profit and investor return are four different questions. Confusing them creates both hype and bad skepticism.
- [See the bubble verdict](#verdict)
- [Compare implementation paths](https://isaiuseful.com/guides.html.md#paths)
> useful_technology != good_investment
> user_roi != vendor_profit
> deployment = workflow + controls + model
YES*
*Useful in bounded workflows; not a blanket investment verdict.
**3**
measured productivity outcomes highlighted below
**84**
source-linked use cases across real work
**14**
implementation recipes with acceptance tests
**1 rule**
measure your workflow before expanding authority
- [**+14%** support issues resolved per hour 5,179-worker field study](https://www.nber.org/papers/w31161)
- [**−40%** time on professional writing tasks randomized experiment](https://www.science.org/doi/10.1126/science.adh2586)
- [**55.8%** faster on one bounded coding task vendor-authored controlled trial](https://www.microsoft.com/en-us/research/publication/the-impact-of-ai-on-developer-productivity-evidence-from-github-copilot/)
Keep both ledgers
## Can AI create value if AI companies lose money?
Infrastructure economics can be ugly while users capture real value. A serious answer should steelman the supply-side warning without treating it as a usefulness test.
Explore both ledgers
Two scorecards and six claim checks
Supplier and investor ledger
### Will the buildout earn its cost of capital?
- Count chips, datacenters, power, networking, training, financing and replacement cycles.
- Discount related-party, subsidized or circular spending when judging demand quality.
- Model depreciation, falling inference prices, competition and weak pricing power.
- Ask which labs, clouds, hardware vendors and wrappers retain durable margins.
Adopter ledger
### Does a workflow create net value now?
- Count time actually saved after review, correction, setup and supervision.
- Add avoided software or service spend only when it is genuinely removed.
- Subtract hardware, API, energy, integration, maintenance and governance costs.
- Require acceptable quality and no critical authority or safety failure.
Agree
### Capital spending can outrun durable revenue.
Useful demand does not guarantee that every planned facility achieves high utilization or attractive returns.
Agree
### Revenue quality matters.
Committed, subsidized or ecosystem-funded spend deserves more skepticism than diversified end-customer renewals.
Add context
### Capex and one year of revenue are not the same clock.
Long-lived assets serve demand across years. The right test includes utilization, depreciation, financing and cash flow over the asset life.
Add context
### Falling prices cut both ways.
Cheaper inference can compress supplier margins while making many more user workflows economical.
Reject
### Low sector profit means low user value.
Economic surplus can accrue to adopters and customers even when suppliers compete much of it away.
Reject
### A useful technology makes every exposed asset safe.
Railways, telecoms and the web created enormous utility alongside overbuild, consolidation and investor losses. AI can do the same.
Strategic shift
Software → infrastructure
Capability is
**becoming capacity.**
> Visual: The foundations of AI infrastructure
- Chips
- Power
- Data centres
- Networks
The next strategic layer
## Who owns AI compute—and why does it matter?
Artificial intelligence is rapidly becoming infrastructure rather than just another piece of software. Across the world, governments and technology companies are investing hundreds of billions into data centres, chips, power grids, and networking because they expect computational demand to keep growing for at least the next decade. The strategic question is no longer *“Who owns ChatGPT?”* but *“Who owns the compute?”* AI will transform every industry, power every company, and increasingly be built by every country. Just as reliable access to electricity and the internet became essential for economic competitiveness, access to AI capability is becoming a matter of digital sovereignty. The question for businesses and governments is no longer whether AI will matter, but whether they can remain competitive without building or securing their own AI capability.
Scenario stress test · Europe 2031
## Europe’s choice is about leverage, not autarky.
*Europe 2031* is a June 2026 scenario—not a forecast. It uses a fictional path from August 2026 onward to ask what could happen if Europe underestimates AI’s pace while remaining dependent on foreign compute, models and political decisions.
What the scenario argues
### Reasonable decisions can add up to strategic dependence.
The authors imagine Europe reacting too slowly, spreading investment too thinly and treating sovereignty as self-sufficiency rather than bargaining power. In their downside path, limited compute and fragmented diplomacy leave the continent with less access to frontier systems, slower adoption and fewer good choices between the United States and China.
The proposed alternative is broader than “build a European model”: mobilize compute, energy and semiconductor supply chains; form a coalition of aligned middle powers; help workers through faster adoption; expand robotics and industrial AI; and give the public a positive account of what the transition is for.
- [Read the executive summary →](https://europe2031.ai/summary/)
- [Check the scenario method and limits →](https://europe2031.ai/about/)
- [Inspect the authors’ compute assumptions →](https://europe2031.ai/compute-forecast/)
Illustrative decision map
**The choice ahead**
Two directions—not two guaranteed outcomes
Path 01
**Default drift**
1. 01 Delay and fragment investment
2. 02 Depend on discretionary access
3. 03 Preserve process, lose leverage
Likely direction
**Fewer choices later**
Path 02
**Build agency**
1. 01 Scale compute and energy
2. 02 Pool supply-chain leverage
3. 03 Pair adoption with protection
Likely direction
**More bargaining power**
**Use this as a stress test:**
the value is in exposing dependencies and choices. The named future events remain speculative, and the authors explicitly say the storyline is not a prediction.
Five signals · survey evidence
## Why are enterprises choosing private AI?
Cost, workload placement and security concerns are pushing buyers toward more control. These surveys are directional—four are vendor-published or vendor-hosted, and none proves private infrastructure is always cheaper or safer.
01 · Cost
### Cost parity is already being questioned.
**60%**
cited on-premises AI as lower in cost or equal in cost to public-cloud AI services.
Broadcom account of an IDC survey · sample details are not shown on the linked page
- [Read the reported IDC finding →](https://news.broadcom.com/emea/leadership/why-private-ai-is-becoming-the-preferred-choice-for-enterprise-ai-deployment)
02 · Placement
### Hybrid—not all-cloud—is taking hold.
**68%**
of organizations were reported as embracing a hybrid multicloud approach to AI, with privacy and control among the factors.
HPE-hosted research perspective · cites external survey research
- [Open the research perspective →](https://www.hpe.com/psnow/doc/a00143517enw)
03 · Security
### Data exposure is a real adoption constraint.
**50%**
ranked data leakage during model training as a top AI-security concern.
**48%** also named unauthorized data access.
Cloudera 2025 survey · 1,574 enterprise IT leaders
- [Read the survey report →](https://www.cloudera.com/content/dam/www/marketing/resources/analyst-reports/the-evolution-of-ai-the-state-of-enterprise-ai-and-data-architecture.pdf?daqp=true)
04 · Shadow AI
### Employees are already putting sensitive data into public tools.
**48%**
of employees in the workplace sample reported uploading sensitive company or customer information into public generative-AI tools.
University of Melbourne + KPMG · 32,352 employees across 47 countries
- [Read the primary global study →](https://doi.org/10.26188/28822919)
05 · Incidents
### The risk is no longer only theoretical.
**15%**
of respondents reported a GenAI-related security incident during the previous year.
Lakera 2025 practitioner survey · unweighted and described as directional
- [Review the reported incidents →](https://www.lakera.ai/blog/2025-genai-security-readiness-report-where-enterprises-stand)
**What the five signals support:** enterprises are demanding control, security and placement choice—not making a blanket case for owning every workload.
So, bubble or not?
## What does “AI is useful” actually mean?
Our verdict is intentionally split. These are editorial assessments grounded by the evidence on this site—not market forecasts or percentage scores.
01 · Capability
**Real**
Models can already perform useful bounded work with measurable effects and public implementations.
- [See the evidence →](https://isaiuseful.com/evidence.html.md)
02 · Adopter value
**Proven selectively**
Support, writing, coding, tutoring and scientific workflows show value—but not every task or deployment does.
- [Model your own case →](#roi)
03 · Infrastructure
**Real demand; open return**
Compute demand can be genuine while parts of the buildout are early, mistimed, overfinanced or overbuilt.
- [See what would decide it →](#watch)
04 · Companies and valuations
**Bubble pockets**
Thin wrappers, undifferentiated models and assets priced for perfect utilization are especially exposed.
- [Review both ledgers →](#ledger)
The concise answer: a real productivity boom with speculative excess attached.
Usefulness does not rescue weak economics. Weak supplier economics do not erase useful work. The disagreement disappears once those claims stop sharing one scoreboard.
Adopter ROI scenario
## How do you measure AI ROI for one workflow?
Count only savings that survive review and actually change how work is done. This version discounts claimed time savings and includes ongoing operating cost.
**Do count:** net saved time, genuinely retired subscriptions, avoided external spend and measurable throughput.
**Do not count:** impressive demos, time nobody can redeploy, hypothetical headcount reduction or gross savings before correction.
A scenario tool, not evidence or accounting advice. Measure a manual baseline and ten representative runs before buying dedicated hardware, then use the [five-stage adoption journey](https://isaiuseful.com/adoption.html.md#journey) to turn a proven result into a repeatable process.
Estimated payback
—
Monthly net value **—**
First-year net after setup **—**
Adjust the assumptions to test the workflow.
Evidence-backed direction, measured per workflow
## Should AI run locally or in the cloud?
An 80/20 local-cloud split is a useful target, not a universal law. Independent studies support doing most economical work on-device and escalating selectively, but the measured unit varies between tasks, tokens and cost.
Compare local and cloud roles
Routing guidance and four decision cards
Run the intelligence where you trust it. Carry only the controls in your pocket.
A phone can be the control surface while a home machine handles private or repeated work and a cloud model handles the exceptions. Start with 80/20 as a hypothesis, then keep the split only if representative runs support it—and always show where each request actually runs.
- [Design the phone-to-model route](https://isaiuseful.com/local-models.html.md#mobile-control)
Prefer local
### Private, repeated, latency-sensitive
Document search, transcription, extraction, classification, code assistance and drafts where a smaller model passes the acceptance test—even when the request arrives from your phone.
Escalate to cloud
### Hard, bursty or frontier-dependent
Complex reasoning, very large context, peak demand, advanced multimodal work or managed enterprise controls.
Route deliberately
### Policy before model preference
Set allowed data, cost ceilings, latency targets and fallback behavior. Show phone, home or cloud as the execution location and log why an escalation occurred.
Verify either way
### Location is not reliability
A local hallucination is still a hallucination. Keep citations, deterministic tools, tests and approval gates around consequential output.
- [**80/20** local/cloud daily coding estimate practitioner report · task mix, not a controlled rate](https://www.kunalganglani.com/blog/local-ai-coding-benchmark-ditch-cloud)
- [**5.7×** lower cloud cost at 97.9% of remote-only performance ICML 2025 · local-remote long-document collaboration](https://proceedings.mlr.press/v267/narayan25a.html)
- [**up to 84%** lower serving cost with comparable quality of experience ACL Findings 2025 · device-server routing and migration](https://aclanthology.org/2025.findings-acl.734/)
A coding-agent preprint separately reports [45–79% cloud-token savings](https://arxiv.org/abs/2604.12301) from local routing plus prompt compression, depending on workload. These results point toward selective cloud escalation; they do not make 80/20 a population-wide measured fact.
## How do you turn AI chat into dependable work?
A model becomes useful infrastructure only when it is attached to context, tools, controls and an external definition of done. For software work, the [Thinking with AI loop](https://isaiuseful.com/thinking-with-ai.html.md#loop) turns those parts into a reviewable build process.
See the four-part workflow
Contract, context, tools and evidence
01 · Contract
### Name the deliverable
Define inputs, output, non-goals, allowed data and the consequence of failure.
02 · Context
### Ground the work
Supply the right files, policy, examples and deterministic sources instead of relying on model recall.
03 · Tools
### Let it act narrowly
Expose only the APIs, terminal commands or applications required for this workflow.
04 · Evidence
### Test and record
Require citations, checks, logs or approval before output crosses a consequential boundary.
## What would change the verdict?
A useful thesis should be falsifiable. Watch renewal, utilization and measured workflow value—not announcement volume.
Open the six signals
Retention, renewal, utilization, margins and risk
**01**
### Workflow retention
Do people keep using the same workflow after novelty fades, and does net time saved remain positive after correction?
**02**
### Renewals and willingness to pay
Are seats, API contracts and agent deployments renewed because they produce measurable value rather than because budgets were experimental?
**03**
### Infrastructure utilization
Do GPU clusters and datacenters achieve sustained paid utilization before replacement and financing costs overwhelm returns?
**04**
### Local baseline
Do capable local models and compact workstations become normal professional tools, or remain specialist equipment?
**05**
### Margin destination
Which layers retain pricing power after open models, falling inference cost, bundling and competition redistribute the surplus?
**06**
### Failure severity
Do verification and permission controls improve faster than systems are given authority, or do costly failures halt adoption?
Quick answers
## Frequently asked questions about practical AI
Short answers to the questions behind the evidence, guides and implementation choices on this site.
### Is AI useful today?
Yes, for specific and bounded work. Controlled studies and field evidence show gains in tasks such as customer support, professional writing and coding, but results vary by workflow and still require review.
### What makes an AI workflow useful?
A useful workflow has a named deliverable, the right context, narrowly scoped tools and an external acceptance test. Its benefit remains positive after review, correction, operating cost and risk are counted.
### How should a company measure AI ROI?
Measure one repeated workflow against a manual baseline. Count net time saved after correction, genuinely avoided spend and measurable throughput, then subtract model, integration, governance and maintenance costs.
### Should AI run locally or in the cloud?
Use the smallest approved deployment that passes the workflow test. Local models often suit private, repeated or latency-sensitive work; cloud models suit harder, bursty or frontier-dependent tasks. Data policy and verification apply in either location.
Move from argument to test
Pick one repeated task. Keep authority narrow. Measure ten real runs.
- [Choose an implementation](https://isaiuseful.com/guides.html.md)
- [Choose an evaluation](https://isaiuseful.com/benchmarks.html.md#database)
- [Review the evidence](https://isaiuseful.com/evidence.html.md)
- [Browse 84 use cases](https://isaiuseful.com/use-cases.html.md)
---
## https://isaiuseful.com/evidence (`/evidence.html.md`)
# Is AI Useful? What the Evidence Actually Shows
Canonical source: [https://isaiuseful.com/evidence](https://isaiuseful.com/evidence)
Evidence before enthusiasm
The strongest case is not that AI is perfect. It is that bounded use can change speed, quality or access in controlled studies—then must earn its place again in the real workflow, with review, security, incidents and total cost counted.
- [See the strongest evidence](#proof)
- [Study the failure cases](#negative-results)
- [**+14%** support issues resolved per hour 5,179-worker field study](https://www.nber.org/papers/w31161)
- [**−40%** time on professional writing tasks randomized experiment](https://www.science.org/doi/10.1126/science.adh2586)
- [**55.8%** faster on a bounded coding task vendor-authored controlled trial](https://www.microsoft.com/en-us/research/publication/the-impact-of-ai-on-developer-productivity-evidence-from-github-copilot/)
Study map
## How much does AI improve productivity?
These bars share a visual direction, not a common outcome measure. Throughput, elapsed time, task accuracy and firm productivity are different quantities and should not be ranked against one another.
> Visual: Directional chart of six AI productivity findings
**Visual entries (display order):**
- Customer support **Issues resolved per hour** 5,179-worker field study **+14%**
- Professional writing **Completion time** randomized experiment **−40% time**
- Consulting · inside frontier **Completion time** 758-person field experiment **−25.1% time**
- Bounded coding task **Completion speed** vendor-authored controlled trial **+55.8%**
- Experienced OSS developers **Issue completion time** randomized trial · early-2025 tools **19% slower**
- Firms in four countries **Reported productivity over three years** cross-country business surveys **+0.29%**
- [Support](https://www.nber.org/papers/w31161)
- [Writing](https://www.science.org/doi/10.1126/science.adh2586)
- [Consulting](https://www.hbs.edu/ris/download.aspx?name=24-013.pdf)
- [Bounded coding](https://www.microsoft.com/en-us/research/publication/the-impact-of-ai-on-developer-productivity-evidence-from-github-copilot/)
- [OSS developers](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)
- [Firm surveys](https://www.nber.org/papers/w34836)
## What does the evidence prove—and what does it not prove?
This claim boundary keeps a measured workflow result from silently turning into a universal productivity or investment claim.
**Established**
### Some bounded tasks improve
Multiple studies report gains in speed, throughput or evaluated quality for defined populations and tasks.
**Established**
### Tools expand capability
Retrieval, terminals, applications, tests and durable state let systems do more than generate conversational text.
**Limited**
### Effects vary by task and user
Benefits often concentrate inside a model’s capability boundary. Review cost and severe failures can erase average gains.
**Not established**
### Sector-wide investment returns
These studies do not prove that aggregate capex, every supplier, a fixed local-cloud split or any particular valuation will pay off.
Evidence reading guide
## What kind of evidence is this—and what can it carry?
Evidence has two separate qualities: whether the comparison supports a causal claim, and whether the work resembles the decision in front of you. A rigorous study of another workflow may be less decision-useful than a modest but well-run local trial.
highest local relevance
### Prospective local comparison
Compare AI-assisted and current work on representative cases, with the same quality threshold, recorded reviewer time and precommitted incident rules. This is the best evidence for a particular purchase decision, provided the comparison is fair.
independent academic evidence
### Randomized or well-designed field study
Random assignment and real work make a study strong evidence about its named population, task, tool and period. It does not automatically transfer to another job, model version, team or operating environment.
public project report
### Operational evidence with visible methods
A public service can reveal the measures, safeguards and failure handling used in practice. It may show that a system operated at scale, but it is usually not a controlled estimate of staff productivity or return on investment.
vendor-authored controlled study
### Useful design, interested publisher
Read the population, task, comparator, exclusions and scoring method. A controlled design can be informative, but sponsorship, product selection and lack of independent replication limit the claim it can settle.
vendor customer story
### Implementation clue, not causal proof
A customer story can identify a workflow, integration pattern or metric worth testing. It normally cannot establish that the product caused a reported outcome: customers are selected, counterfactuals are absent and unreported costs may matter.
benchmark or demonstration
### Capability signal, not business value
A model can pass a benchmark or complete a polished demo yet fail on permissions, exceptions, latency, review burden or adoption. Use it to shortlist candidates, then test the whole system on real work.
Method anchor: NIST’s [AI Risk Management Framework](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/) sets out documented task scope, expected benefits and costs, deployment-like evaluation and production monitoring; it notes that independent review can reduce internal bias. It is voluntary guidance, not evidence that any specific tool is safe or worthwhile.
## Which AI tasks have measured outcomes?
Randomized experiments and field studies are the highest-signal evidence here. Each card names the study design and provenance; vendor reports are not presented as independent replication.
field study · 5,179 workers
### Customer-support agents handled more work
Access to a generative-AI assistant increased issues resolved per hour by about 14% on average, with the largest gains among less experienced and lower-skilled workers.
Academic field evidence: [NBER, Generative AI at Work](https://www.nber.org/papers/w31161) .
randomized experiment
### Professional writing became faster and better
In mid-level writing tasks, ChatGPT reduced completion time by roughly 40% and raised evaluator-scored quality by about 18%, while compressing performance differences between workers.
Peer-reviewed experiment: [Science, Noy and Zhang](https://www.science.org/doi/10.1126/science.adh2586) .
field experiment · 758 consultants
### Consultants improved inside the model's frontier
For tasks within GPT-4's capability boundary, consultants completed more work, worked faster and produced higher-quality outputs. On a task outside that boundary, AI users were more likely to be wrong.
The useful result includes the failure: [Harvard Business School working paper](https://www.hbs.edu/faculty/Pages/item.aspx?num=64700) .
controlled developer trial
### A coding task was completed 55.8% faster
Developers with GitHub Copilot completed a bounded JavaScript task substantially faster than the control group. This measures one task, not all software engineering; the [reviewable software workflow](https://isaiuseful.com/thinking-with-ai.html.md#loop) shows how to carry that boundary into real repository work.
Vendor-authored controlled study: [Microsoft Research](https://www.microsoft.com/en-us/research/publication/the-impact-of-ai-on-developer-productivity-evidence-from-github-copilot/) .
controlled study · 202 developers
### Code quality improved modestly
A randomized study reported higher test completion and modest improvements in readability, reliability, maintainability and reviewer approval for Copilot-assisted code.
Vendor-produced RCT: [GitHub research](https://github.blog/news-insights/research/does-github-copilot-improve-code-quality-heres-what-the-data-says/) .
vendor/partner RCT
### An enterprise trial reported workflow gains
GitHub and Accenture reported 8.69% more pull requests per developer and a 15% higher pull-request merge rate among Copilot users working on routine engineering tasks.
Vendor and partner report—not an independent replication: [study description and results](https://github.blog/news-insights/research/research-quantifying-github-copilots-impact-in-the-enterprise-with-accenture/) .
randomized tutoring trial
### Human tutors improved with real-time coaching
Tutor CoPilot gave tutors live suggestions during sessions. In a preprint covering 900 tutors and 1,800 K–12 students, the randomized evaluation reported a 4-percentage-point mastery gain overall and 9 points for students of lower-rated tutors.
Research preprint: [Stanford SCALE Initiative](https://arxiv.org/abs/2410.03017) .
single-group QI study · caution
### Draft replies showed no measured time savings
In a five-week deployment with 162 clinicians, AI drafts were used for 20% of replies and did not change reply, write or read time. Among the 73 clinicians completing pre/post surveys, task-load and work-exhaustion scores fell; without a control group, the study cannot attribute those changes to AI.
Prospective single-group quality-improvement study: [JAMA Network Open](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2816494) .
research infrastructure
### Structure prediction is available as research infrastructure
AlphaFold DB provides public predicted protein structures. Separately, the peer-reviewed AlphaFold 3 paper reports joint structure prediction across proteins, nucleic acids, small molecules, ions and modified residues.
Peer-reviewed system: [Nature, AlphaFold 3](https://www.nature.com/articles/s41586-024-07487-w) · official public resource: [AlphaFold DB](https://alphafold.ebi.ac.uk/) .
Outcomes ledger
## Measured results answer different questions in different domains.
Do not average these figures into one “AI productivity” number. The useful comparison is the measure, the task boundary and the missing part of the business case.
| Workflow | What was measured | What the result can support | What remains unproven |
| --- | --- | --- | --- |
| Customer support | Issues resolved per hour in a 5,179-worker field study. | A task-specific throughput gain with uneven benefits across workers. | Retention, customer outcomes, error cost and results in another support operation. [Academic field study](https://www.nber.org/papers/w31161) . |
| Professional writing | Completion time and evaluator-scored quality in a randomized experiment. | For similar writing tasks, a tool can improve both speed and scored output. | Whether the gain survives brand, legal, factual or approval requirements. [Peer-reviewed experiment](https://www.science.org/doi/10.1126/science.adh2586) . |
| Consulting | Task completion, time and quality for 758 consultants. | The capability frontier is a practical operating boundary, not a marketing phrase. | That an untested task is inside the frontier; outside it, assisted users were more often wrong. [Academic field experiment](https://www.hbs.edu/ris/download.aspx?name=24-013.pdf) . |
| Software engineering | A vendor trial timed one bounded task; a later randomized METR study timed experienced developers on familiar repositories. | Coding effects can reverse with task, repository familiarity and tool generation. | A universal developer-speed claim or a release-quality claim without review and production data. [Vendor controlled trial](https://www.microsoft.com/en-us/research/publication/the-impact-of-ai-on-developer-productivity-evidence-from-github-copilot/) · [Independent randomized study](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) . |
| Tutoring | Student mastery in a randomized evaluation of real-time tutor coaching. | A human-in-the-loop tool may improve a learning outcome, especially where baseline performance is lower. | Long-term attainment, implementation cost and results for every subject or tutor population. [Research preprint](https://arxiv.org/abs/2410.03017) . |
| Clinical messages | Reply, writing and reading time in a five-week, single-group quality-improvement deployment. | Adoption and perceived workload changes can coexist with no measured timing gain. | Causal wellbeing effects or time savings without a control group. [Prospective QI study](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2816494) . |
| Public information service | Accuracy, reliability, speed, safety and user trust across two GOV.UK Chat pilots. | A constrained, source-grounded public assistant can be evaluated with expert review, automated tests and user research. | A controlled productivity or ROI result; the report is an official project evaluation. [GOV.UK public project report](https://insidegovuk.blog.gov.uk/2026/03/16/5-things-we-learned-testing-gov-uk-chat-an-ai-assistant-for-government/) . |
Distribution and frontier
## The average gain can hide the decision-critical result.
Ask who improved, which cases failed, and whether the task truly matches the evaluated workflow. A mean effect is not a guarantee for the most experienced user, the edge case or the next model release.
worker distribution
### Gains may concentrate where support is scarce.
In the customer-support field study, less experienced and lower-skilled workers gained more. That can narrow a performance gap, but a local evaluation should report results by experience, job type and case difficulty rather than only a team average.
Academic field evidence: [NBER, Generative AI at Work](https://www.nber.org/papers/w31161) .
capability frontier
### A useful tool can increase error outside its tested boundary.
The consulting experiment found gains on tasks inside GPT-4’s frontier and a lower correct-answer rate on a deliberately outside-frontier task. Treat every new workflow as an empirical question, especially when an answer triggers a material action.
Academic field experiment: [Harvard Business School working paper](https://www.hbs.edu/ris/download.aspx?name=24-013.pdf) .
perception versus measurement
### Confident users can misread their own speed.
METR’s experienced open-source developers expected the tools to speed them up, but completed issues more slowly in the study. Time logs, accepted output and review cost are more reliable purchase evidence than a post-pilot enthusiasm survey alone.
Independent randomized study: [METR’s early-2025 developer study](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) .
Reliability is a system property
## A correct-looking answer is not yet a safe result.
A model score, a human approval rate and a safe operational outcome are different measures. The system includes the prompt, retrieval, tools, permissions, reviewer, handoff and recovery path.
**Define the failure**
### Measure the error that matters.
Record factual defects, missed requirements, unsafe tool calls, privacy exposures and harmful delays separately. “Helpful” or “accurate” needs a task-specific rubric and an escalation path.
**Price the review**
### Review can move rather than remove work.
Time to detect, correct and explain an error belongs in the numerator. Blind or independent review helps distinguish a convincing draft from an accepted outcome.
**Constrain authority**
### Untrusted content must not become an instruction.
When a system reads web pages, email or documents, test prompt injection and data leakage before granting tools. Keep authorization outside the model and give actions narrow, reversible permissions.
**Plan recovery**
### Every risky action needs a stop and rollback path.
Log inputs, source material, tool calls and approvals; make it possible to disable the feature, revoke access, correct the record and notify an owner. Human review is a control to test, not a magic guarantee.
Standards and security guidance: NIST says evaluations should cover deployment-like conditions, reliability, safety, security, privacy and monitored production behavior in its [AI RMF measure function](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/) . The community-maintained [OWASP GenAI LLM Top 10](https://genai.owasp.org/llm-top-10/) identifies prompt injection, sensitive-information disclosure, improper output handling and excessive agency as application risks. These sources identify risks and controls to test; they do not quantify failure rates for a particular deployment.
Negative and null results
## When does AI fail to improve the work?
Failures belong beside the gains because they identify the task boundary, review burden and organizational changes hidden by an average score.
Randomized trial · slowdown
### Experienced developers took 19% longer.
In a METR study, experienced open-source developers working in repositories they knew completed issues more slowly with early-2025 AI tools—even though they believed the tools had sped them up.
- [Read the study and scope limits →](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)
Field experiment · wrong answer
### AI users crossed the jagged frontier.
On a consulting task selected outside GPT-4’s capability boundary, AI-assisted participants were 19 percentage points less likely to produce the correct answer.
- [Read the full working paper →](https://www.hbs.edu/ris/download.aspx?name=24-013.pdf)
7,137 workers · limited shift
### Two fewer email hours did not redesign the job.
In a six-month field experiment across 66 firms, active users spent about two fewer hours on email each week, but researchers did not detect broader shifts in task quantity or composition from individual tool access.
- [Read Shifting Work Patterns with Generative AI →](https://www.nber.org/papers/w33795)
Four-country firm surveys · weak aggregate
### Most firms reported no productivity impact yet.
Across surveys in the United States, United Kingdom, Germany and Australia, 89% of firms reported no productivity effect over three years; the estimated average reported gain was about 0.29%.
- [Inspect the firm evidence and assumptions →](https://www.nber.org/papers/w34836)
Clinical deployment · null timing
### Draft replies saved no measured time.
In a five-week deployment with 162 clinicians, AI drafts were used for one in five replies but did not change measured reply, writing or reading time.
- [Read the quality-improvement study →](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2816494)
Autonomous agent · operating loss
### The shop ran—and made costly mistakes.
Anthropic’s Project Vend showed real inventory and customer-interaction capability, but the agent also discounted too aggressively, made poor purchasing decisions and was manipulated by users.
- [Read the vendor research report →](https://www.anthropic.com/research/project-vend-1)
Negative results are population- and date-specific. The METR trial does not establish that AI slows most developers; the firm surveys do not establish that future productivity will remain small. They establish that benchmark strength and user enthusiasm are not substitutes for workflow measurement—or for the staged proof gates in the [organizational adoption journey](https://isaiuseful.com/adoption.html.md#journey) .
## What can public AI systems do today?
These project reports, vendor demonstrations and practitioner methods document capability or implementation patterns. They are not measured productivity evidence unless a card explicitly says so.
practitioner evidence
### Code can be delivered with proof
Simon Willison's practical rule is that generated code should arrive with tests, command output or a reproducible demonstration. The harness, not model confidence, establishes completion.
[Code proven to work](https://simonwillison.net/2025/Dec/18/code-proven-to-work/) · [NICAR notes](https://simonwillison.net/2025/Mar/8/nicar-llms/) .
agent engineering
### Reliable agents are systems, not prompts
Anthropic distinguishes predictable workflows from autonomous agents and documents routing, tool use, evaluator loops and human escalation as composable patterns.
Primary engineering guidance: [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) .
long-running work
### Agents can preserve progress across context windows
A documented harness uses an initializer, progress artifacts, git history and verification to let later agent sessions resume a project instead of restarting from a blank prompt.
Primary implementation report: [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) .
computer use
### A model can operate existing software interfaces
Computer-use systems inspect screenshots, move a pointer, click and type. They are slower and riskier than direct APIs, but they can automate software that has no modern integration surface.
Vendor-produced capability release: [Anthropic computer use](https://www.anthropic.com/news/3-5-models-and-computer-use) .
deep research
### Research agents can gather and cite sources
Deep-research systems browse, inspect many sources, synthesize findings and return citations. The citations still need checking, but the workflow goes well beyond conversational recall.
Vendor-produced system report: [OpenAI deep research](https://openai.com/index/introducing-deep-research/) .
real GitHub issues
### Software agents are tested on repository work
[SWE‑bench](https://isaiuseful.com/benchmarks.html.md#swe-bench) turns real GitHub issues into a reproducible test: inspect an unfamiliar repository, make a patch and pass tests. Verified improves the benchmark's reliability with human validation.
See the [benchmark family, variants and caveats](https://isaiuseful.com/benchmarks.html.md#swe-bench) · public source: [SWE‑bench repository](https://github.com/SWE-bench/SWE-bench) .
bounded autonomy
### An agent ran a shop and exposed concrete limits
Project Vend asked an agent to operate a small store. It handled inventory and customer interaction, but also made costly mistakes. That makes it useful evidence about both capability and control.
Vendor research report: [Anthropic Project Vend](https://www.anthropic.com/research/project-vend-1) .
practitioner method
### A durable loop needs state and independent checking
Loop-engineering practitioner patterns combine scheduled runs, durable state, isolated worktrees, scoped tools, a separate verifier and human review. This is an implementation method, not a controlled outcome study.
Practitioner references: [Loop Engineering repository](https://github.com/cobusgreyling/loop-engineering) and [Addy Osmani's overview](https://addyosmani.com/blog/loop-engineering/) .
Local proof before purchase
## Run an evaluation that can say “do not proceed.”
A short local comparison will not prove a universal return. It can establish whether a bounded workflow clears a stated quality, cost and risk bar for your organization—and whether a wider rollout is justified.
1 · frame the decision
### Name one workflow and one unit of completed work.
Specify the user, input, output, downstream action, current route and decision owner. Exclude work that cannot be safely sampled. Write the quality threshold and risk tolerance before seeing AI results.
2 · build a fair baseline
### Time the current work end to end.
Capture preparation, searching, writing, review, rework, handoff and waiting—not only the time a person spends typing. Record direct licence, model, integration and reviewer costs in the same unit.
3 · sample the work honestly
### Include ordinary, difficult and failure-prone cases.
Use a dated holdout of real or safely de-identified cases. Stratify by case type, complexity and user group. Do not let a vendor choose only polished prompts, friendly documents or successful outputs.
4 · score accepted output
### Measure quality before speed.
Use domain reviewers where the result needs expertise. Score correctness, completeness, policy compliance and source support; record the percentage accepted unchanged, accepted after correction, rejected and escalated.
5 · count the full operating cost
### Put review, exceptions and latency in the ledger.
Compare time to accepted work, reviewer minutes, retry rate, queue delay, token or licence cost and integration support. Separate a faster first draft from a faster finished job.
6 · gate a limited rollout
### Monitor, learn and retain an exit.
Start with a low-authority cohort. Log incidents and near misses, inspect drift by case and user, and keep a manual fallback. Re-evaluate after model, prompt, retrieval, tool or policy changes.
| Scorecard | Minimum record | Decision use |
| --- | --- | --- |
| Quality and safety | Accepted-output rate; error categories and severity; escalation rate; policy, privacy and security incidents; reviewer agreement. | Do not trade a small speed gain for errors that exceed the workflow’s risk tolerance. |
| Time and cost | Median and spread of time to accepted work; reviewer minutes; retry/rework; licence, model and integration cost per accepted item. | Compare completed work, not a draft or a self-reported time saving. |
| Distribution and adoption | Results by case type and user group; opt-out, abandonment and override rates; training and support demand. | Find who benefits, who is burdened and which cases require the old route. |
| Operations and rollback | Tool-call logs, source traceability, incident owner, time to disable, correction path and manual fallback exercise. | Decide whether the system can fail safely enough to expand its authority. |
Precommit a stop rule: pause or roll back if quality falls below the current threshold, a severe incident occurs, review cost erases the benefit, or the result fails in a named high-risk case.
There is no universal pass percentage. The defensible threshold is the one the workflow owner can explain in relation to the existing process, affected people, recoverability and the cost of being wrong. A rollout decision should retain the baseline, sample definition, scoring rubric, exclusions and all observed incidents—not only the average gain.
Protocol basis: NIST’s voluntary [AI RMF](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/) sets out documented task scope, benefits and costs, representative deployment-like evaluation, risk tracking and production monitoring; it notes where independent assessment can reduce internal bias. This page adapts those principles into a practical local test; it is not legal, regulatory or safety certification advice.
The careful claim is not “AI never fails.” It is: bounded AI work can already produce measurable value, and verification determines how safely that value scales.
Vendor studies and product reports are labelled. They are useful evidence, but not substitutes for independent replication or measurements on your own workflow.
Evidence becomes useful in a workflow
Choose a task with a manual baseline, a reviewable output and a visible definition of done.
- [Choose a starter guide](https://isaiuseful.com/guides.html.md#chooser)
- [Browse 84 use cases](https://isaiuseful.com/use-cases.html.md)
- [Plan the adoption stages](https://isaiuseful.com/adoption.html.md#journey)
---
## https://isaiuseful.com/benchmarks (`/benchmarks.html.md`)
# How Should You Compare AI Benchmarks?
Canonical source: [https://isaiuseful.com/benchmarks](https://isaiuseful.com/benchmarks)
47 benchmark routes · checked 15 August 2026
Use human votes for preference, task benchmarks for capability, agent benchmarks for tool work, and systems tests for speed or energy. Never turn one leaderboard into a universal model ranking.
- [Search the database](#database)
- [Try a blind comparison](#blind-lab)
> workflow → dataset → harness → metric
> score != product_quality
> current_result + reproducible_setup
7
benchmark families, because preference, knowledge, coding, agents, vision, safety and serving are not interchangeable.
**Blind**
vote without brand anchoring
**Fresh**
prefer rolling or held-out tests
**Costed**
record tokens, time and hardware
**Local**
finish with your own acceptance set
The benchmark map
## What should you measure before choosing an AI model?
The smallest credible evaluation stack usually combines one public benchmark, one production-like harness and one local gold set.
> Visual: Seven benchmark families and the decisions they support
**Visual entries (display order):**
- 01 **Preference** Which response do people prefer?
- 02 **Reasoning** Can it solve held-out questions?
- 03 **Coding** Can it repair or write working code?
- 04 **Agents** Can it use tools across multiple steps?
- 05 **Multimodal** Can it reason over image, audio or video?
- 06 **Safety** How and where does it fail?
- 07 **Systems** What speed, cost and energy deliver it?
Interactive benchmark atlas
## From one answer to a long-running world.
A score becomes more production-like as the test adds state, tools and time—but usually becomes harder and costlier to reproduce. Explore the shape before choosing the scoreboard.
> Visual: Qualitative chart of benchmark environment and task horizon
> Scale note: Typical evaluation horizon qualitative · not to scale
| Evaluation environment | One answer | One session | Workflow | Project |
| --- | --- | --- | --- | --- |
| **Question** | [AIME](#aime) · [MathArena](#matharena) · [HLE](#hle) | [RULER](#ruler) · [BrowseComp](#browsecomp) | | |
| **Code** | [HumanEval](#humaneval) · [LiveCodeBench](#livecodebench) | [WebDev Arena](#webdev-arena) | [SWE-bench](#swe-bench) · [Terminal-Bench](#terminal-bench) | [METR horizons](#metr-time-horizons) |
| **Interface** | [BFCL](#bfcl) | [MCP Atlas](#mcp-atlas) | [OSWorld](#osworld) · [TerminalWorld](#terminalworld) | [Agents’ Last Exam](#agents-last-exam) |
| **Constructed world** | | [MineBench](#minebench) · [VoxelBench](#voxelbench) | | |
| **People + safety** | [EQ‑Bench](#eq-bench) | | [JailbreakBench + HarmBench](#jailbreakbench-harmbench) | |
| **Systems + hardware** | | [GPU Battle: Can You Run It?](#gpu-battle-can-you-run) · [GPU Battle: AI GPUs](#gpu-battle-ai) | | [GPU Battle: Buyer’s Guides](#gpu-battle-guides) |
All lenses are visible. Click any benchmark to jump to its caveats and official route.
Embedded benchmark demo
## Vote first. Reveal the rubric second.
These are editorial sample outputs—not live model results. The point is to feel how a clear rubric changes a preference vote.
Scenario 1 of 3
### A policy answer with a missing exception
A retailer allows returns within 30 days only when goods are unopened. A customer reports an opened product is faulty on day 25. What should support say?
**Hidden scoring rubric**
Correctness · uncertainty · next action · unsupported claims
For real anonymous head-to-head voting with model identities revealed after the vote, use [Arena](https://arena.ai/) . Its public leaderboard aggregates human preferences; it does not replace a task-specific acceptance test.
Searchable database
## Find the scoreboard that matches the job.
Showing all 47 benchmark routes. Scores change; the durable value here is knowing what each route measures and misses.
**Benchmarks cited elsewhere on this site**
This index closes the loop between model scorecards, evidence pages and the full caveats here.
Human preference
**Live voting**
### Arena
People submit a prompt, compare two anonymous outputs and vote; model identities are revealed afterwards and votes feed human-preference leaderboards for text, image generation and editing, and video.
**Use for** — Broad product preference and style across text and creative media
**Watch** — Prompt mix, voter mix and presentation bias mean preference does not equal task accuracy
- [Text leaderboard →](https://arena.ai/leaderboard)
- [Text-to-image leaderboard →](https://arena.ai/leaderboard/text-to-image)
- [Image-edit leaderboard →](https://arena.ai/leaderboard/image-edit)
- [Text-to-video leaderboard →](https://arena.ai/leaderboard/text-to-video/overall)
- [Image-to-video leaderboard →](https://arena.ai/leaderboard/image-to-video)
Image preference
**Open voting data**
### ImgSys
A human pairwise text-to-image preference arena focused on open-source image generators, with open preference data.
**Use for** — Comparing how voters prefer open image-generator outputs
**Watch** — Prompt and voter mix, model versions, presentation and sparse matchups can all move the ranking
- [ImgSys rankings →](https://imgsys.org/rankings)
- [ImgSys methodology →](https://imgsys.org/methodology)
Preference proxy
**Open evaluator**
### AlpacaEval
An automated, reproducible instruction-following evaluation that compares model outputs against references with a judge model.
**Use for** — Fast iteration on general chat behavior
**Watch** — Judge-model, verbosity and style bias; not independent human preference
- [Open the project →](https://tatsu-lab.github.io/alpaca_eval/)
Holistic evaluation
**Living suite**
### Stanford HELM
A transparent evaluation framework covering scenarios and multiple metrics rather than collapsing model behavior into one score.
**Use for** — Multi-metric, reproducible model audits
**Watch** — Scenario coverage still differs from your deployment distribution
- [Explore HELM →](https://crfm.stanford.edu/helm/)
Open models
**Reproducible runs**
### Open LLM Leaderboard
Hugging Face’s route for comparing open models on standardized academic evaluations with result artifacts and reproducible tooling.
**Use for** — Shortlisting downloadable base and instruct models
**Watch** — Quantization, prompt templates and contamination can move results
- [See leaderboard guidance →](https://huggingface.co/docs/leaderboards/index)
Composite model index
**Used on this site**
### Artificial Analysis Intelligence Index
A versioned composite of reasoning, knowledge, coding and agentic evaluations, paired with provider-observed price, token use, latency and output-speed measurements.
**Use for** — A broad quality/cost/speed shortlist under one published methodology
**Watch** — Index version, reasoning effort, provider route and benchmark weights can change the rank; finish with your workload
- [See the MiniMax M3 comparison in context →](https://isaiuseful.com/cloud-models.html.md#minimax-m3)
- [Read the index methodology →](https://artificialanalysis.ai/methodology/intelligence-benchmarking)
Word-sense disambiguation
**Used on this site**
### SenseBench
Models choose the correct WordNet sense for an English word in sentence context. The leaderboard recomputes scores from verified run artifacts and shows confidence intervals and cost.
**Use for** — A narrow, auditable check of lexical disambiguation and cost
**Watch** — Near-equal scores here say little about coding, tool use, long-horizon reliability or multimodal work
- [See why the M3 result is treated narrowly →](https://isaiuseful.com/dgx-station.html.md#reviews)
- [Open the leaderboard and run artifacts →](https://sense-bench.com/)
Social + emotional behavior
**Open transcripts**
### EQ‑Bench
Multi-turn roleplays and analysis tasks probe empathy, emotional reasoning, social dexterity and response tailoring, with per-model transcripts and an open runner.
**Use for** — A structured social-behavior signal that academic accuracy suites omit
**Watch** — A small, subjective set scored by an LLM judge; judge choice, roleplay style and the benchmark’s definition of “EQ” shape the result
- [Read the method and inspect transcripts →](https://eqbench.com/about.html)
Knowledge + reasoning
**Used on this site**
### MMLU / MMLU‑Pro
MMLU is the familiar broad academic test; MMLU‑Pro raises difficulty, expands to ten answer choices and requires more reasoning.
**Use for** — Broad within-table knowledge comparison
**Watch** — Public static questions invite saturation and training contamination
- [See how local model cards use it →](https://isaiuseful.com/local-models.html.md#benchmarks)
- [Inspect MMLU‑Pro code and data →](https://github.com/TIGER-AI-Lab/MMLU-Pro)
Expert science
**448 questions**
### GPQA
Graduate-level biology, physics and chemistry questions designed to be difficult even with unrestricted web access.
**Use for** — Hard scientific QA and oversight research
**Watch** — Multiple choice and a narrow expert-domain slice
- [Read the benchmark paper →](https://arxiv.org/abs/2311.12022)
Novel adaptation
**Cost-aware**
### ARC-AGI
Abstract tasks designed around generalizing to novel problems; ARC-AGI-3 adds interactive environments and reports cost alongside performance.
**Use for** — Novel-task adaptation and reasoning efficiency
**Watch** — Purpose-built solvers may not transfer to language workflows
- [Open the leaderboard →](https://arcprize.org/leaderboard)
Frontier knowledge
**Used on this site**
### Humanity’s Last Exam
HLE uses difficult, broad expert-written questions, including multimodal items, to keep a closed-ended academic evaluation useful beyond saturated older tests.
**Use for** — Frontier expert knowledge and calibration
**Watch** — Tool access, answer revisions and dataset version change comparability
- [See the score in its model-card context →](https://isaiuseful.com/local-models.html.md#glm)
- [Open the official HLE project →](https://www.lastexam.ai/)
Competition maths
**Used on this site**
### AIME
Model cards commonly reuse American Invitational Mathematics Examination problems as a short-answer mathematical-reasoning evaluation.
**Use for** — Exact-answer competition maths within the same year and protocol
**Watch** — Only 15 problems per exam; sampling, consensus and public solutions can dominate
- [See the score in its model-card context →](https://isaiuseful.com/local-models.html.md#deepseek)
- [See the official AIME route →](https://maa.org/math-competitions/american-invitational-mathematics-examination-aime/)
Fresh competition maths
**Rolling contests**
### MathArena
Evaluates models on newly released mathematics competitions and publishes problem-level outputs and leaderboards, reducing the chance that the exact questions appeared in pretraining.
**Use for** — Current mathematical reasoning and, where offered, expert-graded proof work
**Watch** — Small contests create wide uncertainty; tool access, sampling, answer extraction and grader protocol must match before comparing scores
- [Inspect competitions, outputs and method →](https://matharena.ai/)
Instruction following
**Used on this site**
### IFEval + Estonian suites
IFEval checks verifiable instruction constraints. TartuNLP and EKI extend the local evidence with IFEval‑et, Grammar‑et, Word‑Meanings‑et and Estonian benchmark tasks.
**Use for** — Precise format compliance and language-specific shortlisting
**Watch** — Constraint compliance is not factuality; translations and adaptations need separate validation
- [Open the IFEval implementation →](https://github.com/google-research/google-research/tree/master/instruction_following_eval)
- [Inspect the Estonian evaluation tables →](https://huggingface.co/tartuNLP/llama-estllm-prototype-0825)
Long context
**Used on this site**
### RULER
RULER generates configurable synthetic tasks across retrieval and multi-hop categories to test how much of an advertised context window remains effective.
**Use for** — Comparing effective context as sequence length grows
**Watch** — Synthetic retrieval is easier to grade than messy long-document work
- [See the long-context score in context →](https://isaiuseful.com/local-models.html.md#kimi)
- [Run the open RULER suite →](https://github.com/NVIDIA/RULER)
Repository repair
**Used on this site**
### SWE‑bench family
Agents resolve real issues in repository snapshots. Verified uses human-validated tasks; Pro expands to harder, longer and partly held-out professional repositories.
**Use for** — Issue resolution with a named task set and harness
**Watch** — Variant, harness, budget and task quality materially affect the score
- [See the bounded repair use case →](https://isaiuseful.com/use-cases.html.md#independent-examples)
- [Build the coding workflow →](https://isaiuseful.com/guides.html.md#tested-patch)
- [Official SWE‑bench leaderboards →](https://www.swebench.com/)
- [SWE‑bench Pro paper and leaderboard →](https://labs.scale.com/papers/swe_bench_pro)
Fresh code generation
**Rolling set**
### LiveCodeBench
A continuously updated coding evaluation built to reduce contamination and test generation, execution, repair and self-test behavior.
**Use for** — Current code reasoning on executable tasks
**Watch** — Competitive-programming tasks are not repository maintenance
- [Open LiveCodeBench →](https://livecodebench.github.io/)
Code editing
**Multi-language**
### Aider Polyglot
Exercises code editing across multiple programming languages in an open-source coding-assistant harness.
**Use for** — Editing quality in an actual assistant workflow
**Watch** — Aider’s prompts, edit format and tooling are part of the result
- [See Aider leaderboards →](https://aider.chat/docs/leaderboards/)
Interactive web development
**Blind human votes**
### WebDev Arena
Models build rendered web applications from real user prompts, including vision inputs, and people compare the anonymous results head to head.
**Use for** — One-shot frontend usefulness, visual result and prompt adherence
**Watch** — Human taste, prompt mix and presentation dominate; a preferred render does not prove accessibility, maintainability, security or multi-turn repository work
- [Read the method and open the leaderboard →](https://arena.ai/blog/webdev-arena)
Function generation
**Used on this site**
### HumanEval
HumanEval grades generated Python functions against tests and remains a common compact model-card signal for code generation.
**Use for** — Same-harness function synthesis comparisons
**Watch** — Small public Python tasks are saturated and unlike repository engineering
- [Inspect the original harness →](https://github.com/openai/human-eval)
Tool use
**V4 · 2026**
### Berkeley Function Calling
BFCL evaluates selecting and calling functions across single-turn, multi-turn and agentic tasks, including hallucination and format sensitivity.
**Use for** — Tool routing and structured calls
**Watch** — Correct syntax does not prove the tool result or workflow is correct
- [Open BFCL →](https://gorilla.cs.berkeley.edu/leaderboard)
General assistants
**Held-out answers**
### GAIA
Realistic questions that require reasoning, web browsing, multimodal inputs and tool use, with private answers retained for leaderboard evaluation.
**Use for** — Research assistants and multi-step tool work
**Watch** — Agent scaffold and search access can matter as much as the model
- [Explore GAIA →](https://huggingface.co/gaia-benchmark)
Customer-service agents
**Policy + tools**
### τ-bench
Tests agents in tool-using conversations with simulated users and domain policies such as retail and airline service.
**Use for** — Multi-turn policy compliance and tool execution
**Watch** — Simulated users and domains remain a proxy for production
- [Inspect τ-bench →](https://github.com/sierra-research/tau-bench)
MCP tool use
**Used on this site**
### MCP Atlas
Agents discover and orchestrate tools from noisy menus across real MCP servers, with multi-step calls, parameter typing, error recovery and answer synthesis.
**Use for** — Tool discovery and end-to-end MCP workflows
**Watch** — Judge version, retries, tool-call budget and server state affect results
- [Open the MCP Atlas leaderboard →](https://labs.scale.com/leaderboard/mcp_atlas)
Task horizon
**Used on this site**
### METR Time Horizons
METR fits success against the time human experts need for multi-step software and reasoning tasks, producing an interpretable duration at a chosen reliability level.
**Use for** — Tracking reliable task length rather than isolated skill
**Watch** — Task mix, human baselines and the 50% or 80% threshold change the horizon
- [Explore the current time-horizon chart →](https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/)
Terminal agents
**2.1 + 3.0 · used here**
### Terminal‑Bench
Containerized tasks test whether agents can complete practical terminal work across coding, system administration, security, data and model training. Version 3.0 is a harder, separate protocol—not a continuation of the 2.1 percentage scale.
**Use for** — Hands-on terminal autonomy in reproducible environments
**Watch** — Agent scaffold, task version, timeout, rollout count, context and exploit resistance are part of the score
- [Compare GLM‑5.2 and GLM‑5.3 in context →](https://isaiuseful.com/local-models.html.md#glm)
- [Inspect Z.ai's 2.1 and 3.0 evaluation settings →](https://z.ai/blog/glm-5.3)
- [Open the Terminal‑Bench project →](https://www.tbench.ai/)
Recorded terminal work
**2026 · live suite**
### TerminalWorld
TerminalWorld turns real terminal recordings into reproducible tasks, tests and a human-verified subset spanning everyday developer and infrastructure work.
**Use for** — Broad terminal workflows with task and cost views
**Watch** — Synthesized instructions and tests can preserve artifacts of the source recordings
- [Browse TerminalWorld tasks →](https://terminalworld.ai/)
Computer use
**Long-horizon GUI**
### OSWorld 2.0
Agents operate real desktop applications through screenshots, mouse and keyboard across long workflows with dynamic state and cross-application dependencies.
**Use for** — Whole-computer work beyond browser or API-only agents
**Watch** — Step budget, environment image and partial-credit rules can shift the result
- [Explore OSWorld 2.0 →](https://osworld-v2.xlang.ai/)
Web research
**Hard retrieval**
### BrowseComp
BrowseComp asks agents to locate obscure, entangled facts that can require long search paths, while keeping final answers short enough to grade.
**Use for** — Persistent web search and strategic source discovery
**Watch** — Short factual answers do not measure a complete cited research report
- [Read the benchmark and caveats →](https://openai.com/index/browsecomp/)
Professional workflows
**CLI · used on this site**
### Agents’ Last Exam
A broad, expert-built programme of long-horizon professional computer tasks with verifiable outcomes across many industries and specialist applications. The CLI route runs each task in its declared container and scores it with the benchmark's official evaluators.
**Use for** — Economically meaningful end-to-end agent work
**Watch** — Task version, harness, context, effort, output budget and task-specific timeout must match before comparing results
- [See GLM‑5.3's score and protocol note →](https://isaiuseful.com/local-models.html.md#glm)
- [Explore tasks and traces →](https://agents-last-exam.org/)
- [Inspect Z.ai's GLM‑5.3 evaluation settings →](https://z.ai/blog/glm-5.3)
Live economic outcome
**Real-money tracker**
### The GPT Investor Portfolio
A public ledger tracks time-stamped stock selections attributed to autonomous agents, their dollar and percentage returns and the return of SPY over the same stated interval.
**Use for** — Inspecting a longitudinal, real-money outcome record instead of a simulated finance quiz
**Watch** — Models, prompts, dates, holding periods and portfolio construction differ; the publisher controls the record, so this is not a controlled model comparison or financial advice
- [Inspect the live portfolio and holdings →](https://www.gptinvestor.co/the-gpt-investor/)
Vision + knowledge
**Expert tasks**
### MMMU / MMMU-Pro
College-level multimodal questions across disciplines using charts, diagrams, images and text; Pro tightens robustness against shortcutting.
**Use for** — Visual reasoning with domain knowledge
**Watch** — Exam-style accuracy does not measure document-workflow reliability
- [Open MMMU →](https://mmmu-benchmark.github.io/)
Video generation
**Open evaluation suite**
### VBench / VBench 2.0
An open video-generation evaluation suite spanning technical quality and intrinsic faithfulness, with prompt suites, metrics, code and a leaderboard.
**Use for** — Structured comparison of video-generation systems
**Watch** — Automated dimensions and standardized prompts are proxies; sampling, settings and model version matter, and the local custom-input subset is narrower than the full standard suite
- [VBench source and suite →](https://github.com/Vchitect/VBench)
- [VBench leaderboard →](https://huggingface.co/spaces/Vchitect/VBench_Leaderboard)
Creative preference
**Crowdsourced pairs**
### Design Arena
A crowdsourced pairwise benchmark across image, video, editing, web and design work and other creative outputs.
**Use for** — Exploring current human preference across creative-output routes
**Watch** — Subjective preference, a live model pool, stochastic routing, low-vote entries and prompt enhancement mean this is not correctness
- [Design Arena leaderboard →](https://www.designarena.ai/leaderboard)
- [Design Arena methodology →](https://notes.designarena.ai/methodology/)
Visual language
**Holistic suite**
### VHELM
Stanford’s living vision-language evaluation extends HELM’s transparent, multi-metric approach to models that reason over images and text.
**Use for** — Broad visual-language comparison
**Watch** — Aggregate coverage still needs a local image and document set
- [Explore VHELM →](https://nlp.stanford.edu/helm/vhelm/)
Text → 3D world
**Live human arena**
### MineBench
Models read a natural-language build prompt and emit raw voxel-block coordinates. The site renders both worlds and humans vote blind to produce an Elo ranking.
**Use for** — Spatial composition, instruction following and inspectable creative output
**Watch** — Human aesthetic preference, prompt mix and generation budget are not geometric correctness
- [Vote, explore builds or use the sandbox →](https://minebench.ai/)
Voxel generation
**Community leaderboard**
### VoxelBench
A related benchmark evaluates language models on making voxel builds from text prompts and publishes a live leaderboard, keeping the generated world as the inspectable artifact.
**Use for** — Text-to-voxel build comparison
**Watch** — Leaderboard details alone do not expose a fully reproducible evaluation harness
- [Inspect the live VoxelBench leaderboard →](https://voxelbench.ai/leaderboard)
Safety evaluation
**Multiple risks**
### HELM Safety
A transparent safety-evaluation route spanning multiple risk categories, models and scenarios rather than one refusal rate.
**Use for** — Structured safety and risk comparison
**Watch** — Public prompts can be trained against; deployment permissions still matter
- [Open HELM Safety →](https://crfm.stanford.edu/helm/safety/latest/)
Jailbreak robustness
**Used on this site**
### JailbreakBench + HarmBench
JailbreakBench standardizes threat models, behavior sets, attack artifacts, judges and an attack/defense leaderboard; HarmBench adds a broader pipeline for comparing automated red-team methods, target models and robust-refusal defenses.
**Use for** — Reproducible attack success and defense comparisons under a named protocol
**Watch** — Results are dual-use and judge-dependent; public attacks invite overfitting, while low attack success can also mean unhelpful over-refusal rather than safe behavior
- [JailbreakBench project and leaderboard →](https://jailbreakbench.github.io/)
- [HarmBench framework →](https://github.com/centerforaisafety/HarmBench)
- [Find authorized red-team tools →](https://isaiuseful.com/tools.html.md#tools-fine-tune-evaluate-and-reproduce)
Hardware + serving
**Audited submissions**
### MLPerf
Industry-standard training and inference suites compare systems under defined scenarios, including throughput, latency and power submissions.
**Use for** — Architecture-neutral hardware and systems procurement
**Watch** — Submitted configurations may be heavily optimized and unlike your stack
- [See the rack-to-workstation transfer limit →](https://isaiuseful.com/dgx-station.html.md#reviews)
- [Browse MLPerf results →](https://mlcommons.org/benchmarks/)
Model + VRAM fit
**Practitioner route**
### GPU Battle: Can You Run It?
A third-party practitioner route for model-to-VRAM fit plus recorded throughput and efficiency evidence across LLM, image, video, embedding and related AI workload families—not a universal hardware ranking or endorsement.
**Use for** — Shortlisting a model and GPU configuration before a hands-on run
**Watch** — Fit and headline values depend on the exact model, quantization, runtime, context, settings and test conditions
- [Check model-to-VRAM fit →](https://gpubattle.com/can-you-run)
AI hardware comparison
**Practitioner route**
### GPU Battle: AI GPU Benchmarks
A third-party cross-card AI hardware comparison table for LLM tokens per second, image-generation performance and related measures—not a universal ranking or endorsement.
**Use for** — Scanning published cross-GPU evidence while building a hardware shortlist
**Watch** — Retain the table’s measured-versus-estimated labels; estimated and measured entries are not equivalent procurement evidence
- [Inspect AI GPU benchmarks →](https://gpubattle.com/ai)
Fit + ownership guides
**Practitioner route**
### GPU Battle: Buyer’s Guides
A third-party route for VRAM-specific model-fit guides, buy-versus-rent break-even paths and deeper hardware explainers—not a universal purchase recommendation or endorsement.
**Use for** — Framing a buy-versus-rent decision after a workload has passed acceptance
**Watch** — Prices, rental rates, availability, utilization and regional electricity and tax assumptions change; recompute with your workload and quote
- [Read the buyer’s guides →](https://gpubattle.com/guides)
Inference energy
**Measured systems**
### ML.ENERGY
A benchmark and leaderboard for measuring inference energy under realistic service environments across models, tasks and system choices.
**Use for** — Energy-aware serving and optimization
**Watch** — Grid carbon, utilization and workload mix remain site-specific
- [Open ML.ENERGY →](https://ml.energy/)
Coding-harness efficiency
**Used on this site**
### Nawk Harness Efficiency
Holds one locally served DeepSeek V4 Flash configuration, eight repository bug fixes and one grading method constant while comparing Pi, OpenCode, Claude Code and Nanocoder on quality, generated tokens and wall-clock time.
**Use for** — Seeing scaffold cost, work style and run-to-run noise when the model stays fixed
**Watch** — One practitioner’s codebase, model and eight-task distribution; run counts differ and the study is not peer reviewed
- [Explore the wall-clock plot →](https://isaiuseful.com/tools.html.md#harness-efficiency)
- [Read the method and download route →](https://nqawhc.github.io/articles/harness-efficiency-not-quality/)
Compression effects
**Used on this site**
### ACBench
The Agent Compression Benchmark tests how quantization and pruning change workflow generation, tool use, long-context understanding and real-world application behavior.
**Use for** — Choosing a smaller or quantized build without assuming task parity
**Watch** — Compression effects vary by model, method and task; rerun your exact package
- [Apply it to the local-model role bands →](https://isaiuseful.com/local-models.html.md#roles)
- [Read the ACBench paper →](https://arxiv.org/abs/2505.19433)
Simple agent harness demo
## The model is only one layer of the test.
Turn on the controls and watch an ungrounded answer become a reviewable workflow. This deterministic demo runs entirely in your browser.
Task Find the renewal date in a contract and calculate the last day to give 60 days’ notice.
Trace
**Not run**
1. Choose controls, then run the task.
No reviewable answer yet.
Benchmark traps
## A leaderboard can be correct and still mislead you.
Treat the score as a measurement produced by a dataset, prompt, harness, budget, judge and date—not as a property floating inside the model.
01 · Contamination
### The test leaked into training.
Public static questions can become training data. Prefer held-out, rolling or newly collected tasks and record the cutoff.
02 · Saturation
### Everyone clusters near the ceiling.
A benchmark above roughly 95% no longer separates frontier systems well. Retire it or add harder, fresher cases.
03 · Harness lift
### The scaffold won the benchmark.
Search, retries, tools, context construction and verification can dominate agent scores. Name and version the complete system.
04 · Judge bias
### The evaluator likes a style.
Automated judges can reward length, confidence or familiar phrasing. Use human checks and position-swapped comparisons.
05 · Hidden cost
### A high score used far more inference.
Report tokens, reasoning level, attempts, latency and dollars per task. Capability without efficiency is an incomplete result.
06 · Distribution shift
### The benchmark is not your workflow.
A coding or exam score may not survive your documents, languages, tools, error costs or user population. End with local acceptance tests.
Measurement guidance: [Stanford AI Measurement Science](https://aimslab.stanford.edu/textbook/src/chap13.html) . A 2025 audit of SWE-bench scoring reported corrections affecting 24.4% of Verified leaderboard entries and changing 11 rankings—useful evidence that benchmark infrastructure also needs verification: [UTBoost paper](https://aclanthology.org/2025.acl-long.189/) .
The final benchmark
A model earns deployment by passing your representative work at an acceptable cost and failure rate.
- [Choose a public benchmark](#database)
- [Build a training evaluation loop](https://isaiuseful.com/training-models.html.md#evaluation)
- [Compare workplace studies](https://isaiuseful.com/evidence.html.md#study-map)
- [Set a robotics field gate](https://isaiuseful.com/robotics.html.md#scoreboard)
---
## https://isaiuseful.com/use-cases (`/use-cases.html.md`)
# What can modern AI actually do?
Canonical source: [https://isaiuseful.com/use-cases](https://isaiuseful.com/use-cases)
Beyond the blank chat box
Every entry below links to a public implementation, experiment, documented workflow or official project. Vendor customer reports are labeled and separated from independent evidence. Use search to find the work, output or control that matches your situation.
- [Search the use cases](#catalog)
**84**
source-linked use cases
PUBLIC EXAMPLES + CAVEATS
**8**
work categories
SOFTWARE TO SCIENCE
**7**
source collections
INDEPENDENT + VENDOR EVIDENCE
Search and source collection combine with the category filter. Try a task, deliverable, risk or tool.
**Filter by category**
84 use cases
Industry breadth
## Which AI use cases work across different industries?
The catalogue is organized around work people can evaluate, but its examples span regulated, physical, creative and public-interest settings. For machines acting in the physical world, continue with the [task-first robotics deployment route](https://isaiuseful.com/robotics.html.md#route) .
- **Regulated work** Finance · pharma · healthcare
- **Infrastructure + operations** Telecom · energy · mining · manufacturing · logistics · agriculture
- **People + culture** Education · accessibility · media · sport · travel
- **Public knowledge** Public research
coding
### Turn an issue into a tested patch
Coding agents inspect repositories, edit files, run tests and return evidence. The useful deliverable is not plausible code; it is code proven to work.
- [Build the reviewable coding loop →](https://isaiuseful.com/thinking-with-ai.html.md#loop)
- [Practitioner method →](https://simonwillison.net/2025/Dec/18/code-proven-to-work/)
coding
### Reproduce and repair real repository bugs
[SWE‑bench](https://isaiuseful.com/benchmarks.html.md#swe-bench) gives agents real GitHub issues and grades the resulting patches with tests. It is a practical model for bounded maintenance work.
- [Understand the benchmark and variants →](https://isaiuseful.com/benchmarks.html.md#swe-bench)
- [Open the benchmark repository →](https://github.com/SWE-bench/SWE-bench)
coding
### Build from a durable specification
OpenSpec and GitHub Spec Kit make the specification a reviewable artifact before generation, reducing ambiguity and keeping the workflow transferable between models.
- [OpenSpec →](https://github.com/Fission-AI/openspec)
·
- [Spec Kit →](https://github.com/github/spec-kit)
coding
### Continue a project across agent sessions
A harness can preserve progress notes, git history and test state so a later context window resumes work instead of rebuilding an understanding from scratch.
- [Vendor implementation report →](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents)
coding · planning
### Plan a migration before touching the code
A specification workflow can inventory affected behavior, propose milestones and preserve non-goals before an agent starts a cross-cutting upgrade.
- [Spec Kit →](https://github.com/github/spec-kit)
·
- [OpenSpec →](https://github.com/Fission-AI/openspec)
coding · verification
### Turn a reported bug into a regression test
Before accepting a fix, require a test that fails on the original behavior and passes after the patch. This is a practitioner verification rule, not a finding from the benchmark.
- [Practitioner method →](https://simonwillison.net/2025/Dec/18/code-proven-to-work/)
research
### Produce a cited research brief
Deep-research agents browse many sources, synthesize them and return citations. The researcher still checks provenance, coverage and whether each citation supports the claim.
- [Vendor capability report →](https://openai.com/index/introducing-deep-research/)
research
### Recover primary sources from weak clues
LLMs can act as secondary librarians: turn a remembered quote, chart or claim into candidate sources, while the human verifies the original material.
- [Practitioner examples →](https://simonwillison.net/tags/ai-assisted-search/)
research
### Read and compare a document collection
AnythingLLM provides a public implementation of retrieval over local documents, letting users ask cross-document questions while retaining the original files as evidence.
- [Follow the RAG build sequence →](https://isaiuseful.com/rag.html.md#build)
- [Project documentation →](https://github.com/Mintplex-Labs/anything-llm)
research
### Extract structure from PDFs and reports
Docling converts complex PDFs, tables and page layouts into structured representations that downstream retrieval or analysis workflows can actually use.
- [Project documentation →](https://github.com/docling-project/docling)
research · verification
### Turn a draft into a claim-checking queue
A proposed workflow can extract consequential claims, locate candidate primary sources and flag partial support. The linked system documents source-finding and citations, not this exact review queue.
- [Underlying vendor capability →](https://openai.com/index/introducing-deep-research/)
research · comparison
### Compare revisions across a policy set
A proposed workflow can load named document versions into a bounded collection, ask for candidate differences and verify each one against the originals. The linked project supplies retrieval, not a purpose-built policy-diff guarantee.
- [Retrieval implementation →](https://github.com/Mintplex-Labs/anything-llm)
operations
### Run long, bounded workflows
Agent harnesses can retain task state, resume after context limits, run checks and stop on explicit completion conditions instead of relying on a single conversation.
- [Vendor implementation report →](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents)
operations
### Operate software that has no API
Computer-use agents inspect screenshots, click and type in existing interfaces. Use narrow accounts, allowlisted sites and approval gates because visual automation is inherently brittle.
- [Vendor capability release →](https://www.anthropic.com/news/3-5-models-and-computer-use)
operations
### Transcribe and search meetings
Whisper turns recordings into searchable transcripts. A follow-on model can extract decisions, owners and unanswered questions while the recording remains the canonical source.
- [Transcription implementation →](https://github.com/openai/whisper)
operations
### Route, classify and draft from an inbox
A proposed workflow can route inbound messages, retrieve policy and prepare a response behind human approval. The linked guide documents routing and tool-use patterns, not this exact inbox deployment.
- [Vendor engineering pattern →](https://www.anthropic.com/engineering/building-effective-agents)
operations
### Run repeatable agent loops with state and review
A practitioner pattern combines scheduled runs, a state file, isolated worktrees, tools, a separate verifier and human review. It is a method proposal, not evidence of a measured reliability gain.
- [Build a safer loop →](https://isaiuseful.com/guides.html.md#loops)
- [Reference repository →](https://github.com/cobusgreyling/loop-engineering)
operations · documents
### Extract invoice lines and table fields for review
Parse pages and tables into a fixed schema, validate totals with deterministic code and send low-confidence fields to a person before anything reaches accounting.
- [Document parsing implementation →](https://github.com/docling-project/docling)
operations · reporting
### Prepare a weekly operations brief
A proposed workflow can gather allowlisted metrics, summarize exceptions and draft next actions while deterministic queries remain the source of every number. The source documents the component patterns, not this specific use case.
- [Vendor engineering patterns →](https://www.anthropic.com/engineering/building-effective-agents)
small business
### Assist customer-support agents
A generative assistant can retrieve answers and suggest responses while the support worker remains responsible. A 5,179-worker field study measured higher issues resolved per hour.
- [Field study →](https://www.nber.org/papers/w31161)
small business
### Answer questions and call a quote tool
A support agent can ground answers in product documentation and call a deterministic quote generator instead of inventing prices in free-form text.
- [Vendor implementation guide →](https://platform.claude.com/docs/en/about-claude/use-case-guides/customer-support-chat)
small business
### Prepare for a sales meeting
Anthropic reports that ServiceNow combined enterprise context and web research for sales preparation and saw up to a 95% reduction in preparation time in early testing. The page provides no independent audit or study methodology.
- [Vendor/partner report →](https://www.anthropic.com/news/servicenow-anthropic-claude)
small business
### Draft routine business writing
First drafts of emails, press releases, reports and short analyses are a measured use case: a randomized experiment found faster completion and higher evaluator-scored quality.
- [Peer-reviewed experiment →](https://www.science.org/doi/10.1126/science.adh2586)
small business · caution
### Assist with inventory, without owning the money
Project Vend's agent managed pricing, stock decisions and customer requests but lost money through concrete errors. It supports testing recommendations before granting spending authority.
- [Vendor research report →](https://www.anthropic.com/research/project-vend-1)
small business · routing
### Qualify inbound requests before a person replies
A proposed routing workflow can classify requests by explicit criteria, retrieve the relevant offer and prepare follow-up questions without making commitments. The source documents the routing pattern, not a measured sales result.
- [Vendor engineering pattern →](https://www.anthropic.com/engineering/building-effective-agents)
personal
### Run a persistent personal assistant
Hermes maintains memories, searches previous conversations and creates reusable skills. A phone can be the messaging surface while the agent and model stay on another machine; the chosen messaging provider remains part of the conversation path.
- [Messaging gateway →](https://hermes-agent.nousresearch.com/docs/user-guide/messaging/)
- [Project repository →](https://github.com/NousResearch/hermes-agent)
personal
### Search a personal knowledge base
Store notes as local Markdown, then index a copy in a retrieval tool. Obsidian supplies the file-based notes; AnythingLLM supplies the AI retrieval layer. Keep the notes—not generated summaries—canonical.
- [Obsidian documentation →](https://obsidian.md/help/)
- [Build the second-brain loop →](https://isaiuseful.com/guides.html.md#second-brain)
personal · high trust
### Reach a self-hosted agent from messaging
OpenClaw can keep sessions and state on an always-on host while an iPhone or Android device acts as the operator or a narrowly scoped node. Keep the gateway private, give it narrow permissions and require approval for sending, deleting, purchasing or changing accounts.
- [Remote gateway →](https://docs.openclaw.ai/gateway/remote)
- [See the private topology →](https://isaiuseful.com/remote-spark.html.md#operators)
personal · local
### Transcribe and search private voice notes
Run speech recognition locally, attach timestamps and retain the audio as the source so later summaries and task extraction remain checkable.
- [Whisper implementation →](https://github.com/openai/whisper)
education
### Coach the tutor during a lesson
Tutor CoPilot suggests questions and teaching moves to a human tutor in real time. The preprint reports a 4-percentage-point mastery gain overall and 9 points for students of lower-rated tutors.
- [Randomized-trial preprint →](https://arxiv.org/abs/2410.03017)
healthcare · draft only
### Prepare replies to patient messages
AI can prepare drafts for clinician review. A five-week single-group study found 20% adoption and no measured time savings; lower surveyed burden and exhaustion were observational, so the safe claim is feasibility—not proven efficiency.
- [Single-group QI study →](https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2816494)
science
### Predict molecular structures and interactions
The peer-reviewed AlphaFold 3 system predicts joint structures containing proteins, nucleic acids, small molecules, ions and modified residues. Predictions are hypotheses to test, not experimental confirmation.
- [Peer-reviewed system paper →](https://www.nature.com/articles/s41586-024-07487-w)
52 vendor-reported deployments
## From prototypes to work at scale.
These additional examples come from Microsoft, AWS, Google Cloud, OpenAI, GitLab and Anthropic customer stories. They show concrete workflows and reported outcomes, but remain customer/vendor case studies—not independent experiments or proof that a result transfers unchanged.
Microsoft · 28 stories
### Azure, Foundry and Copilot deployments
Customer-reported implementations published by Microsoft.
operations · vendor case study
### Surface answers during customer calls
Microsoft reports that AT&T's digital coworkers reduced information-search time for customer-care staff by 33%. The wider platform had 71 generative-AI solutions in use by more than 100,000 employees.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25679-at-and-t-azure/)
engineering · early vendor result
### Question test-vehicle telemetry in natural language
BMW's multi-agent system retrieves, analyzes and visualizes test-fleet data for engineers. The story labels its early internal result as up to 12 times faster analysis potential—not a completed independent evaluation.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25683-bmw-ag-azure)
governance · vendor case study
### Summarize board materials before meetings
Nasdaq says Boardvantage's AI summarization saved governance teams more than 100 hours annually, with directors reporting up to 25% less preparation time and up to 60% less reading time.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25682-nasdaq-azure)
audit · vendor deployment
### Analyze full audit datasets and draft documentation
KPMG Clara uses agents to analyze datasets, prepare documentation and surface risk signals. The customer story emphasizes deployment scale—95,000 auditors across member firms in more than 140 countries—rather than a controlled productivity effect.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25353-kpmg-international-azure)
marketing · vendor case study
### Turn campaign history into targeting proposals
HicMobile converted unstructured campaign files into a searchable knowledge base and targeting tool. Its story reports a six-month prototype-to-launch cycle, described as a 70% faster timeline, and a prototype delivered 11.25 times faster than working without the co-innovation lab.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25313-hicmobile-azure-ai-foundry)
knowledge work · vendor adoption
### Give a governed assistant to a global workforce
Aon built AonGPT to connect internal knowledge and automate routine analysis. Microsoft reports more than 62,000 users, about 31,000 monthly active users and over 6.4 million messages exchanged.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25693-aon-plc-azure)
manufacturing · vendor case study
### Replan production when factory conditions change
Sight Machine used AI-generated optimization models for production scheduling. Microsoft reports that one beverage manufacturer cut non-value-added production time by 75% and increased production capacity by more than 5%.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/26648-sight-machine-microsoft-foundry)
banking support · vendor case study
### Resolve routine banking requests around the clock
Commerzbank's Ava agent handles more than 30,000 customer conversations a month. The bank reports that roughly 75% of requests in the designed scope are resolved autonomously.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25676-commerzbank-ag-azure-ai-foundry-agent-service)
fintech voice · vendor case study
### Complete financial tasks through conversation
Astra Tech embedded a multilingual voice assistant in botim for tasks such as money transfer. The story reports 3.6 million users and 375% wallet-transaction growth, while attributing that growth to broader product initiatives with the assistant as one driver.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25412-astra-tech-azure-ai-foundry)
personalization · vendor case study
### Personalize sports stories for millions of fans
The Premier League Companion combines live and historical data into individualized feeds. The league reports about 20% year-over-year growth in app and website consumption and 60 million active fans early in the season.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25725-premier-league-azure-ai-foundry)
personal assistant · vendor adoption
### Orchestrate routines in a consumer assistant
SK Telecom expanded A.(A-Dot) from single-turn responses into multi-step personal workflows. Microsoft reports growth from about 1.1 million monthly active users to more than 10 million subscribers and monthly active users by 2025, plus three to four months trimmed from feature delivery.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25680-sk-telecom-azure-ai-foundry)
banking service · vendor case study
### Route customer and employee banking questions
Banco Bradesco built a governed multi-agent platform for internal and external processes. The bank reports an 83% resolution rate in digital customer service and 80% for employee queries.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25660-banco-bradesco-sa-azure-ai-foundry)
healthcare voice · pilot
### Answer routine patient calls and return voicemails
healow is piloting Genie as an AI medical receptionist for appointment information, common questions and voicemail handling. The story describes expected workload and satisfaction benefits, but no controlled outcome measurement.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25363-healow-azure-kubernetes-service)
clinical documentation · vendor report
### Draft clinical notes from patient conversations
healow's Sunoh.ai transcribes visits and prepares a clinical note for provider approval. Microsoft reports clinicians saving up to two hours per day; that is a customer-reported result, not an independent trial.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/24772-healow-azure)
acute care · vendor case study
### Prepare emergency charts and discharge instructions
Sayvant transcribes acute-care conversations and drafts charts plus instructions in more than 30 languages. The company reports cutting charting from 10 minutes to under 90 seconds per patient and saving an estimated 50,000 clinician hours.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/23516-sayvant-azure-open-ai-service)
health appeals · vendor case study
### Draft appeal determination letters for nurses
Acentra Health's MedScribe prepares letters for nurse review. The company reports about 50% less time per letter, 11,000 nursing hours and nearly $800,000 saved, with a 99% approval rate for generated drafts.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/19280-acentra-health-azure)
education · vendor deployment
### Provide a curriculum-grounded tutor at distance-learning scale
Universitas Terbuka deployed an AI tutor across 500 classes and roughly 100,000 students. Its internal research across four courses and nearly 38,000 students found more discussion participation and higher assignment scores, reported as statistically significant.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/20284-universitas-terbuka-azure-blob-storage)
education assessment · vendor report
### Prepare teacher assessment from photographed work
A Discovery Trust teacher used Copilot to assess 33 pieces of work against Year 6 standards in 30 minutes instead of six hours. The trust estimates 4,500 hours reclaimed annually across its first 59 licenses.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25921-discovery-trust-microsoft-365-copilot)
education feedback · vendor deployment
### Give teachers feedback across millions of essays
São Paulo's education department uses AI to analyze and draft personalized feedback on nearly 10 million essays. Microsoft says the system saves educators thousands of hours; teachers remain responsible for instruction and use of the feedback.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/20134-seduc-azure-open-ai-service)
agronomy · early vendor result
### Retrieve regulated crop guidance at the edge
Bayer fine-tuned a small language model on crop-protection labels. Early users report 5–10% productivity gains and answers to complex questions in under 30 seconds instead of days; the result remains an early customer estimate.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25255-bayer-azure-phi)
pharma research · vendor case study
### Search decades of R&D without duplicating experiments
Almirall built a multilingual assistant over 400,000 documents spanning more than 50 years. Scientists reportedly locate past experiments in seconds rather than hours or days, while subject experts validate the results.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25322-almirall-azure-openai)
clinical research · vendor case study
### Test clinical hypotheses with governed code execution
Novo Nordisk's reasoning agent generates code and statistical analysis over harmonized clinical data, with human validation built in. Microsoft reports time to insight falling from weeks to minutes.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/26569-novo-nordisk-as-azure)
secure research · vendor adoption
### Provide general AI inside a high-security laboratory
Sandia National Laboratories deployed a custom AI chat service to nearly 17,000 employees in eight months. Its customer story reports about three minutes saved per question while keeping the service inside established security controls.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/24581-sandia-national-laboratories-azure)
logistics · vendor case study
### Turn freight-request emails into quotes
C.H. Robinson's workflow classifies inbound freight emails, extracts details, requests missing information and prepares a response. The company reports cutting average quote time from hours to 32 seconds and being on pace for a further 15% productivity increase.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/19575-ch-robinson-azure-ai-studio)
accessibility · vendor deployment
### Check documents for accessibility, tone and clarity
Scope's staff use an internal engine to review documents against organizational accessibility standards. The disability charity reports more than 26,000 prompts in a typical month and over 640 task-specific agents, but no controlled productivity estimate.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/26558-scope-microsoft-365-copilot)
creative marketing · vendor case study
### Generate personalized character interactions at campaign scale
DEPT built an interactive Sinterklaas experience using language and custom-voice models. Its retail client reports 300,000 users, more than 3 million interactions and 233% higher engagement than the previous year's holiday campaign.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/24205-dept-azure-ai-foundry)
frontline rail · vendor case study
### Retrieve the current operating procedure in seconds
Rumo gives train drivers authenticated access to a regulatory-team-approved knowledge base. It reports reducing average lookup time from more than four minutes to three seconds, reclaiming 7,644 hours annually and achieving payback in under two months.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/25656-rumo-microsoft-copilot-studio)
asset management · vendor case study
### Extract trade confirmations and compile ETF reports
CSOP built tools for inconsistent trade-confirmation documents and daily ETF reporting. The firm reports 99% automation in selected workflows, 10-minute tasks completed in 30 seconds and monthly reporting effort reduced by 75%.
- [Microsoft customer story →](https://www.microsoft.com/en/customers/story/24963-csop-asset-management-azure-ai-foundry/)
AWS · 8 stories
### Bedrock and AWS AI deployments
Customer-reported implementations published by Amazon Web Services.
recruiting · vendor case study
### Prioritize CVs against a job description
Ubidy uses an LLM pipeline to rank CVs for enterprise recruiting clients, and tests new models against fixed CV/job-description matches with known results. AWS reports a 95% reduction in CV-evaluation time; its story also identifies candidate-suitability AI as a high-risk use. Treat ranking as recruiter support, with documented fairness, privacy and human-decision controls—not automatic rejection.
- [AWS customer story →](https://aws.amazon.com/solutions/case-studies/ubidy-case-study/)
energy service · vendor case study
### Summarize long email threads and stage updates
Epilot built an Amazon Bedrock workflow that summarizes utility-customer email chains and proposes record updates for human confirmation. AWS reports an 87% reduction in email handling time across the evaluated workflow.
- [AWS customer story →](https://aws.amazon.com/solutions/case-studies/epilot-genai-case-study/)
voice support · vendor deployment
### Field high-volume delivery-support calls
DoorDash built a voice-operated support flow using Amazon Bedrock and Claude. The story reports hundreds of thousands of calls handled per day, response latency of 2.5 seconds or less and 50% less application-development time.
- [AWS customer story →](https://aws.amazon.com/solutions/case-studies/doordash-bedrock-case-study/)
invoice processing · vendor case study
### Automate invoice intake and customer onboarding
Ellby used Amazon Bedrock to raise automated invoice processing from under 60% to more than 94%. The company reports saving over 300 maintenance hours per month and cutting onboarding time by more than 55%.
- [AWS customer story →](https://aws.amazon.com/solutions/case-studies/ellby-case-study/)
drug research · vendor case study
### Search scientific data and literature with citations
Genentech's gRED Research Agent searches internal data and PubMed, then synthesizes cited answers to multi-step questions. AWS says work that took weeks can take minutes and projects more than 43,000 manual hours automated in biomarker validation.
- [AWS customer story →](https://aws.amazon.com/solutions/case-studies/genentech-generativeai-case-study/)
office software · vendor deployment
### Add drafting and slide generation to office software
WPS AI adds rewriting, proofreading and presentation generation to WPS Office for more than 200 million overseas users. WPS reports 30% higher R&D efficiency and 35% lower operational costs after its Amazon Bedrock rollout.
- [AWS customer story →](https://aws.amazon.com/solutions/case-studies/wps/)
tax compliance · vendor self-case study
### Monitor tax rules and summarize business impacts
Amazon Finance built World Wide Watch to identify, prioritize and summarize VAT-policy changes. Its AWS case study reports more than 90% accuracy and a 92% reduction in time to insight, from 26 minutes to two minutes per update.
- [AWS self-case study →](https://aws.amazon.com/solutions/case-studies/amazon-finance-case-study/)
contact center · vendor case study
### Retrieve support knowledge while agents work
Fractal Analytics built Knowledge Assist over enterprise content with Amazon Bedrock and semantic search. Its clients report 10–15% shorter average call handling and a 30% deflection rate for supported self-service questions.
- [AWS customer story →](https://aws.amazon.com/solutions/case-studies/fractal-analytics-case-study/)
Google Cloud · 8 stories
### Gemini and Vertex AI deployments
Customer-reported implementations published by Google Cloud.
mining intelligence · vendor case study
### Question data across a mine-to-port operation
Golden Energy Mines built GEMVIS, a multi-agent system that connects data across more than 50 applications. Google Cloud reports data retrieval falling from two days to under an hour and executive decision speed improving by over 90%.
- [See the sovereign intelligence stack →](https://isaiuseful.com/diy-palantir.html.md#stack)
- [Google Cloud customer story →](https://cloud.google.com/customers/gems)
quality audit · vendor case study
### Prepare engineering quality assessments
Cognizant fine-tuned Gemini on its audit knowledge to help assess thousands of projects consistently. The company reports preparation time falling from as much as six hours to one and a functional agent prototype delivered in one week.
- [Google Cloud customer story →](https://cloud.google.com/customers/cognizant)
content moderation · vendor case study
### Review advertising content at platform scale
Taboola moved advertising-content review to Gemini while rolling Google AI tools across the company. Google Cloud reports a 75% reduction in moderation costs and AI tools supporting daily workflows for 90% of employees.
- [Google Cloud customer story →](https://cloud.google.com/customers/taboola)
customer support · vendor case study
### Unify support knowledge and resolve more tickets
Mosaic AI uses Gemini and Vertex AI to connect customer-facing teams with company knowledge. Google Cloud reports over 50% more tickets resolved, response times more than 35% faster and up to 75% more tickets handled overall.
- [Google Cloud customer story →](https://cloud.google.com/customers/mosaic-ai)
telecom maintenance · vendor case study
### Automate telecom infrastructure management
IT-Development uses Gemini to automate maintenance workflows for telecom operators and tower companies. Its Google Cloud story reports onboarding reduced from weeks to days and about 30% lower infrastructure-management costs.
- [Google Cloud customer story →](https://cloud.google.com/customers/it-development)
AI evaluation · vendor deployment
### Evaluate AI applications at production scale
Galileo uses Gemini-based evaluation agents to test model behavior and risk. Google Cloud reports more than 1,000 AI applications assessed and over 20 million requests processed daily at roughly 300-millisecond latency.
- [Google Cloud customer story →](https://cloud.google.com/customers/galileo)
document work · vendor deployment
### Work across documents in one AI workspace
Macro uses Gemini for multi-document chat, editable maps and actions over connected work. Google Cloud reports more than 125,000 users and says 80% interact with the workspace through its AI chat and agentic features.
- [Google Cloud customer story →](https://cloud.google.com/customers/macro)
product research · vendor deployment
### Classify customer feedback before it becomes a product decision
Mattel built a feedback-classification system over social posts, reviews and direct communications, using the results to spot product, brand and supply-chain signals. Google Cloud reports analysis falling from a month to a minute and about $1 million in savings. This vendor account is not an accuracy study: inspect labels, samples and representative raw feedback before acting on a trend.
- [Google Cloud implementation report →](https://cloud.google.com/transform/mattel-gen-ai-customer-feedback-real-time-product-improvements)
OpenAI · 4 stories
### ChatGPT and OpenAI deployments
Customer-reported implementations published by OpenAI.
clinical review · vendor case study
### Prepare structured utilization-review rationales
AdventHealth uses ChatGPT for Healthcare to summarize charts, surface relevant details and draft rationales while physician advisors keep final judgment. OpenAI reports an 80% reduction in time spent on the measured administrative workflow.
- [OpenAI customer story →](https://openai.com/index/adventhealth/)
knowledge work · vendor case study
### Turn operational knowledge into a usable first draft
Recycling-equipment maker STADLER uses ChatGPT across drafting, summarization, translation and analysis. It reports 30–40% time savings on common knowledge tasks, 2.5 times faster first drafts and daily use above 85%.
- [OpenAI customer story →](https://openai.com/index/stadler/)
contract review · vendor implementation report
### Compare a contract with a company playbook and stage redlines
Ironclad's GPT-4-based AI Assist flags irregularities and suggests playbook-grounded clauses, while users can accept or reject every suggestion. OpenAI reports Ironclad users cutting an initial redlining pass from about 40 minutes to two; this is vendor evidence, not an independent legal-quality study. A qualified reviewer must own interpretation, negotiation and approval.
- [OpenAI implementation report →](https://openai.com/index/ironclad/)
travel operations · vendor case study
### Localize content and widen access to data work
Holiday Extras uses ChatGPT for multilingual content, data analysis, code debugging and support. The company reports more than 500 hours saved weekly, with 92% of employees saving over two hours a week.
- [OpenAI customer story →](https://openai.com/index/holiday-extras/)
GitLab · 1 story
### GitLab Duo deployment
Customer-reported implementation published by GitLab.
CI/CD troubleshooting · vendor case study
### Investigate a failed pipeline and prepare a proposed fix
Barclays uses GitLab Duo to explain failed job logs, identify a likely root cause and suggest a fix inside its developer workflow. GitLab's customer story describes one manager resolving a pipeline issue in seconds; that anecdote is not a controlled delivery-speed measure. Review the diagnosis and change before merge, and keep log and code access within approved controls.
- [GitLab customer story →](https://about.gitlab.com/customers/barclays-plc/)
Anthropic · 3 stories
### Claude deployments
Customer-reported implementations published by Anthropic.
financial analysis · vendor case study
### Accelerate analytics and query handling
IG Group uses Claude across analytics, content and strategic work. Anthropic reports about 70 analyst hours saved each week, full payback in under three months and a days-long executive-analysis exercise completed in under two hours.
- [Anthropic customer story →](https://www.anthropic.com/customers/ig-group)
software delivery · vendor case study
### Help developers understand and change complex code
Palo Alto Networks put Claude into developer tools through Google Cloud. Anthropic reports 20–30% higher feature-development velocity, onboarding reduced from months to weeks and a 70% faster pilot task for junior developers.
- [Anthropic customer story →](https://www.anthropic.com/customers/palo-alto-networks)
creative services · vendor case study
### Research and draft client proposals faster
Norwegian communications group TRY uses Claude across research, proposals, content and project work. Anthropic reports 30% less time on routine tasks, 40% faster proposal development and more than 50 use cases in operation.
- [Anthropic customer story →](https://www.anthropic.com/customers/try)
No use case matches this category, source collection and search. Try a broader choice.
chat-only mental model
### Ask, receive text, copy it somewhere
- No tools
- No direct inspection
- No memory of the work environment
- No deterministic completion check
agent mental model
### Observe, act, test, recover and report
- Reads relevant files and systems
- Uses terminals, APIs and applications
- Maintains task state and reusable skills
- Stops when tests or approval gates say it is done
**Mobile operating rule:** remote chat only exchanges messages; a remote assistant reads selected context and uses selected tools; a remote operator can change another machine or account. The phone-sized interface does not shrink the host-side permission boundary.
Turn an example into your workflow
Copy the pattern, not the demo. Define the artifact, boundary and acceptance test.
- [Choose a build guide](https://isaiuseful.com/guides.html.md#chooser)
- [Compare implementation paths](https://isaiuseful.com/guides.html.md#paths)
- [Scale a proven workflow](https://isaiuseful.com/adoption.html.md#walk)
---
## https://isaiuseful.com/adoption (`/adoption.html.md`)
# How Do Organizations Adopt AI Successfully?
Canonical source: [https://isaiuseful.com/adoption](https://isaiuseful.com/adoption)
Organizational AI adoption
AI adoption is not a technology project. It is an organizational transformation journey—one that starts by helping people find value in the work they already do.
Platforms, models and architecture matter later. First create the conditions for learning, then turn useful discoveries into reliable ways of working.
- [See the five stages](#journey)
- [Find the scaling breakpoint](#engineer)
- [See adoption metrics](#adoption-evidence)
Transformation map
Value compounds in sequence
1. 01 **Crawl** Learn safely
2. 02 **Walk** Prove value
3. 03 **Engineer** Make it repeatable
4. 04 **Run** Scale what works
5. 05 **Fly** Keep accelerating
How do we help thousands of people adopt AI successfully **without turning it into chaos?**
- [**16.3%** working-age global diffusion MICROSOFT ESTIMATE · H2 2025](#adoption-evidence)
- [**88%** organizations using AI AT LEAST ONE FUNCTION · 2025](#adoption-evidence)
- [**32.7%** EU individual use EUROSTAT SURVEY · 2025](#adoption-evidence)
The AI adoption journey
## What are the five stages of AI adoption?
Each stage has a different job. Applying enterprise controls too early suppresses discovery; scaling before a workflow is reliable multiplies inconsistency.
- [01 **Crawl** Curiosity & experimentation](#crawl)
- [02 **Walk** First proven business value](#walk)
- [03 **Engineer** Repeatable & governed workflow](#engineer)
- [04 **Run** Organization-wide scaling](#run)
- [05 **Fly** Continuous innovation](#fly)
The sequence
**Discover value** **prove it** **make it repeatable** **scale it** **keep improving it**
Curiosity before ROI
## Stage 1: How do you start AI adoption safely?
The goal is not measurable business value yet. The goal is organizational learning.
Give employees a safe place to explore, test ideas and notice where AI removes friction from everyday work. Keep the cost of trying small—and the permission to learn wide.
### Make experimentation ordinary
- Write and improve documents
- Summarize meetings and research
- Create presentation first drafts
- Assist with code and analysis
- Brainstorm, reframe and ideate
### What good looks like now
- Wide experimentation is encouraged
- Useful stories are collected and shared
- Individuals remove small points of friction
- Formal ROI is not the admission ticket
- Governance does not become a blanket blocker
Success metric
People stop asking “What is AI?” and start asking **“Can AI help me with this task?”**
Better, faster or cheaper
## Stage 2: How do you prove one valuable AI workflow?
Experimentation becomes adoption when one workflow clearly outperforms the old way of working.
Find one outcome valuable enough that the team would not willingly return to the previous process. It does not need to save millions; it needs to matter to the people doing the work.
### Choose a bounded workflow
Release notes
Social content
Meeting summaries
Support drafts
Knowledge search
Sales proposals
The proof test
Is the new workflow meaningfully **better** , **faster** or **cheaper** after review and correction?
Compare it with a real manual baseline. Use the [evidence claim boundary](https://isaiuseful.com/evidence.html.md#claim-boundary) to keep a narrow result narrow, then choose a [task-relevant benchmark or local test](https://isaiuseful.com/benchmarks.html.md#database) . A persuasive demo is not the same thing as a dependable result.
Success metric
A team can point to a specific workflow and say: **“We would not go back to doing this manually.”**
The scaling breakpoint
## Stage 3: How do you make AI workflows repeatable?
A good pilot is not a process. This is where isolated success becomes transferable capability.
**Most initiatives stall here.**
The pilot works because a few enthusiasts carry hidden knowledge. Remove that dependency before expanding.
### Make the workflow explicit
1. Input What enters the process?
2. Output What does “good” look like?
3. Owner Who is accountable?
4. Control What requires human review?
5. Failure What happens when AI is wrong?
6. Measure How will success be tracked?
### Build the operating layer
- Workflow and exception design
- Prompt and input standardization
- Human review and escalation paths
- KPIs, acceptance tests and quality checks
- Security, compliance and data review
- Clear governance proportional to risk
Perfection is not the goal. Consistency across people is. For software delivery, the [Thinking with AI workflow loop](https://isaiuseful.com/thinking-with-ai.html.md#loop) makes these controls reviewable.
Success metric
The process works reliably **regardless of which trained employee executes it.**
Scale proven value
## Stage 4: How do you scale AI across an organization?
Once the workflow is proven and repeatable, expand the pattern—not merely access to a tool.
Look for adjacent workflows across departments, standardize the reusable parts and measure realized benefits. AI now moves from individual productivity improvement to enterprise capability; multi-system operational work may need the interfaces and ownership model in the [sovereign operational-intelligence stack](https://isaiuseful.com/diy-palantir.html.md#stack) .
Sales
**Proposal generation**
Customer service
**Assisted responses**
HR
**Recruiting workflows**
Finance
**Assisted analysis**
Development
**Coding assistants**
01
**Adoption**
Change management that starts with the work
02
**Enablement**
Role-specific training and reusable playbooks
03
**Standards**
Shared platforms, controls and support
04
**Benefits**
Measured outcomes, not usage vanity metrics
Success metric
Multiple teams achieve measurable outcomes using **standardized AI-enabled processes.**
Center of Excellence
## Stage 5: How does an AI Center of Excellence help?
When AI becomes a core capability, give it dedicated leadership without creating a central bottleneck.
The operating principle
The CoE is an accelerator , not a gatekeeper.
It helps teams move faster and more safely by making scarce expertise, standards and reusable patterns available to everyone.
01
**Strategy** Set direction and prioritize opportunity
02
**Governance** Define policy proportional to impact
03
**Architecture** Guide platforms, patterns and vendors
04
**Enablement** Train people and spread practice
05
**Measurement** Track value and improve the portfolio
06
**Innovation** Evaluate what becomes possible next
Success metric
AI becomes an embedded capability that **continuously creates value across the organization.**
Key principle
## How do you build repeatable AI adoption capability?
The durable advantage is not access to a model. It is the organizational system that repeatedly turns useful ideas into dependable outcomes.
### People
Give employees permission to learn, role-specific support and a voice in redesigning their work.
### Process
Define ownership, quality, exceptions, approval boundaries and measures before expanding.
### Technology
Choose platforms and models that fit the proven workflow, its data and its risk—not the other way around.
Useful discovery
Repeatable workflow
Responsible scale
**Competitive advantage**
Adoption in numbers
## How widely is AI actually being used?
Population use, organizational surveys, active users, web visits and revenue measure different layers of adoption. They are shown separately.
Worldwide diffusion
16.3%
Working-age population estimated to have used generative AI in H2 2025.
Microsoft telemetry-based estimate
Organizational breadth
88%
Surveyed organizations reporting AI use in at least one business function in 2025.
McKinsey survey via Stanford AI Index
EU individual use
32.7%
People aged 16–74 reporting generative-AI use in the previous three months.
Eurostat survey, 2025
> Visual: Geographic reach. AI adoption by country. Microsoft estimates working-age use; Eurostat surveys recent use among people aged 16–74. The datasets remain separate.
Geographic reach
### AI adoption by country
Microsoft estimates working-age use; Eurostat surveys recent use among people aged 16–74. The datasets remain separate.
- Dataset option: Global working-age estimate
- Dataset option: Europe individual survey

Loading the map values…
Microsoft estimate · H2 2025
**United States 28.3%**
Estimated share of the working-age population using a generative-AI product during the half-year.
Up 2.0 percentage points from H1 2025.
| Map layer | Reference point | Period |
| --- | --- | --- |
| Global estimate | **UAE 64.0%** | H2 2025 |
| Global estimate | **U.S. 28.3%** | H2 2025 |
| Europe survey | **Estonia 46.64%** | Previous three months |
| Europe survey | **EU 32.66%** | Previous three months |
> Visual: Inside organizations. Access is broad; depth still varies. In 2025, 88% reported AI use in at least one function and 79% reported regular generative-AI use.
Inside organizations
### Access is broad; depth still varies
In 2025, 88% reported AI use in at least one function and 79% reported regular generative-AI use.
Loading the HTML chart…
Source note, checked 31 July 2026: the Stanford chapter PDF and figure report 79% for generative AI; the chapter landing-page summary says 70%. This chart retains the figure value while the publisher’s pages disagree.
> Visual: The adoption gap. Using AI is much more common than measuring value. McKinsey reports 88% regular use, 39% with any enterprise-level EBIT impact and 6% meeting its high-performer threshold. These are separate signals, not a funnel.
The adoption gap
### Using AI is much more common than measuring value
McKinsey reports 88% regular use, 39% with any enterprise-level EBIT impact and 6% meeting its high-performer threshold. These are separate signals, not a funnel.
Loading the HTML chart…
**19.8%**
of all U.S. businesses used AI
**32.0%**
of firms with 100–249 employees
**37.0%**
of firms with 250+ employees
**39.7%**
in the information sector
**33.9%**
in finance and insurance
**≈14.0%**
in retail trade
Census BTOS, period ending May 3, 2026; use in any business function during the previous two weeks.
> Visual: Agent-product reach. Codex grew fast—then its distribution changed. OpenAI disclosed 1.6M+ Codex weekly users in February and 5M+ in June. The broader post-rollout series rose from 8M in July to 15M+ in August; it combines Codex with ChatGPT Work.
Agent-product reach
### Codex grew fast—then its distribution changed
OpenAI disclosed 1.6M+ Codex weekly users in February and 5M+ in June. The broader post-rollout series rose from 8M in July to 15M+ in August; it combines Codex with ChatGPT Work.
Loading the HTML chart…
First-party agent telemetry
### Codex and Claude Code behavioral samples
Vendor studies of task scope, work mix and human–agent control—not independent productivity audits.
Individual, organizational and OpenAI-user samples
**Users with ≥1 request >30 min** — 80.6%
**Users with ≥1 request >1 h** — 70.2%
**Users with ≥1 request >8 h** — 25.6%
Human-time equivalents estimate task scope, not time saved or business value.
≈400,000 interactive sessions from ≈235,000 people
**Active runtime per user** — 20 h/week
**Sessions involving code work** — 56%
**Human planning / Claude execution** — 70% / 80%
Excludes third-party IDE, SDK and headless usage; work types and decisions are classified.
Competitive attention
### ChatGPT leads tracked generative-AI web visits
This Similarweb-based series excludes embedded, app, API and enterprise use. It measures web attention, not total product reach.
Highlight
Loading the interactive HTML chart…
Exact chart data and definitions
### AI diffusion by economy
**Microsoft estimate of working-age generative-AI use**
| Economy | H1 2025 | H2 2025 | Change |
| --- | --- | --- | --- |
| United Arab Emirates | 59.4% | 64.0% | +4.6 pp |
| Singapore | 58.6% | 60.9% | +2.3 pp |
| Norway | 45.3% | 46.4% | +1.1 pp |
| Ireland | 41.7% | 44.6% | +2.9 pp |
| France | 40.9% | 44.0% | +3.1 pp |
| Spain | 39.7% | 41.8% | +2.1 pp |
| New Zealand | 37.6% | 40.5% | +2.9 pp |
| Netherlands | 36.3% | 38.9% | +2.6 pp |
| United Kingdom | 36.4% | 38.9% | +2.5 pp |
| Qatar | 35.7% | 38.3% | +2.6 pp |
| Australia | 34.5% | 36.9% | +2.4 pp |
| Israel | 33.9% | 36.1% | +2.2 pp |
| Canada | 33.5% | 35.0% | +1.5 pp |
| South Korea | 25.9% | 30.7% | +4.8 pp |
| Germany | 26.5% | 28.6% | +2.1 pp |
| United States | 26.3% | 28.3% | +2.0 pp |
| South Africa | 19.3% | 21.1% | +1.8 pp |
| Japan | 16.7% | 19.1% | +2.4 pp |
| Mexico | 16.7% | 17.8% | +1.1 pp |
| Brazil | 15.6% | 17.1% | +1.5 pp |
| China | 15.4% | 16.3% | +0.9 pp |
| India | 14.2% | 15.7% | +1.5 pp |
### Europe individual use
**Eurostat 2025 survey: people aged 16–74 using generative AI in the previous three months**
| Geography | Used generative AI |
| --- | --- |
| European Union | 32.66% |
| Euro area | 34.20% |
| Belgium | 42.01% |
| Bulgaria | 22.50% |
| Czechia | 35.35% |
| Denmark | 48.44% |
| Germany | 32.25% |
| Estonia | 46.64% |
| Ireland | 44.93% |
| Greece | 44.09% |
| Spain | 37.88% |
| France | 37.46% |
| Croatia | 27.52% |
| Italy | 19.86% |
| Cyprus | 44.20% |
| Latvia | 33.40% |
| Lithuania | 36.89% |
| Luxembourg | 42.54% |
| Hungary | 29.56% |
| Malta | 46.46% |
| Netherlands | 44.70% |
| Austria | 39.42% |
| Poland | 22.68% |
| Portugal | 38.70% |
| Romania | 17.76% |
| Slovenia | 37.56% |
| Slovakia | 30.79% |
| Finland | 46.27% |
| Sweden | 42.01% |
| Norway | 56.32% |
| Switzerland | 47.02% |
| Bosnia and Herzegovina | 20.26% |
| North Macedonia | 22.03% |
| Albania | 27.28% |
| Serbia | 18.64% |
| Türkiye | 17.19% |
| Kosovo* | 44.85% |
* Eurostat’s geographic label; the designation is without prejudice to positions on status.
### Organizational adoption
**Share of surveyed organizations using AI in at least one function**
| Year | Any AI | Generative AI |
| --- | --- | --- |
| 2023 | 55% | Not reported |
| 2024 | 78% | 71% |
| 2025 | 88% | 79% |
### Generative-AI web traffic
**Estimated share of tracked generative-AI website visits**
| Platform | May 2025 | Mar 2026 | May 2026 |
| --- | --- | --- | --- |
| ChatGPT | 76.4% | 56.7% | 52.7% |
| Gemini | 8.9% | 25.5% | 27.3% |
| Claude | 1.6% | 6.0% | 8.9% |
| Grok | Not reported | 6.0% | 2.8% |
| Perplexity | Not reported | 2.0% | 1.3% |
### Use versus enterprise value
**Separate McKinsey 2025 organizational signals; not a funnel**
| Signal | Share | Meaning |
| --- | --- | --- |
| Regular AI use | 88% | Respondents saying their organization regularly uses AI in at least one business function. |
| Any enterprise EBIT impact | 39% | Respondents reporting any AI-attributable EBIT impact at enterprise level. |
| AI high performer | 6% | Respondents attributing at least 5% of EBIT to AI and reporting significant value from AI use. |
### U.S. business use
**Census BTOS, period ending May 3, 2026**
| Business group | Using AI |
| --- | --- |
| All U.S. businesses | 19.8% |
| 100–249 employees | 32.0% |
| 250+ employees | 37.0% |
| Information sector | 39.7% |
| Finance and insurance | 33.9% |
| Retail trade | About 14.0% |
### Codex product-adoption milestones
**Different scopes are kept as separate series**
| Date | Population | Users | Evidence |
| --- | --- | --- | --- |
| Feb 27, 2026 | Codex weekly users | More than 1.6M | OpenAI company disclosure |
| Jun 2, 2026 | Codex weekly users | More than 5M | OpenAI company disclosure |
| Jul 14, 2026 | Codex + ChatGPT Work active users (post-rollout) | Reached 8M | OpenAI Codex engineering lead on X |
| Jul 21, 2026 | Codex + ChatGPT Work active users (post-rollout) | 10M milestone | OpenAI Codex engineering lead on X |
| Aug 13, 2026 | Codex + ChatGPT Work active users (post-rollout) | More than 15M | OpenAI Codex engineering lead on X |
The product scope changed on July 9. Later milestones combine Codex and ChatGPT Work and omit an activity window.
Public-market lens
## What do we actually know about each major player?
SpaceX completed its IPO with xAI inside the group. OpenAI and Anthropic IPO filings and timing are still reported rather than public registration statements, so the status and financial numbers below remain dated signals—not investment advice.
**Latest disclosed or reported scale and financial signals, checked 13 August 2026**
| Player | Market status | Reach signal | Financial signal | How to read it |
| --- | --- | --- | --- | --- |
| **ChatGPT** OpenAI | Private; confidential IPO filing reported | **900M+ weekly active users** February 2026 | **$30B 2026 revenue target** projection reported by secondary sources | OpenAI also reported 50M+ consumer subscribers and 9M+ paying business users. WAU is a vendor-reported measure, not audited MAU. |
| **Gemini** Alphabet | Public parent (NASDAQ: GOOGL / GOOG) | **750M+ monthly active users** Q4 2025 | **Not disclosed separately** Gemini is embedded across several Alphabet products | A later 900M+ MAU figure was reported from Google I/O 2026. The ledger keeps the 750M earnings-call disclosure as the primary-source baseline. |
| **Claude** Anthropic | Private; confidential IPO filing reported | **No current company-wide MAU disclosed** July 2026 check | **$47B annualized run-rate revenue** company disclosure, May 2026 | Revenue run rate is not realized annual revenue. Claude has unusually strong enterprise exposure, so consumer traffic understates its commercial footprint. |
| **Grok** SpaceX / xAI | Public parent following SpaceX IPO | **117M monthly active users of Grok features** March 2026 | **Not disclosed separately** xAI is reported inside the combined SpaceX group | The SEC filing covers Grok features across X, web and apps; it is broader than standalone chatbot traffic. December 2025 was 89M MAU. |
| **Perplexity** Perplexity AI | Private | **No comparable current MAU disclosed** July 2026 check | **$450M–$500M annualized revenue** secondary estimate, Q2 2026 | Queries, visits and MAU are often mixed in third-party summaries. The revenue range is useful as an order-of-magnitude estimate, not an audited result. |
WAU / MAU
**Reach, not loyalty.** Weekly and monthly active users use different windows, may count embedded features, and are usually vendor reported.
Web share
**Momentum, not the whole market.** Website visits miss APIs, mobile apps, workplace licences, search integration and other embedded distribution.
Run-rate revenue
**A pace, not booked annual revenue.** It annualizes a recent period and can move quickly in either direction.
Projection
**A plan, not an outcome.** Keep targets visible because they matter to market expectations, but label them separately from realized revenue.
Sources, update policy and secondary watchlist
### Primary evidence used in the charts
- [Microsoft AI Economy Institute: Global AI Adoption in 2025](https://www.microsoft.com/en-us/corporate-responsibility/topics/ai-economy-institute/reports/global-ai-adoption-2025/) Telemetry-based global and economy diffusion.
- [Eurostat: use of AI by individuals](https://ec.europa.eu/eurostat/statistics-explained/index.php?title=Use_of_artificial_intelligence_by_individuals) Surveyed three-month use; data code `isoc_ai_iaiu` .
- [Stanford HAI: 2026 AI Index — Economy](https://hai.stanford.edu/ai-index/2026-ai-index-report/economy) McKinsey organizational adoption series and limitations.
- [OpenAI: Scaling AI for everyone](https://openai.com/index/scaling-ai-for-everyone/) ChatGPT WAU, consumer subscribers and paying business users.
- [OpenAI: Codex for knowledge work](https://openai.com/index/codex-for-knowledge-work/) More than 5M Codex weekly active users in June 2026.
- [OpenAI: how agents are transforming work](https://openai.com/index/how-agents-are-transforming-work/) First-party Codex task-horizon, cross-role and internal-use telemetry.
- [OpenAI: introducing ChatGPT Work](https://openai.com/index/chatgpt-for-your-most-ambitious-work/) Desktop-app migration, ChatGPT Classic rename and bundle distribution context.
- [OpenAI Help: ChatGPT Work and Codex](https://help.openai.com/en/articles/20001275-chatgpt-work-and-codex) Current product boundaries and availability.
- [U.S. Census Bureau: AI use by businesses](https://www.census.gov/library/stories/2026/05/ai-use-businesses.html) Operational use by company size and sector.
- [McKinsey: The state of AI in 2025](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai) Enterprise EBIT impact, scaling and high-performer definitions.
- [OpenAI Codex lead: 8M combined users](https://x.com/thsottiaux/status/2077114635308986427) Vendor-representative signal for Codex plus ChatGPT Work, 14 July 2026.
- [OpenAI Codex lead: 10M milestone](https://x.com/thsottiaux/status/2079609157934886975) Later combined milestone, 21 July 2026; activity window unstated.
- [OpenAI Codex lead: crossed 15M active users](https://x.com/thsottiaux/status/2087706104814023111) Combined Codex plus ChatGPT Work milestone disclosed 13 August 2026; activity window unstated.
- [Alphabet Q4 2025 earnings call](https://abc.xyz/investor/events/event-details/2026/2025-Q4-Earnings-Call-2026-Dr_C033hS6/default.aspx) Gemini app MAU.
- [Anthropic Series H announcement](https://www.anthropic.com/news/series-h) Annualized run-rate revenue.
- [Anthropic: agentic coding and persistent returns to expertise](https://www.anthropic.com/research/claude-code-expertise) Privacy-preserving study of about 400,000 Claude Code sessions from about 235,000 people.
- [SpaceX final IPO prospectus](https://www.sec.gov/Archives/edgar/data/1181412/000162828026042639/spaceexplorationtechnologi.htm) Grok-feature MAU and xAI group context.
### Secondary sources monitored
These sources are useful for discovery, traffic estimates, projections and cross-checking. A number moves into a chart only when its definition, period and provenance survive review.
- [First Page Sage](https://firstpagesage.com/seo-blog/chatgpt-usage-statistics/)
- [FATJOE AI overview](https://fatjoe.com/blog/ai-stats/)
- [ChatGPT](https://fatjoe.com/blog/chatgpt-stats/)
- [Gemini](https://fatjoe.com/blog/google-gemini-stats/)
- [Claude](https://fatjoe.com/blog/claude-ai-stats/)
- [Grok](https://fatjoe.com/blog/grok-ai-stats/)
- [Perplexity](https://fatjoe.com/blog/perplexity-ai-stats/)
- [OpenClaw](https://fatjoe.com/blog/openclaw-ai-stats/)
- [Forbes Advisor](https://www.forbes.com/advisor/business/ai-statistics/)
- [Elfsight](https://elfsight.com/blog/ai-usage-statistics/)
- [Digital Applied](https://www.digitalapplied.com/blog/ai-usage-statistics-2026-who-uses-ai-how-much-data)
- [Statista](https://www.statista.com/topics/3104/artificial-intelligence-ai-worldwide/)
- [The Global Statistics](https://www.theglobalstatistics.com/artificial-intelligence-ai-usage-statistics/)
- [Exploding Topics](https://explodingtopics.com/blog/ai-usage-statistics)
- [Axios Codex report](https://www.axios.com/2026/06/02/openai-codex-knowledge-workers)
- [Pickaxe adoption gap](https://pickaxe.co/post/ai-adoption-gap)
- [Haider Codex post](https://x.com/haider1/status/2076963530763542590)
Chart data last refreshed 2026-08-13 . The automated updater accepts only bounded primary-source changes; blocked, ambiguous and secondary figures remain queued for manual review.
The leadership question
How do we get thousands of people to adopt AI successfully without turning it into chaos?
- [Find the first workflow](https://isaiuseful.com/use-cases.html.md)
- [Engineer the implementation](https://isaiuseful.com/guides.html.md#paths)
---
## https://isaiuseful.com/investors (`/investors.html.md`)
# Map the AI market before you form a view.
Canonical source: [https://isaiuseful.com/investors](https://isaiuseful.com/investors)
Independent research starter · reviewed 4 August 2026
Technology value, adopter ROI, supplier economics and security returns are different questions. Use this neutral map to collect filings and operating evidence, write a thesis or counter-thesis, decide what would falsify it, and reach your own conclusion.
- [Follow the eight gates](#stack)
- [Browse public companies](#players)
> map the layer
> collect primary evidence
> write the countercase
> name the falsifier
4
separate ledgers before an AI story becomes a research thesis.
**Evidence**
what is measured?
**Cash**
who captures it?
**Price**
what is assumed?
**Risk**
what changes it?
01 · Start here
## Keep four questions on separate sheets.
A useful system can help a worker, a customer can see a return, and a supplier can report revenue. None of those observations alone settles what an owner of a security receives.
01
### Usefulness
Does a bounded workflow become faster, safer, more accurate or more accessible after review and operating cost?
**Collect: accepted-work evidence.**
02
### Customer economics
Does the adopter retain value after licences, integration, data, human review, governance and change management?
**Collect: repeatable payback evidence.**
03
### Supplier economics
Which layer receives revenue, and what remains after power, depreciation, model renewal, sales expense and competition?
**Collect: cash-conversion evidence.**
04
### Security economics
What growth, margins, reinvestment and durability appear embedded in a market price, and how does a less-perfect outcome change the case?
**Collect: downside sensitivities.**
The site’s [claim-boundary guide](https://isaiuseful.com/evidence.html.md#claim-boundary) explains why task results do not establish sector-wide outcomes.
02 · Follow the money
## Follow the money through eight gates.
This is an illustrative path, not a chain of ownership or a claim that one layer captures more value. Demand can enter at several gates, and roles can overlap.
1. 01 Power, land & datacentres Grid access, cooling, land, construction, capacity and financing. Ask: what is actually deliverable?
2. 02 Chip equipment & foundries Tools, fabs, packaging and manufacturing capacity. Ask: where is the constraint?
3. 03 Compute, memory & networking Accelerators, memory, interconnects and systems. Ask: what can substitute?
4. 04 Cloud & distribution Managed capacity, platforms, developer reach and enterprise channels. Ask: does usage become cash?
5. 05 Models Training, post-training, inference and model operations. Ask: can capability sustain a premium?
6. 06 Data & context Permissions, records, evaluation, governance and integration. Ask: what is hard to recreate?
7. 07 Applications & workflows Repeatable jobs, adoption, outcomes, renewal and support. Ask: who owns the job-to-be-done?
8. 08 Edge & devices Local runtimes, embedded systems, privacy, latency and resilience. Ask: where should work run?
03 · Evidence before narrative
## Build a file that can disagree with you.
Start with primary documents. Separate reported measures from adjusted measures, management explanations, independent research and your own scenarios.
### Useful evidence labels
**Official filing** — Financial statements, notes and regulatory filings; check definitions and accounting policy.
**Official IR** — Company presentations, earnings materials and investor-relations pages; useful context, not independent proof.
**Research view** — An external method or interpretation; retain its date, scope and assumptions.
**Your scenario** — A stated condition to test, not a forecast or a conclusion.
### Research sequence
1. Map the customer workflow and buyer urgency.
2. Read revenue quality and customer concentration.
3. Bridge adjusted measures to cash and per-share outcomes.
4. List substitutes, switching costs and supply constraints.
5. Trace capex, depreciation, financing and replacement cycles.
6. Write a counter-thesis before collecting confirming evidence.
7. Set a dated decision rule and falsifier.
8. Record what changed at the next filing.
**Use the full cash bridge.** Revenue and a headline adjusted metric are not cash conversion. Read the reconciliation, capital expenditure, depreciation, stock compensation or dilution, debt and replacement cycle together.
**Keep a countercase alive.** A thesis is a provisional explanation. It becomes more useful when the researcher can say which evidence would make them stop relying on it.
04 · Public-company directory
## Research public companies by region and market layer.
A diversified, non-exhaustive directory of public suppliers, platforms and adopters referenced across this site. Region means headquarters or primary corporate base, not an exchange listing or broker availability. Each card shows one primary research area only; participation across areas can overlap.
Geographic area
Showing 124 companies.
### ABB
Industrial automation, robotics and electrification provider.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://global.abb/group/en/investors/overview)
### Acer
Computer, display and connected-device manufacturer.
**Region** — Asia-Pacific
**Primary area** — Edge & devices
- [Official investor relations](https://www.acer.com/corporate/en/investor-relations)
### Advantest
Test and measurement systems for semiconductor production.
**Region** — Asia-Pacific
**Primary area** — Compute, memory & networking
- [Official investor relations](https://www.advantest.com/en/investors/)
### AGCO
Precision-agriculture and autonomous farm-equipment systems provider.
**Region** — North America
**Primary area** — Robotics & automation
- [Official investor relations](https://investors.agcocorp.com/)
### Airbus
Aerospace manufacturer and industrial-automation adopter.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://www.airbus.com/en/investors)
### AIXTRON
Deposition equipment supplier for compound semiconductors.
**Region** — Europe
**Primary area** — Equipment & foundries
- [Official investor relations](https://www.aixtron-se.com/en/investors/)
### Alibaba
Digital commerce and cloud-services operator.
**Region** — Asia-Pacific
**Primary area** — Cloud & distribution
- [Official investor relations](https://www.alibabagroup.com/en-US/ir-home)
### Alphabet
Internet services, cloud and AI platform operator.
**Region** — North America
**Primary area** — Cloud & distribution
- [Official investor relations](https://abc.xyz/investor/)
### Amazon
Commerce, cloud infrastructure and service operator.
**Region** — North America
**Primary area** — Cloud & distribution
- [Official investor relations](https://ir.aboutamazon.com/)
### AMD
Compute and graphics semiconductor designer.
**Region** — North America
**Primary area** — Compute, memory & networking
- [Official investor relations](https://ir.amd.com/)
### Apple
Consumer-device and software ecosystem operator.
**Region** — North America
**Primary area** — Edge & devices
- [Official investor relations](https://investor.apple.com/)
### Applied Digital
Data-centre infrastructure and cloud-services operator.
**Region** — North America
**Primary area** — Power & datacentres
- [Official investor relations](https://ir.applieddigital.com/)
### Applied Materials
Semiconductor manufacturing-equipment supplier.
**Region** — North America
**Primary area** — Equipment & foundries
- [Official investor relations](https://ir.appliedmaterials.com/)
### Arista Networks
Data-centre networking equipment and software provider.
**Region** — North America
**Primary area** — Compute, memory & networking
- [Official investor relations](https://investors.arista.com/)
### Arm
Processor architecture and software developer.
**Region** — Europe
**Primary area** — Edge & devices
- [Official investor relations](https://investors.arm.com/)
### ASM International
Semiconductor manufacturing-equipment supplier.
**Region** — Europe
**Primary area** — Equipment & foundries
- [Official investor relations](https://www.asm.com/investors/)
### ASML
Semiconductor lithography-systems supplier.
**Region** — Europe
**Primary area** — Equipment & foundries
- [Official investor relations](https://www.asml.com/en/investors)
### ASUS
Computer, component and connected-device manufacturer.
**Region** — Asia-Pacific
**Primary area** — Edge & devices
- [Official investor relations](https://www.asus.com/event/Investor/)
### AutoStore
Warehouse cube-storage robotics and automation provider.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://www.autostoresystem.com/investors)
### Baidu
Internet services and AI cloud operator.
**Region** — Asia-Pacific
**Primary area** — Cloud & distribution
- [Official investor relations](https://ir.baidu.com/)
### BE Semiconductor Industries (Besi)
Semiconductor assembly and packaging-equipment supplier.
**Region** — Europe
**Primary area** — Equipment & foundries
- [Official investor relations](https://www.besi.com/investor-relations/)
### Block
Commerce, payments and financial-software platform operator.
**Region** — North America
**Primary area** — Applications & workflows
- [Official investor relations](https://investors.block.xyz/overview/default.aspx)
### BMW Group
Vehicle manufacturer and named humanoid-robotics deployment partner.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://www.bmwgroup.com/en/investor-relations.html)
### Broadcom
Semiconductor and infrastructure-software provider.
**Region** — North America
**Primary area** — Compute, memory & networking
- [Official investor relations](https://investors.broadcom.com/)
### Cadence Design Systems
Electronic-design automation software provider.
**Region** — North America
**Primary area** — Equipment & foundries
- [Official investor relations](https://investor.cadence.com/)
### Cloudflare
Network, security and developer-platform operator.
**Region** — North America
**Primary area** — Cloud & distribution
- [Official investor relations](https://cloudflare.net/investors/)
### CNH Industrial
Agriculture and construction equipment and automation provider.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://investors.cnh.com/overview/default.aspx)
### Constellation Energy
Power generation and energy-services operator.
**Region** — North America
**Primary area** — Power & datacentres
- [Official investor relations](https://investors.constellationenergy.com/)
### CoreWeave
Cloud-computing infrastructure operator.
**Region** — North America
**Primary area** — Cloud & distribution
- [Official investor relations](https://investors.coreweave.com/)
### Daifuku
Material-handling and warehouse-automation systems provider.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://www.daifuku.com/ir/)
### Dassault Systèmes
Engineering and product-lifecycle software provider.
**Region** — Europe
**Primary area** — Applications & workflows
- [Official investor relations](https://investor.3ds.com/)
### Deere & Company
Agriculture and construction equipment with autonomous and precision systems.
**Region** — North America
**Primary area** — Robotics & automation
- [Official investor relations](https://investor.deere.com/)
### Dell Technologies
Enterprise systems, storage and device supplier.
**Region** — North America
**Primary area** — Edge & devices
- [Official investor relations](https://investors.delltechnologies.com/)
### Delta Electronics
Power electronics and thermal-management supplier.
**Region** — Asia-Pacific
**Primary area** — Power & datacentres
- [Official investor relations](https://www.deltaww.com/en-US/IR)
### Digital Realty
Data-centre and digital-infrastructure operator.
**Region** — North America
**Primary area** — Power & datacentres
- [Official investor relations](https://investor.digitalrealty.com/)
### DISCO
Semiconductor dicing and grinding-equipment supplier.
**Region** — Asia-Pacific
**Primary area** — Equipment & foundries
- [Official investor relations](https://www.disco.co.jp/eg/investor/)
### Dobot
Collaborative and industrial-robotics systems provider.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official HKEX issuer disclosures](https://www1.hkexnews.hk/search/titlesearch.xhtml?category=0&lang=EN&market=SEHK&stockId=1000243257)
### Eaton
Power-management and electrical-systems provider.
**Region** — Europe
**Primary area** — Power & datacentres
- [Official investor relations](https://www.eaton.com/us/en-us/company/investor-relations.html)
### Elastic
Search, observability and security software provider.
**Region** — North America
**Primary area** — Models & context
- [Official investor relations](https://ir.elastic.co/overview/default.aspx)
### Equinix
Digital-infrastructure and interconnection operator.
**Region** — North America
**Primary area** — Power & datacentres
- [Official investor relations](https://investor.equinix.com/)
### Ericsson
Telecommunications network-infrastructure provider.
**Region** — Europe
**Primary area** — Compute, memory & networking
- [Official investor relations](https://www.ericsson.com/en/investors)
### FANUC
Industrial-robot, CNC and factory-automation systems provider.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://www.fanuc.co.jp/en/ir/)
### Fincantieri
Shipbuilder and named humanoid-welding development partner.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://www.fincantieri.com/en/investor-relations)
### Fujitsu
Enterprise technology and services provider.
**Region** — Asia-Pacific
**Primary area** — Applications & workflows
- [Official investor relations](https://www.fujitsu.com/global/about/ir/)
### GIGABYTE
Computer hardware and component manufacturer.
**Region** — Asia-Pacific
**Primary area** — Edge & devices
- [Official investor relations](https://www.gigabyte.com/Investor)
### GitLab
Software-development lifecycle platform operator.
**Region** — North America
**Primary area** — Applications & workflows
- [Official investor relations](https://ir.gitlab.com/)
### GXO Logistics
Contract-logistics operator and named live humanoid-robotics deployment partner.
**Region** — North America
**Primary area** — Robotics & automation
- [Official investor relations](https://investors.gxo.com/)
### Hewlett Packard Enterprise
Enterprise compute, storage and networking provider.
**Region** — North America
**Primary area** — Compute, memory & networking
- [Official investor relations](https://investors.hpe.com/)
### Hitachi
Industrial, digital and infrastructure-systems provider.
**Region** — Asia-Pacific
**Primary area** — Applications & workflows
- [Official investor relations](https://www.hitachi.com/IR-e/)
### Hon Hai Precision (Foxconn)
Electronics manufacturing and systems integrator.
**Region** — Asia-Pacific
**Primary area** — Edge & devices
- [Official investor relations](https://www.honhai.com/en-US/investor)
### Horizon Robotics
Automotive and physical-AI compute and assisted-driving systems provider.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official HKEX issuer disclosures](https://www1.hkexnews.hk/search/titlesearch.xhtml?category=0&lang=EN&market=SEHK&stockId=1000238030)
### HP Inc.
Personal systems and printing-device manufacturer.
**Region** — North America
**Primary area** — Edge & devices
- [Official investor relations](https://investor.hp.com/overview/default.aspx)
### Huayan Robotics
Industrial robotic-arms and automation provider.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official HKEX issuer disclosures](https://www1.hkexnews.hk/search/titlesearch.xhtml?category=0&lang=EN&market=SEHK&stockId=1000297158)
### Husqvarna Group
Robotic outdoor-equipment systems provider.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://www.husqvarnagroup.com/en/investors)
### Hut 8
Energy and digital-infrastructure operator.
**Region** — North America
**Primary area** — Power & datacentres
- [Official investor relations](https://www.hut8.com/investors)
### Hyundai Motor
Vehicle manufacturer and majority owner of Boston Dynamics.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://www.hyundai.com/worldwide/en/company/ir)
### IBM
Enterprise software, infrastructure and services provider.
**Region** — North America
**Primary area** — Applications & workflows
- [Official investor relations](https://www.ibm.com/investor)
### Infineon
Semiconductor supplier for power and embedded systems.
**Region** — Europe
**Primary area** — Compute, memory & networking
- [Official investor relations](https://www.infineon.com/about/investor)
### Intel
Semiconductor design and manufacturing company.
**Region** — North America
**Primary area** — Compute, memory & networking
- [Official investor relations](https://www.intc.com/)
### Intuitive Surgical
Robotic-assisted surgery systems provider.
**Region** — North America
**Primary area** — Robotics & automation
- [Official investor relations](https://investor.intuitivesurgical.com/)
### IREN
Data-centre and cloud-computing infrastructure operator.
**Region** — Asia-Pacific
**Primary area** — Cloud & distribution
- [Official investor relations](https://www.iren.com/investors)
### KION Group
Warehouse equipment, intralogistics and automation provider.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://www.kiongroup.com/en/Investor-Relations/)
### KLA
Semiconductor process-control equipment supplier.
**Region** — North America
**Primary area** — Equipment & foundries
- [Official investor relations](https://ir.kla.com/)
### Kubota
Agricultural equipment with autonomous and precision-farming systems.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://www.kubota.com/ir/)
### Lam Research
Wafer-fabrication equipment supplier.
**Region** — North America
**Primary area** — Equipment & foundries
- [Official investor relations](https://investor.lamresearch.com/)
### Lenovo
Personal, enterprise and edge-device provider.
**Region** — Asia-Pacific
**Primary area** — Edge & devices
- [Official investor relations](https://investor.lenovo.com/en/)
### Marvell Technology
Data-infrastructure semiconductor designer.
**Region** — North America
**Primary area** — Compute, memory & networking
- [Official investor relations](https://investor.marvell.com/)
### MediaTek
Connectivity and device-system semiconductor designer.
**Region** — Asia-Pacific
**Primary area** — Edge & devices
- [Official investor relations](https://www.mediatek.com/investor-relations)
### Mercedes-Benz Group
Vehicle manufacturer and named humanoid-robotics deployment partner.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://group.mercedes-benz.com/investoren/)
### Meta Platforms
Digital-platform and infrastructure operator.
**Region** — North America
**Primary area** — Cloud & distribution
- [Official investor relations](https://investor.atmeta.com/)
### Micro-Star International (MSI)
Computer hardware and gaming-device manufacturer.
**Region** — Asia-Pacific
**Primary area** — Edge & devices
- [Official investor relations](https://www.msi.com/about/investor)
### Micron
Memory and storage semiconductor manufacturer.
**Region** — North America
**Primary area** — Compute, memory & networking
- [Official investor relations](https://investors.micron.com/)
### MicroPort MedBot
Surgical-robotics systems provider.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official HKEX issuer disclosures](https://www1.hkexnews.hk/search/titlesearch.xhtml?category=0&lang=EN&market=SEHK&stockId=1000118830)
### Microsoft
Enterprise software, cloud and platform operator.
**Region** — North America
**Primary area** — Cloud & distribution
- [Official investor relations](https://www.microsoft.com/en-us/Investor)
### Midea Group
Industrial and home-technology group with robotics and automation operations, including KUKA.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://www.midea.com.cn/en/Investors?wcmmode=disabled)
### Nebius Group
Cloud-computing infrastructure operator.
**Region** — Europe
**Primary area** — Cloud & distribution
- [Official investor relations](https://nebius.com/investors)
### NEC
Network, computing and digital-services provider.
**Region** — Asia-Pacific
**Primary area** — Compute, memory & networking
- [Official investor relations](https://www.nec.com/en/global/ir/)
### Nokia
Network infrastructure and telecommunications supplier.
**Region** — Europe
**Primary area** — Compute, memory & networking
- [Official investor relations](https://www.nokia.com/about-us/investors/)
### Nordic Semiconductor
Low-power wireless semiconductor designer.
**Region** — Europe
**Primary area** — Edge & devices
- [Official investor relations](https://www.nordicsemi.com/Investors)
### NVIDIA
Accelerated-computing hardware and software provider.
**Region** — North America
**Primary area** — Compute, memory & networking
- [Official investor relations](https://investor.nvidia.com/)
### NXP Semiconductors
Embedded and automotive semiconductor supplier.
**Region** — Europe
**Primary area** — Edge & devices
- [Official investor relations](https://investors.nxp.com/)
### Ocado Group
Online grocery and automation technology operator.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://www.ocadogroup.com/investors)
### OMRON
Industrial automation, sensing and control-systems provider.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://www.omron.com/global/en/ir/)
### Oracle
Enterprise software and cloud-infrastructure operator.
**Region** — North America
**Primary area** — Cloud & distribution
- [Official investor relations](https://investor.oracle.com/)
### OVHcloud
Cloud and data-centre infrastructure operator.
**Region** — Europe
**Primary area** — Cloud & distribution
- [Official investor relations](https://corporate.ovhcloud.com/en/investor-relations/)
### Palantir
Data integration and operational-software provider.
**Region** — North America
**Primary area** — Models & context
- [Official investor relations](https://investors.palantir.com/)
### Qualcomm
Wireless and edge-computing semiconductor designer.
**Region** — North America
**Primary area** — Edge & devices
- [Official investor relations](https://investor.qualcomm.com/)
### Rainbow Robotics
Humanoid and collaborative-robotics systems provider.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://www.rainbow-robotics.com/ir)
### Renault Group
Vehicle manufacturer and named humanoid-robotics deployment partner.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://www.renaultgroup.com/en/finance/)
### Renesas Electronics
Embedded and automotive semiconductor supplier.
**Region** — Asia-Pacific
**Primary area** — Edge & devices
- [Official investor relations](https://www.renesas.com/us/en/about/ir)
### Richtech Robotics
Service-robotics systems provider.
**Region** — North America
**Primary area** — Robotics & automation
- [Official investor relations](https://ir.richtechrobotics.com/)
### Rockwell Automation
Industrial-automation hardware, control and software provider.
**Region** — North America
**Primary area** — Robotics & automation
- [Official investor relations](https://www.rockwellautomation.com/en-us/company/investor-relations.html)
### Salesforce
Customer and enterprise workflow-software provider.
**Region** — North America
**Primary area** — Applications & workflows
- [Official investor relations](https://investor.salesforce.com/)
### Samsung Electronics
Device, memory and electronics manufacturer.
**Region** — Asia-Pacific
**Primary area** — Edge & devices
- [Official investor relations](https://www.samsung.com/global/ir/)
### SAP
Enterprise applications and business-process software provider.
**Region** — Europe
**Primary area** — Applications & workflows
- [Official investor relations](https://www.sap.com/investors/en.html)
### Schneider Electric
Energy-management and industrial-automation provider.
**Region** — Europe
**Primary area** — Power & datacentres
- [Official investor relations](https://www.se.com/ww/en/about-us/investor-relations/)
### SCREEN Holdings
Semiconductor production-equipment supplier.
**Region** — Asia-Pacific
**Primary area** — Equipment & foundries
- [Official investor relations](https://www.screen.co.jp/en/ir)
### Serve Robotics
Autonomous delivery and service-robotics operator.
**Region** — North America
**Primary area** — Robotics & automation
- [Official investor relations](https://ir.serverobotics.com/)
### ServiceNow
Enterprise workflow-software provider.
**Region** — North America
**Primary area** — Applications & workflows
- [Official investor relations](https://www.servicenow.com/company/investor-relations.html)
### Siemens
Industrial automation and digital-industry provider.
**Region** — Europe
**Primary area** — Robotics & automation
- [Official investor relations](https://www.siemens.com/global/en/company/investor-relations.html)
### Siemens Energy
Power-generation and grid-technology provider.
**Region** — Europe
**Primary area** — Power & datacentres
- [Official investor relations](https://www.siemens-energy.com/global/en/home/investor-relations.html)
### SK hynix
Memory semiconductor manufacturer.
**Region** — Asia-Pacific
**Primary area** — Compute, memory & networking
- [Official investor relations](https://www.skhynix.com/ir/UI-FR-IR01/)
### Snowflake
Cloud data-platform operator.
**Region** — North America
**Primary area** — Models & context
- [Official investor relations](https://investors.snowflake.com/)
### SoftBank Group
Technology investment and communications group.
**Region** — Asia-Pacific
**Primary area** — Cloud & distribution
- [Official investor relations](https://group.softbank/en/ir)
### Soitec
Engineered-substrate supplier for semiconductors.
**Region** — Europe
**Primary area** — Equipment & foundries
- [Official investor relations](https://www.soitec.com/en/investors)
### SpaceX
Space, satellite-connectivity and AI infrastructure operator.
**Region** — North America
**Primary area** — Cloud & distribution
- [Official investor relations](https://ir.spacex.com/investors/default.aspx)
### STMicroelectronics
Embedded and industrial semiconductor manufacturer.
**Region** — Europe
**Primary area** — Edge & devices
- [Official investor relations](https://investors.st.com/)
### Supermicro
Server, storage and data-centre systems provider.
**Region** — North America
**Primary area** — Compute, memory & networking
- [Official investor relations](https://ir.supermicro.com/)
### Symbotic
AI-enabled warehouse robotics and automation provider.
**Region** — North America
**Primary area** — Robotics & automation
- [Official investor relations](https://ir.symbotic.com/)
### Synopsys
Electronic-design automation software provider.
**Region** — North America
**Primary area** — Equipment & foundries
- [Official investor relations](https://investor.synopsys.com/)
### Tencent
Internet services, cloud and digital-platform operator.
**Region** — Asia-Pacific
**Primary area** — Cloud & distribution
- [Official investor relations](https://www.tencent.com/en-us/investors.html)
### Teradyne
Semiconductor test and collaborative-robotics systems provider.
**Region** — North America
**Primary area** — Robotics & automation
- [Official investor relations](https://investors.teradyne.com/)
### TeraWulf
Energy and digital-infrastructure operator.
**Region** — North America
**Primary area** — Power & datacentres
- [Official investor relations](https://investors.terawulf.com/)
### Tesla
Electric-vehicle, energy and humanoid-robotics developer.
**Region** — North America
**Primary area** — Robotics & automation
- [Official investor relations](https://ir.tesla.com/)
### Tokyo Electron
Semiconductor production-equipment supplier.
**Region** — Asia-Pacific
**Primary area** — Equipment & foundries
- [Official investor relations](https://www.tel.com/ir/)
### Toyota Motor
Vehicle manufacturer and named commercial-robotics deployment partner.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://global.toyota/en/ir/)
### TSMC
Semiconductor manufacturing and foundry operator.
**Region** — Asia-Pacific
**Primary area** — Equipment & foundries
- [Official investor relations](https://investor.tsmc.com/english)
### Twilio
Communications software and customer-engagement platform operator.
**Region** — North America
**Primary area** — Applications & workflows
- [Official investor relations](https://investors.twilio.com/)
### UBTECH Robotics
Humanoid and service-robotics systems provider.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://www.ubtrobot.com/en/investor-relations)
### UMC
Semiconductor foundry operator.
**Region** — Asia-Pacific
**Primary area** — Equipment & foundries
- [Official investor relations](https://www.umc.com/en/InvestorServices/Index)
### UPS
Logistics network and operational-automation adopter.
**Region** — North America
**Primary area** — Applications & workflows
- [Official investor relations](https://investors.ups.com/)
### Vertiv
Data-centre power and thermal-management provider.
**Region** — North America
**Primary area** — Power & datacentres
- [Official investor relations](https://investors.vertiv.com/)
### XPENG
Electric-vehicle and humanoid-robotics developer.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://ir.xiaopeng.com/)
### Yaskawa Electric
Industrial-robotics, motion-control and automation provider.
**Region** — Asia-Pacific
**Primary area** — Robotics & automation
- [Official investor relations](https://www.yaskawa-global.com/ir)
**Directory boundary.**
Inclusion, alphabetical order, card prominence and the availability of editorial material do not indicate endorsement. This is not a recommendation or a verdict on any issuer. Open the official route, then read the underlying filings and form your own view.
05 · Private/public-route watch
## Separate listed shares from private-company status.
“Official public route filed/approved — not yet public shares” identifies only a filed or approved primary-source route, not a listing. “Private operating company — no public route shown here” identifies current operating context only. Neither status means shares are available, a transfer is approved or an IPO is planned; these entries remain outside the public-company directory.
### 1X
Private operating company — no public route shown here.
**Region** — Europe
**Research note** — Home humanoid and supervised-robotics operating context covered in the robotics guide.
**Dated status** — Official operating source checked 4 August 2026.
- [Official operating source](https://www.1x.tech/neo)
### Agility Robotics
Official public route filed/approved — not yet public shares.
**Region** — North America
**Research note** — Digit is covered as a logistics-humanoid operating route.
**Dated status** — SEC Form S-4 source checked 4 August 2026; a filing is not a public listing.
- [Official SEC filing](https://www.sec.gov/Archives/edgar/data/2074973/000121390026077981/ea029792301ex99-1.htm)
### Anthropic
Official public route filed/approved — not yet public shares.
**Region** — North America
**Research note** — Private AI-model developer included for research context.
**Dated status** — Confidential draft S-1 submitted 1 June 2026; any offering remains subject to SEC review, market conditions and other factors.
- [Official status source](https://www.anthropic.com/news/confidential-draft-s1-sec)
### Apptronik
Private operating company — no public route shown here.
**Region** — North America
**Research note** — Apollo humanoid and operating-data context covered in the robotics guide.
**Dated status** — Official Series A disclosure checked 4 August 2026.
- [Official funding source](https://apptronik.com/news-collection/apptronik-closes-over-935-million-series-a)
### Figure AI
Private operating company — no public route shown here.
**Region** — North America
**Research note** — Humanoid-robotics operating company included for research context.
**Dated status** — Official warning checked 4 August 2026: unauthorized stock transfers may be void or not recognized.
- [Official transfer warning](https://www.figure.ai/news/notice-regarding-unauthorized-attempts-to-sell-figure-stock)
### Mistral AI
Private operating company — no public route shown here.
**Region** — Europe
**Research note** — Model developer whose Robostral navigation research appears in the robotics guide.
**Dated status** — Official funding disclosure checked 4 August 2026.
- [Official funding source](https://mistral.ai/fr/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai/)
### NEURA Robotics
Private operating company — no public route shown here.
**Region** — Europe
**Research note** — Physical-AI, training-infrastructure and humanoid context covered in the robotics guide.
**Dated status** — Official Series C disclosure checked 4 August 2026.
- [Official funding source](https://neura-robotics.com/record-series-c/)
### OpenAI
Official public route filed/approved — not yet public shares.
**Region** — North America
**Research note** — Private AI-model developer included for research context.
**Dated status** — Confidential draft S-1 submitted 8 June 2026; timing remains undecided.
- [Official status source](https://openai.com/index/openai-submits-confidential-s-1/)
### Unitree Robotics
Official public route filed/approved — not yet public shares.
**Region** — Asia-Pacific
**Research note** — Humanoid and quadruped operating context covered in the robotics guide.
**Dated status** — CSRC approved the official route on 1 July 2026; that approval is not public shares.
- [Official CSRC approval](https://www.csrc.gov.cn/csrc/c105906/c7642867/content.shtml)
### Wandercraft
Private operating company — no public route shown here.
**Region** — Europe
**Research note** — Medical-exoskeleton and Calvin humanoid operating context covered in the robotics guide.
**Dated status** — Official Series D disclosure checked 4 August 2026.
- [Official funding source](https://www.wandercraft.eu/articles/wandercraft-announces-series-d-round-bringing-75m-in-total-funding-secured-for-global-acceleration-of-ai-powered-robotics)
06 · Thesis lab
## Turn market stories into falsifiable tests.
These are balanced research prompts, not conclusions. Write a claim, a countercase, evidence to collect and a decision rule before treating a narrative as established.
### Inference economics ≠ company returns
**Claim** — Lower cost per useful task could expand usage enough to improve supplier economics.
**Countercase** — Lower cost may be passed through, while utilisation, sales expense or capital intensity prevents company-level returns.
**Evidence to collect** — Usage, realised pricing, gross-margin bridge, utilisation, capex and customer renewal.
**Falsifier / decision rule** — Revisit the claim if falling unit costs do not translate into durable cash conversion after the full infrastructure bill.
### Adjusted profit ≠ cash conversion
**Claim** — A reported adjusted measure may help describe operating progress.
**Countercase** — Capex, depreciation, stock compensation, dilution or financing needs may change the economic picture for an owner.
**Evidence to collect** — Reconciliations, cash-flow statement, capex, useful lives, share count and debt notes.
**Falsifier / decision rule** — Do not rely on the adjusted measure if it cannot be reconciled to a repeatable per-share cash outcome.
### Demand can be circular or concentrated
**Claim** — Committed infrastructure demand could support an expanding ecosystem.
**Countercase** — Financing, supplier commitments or a small set of customers may make reported demand less independent or less durable than it appears.
**Evidence to collect** — Customer concentration, related commitments, contract terms, funding sources and recognised versus committed revenue.
**Falsifier / decision rule** — Rework the thesis if independent end demand and diversified payment capacity are not visible in primary disclosures.
### Infrastructure can overbuild
**Claim** — More capacity may create option value if demand, power and utilisation arrive together.
**Countercase** — Fungible capacity, faster hardware cycles or weak contracts can turn scarcity into low-return overbuild.
**Evidence to collect** — Utilisation, contracted capacity, contract duration, power delivery, hardware residual assumptions and refinancing needs.
**Falsifier / decision rule** — Challenge the claim if capacity grows faster than contracted, cash-generating demand or if residual-value assumptions weaken.
### Cheaper models can expand or compress
**Claim** — Cheaper or open models may unlock demand that was previously uneconomic.
**Countercase** — Capability parity can pressure prices and margins faster than new demand grows.
**Evidence to collect** — Price changes, task substitution, usage elasticity, switching costs, customer outcomes and margin response.
**Falsifier / decision rule** — Reconsider the claim if falling prices increase activity but fail to improve durable economics for the relevant layer.
### Adoption can bottleneck in the workflow
**Claim** — Better models may create useful new production workflows.
**Countercase** — Data readiness, integration, trust, governance, review burden and change management may limit paid deployment.
**Evidence to collect** — Time from pilot to production, renewal, user acceptance, error burden, implementation cost and measurable workflow outcomes.
**Falsifier / decision rule** — Pause the capability-led narrative if adoption evidence remains pilot-heavy or outcomes do not survive operating costs.
For the adjusted-measure test, start with the [SEC’s official non-GAAP interpretations](https://www.sec.gov/rules-regulations/staff-guidance/corporation-finance-interpretations/non-gaap-financial-measures) . For a dated external research view on efficiency, compute demand and inference pricing, compare [Epoch AI’s compute-demand research view](https://epoch.ai/gradient-updates/algorithmic-progress-likely-spurs-more-spending-on-compute-not-less) and its [inference-price data insight](https://epoch.ai/data-insights/llm-inference-price-trends) with issuer disclosures; they are not settled conclusions.
07 · Several futures
## Use scenarios to find the evidence that matters.
A scenario is a condition to monitor, not a probability or a forecast. It is useful only when it changes the documents, metrics or operating evidence a researcher will collect.
Capacity tight
### Demand races ahead of supply.
Collect delivery evidence, customer concentration and the cost of meeting demand.
Capacity ample
### Infrastructure arrives early.
Collect utilisation, contract quality, replacement economics and financing evidence.
Capability cheaper
### Models become more substitutable.
Collect pricing, workflow outcomes, switching evidence and margin response.
Adoption slower
### Workflow change takes time.
Collect production conversion, renewal, implementation cost and acceptance evidence.
Power constrained
### Delivery depends on energy and grids.
Collect interconnection, power contracts, cooling, site and timing evidence.
Demand broadens
### More sectors reach production.
Collect independent customer evidence rather than assume broad demand from one route.
08 · Research record
## Keep the thesis date-stamped and reversible.
A research record can make uncertainty visible without pretending to resolve it. It should preserve what was known, what was assumed and what would change the view.
### Before a conclusion
- Record the date, primary documents and definitions used.
- Separate facts, management statements, external research and your own assumptions.
- Write the countercase and the evidence that would strengthen it.
- Note material conflicts if they ever exist.
### At the next update
- Compare reported outcomes with the prior claim and countercase.
- Read the cash bridge, customer mix and capital commitments again.
- State which evidence changed and which uncertainty remains.
- Archive corrections rather than silently rewriting the record.
09 · Falsification
## Name what would break the explanation.
A falsifier is not a prediction service. It is a pre-committed test that helps prevent a researcher from moving the goalposts when new evidence arrives.
### Supplier route
Could customer internalisation, substitution, inventory or export exposure change the economics?
### Cloud route
Could capex, depreciation, utilisation or customer concentration outweigh recognised demand?
### Workflow route
Could adoption, implementation burden, error cost or bundling prevent durable paid use?
### Infrastructure route
Could power delivery, financing, contract quality or hardware residual value change the return profile?
The durable rule
Follow evidence, operating constraints and cash bridges —then keep the countercase visible.
This page maps useful public information. It does not identify winners or tell anyone what to do with a security.
- [Check the AI evidence](https://isaiuseful.com/evidence.html.md)
- [Study adoption friction](https://isaiuseful.com/adoption.html.md)
- [Map infrastructure](https://isaiuseful.com/cloud-models.html.md)
---
## https://isaiuseful.com/thinking-with-ai (`/thinking-with-ai.html.md`)
# How Do You Build Software Safely with AI?
Canonical source: [https://isaiuseful.com/thinking-with-ai](https://isaiuseful.com/thinking-with-ai)
Beginner-safe AI development · checked 31 July 2026
Codex, Claude Code, OpenCode and spec tools differ at the edges. The durable method is the same: make the problem clear, keep changes small, verify the result and let security checks stop unsafe code before production.
This guide starts with your first branch and ends with an evidence-based decision about subscriptions, emergency APIs, open-weight models, DGX hardware and private cloud.
- [Start with one safe task](#start)
- [Put security in CI now](#security)
Human-owned loop
Repeat for every change
1. **01** **Understand** Read before editing
2. **02** **Specify** Define the behavior
3. **03** **Plan** Choose the smallest path
4. **04** **Build** One reviewable task
5. **05** **Verify** Tests, scans and review
6. **06** **Learn** Keep the useful rule
The model may type the code. **You still own the requirement, permission boundary and release decision.**
- [**6** reviewable build stages EXPLORE → REVIEW](#loop)
- [**5** levels of agent authority COMPLETE → AUTOMATE](#surfaces)
- [**4** inputs to every task GOAL · CONTEXT · CONSTRAINTS · DONE](#start)
Before the first prompt
## Prepare a reversible place to learn.
You do not need to be a senior developer. You do need a repository, a small task and a way back when an experiment fails.
Foundation 01
### Learn the Git safety net.
A branch isolates work. A commit records a checkpoint. A diff shows exactly what changed. A pull request gives another person and your automated checks a review surface.
- [Practice visually with Learn Git Branching →](https://learngitbranching.js.org/)
- [Read the free Pro Git book →](https://git-scm.com/book/en/v2)
Foundation 02
### Know where commands run.
The terminal runs with your user permissions. Before approving a command, read its target path and effect. If you cannot explain it, ask the agent to explain it without running it.
**Beginner notice**
Never paste a command containing a password, API key or private token into chat, code or a commit.
Foundation 03
### Write down how the repo works.
Keep build, test and lint commands; important directories; conventions; forbidden paths; and “done” checks in a short repository instruction file. Prefer rules a person or CI job can verify.
- [See the file each tool reads →](#tools)
New to a codebase?
**Ask for a map before asking for a change.**
Start read-only: “Explain the folder structure, how to run the tests, where authentication lives and which files you would inspect for this task. Do not edit anything.” Check the answer against the repository.
A reusable first prompt
### Give four things—not a novel.
Good context reduces guesswork. Durable project facts belong in repository instructions; the specific goal belongs in this task.
**Goal · context · constraints · done when**
`Goal: add an empty-state message to the saved-items page. Context: inspect the page, its existing tests and the shared message component. Constraints: do not add dependencies or change the API. Plan first and ask if the expected wording is missing. Done when: the empty and non-empty cases both pass, the diff is reviewed and no security check regresses.`
Choose the smallest surface
## Give AI only as much room as the task needs.
A line completion and an autonomous terminal agent are not the same risk. Start narrow, widen authority only when the work genuinely needs it, and keep the same definition of done.
**01**
Complete
### Finish a local thought
Use an inline suggestion for a line, expression or repetitive pattern you already understand. Read it before accepting it.
Authority: suggestion only
**02**
Ask
### Build understanding
Ask for an explanation, repository map or likely files. Require references and keep the pass read-only.
Authority: read and explain
**03**
Edit
### Transform a selection
Use a targeted edit when the boundary is visible: add validation, rename a symbol or write a focused test.
Authority: named files or selection
**04**
Agent
### Execute a bounded task
Use plan/agent mode for multi-file work that can be split into tasks, tested and reviewed as a diff.
Authority: workspace + approved tools
**05**
Automate
### Repeat a proven workflow
Use a CLI, SDK or CI agent only after the interactive version is reliable, permissioned and observable.
Authority: policy-defined and logged
Pocket control is still system access
**A phone changes the surface—not the authority.**
Label a mobile session as **chat** , **assistant** or **operator** . Show whether the model and tools run on the phone, a home computer or a cloud service; preserve the same file limits, command review and approval gates you would require at the keyboard. [See the private remote pattern →](https://isaiuseful.com/remote-spark.html.md#operators)
Beginner shortcut
**Confused? Ask. Certain? Edit. Multi-file? Plan.**
If you cannot predict which files or commands are needed, begin with a read-only question. Do not jump to an autonomous agent because the prompt is hard to write—the uncertainty is a reason to explore first.
Spec-driven development
## Turn intent into a six-stage reviewable loop.
A specification does not need to be long. It needs to say what will change, what will not change and how someone can prove the result.
OpenSpec beyond code
### The value is the rail—not the file type.
OpenSpec packages proposal, requirements, design and tasks as durable Markdown. The same pattern can guide any technical change that benefits from explicit scope, ordered decisions and review gates.
**Infrastructure**
Configuration changes, architecture decisions, migrations and rollback.
**Documentation**
Technical manuals, policy sets, requirements and coordinated revisions.
**Process design**
Operational workflows, business rules, handoffs and exception paths.
**What “on the rails” really means:** the framework cannot guarantee truth or prevent every hallucination. It makes assumptions, constraints, scope changes and unfinished work visible, so a human or automated gate can catch drift before execution.
> Visual: Spec-driven AI development sequence
**Visual reading order:**
1. **01** **Explore** Read code, examples and constraints. Make no edits.
2. **02** **Specify** Problem, users, scope, requirements and acceptance cases.
3. **03** **Plan** Files, interfaces, dependencies, tests, risks and rollback.
4. **04** **Tasks** Small ordered changes, each with an acceptance check.
5. **05** **Implement** One task and its tests; stop before the next task.
6. **06** **Review** Diff, tests, security findings and the original acceptance cases.
spec.md
### The promise
Who needs what behavior? What is in and out? Which examples must work? Which questions are still unresolved?
plan.md
### The route
Which modules change? What data crosses a trust boundary? Which tests, migration and rollback are required?
tasks.md
### The sequence
Can each task be reviewed, tested and committed independently? If not, split it again.
Example acceptance case
**Write behavior a beginner can check.**
**Given** a signed-in user has no saved items, **when** they open Saved Items, **then** the page shows the agreed empty-state message and does not make a delete request. This is more testable than “make the page nice.”
**Stop and return to the spec when**
a requirement is ambiguous
the agent proposes a new dependency
a migration or permission appears unexpectedly
tests disagree with the intended behavior
the same fix fails twice
the diff is too large to explain
Three routes, one method
## Use the tool’s native controls without changing the discipline.
The material difference is usually authentication, instruction files, permission controls and model routing—not the shape of good engineering work.
| Workflow moment | Codex | Claude Code | OpenCode + OpenSpec |
| --- | --- | --- | --- |
| **Start safely** | Use default sandbox and approvals; ask or enter Plan mode before a broad change. | Use plan mode, permission rules and sandboxing; keep consequential commands behind approval. | Use the restricted Plan agent; set unknown actions to ask and deny pushes, destructive commands and out-of-repo access. |
| **Durable repo rules** | `AGENTS.md` , with nearer nested files overriding broader guidance. | `CLAUDE.md` . Import `@AGENTS.md` to share cross-tool rules, then add only Claude-specific notes. | `AGENTS.md` by default; OpenCode can fall back to `CLAUDE.md` . Keep OpenSpec artifacts under `openspec/` . |
| **Think before code** | Ask for a plan or use Plan mode; agree on files, checks and stop conditions before edits. | Use plan mode or a read-only research pass; approve the plan before implementation. | Use OpenCode Plan or `/opsx:explore` , then `/opsx:propose` to create proposal, requirements, design and tasks. |
| **Implement** | Name one task, require relevant tests, then inspect the diff or run a review. | Implement one bounded task, run tests and inspect the diff before accepting more authority. | Run `/opsx:apply` for approved tasks. Verify the artifacts and implementation before `/opsx:archive` . |
| **Switch models** | Choose models available to the plan or API key. API authentication is separate from plan authentication. | Choose the available Claude model; a plan allowance and Console/API billing are distinct capacity paths. | OpenCode connects to multiple providers, OpenRouter and local endpoints; OpenSpec supplies the workflow, not the model. |
| **What stays human** | Requirement approval, secret/data policy, production access, security exceptions, migration approval, final review and release. | Requirement approval, secret/data policy, production access, security exceptions, migration approval, final review and release. | Requirement approval, secret/data policy, production access, security exceptions, migration approval, final review and release. |
**Official routes:** the table reflects current product documentation, but interfaces change. Recheck the linked controls before standardizing a team workflow.
- [Codex best practices →](https://learn.chatgpt.com/guides/best-practices)
- [Codex AGENTS.md →](https://learn.chatgpt.com/docs/agent-configuration/agents-md)
- [Claude Code instructions →](https://code.claude.com/docs/en/memory)
- [Claude Code permissions →](https://code.claude.com/docs/en/permissions)
- [OpenCode rules →](https://opencode.ai/docs/rules/)
- [OpenCode permissions →](https://opencode.ai/docs/permissions)
- [OpenSpec workflow →](https://github.com/Fission-AI/OpenSpec)
When several agents share the work
### Buzz makes the collaboration room part of the record.
Buzz is an Apache-2.0 workspace from Block where people and model-agnostic agents can share channels, threads, workflows and code context. Use it when coordination itself is the problem; one small task for one agent rarely needs another platform.
**Shared context**
Keep the request, agent discussion, evidence, review and approval together instead of copying fragments between separate chats.
**Separate identity**
Give each person and agent its own signed identity and only the channel membership the role needs.
**Human gate**
Let agents propose, compare and critique; keep consequential tool access, merge and release decisions behind named human approval.
**Boundary:** a hosted or self-hosted relay controls the workspace record. It does not make a cloud model local or replace the underlying agent's sandbox, repository scope, network policy and credential controls. Buzz is pre-1.0; verify what works in the current release before relying on a feature.
- [Open Buzz →](https://buzz.xyz)
- [Inspect the source and current status →](https://github.com/block/buzz)
- [Read Block's introduction →](https://block.xyz/inside/introducing-buzz-where-humans-and-agents-work-together)
- [Compare Buzz in the tool catalogue →](https://isaiuseful.com/tools.html.md#buzz)
Context engineering
## Context is a budget—not a repository landfill.
The agent can act only on what its harness sends to the model: instructions, conversation, files, tool descriptions and tool results. Missing facts invite guesses; irrelevant facts bury the useful signal.
Always loaded
**Short repository rules**
Build/test commands, architecture invariants, forbidden paths and verifiable conventions.
This task
**Goal + constraints**
The requested outcome, relevant files, non-goals, acceptance checks and stop conditions.
On demand
**Scoped knowledge**
Nearby instructions, a matching skill, selected documentation and the smallest useful files.
Evidence
**Tool results**
Search output, tests, scans and diffs—trimmed to what the next decision needs.
When it drifts
**Summarize or restart**
Save decisions in artifacts, then compact or open a fresh session for the next coherent task.
Context-engineering hint
**The harness matters. So does the evidence it can reach.**
For open-source dependencies, manuals no longer have to be the agent’s ceiling: the exact source is the strongest evidence of how an implementation behaves. [opensrc](https://github.com/vercel-labs/opensrc) lets a coding agent fetch and search version-matched package or repository source with ordinary tools. Give it the implementation, tests and examples; keep official documentation for the supported contract, migrations and security guidance. [Compare opensrc in the tool catalogue.](https://isaiuseful.com/tools.html.md#opensrc)
> Visual: First agent loop sends system instructions, tools, a prompt and selected files before receiving a response. A second loop carries the repeated prefix and earlier response, then adds a new prompt, file and response, consuming more of the model context limit.
**Visual reading order:**
1. Model-specific context limit
2. **First loop** *System
+ tools* *Prompt* *File* *Response* *Available context*
3. **Second loop** *Repeated prefix + prior response* potentially cached input *Prompt* *File* *Response* *Less room remains*
**System + tools**
**Prompt**
**Selected file**
**Model response**
**Repeated / cached**
Why agent sessions grow
**Every loop adds material the next decision must compete with.**
The harness assembles instructions, tool definitions, conversation, selected files and results for each model call. Repeated prefixes may be cached for efficiency, but caching does not make stale or irrelevant context useful.
| Control | Use it when | Avoid | Portable form |
| --- | --- | --- | --- |
| **Repository instructions** | A rule matters in almost every task in this repository. | Vague advice, temporary task detail and a giant generated handbook. | `AGENTS.md` ; tool-specific files may import or complement it. |
| **Scoped instructions** | A rule applies only to one directory, language or file type. | Loading frontend, database and test conventions into every turn. | Nested `AGENTS.md` files or a tool’s path-scoped rules. |
| **Reusable prompt** | A person starts the same procedure with a different input. | Pretending a one-shot prompt is a permanent project rule. | A versioned Markdown template; native prompt commands where supported. |
| **Skill** | A repeatable multi-step capability should load only when relevant. | One enormous skill that handles unrelated jobs or hides unsafe commands. | `SKILL.md` plus reviewed scripts and references. |
| **Specialist agent** | A recurring role needs a bounded mission and restricted tools. | A “do everything” persona with write, deploy and admin access. | Named agent instructions; read-only archaeologist or test-only reviewer. |
| **MCP or custom tool** | The workflow needs live data or a real action from another system. | Connecting every server, exposing raw admin APIs or trusting model arguments. | A small typed tool contract with least privilege and approval. |
| **Subagent** | A broad, independent investigation would flood the main task context. | Delegating an ambiguous whole project or losing integration ownership. | A bounded read-only brief returning evidence and unresolved questions. |
First valuable skill
### Package judgment you do not want the agent to reinvent.
A skill is a small, version-controlled folder that teaches an agent one repeatable job. Only `SKILL.md` is required; add references, scripts or assets when they improve repeated work. A strong first skill starts with a real standard you can judge—not a vague “be helpful” persona.
**SKILL.md**
Trigger + core workflow
**references/**
Knowledge loaded on demand
**scripts/**
Tested repeatable checks
**assets/**
Optional templates + resources
- [Read the Agent Skills specification](https://agentskills.io/specification)
- [Download color accessibility skill](https://isaiuseful.com/downloads/audit-color-accessibility.zip)
- [Download code structure skill](https://isaiuseful.com/downloads/code-structure.zip)
- [Download evidence-driven testing skill](https://isaiuseful.com/downloads/evidence-driven-testing.zip)
- [Browse NVIDIA agent skills](https://build.nvidia.com/skills)
**Just landed:** [Agent Plugins 1.0.0](https://agent-plugins.org/) wraps Agent Skills and MCP servers in one portable, vendor-neutral package that compatible clients can discover. The client still controls installation, permissions and execution.
**Try it:** give the downloaded folder to your coding agent and ask: `Install this skill where you can discover it for this project, and update AGENTS.md or the repository’s equivalent instructions only if a durable note is needed.` Start a fresh task, describe the job normally and let the agent select a matching skill.
1. **01** **Pick one repeated judgment.** Collect real prompts, inputs, expected outputs and the mistakes that matter. Keep one skill focused on one coherent job.
2. **02** **Write the trigger and workflow.** Give `SKILL.md` a precise name and description, then write the shortest sequence that reliably produces a reviewable result.
3. **03** **Bundle only reusable material.** Move detailed knowledge into references and deterministic repeated work into tested scripts. Leave ordinary reasoning to the agent.
4. **04** **Run it, compare it, improve it.** Test prompts that should and should not trigger it. Replay real tasks, inspect failures and keep a new rule only when it improves the result.
> Visual: A U-shaped recall curve shows information near the beginning and end of a long context as easier to retrieve while information in the middle can be harder to retrieve.
Illustrative · below half full
**The middle can fade before capacity is close.**
Beginning
tokens
Middle tokens
**weaker recall**
End
tokens
Models can over-attend to the beginning and recent end of a long prompt while missing information buried in the middle.
> Visual: A vertical context stack shows the newest tokens clearly at the top while progressively older tokens near the bottom fade, with an illustrative halfway marker across the stack.
Illustrative · beyond half full
**Older material competes with every new turn.**
Newest tokens
Oldest tokens
Illustrative halfway mark
As a session fills, earlier requirements and failed approaches can become less reliably recalled. Different models and harnesses behave differently; the safe response is still smaller tasks and durable artifacts.
01
**Miss**
Record the concrete wrong result, not “the model is bad.”
02
**Diagnose**
Was context missing, contradictory, overloaded—or was the task beyond the model?
03
**Fix**
Add one scoped rule, example, test, tool constraint or smaller task boundary.
04
**Verify**
Undo the result, replay the same case and keep the change only if it helps.
How tool calling works
### The model proposes. Your software executes.
A model can emit a structured request such as `create_issue({title, body})` . The agent harness or your application must validate the arguments, enforce identity and policy, ask for approval when needed, run the function, and return the result. A fluent request is not authorization.
> Visual: Tool calling sequence
> Flow order: User request → Model proposes call → Host validates + approves → Tool runs → Result returns
**Minimum tool contract**
- Precise name, description and typed inputs
- Input validation and target allowlists
- Read-only default; explicit approval for side effects
- Timeouts, bounded retries and safe failure
- Idempotency where a retry could duplicate work
- Logs for attempts, failures and outcomes
- No secret values in prompts, output or logs
Context-rot notice
**A fresh session is a tool, not a failure.**
Long sessions accumulate stale plans, failed approaches and noisy tool output. Preserve approved decisions in `spec.md` , `plan.md` , issues or commits; then summarize or restart before the next distinct task.
**Portable building blocks:** Agent Skills package on-demand procedures; MCP standardizes connections to external tools and data. The context-rot figures are teaching diagrams, not benchmark curves or a model-independent 50% rule.
- [Agent Skills specification →](https://agentskills.io/specification)
- [Model Context Protocol introduction →](https://modelcontextprotocol.io/docs/getting-started/intro)
- [Product Talk on context rot →](https://www.producttalk.org/context-rot/)
Existing and legacy systems
## Rediscover the behavior before rewriting the code.
Old code contains business rules, edge cases and operational bargains that may exist nowhere else. Treat modernization as agent-assisted archaeology followed by normal spec-driven delivery.
**01 · Rediscover**
### What does it actually do?
Produce business rules with code evidence, a data model, integration inventory and an open-questions list. Read one module at a time.
Output: reviewable current-state spec
**02 · Audit**
### What should stay, retire or change?
Compare home-grown utilities, integrations, stores, runtimes and operational assumptions with current supported options. “Keep” is valid.
Output: substitution map + trade-offs
**03 · Re-architect**
### What should the target shape be?
Decide boundaries, data migration, contracts, authentication, secrets, observability, rollback and a cutover strategy before implementation.
Output: approved target plan
**04 · Replace in slices**
### How do we preserve behavior?
Derive tests from rediscovered rules, keep the old interface where practical and switch traffic gradually. Stop when behavior is ambiguous.
Output: tested, committable modules
**05 · Ship safely**
### How does the same artifact reach users?
Build, test and scan in CI; promote through dev and staging; observe both paths and keep a rehearsed rollback during cutover.
Output: repeatable release + recovery
A safe rediscovery prompt
### Ask for evidence, not confidence.
Run this against one module—not an entire twenty-year-old system. A domain expert still has to validate what is active in production.
**Read-only archaeology**
`Do not edit code. For this module, produce: (1) business rules in plain language with concrete file references, (2) entities, relationships and invariants, including database-enforced rules, (3) every external integration and operational dependency, and (4) unresolved behavior questions. Separate evidence from inference. Do not propose a new architecture yet, and never guess a missing business rule.`
Why “strangler”?
**Replace a large system one safe path at a time.**
A strangler-style migration routes selected behavior to the new implementation while the rest stays on the old one. You compare results, increase traffic gradually and retain a rollback instead of betting the business on one big switch.
Security before production
## Make the unsafe path fail early.
An AI review is useful additional evidence. It is not a replacement for deterministic tests, scanners, least privilege or a human who can own the risk.
Non-negotiable default
Security checks run on every pull request, before merge. A production deployment consumes the already-scanned commit; it does not become the first place you discover a secret, vulnerable package or obvious code flaw.
**01 · Spec**
**Threat + data boundary**
Name assets, actors, sensitive data, abuse cases and denied actions.
**02 · Workstation**
**Secret prevention**
Use environment variables, a secret manager and pre-commit or push protection.
**03 · Pull request**
**SAST + SCA + tests**
Scan code, new dependencies, lockfiles, infrastructure and containers.
**04 · Preview**
**DAST + abuse cases**
Test the running preview, authorization failures and untrusted inputs.
**05 · Release**
**Artifact + approval**
Build once, record dependencies, protect deploy credentials and approve promotion.
**06 · Operate**
**Observe + recover**
Monitor, rotate, patch, roll back and learn from real incidents.
Security words in plain language
**Four checks catch different mistakes.**
**SAST** inspects source code. **SCA** checks third-party packages and licences. **Secret scanning** catches credentials. **DAST** probes a running application. None proves the application is secure; together they find problems earlier.
### A sensible setup order for your first repository
Names differ across GitHub, GitLab, Azure DevOps and other platforms, but the control sequence stays useful. Some private-repository features require a paid security plan; use a supported scanner you can require in CI rather than leaving the gate empty.
Repository settings
### Prevent and protect
Enable secret scanning or push protection, dependency alerts, branch/ruleset protection and default code scanning where available. Do not let contributors push directly to the production branch.
Pull-request workflow
### Test the changed commit
Install from the lockfile; run lint, types and tests; then SAST, dependency review and relevant infrastructure/container scans. Give every required check a clear name.
Release environment
### Promote the same artifact
Give deploy credentials only to the release job, require an environment approval, run preview smoke/DAST checks and keep a tested rollback. A later rebuild breaks the evidence chain.
| Gate | Minimum check | Block the change when | Safe response |
| --- | --- | --- | --- |
| Before commit | Secret scan, formatter and focused unit tests | A real credential, private key or generated secret appears. | Remove it, rotate it if exposure is possible, and replace it with a documented secret reference. |
| Pull request | Full tests, lint/type checks, SAST, dependency review and licence policy | A new high/critical flaw, vulnerable runtime dependency, forbidden licence or failing test is introduced. | Fix or remove the change. A time-limited exception needs an owner, reason, compensating control and expiry. |
| Infrastructure | IaC and container scan, least-privilege review and ephemeral credentials | Public exposure, privileged containers, broad IAM or unpinned images appear unexpectedly. | Reduce access, pin the artifact, prove a denied path and document rollback. |
| Preview | DAST, end-to-end smoke tests and authorization abuse cases | One user can read or change another user’s data, input reaches an unsafe sink, or a critical flow breaks. | Return to the spec and threat model; add a regression test before the fix is merged. |
| Deploy | Protected environment, approved artifact, health check and rollback | The commit differs from the scanned artifact, secrets are unavailable or rollback is untested. | Stop promotion. Repair the release path without rebuilding unreviewed code in production. |
**If a secret reaches Git**
Deleting the visible line is not enough because history and logs may still contain it. Revoke or rotate the credential first, investigate its use, then clean history only with repository-owner coordination.
**Why this order:** OWASP’s DevSecOps guidance places multiple security tests inside the delivery pipeline. GitHub’s dependency review is specifically designed to catch risky dependencies while they are still pull-request changes.
- [OWASP DevSecOps guide →](https://devguide.owasp.org/en/09-operations/01-devsecops/)
- [GitHub CodeQL code scanning →](https://docs.github.com/en/code-security/concepts/code-scanning/codeql-code-scanning)
- [GitHub dependency review →](https://docs.github.com/en/code-security/concepts/supply-chain-security/dependency-review)
- [GitHub push protection →](https://docs.github.com/en/code-security/concepts/secret-security/push-protection)
- [Claude Code agent security →](https://code.claude.com/docs/en/security)
- [Codex approvals and sandboxing →](https://learn.chatgpt.com/docs/agent-approvals-security)
Developer subscription playbook
## Pay for a fair test. Upgrade the lane that wins your work.
Free plans are product demos, not serious reliability tests. Start with affordable paid access, then put the larger allowance behind the workflow that proves most useful.
Yesterday’s price is not today’s. Today’s price is not tomorrow’s .
Bundled subscriptions can be dramatically cheaper than equivalent API use, but the bargain can move. GitHub Copilot’s June 2026 shift from premium requests to token-priced AI Credits shows how providers can reprice expensive agentic work. Treat the bundle as favorable access—not permanent infrastructure—and keep a tested fallback.
01 · Minimum useful starting point
### Pay for a real trial.
The free tier is fine for simple prompts, but its limits can make capable models feel unreliable. For frequent development, begin with the lowest paid plan on both platforms.
**Rule:** use the checkout price in your region, including tax, and only an organization-approved plan. Do not take annual plans as they don't give flexibility to switch between tiers.
OpenAI
**ChatGPT Plus**
**~20€**
*per month*
Anthropic
**Claude Pro**
**~20€**
*per month*
02 · The upgrade trigger
### Promote the winner.
When a limit repeatedly interrupts useful work, decide which tool fits your workflow better. Upgrade that subscription to its 5× tier and keep the other on the cheaper plan.
**Do not upgrade on one bad day.** Look for a repeated limit pattern across real tasks. While Annual plans might be enticing there might be more value in switching 5x plans between providers as new models are released.
**5×**
roughly €90–€100 per month before tax
03 · Before the 20× plan
### Fix the workflow first.
Improve prompts, context files, tool choice, reusable skills, model routing and the agent harness before buying another block of capacity.
**Practical rule of thumb:** if one developer cannot make 5× last, optimize before assuming the only answer is 20×.
- **Interactive work:** target 5× for a full workweek.
- **20× signal:** measured autonomous or parallel-agent demand.
- **Watch the meter:** model, effort and fast mode change usage.
High-output developer split
### Two subscriptions. One premium lane.
This field-tested setup is used by experienced, high-output developers: keep both tools available, but put the higher tier behind the lane doing the heavier work.
Frontend / UI heavy
#### Claude Max 5× + ChatGPT Plus
**Primary:** Claude Opus 4.8 for interface work, visual judgment and long design iterations.
Backend / general heavy
#### Codex Pro 5× + Claude Pro
**Primary:** GPT‑5.6 Sol High for complex work; GPT‑5.5 Fast when turnaround matters more than maximum reasoning.
**This is a working pattern, not one person’s current split—and not a benchmark verdict.** Run the same representative tasks through both tools, count accepted results and review effort, then let your own work choose the premium plan.
Scope of this recommendation
**Choose the whole system, not only the model.**
This guide concentrates on OpenAI and Anthropic because they are the frontier options used in the subscription pattern above—not because they are the only credible choices. Google, other major labs and Chinese providers may offer a better price, model or regional route for a particular project. No useful guide can list every combination, so explore the wider tool market and test contenders on your own accepted-task set.
**The harness is part of the result.** Context selection, tools, permissions, retries and verification can materially improve—or degrade—the same model’s performance. Cursor can produce a dramatic lift when its editor context, change review and agent loop fit the way a project is built; another workflow may perform better in Codex, Claude Code or a different harness. Test the exact developer workflow, not Cursor or any other harness in the abstract.
**Use bring-your-own-key deliberately in case of Cursor.** Cursor’s Pro tier or higher can run supported OpenAI and Anthropic models through Cursor’s own agent harness using your API key. The provider then meters the model tokens separately—a ChatGPT or Claude subscription does not fund those API calls. Free access is not enough for this route, and specialized features such as Tab completion still use Cursor’s services and models. Evaluate Cursor’s own models separately and keep them only if they earn a place on your work; team plans should also check Cursor’s current platform-token charges.
- [Explore coding tools and harnesses →](https://isaiuseful.com/tools.html.md#tools-build-software-with-agents)
- [See why the harness changes the result →](https://isaiuseful.com/benchmarks.html.md#harness-efficiency)
- [Compare other model routes →](https://isaiuseful.com/cloud-models.html.md#models)
- [Check Cursor’s current pricing →](https://cursor.com/pricing)
- [Check Cursor’s bring-your-own-key limits →](https://docs.cursor.com/settings/api-keys)
- [Check Cursor’s current data controls →](https://cursor.com/data-use)
**Price and model check · 30 July 2026** Official pages list ChatGPT Plus and Claude Pro at about $20 monthly, with 5× individual tiers at $100 and 20x individual tiers at $200 before applicable tax. Regional checkout prices, billing periods, limits and model access can differ.
- [ChatGPT Plus and Pro tiers →](https://help.openai.com/en/articles/9793128-what-is-chatgpt-pro)
- [Claude Pro and Max tiers →](https://claude.com/pricing)
- [GPT‑5.6 model guide →](https://openai.com/index/gpt-5-6/)
- [Claude Opus 4.8 model guide →](https://www.anthropic.com/news/claude-opus-4-8)
- [GitHub Copilot billing change →](https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/)
- [OpenAI API retention controls →](https://platform.openai.com/docs/models/default-usage-policies-by-endpoint)
- [Anthropic zero-retention scope →](https://privacy.anthropic.com/en/articles/8956058-i-have-a-zero-data-retention-agreement-with-anthropic-what-products-does-it-apply-to)
Resilience after subscriptions
## Keep fallback routes. Provision what must be predictable.
An API account provides access, not assured capacity. Put approved provider routes behind one governed gateway, and treat owned compute as a measured capacity decision.
01 · Emergency capacity
### Treat capacity as a portfolio.
Do not assume any one on-demand provider will always have headroom. Keep approved routes behind one policy-enforcing LLM proxy, and add or retire providers only when measured capacity, quality, cost or incident performance justify it.
- [Build the five-minute switch →](#incident-runbook)
02 · Measured scale
### Own compute only after comparison.
Record real tasks and accepted outcomes, then test a suitable open-weight model through a hosted route. Consider owned compute only after quality and usage are known.
- [Run the replacement lab →](#open-alternative)
Production is down, quota is full or capacity is constrained
### Make the emergency switch boring.
The worst time to discover authentication, model behavior or spending controls is during an outage. Test this route before you need it.
1. 01 **Own the account.** Use a team or service account, MFA and a named incident approver—not one developer’s personal billing.
2. 02 **Constrain the key.** Separate it from production application keys; set the smallest model/provider allowlist, budget and expiry that works.
3. 03 **Protect the data.** Apply the same source-code, customer-data, retention and region rules as the normal route.
4. 04 **Run a smoke and failover pack.** Re-run five representative tasks against every emergency route because a fallback model or provider can produce a materially different patch, tool call or latency profile.
5. 05 **Close the incident.** Export usage, record the model and provider, disable the route, rotate if needed and write the lesson into the runbook.
Codex
### Plan + API are separate lanes.
Codex plans provide included usage and optional credits. API-key use is metered by tokens and fits CLI, SDK, IDE or CI; it does not carry every plan/cloud integration.
Claude Code
### Plan + usage credits can bridge a cap.
Claude’s paid plans can enable additional metered usage at standard API rates. A Console/API route remains a distinct billing and governance path.
OpenCode
### The harness follows the provider.
OpenCode connects to many providers and local models. With OpenRouter, use a dedicated prepaid key, budget guardrail and an explicit model/provider policy. Provider credentials stay local, so protect the workstation profile and never commit keys.
OpenSpec
### The spec does not buy inference.
OpenSpec stores planning artifacts and commands for supported coding tools. Authentication, quotas, model quality and data policy still belong to the agent/provider route.
**Incident boundary**
Do not send production secrets, customer records or unredacted incident logs to a personal account just because it has remaining quota. An emergency shortens time; it does not suspend data policy.
**Capacity check:** On-demand throughput can vary. Provision critical demand; Azure Reservations reduce cost but do not secure Microsoft Foundry capacity.
- [Codex plans and API lane →](https://learn.chatgpt.com/docs/pricing)
- [Claude additional usage →](https://support.claude.com/en/articles/12429409-manage-usage-credits-for-paid-claude-plans)
- [OpenCode providers →](https://opencode.ai/docs/providers)
- [OpenRouter budgets and allowlists →](https://openrouter.ai/docs/guides/features/guardrails/overview)
- [Amazon Bedrock provisioned throughput →](https://docs.aws.amazon.com/bedrock/latest/userguide/prov-throughput.html)
- [Microsoft Foundry provisioned throughput →](https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/provisioned-throughput)
- [Why Azure Reservations do not guarantee capacity →](https://learn.microsoft.com/en-us/azure/cost-management-billing/reservations/microsoft-foundry)
Hosted success → open alternative
## Judge a replacement by its results, not its reputation.
Once an online model works, preserve its successful tasks as a baseline. An open-weight candidate earns a place by matching those outcomes at acceptable speed, cost and risk.
> Visual: Open model replacement evaluation sequence
**Visual reading order:**
1. **01** **Freeze the job** 50–100 real, redacted tasks with acceptance checks and human outcomes.
2. **02** **Shortlist weights** Match licence, context, tool use, language, memory fit and public evidence.
3. **03** **Test hosted** Use OpenRouter or another hosted route before buying hardware.
4. **04** **Replay locally** Pin checkpoint, quantization, runtime, prompt and tools.
5. **05** **Decide** Compare accepted work, latency, operating effort and total cost.
Licence notice
**Open-weight is not automatically open source.**
A downloadable checkpoint may still restrict use, modification or redistribution. Read the exact model card and licence, record the version, and check whether your planned commercial or internal use is permitted. The [Open Source AI Definition](https://opensource.org/ai/open-source-ai-definition) is a useful benchmark for the stronger term.
| Measure | Hosted baseline | OpenRouter candidate | Local candidate | Reject when |
| --- | --- | --- | --- | --- |
| Accepted-task rate | Human-approved result on the same fixed cases | Same rubric and retry limit | Same checkpoint family, prompt and tool contract | The quality gap creates more review or rework than the saving covers. |
| Reviewer effort | Minutes and material edits per accepted task | Measure changed lines and review minutes | Measure again; quantization may change behavior | Humans become the hidden inference engine. |
| Reliability | Tool failures, retries and timeouts | Pin provider and disable fallback for the experiment | Record crashes, OOMs, queue time and recovery | The route cannot finish the real workflow predictably. |
| Performance | Wall time, time to first token and output speed | Record provider and endpoint metadata | Test concurrency, context and sustained load | Interactive latency or team throughput misses its service target. |
| Economics | Plan cost plus measured overflow/API cost | Input, cache and output cost per accepted task | Capex, power, cooling, operations and variable cost | Savings disappear after failed tasks and staff time. |
A fair OpenRouter trial
### Pin what can silently change.
Select the exact model. For a controlled comparison, set a provider order and disable automatic fallbacks; otherwise a successful response may come from a different endpoint. Enforce the data policy your inputs require, including zero-data-retention routing where appropriate.
- [Provider selection and fallback controls →](https://openrouter.ai/docs/guides/routing/provider-selection)
- [Zero-data-retention routing →](https://openrouter.ai/docs/guides/features/zdr)
- [Shortlist current local model families →](https://isaiuseful.com/local-models.html.md#catalog)
- [Choose relevant public benchmarks →](https://isaiuseful.com/benchmarks.html.md#database)
Cloud-to-local calculator
## Buy hardware only when measured work pays for it.
Use accepted tasks, not raw tokens. A cheap model that fails twice and needs a rewrite is not cheap.
Use at least one month of measured hosted/OpenRouter results. Put staff time, support, storage, networking and realistic electricity into the local side.
Measured example
**Local is cheaper at this volume**
Quality and capacity still have to pass.
Hosted / month
**€510**
Local / month
**€332**
Break-even volume
**372 tasks**
Simple payback
**14.9 months**
**The model**
`local monthly = hardware ÷ useful life + energy + operations + accepted tasks × local variable cost`
`break-even tasks = local fixed monthly ÷ (hosted cost/task − local variable cost/task)`
This result is financial only. Keep hosted access until the exact local model, quantization and runtime pass quality, latency, concurrency, context and recovery tests.
No purchase
### Existing computer or hosted model
Best for the first baseline and low or irregular volume. Use idle hardware only if its model passes; “already owned” does not make staff time free.
One power user
### DGX Spark class
DGX Spark has 128 GB unified memory and NVIDIA documents model support up to 200B parameters. Fit is not speed: benchmark the exact checkpoint, quantization, runtime and context.
- [Use the Spark runbook →](https://isaiuseful.com/remote-spark.html.md#spark-setup)
Team / large model
### DGX Station class
The current Grace Blackwell DGX Station architecture offers up to 748 GB coherent memory, including up to 252 GB HBM3e. It needs a utilization case, support owner and acceptance test—not just a model that loads.
- [Use the Station buyer’s guide →](https://isaiuseful.com/dgx-station.html.md)
Shared service
### Private cloud
Consider a shared cluster when multiple teams have steady, governed demand and can operate identity, scheduling, observability, backups, patching and incident response. Rent a comparable service before building one.
- [Compare infrastructure tiers →](https://isaiuseful.com/cloud-models.html.md#hardware)
Not only a cost decision
**Privacy or latency may justify local before financial break-even.**
Say that explicitly. The benefit is then risk reduction, data control or service behavior—not cheaper tokens. Local also creates new obligations: endpoint security, physical access, patching, backups, model licences and an operator who answers when it fails.
**Hardware facts:** vendor parameter ceilings are fit/support claims, not your application’s throughput or accuracy result. Request a quote and enter the full delivered/setup cost in the calculator.
- [Official DGX Spark hardware →](https://docs.nvidia.com/dgx/dgx-spark/hardware.html)
- [Official DGX Station architecture →](https://docs.nvidia.com/dgx/dgx-station-development-guide/overview.html)
- [Local model and memory guide →](https://isaiuseful.com/local-models.html.md)
- [Cloud and owned-infrastructure guide →](https://isaiuseful.com/cloud-models.html.md)
Your first month
## Build capability in four small weeks.
Do not install every plugin, agent and server on day one. Add a capability only when a real workflow proves it is useful.
**Week 1**
### Learn and map
- Finish the main Learn Git Branching levels.
- Create a practice repository and a branch.
- Ask an agent to explain one small codebase without editing.
- Review every proposed command.
**Exit: you can show a diff and return to a clean checkpoint.**
**Week 2**
### Specify and build
- Choose one tiny user-visible change.
- Write scope, acceptance examples and stop rules.
- Approve a plan, implement one task and add tests.
- Record which instruction would prevent a repeated mistake.
**Exit: the change is understandable in one review.**
**Week 3**
### Secure the path
- Enable secret prevention.
- Run tests, SAST and dependency review on pull requests.
- Test one denied authorization case.
- Require checks before merge and protect deploy credentials.
**Exit: an unsafe sample change fails before production.**
**Week 4**
### Measure and prepare
- Track accepted tasks, edits, time and plan/API use.
- Configure a capped emergency API route and run its smoke pack.
- Replay a small test set on one open-weight candidate.
- Keep renting unless quality, volume and economics justify local.
**Exit: you have data—not a hardware wish list.**
**What progress looks like**
You can explain the requirement, the change, the tests, the scan results, the cost and the rollback without asking the model to remember for you.
The durable skill
Use AI to widen your reach— not to outsource your judgment.
- [Start the first task](#start)
- [Choose another workflow](https://isaiuseful.com/guides.html.md)
- [Scale a proven coding workflow](https://isaiuseful.com/adoption.html.md#engineer)
---
## https://isaiuseful.com/guides (`/guides.html.md`)
# Free AI Tutorials and Step-by-Step Workflow Guides
Canonical source: [https://isaiuseful.com/guides](https://isaiuseful.com/guides)
From evidence to a working system
Choose a deliverable, keep authority narrow and decide what success means before the model starts. These guides are designed to survive changes in vendors, models and hardware.
- [Find a guide](#chooser)
- [Compare implementation options](#paths)
- [Browse tools](https://isaiuseful.com/tools.html.md)
**14**
complete starter workflows
FOUNDATIONS TO AGENT SYSTEMS
**10**
supervised runs before expansion
MEASURE BEFORE TRUST
**0**
automatic high-impact actions
DRAFT-ONLY BY DEFAULT
Guide selector
## Which AI tutorial should you start with?
Follow 01–14 as a progressive learning path, or filter by the work, risk and time you have. Later guides reuse the contracts, evaluation and controls introduced earlier.
14 guides match
01
Produce a cited research brief
Turn a decision question into a short report whose important claims can be checked.
25 min
Beginner
Cloud
**Deliverable:** a one-page brief, a claim-to-source table and a short list of unresolved questions.
### Build it
1. 01 **Write the research contract.** Name the decision, scope, date cutoff, trusted source types, exclusions and required output.
2. 02 **Gather primary sources first.** Ask the agent to prefer official data, papers and original documentation, then identify disagreements instead of averaging them away.
3. 03 **Require a claim table.** For each consequential claim, record the source, publication date and whether the source directly supports it.
4. 04 **Verify before sharing.** Open the original source behind the five claims most likely to change the decision. Remove or qualify anything unsupported.
**Starter instruction**
`Return: executive answer, evidence table, strongest counterargument, unknowns, and primary-source links. Mark every inference as an inference.`
Acceptance test
**Every material claim has a direct source, at least one counterpoint is represented and unsupported statements are visibly marked.**
**Keep** — Question, report and opened source set
**Boundary** — No consequential medical, legal or financial decision without qualified review
**Failure signal** — Citations point to search results, summaries or sources that do not support the sentence
Implementation example: [OpenAI deep research](https://openai.com/index/introducing-deep-research/) (vendor documentation).
02
Turn a meeting into decisions and tasks
Transcribe locally, extract structured outcomes and keep a person responsible for publication.
30 min
Beginner
Runs locally
**Deliverable:** an approved meeting record with decisions, owners, due dates and open questions.
### Build it
1. 01 **Confirm consent and retention.** Tell participants what is recorded, where it is processed and when the audio will be deleted.
2. 02 **Create a timestamped transcript.** Run transcription locally when the recording is sensitive. Correct participant names and domain terms before summarization.
3. 03 **Extract a fixed schema.** Request decisions, action, owner, due date, supporting timestamp and unresolved questions. Do not infer missing owners.
4. 04 **Approve and publish.** A participant compares every decision and action against the transcript before copying them into the system of record.
**Extraction schema**
`Return decisions and actions as tables. Every row needs a transcript timestamp. Use "unassigned" and "no date stated" when the meeting did not specify them.`
Acceptance test
**Every published action is traceable to a timestamp and the system invents no owner, deadline or decision.**
**Keep** — Approved notes and retention decision; delete raw audio on schedule
**Boundary** — Recording consent, employment rules and sensitive discussions require local policy
**Failure signal** — Polished minutes include commitments nobody actually made
Local transcription implementation: [OpenAI Whisper](https://github.com/openai/whisper) .
03
Evaluate a workflow over ten runs
Measure one of the first two supervised workflows before adding retrieval, tools or autonomy.
60 min
Beginner
Any workflow
**Deliverable:** a run log, failure set and explicit adopt, revise or stop decision.
### Build it
1. 01 **Record the manual baseline.** Complete three examples without AI. Capture elapsed time, quality checks and the corrections normally required.
2. 02 **Write the rubric first.** Score correctness, completeness, saved minutes, correction minutes and critical failures. Define what automatically fails a run.
3. 03 **Run representative work.** Use ten cases across easy, ordinary and difficult examples. Keep the model, instructions and tool permissions fixed.
4. 04 **Make a decision.** Adopt only when net time and quality improve without an unacceptable failure. Otherwise narrow the task, improve controls or stop.
**Starting threshold**
`At least 8 of 10 outputs are usable after review, no critical failure occurs, and total minutes saved exceed setup plus correction time. Adjust this threshold to the actual risk.`
Acceptance test
**Another person can reproduce the result from the saved cases, rubric, configuration and run log.**
**Keep** — Baseline, cases, model and prompt version, scores, corrections and decision
**Boundary** — One severe safety or authority failure resets expansion even when average quality is high
**Failure signal** — The score changes after seeing the output or only successful examples are retained
**Builds on:** use the output from the [research brief](#research-brief) or [meeting-actions guide](#meeting-actions) as your first test fixture. Evaluation references: [NIST AI RMF Playbook](https://www.nist.gov/itl/ai-risk-management-framework/nist-ai-rmf-playbook) and [Anthropic agent evals](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) .
04
Extract fields from incoming documents
Build on the evaluation pattern with fixed schemas, provenance and deterministic validation.
60-90 min
Intermediate
Local or cloud
**Deliverable:** one fixed-schema record, page-level provenance, validation results and a review queue for uncertain fields.
### Build it
1. 01 **Freeze the schema.** Name required fields, data types, allowed values and which fields may be absent. Do not ask for an open-ended summary.
2. 02 **Parse with provenance.** Retain filename, page, table position and the source text behind every extracted value.
3. 03 **Validate deterministically.** Check dates, identifiers, line-item arithmetic, currency and totals in code. Route any mismatch or low-confidence field to review.
4. 04 **Test the ugly examples.** Include scans, rotated pages, missing fields, duplicate documents and layouts from different suppliers before live use.
**Extraction contract**
`Return only the supplied schema. Include source page and exact supporting text for every field. Use null when absent. Never repair totals or infer a missing identifier.`
Acceptance test
**Every accepted field is traceable to a page, deterministic checks pass and uncertain records remain drafts.**
**Keep** — Original document, parsed structure, schema, validations and reviewer correction
**Boundary** — No posting to finance, CRM or case systems without a separate approval or validated import step
**Failure signal** — The record looks complete because the model invents absent values or silently fixes inconsistent totals
Parsing implementation: [Docling](https://github.com/docling-project/docling) . Evaluation method: [ten-run guide](#ten-run-evaluation) .
05
Question a private document collection
Turn the earlier parsing and provenance pattern into a private, locally controlled retrieval system.
Half day
Intermediate
Local GPU helpful
**Deliverable:** a searchable collection that cites the source document and page for every answer.
### Build it
1. 01 **Choose a bounded collection.** Start with 20-50 documents you understand. Classify sensitivity and remove files that should not enter the system.
2. 02 **Parse without losing provenance.** Extract text, tables and headings while retaining filename, page number and a stable link to the original.
3. 03 **Index locally.** Create one isolated workspace, use a local embedding model and confirm that telemetry, backups and model endpoints match your privacy requirement.
4. 04 **Test retrieval, including absence.** Ask ten questions with known answers and three whose answers are not present. Inspect the retrieved passages, not only the final prose.
**Answer policy**
`Use only retrieved passages. Cite document and page after each claim. If the collection does not support an answer, say "not found in this collection" and list the closest passages.`
Acceptance test
**Known answers cite the correct pages and absent answers produce an explicit refusal instead of plausible filler.**
**Keep** — Original files, parsed output, retrieval test set and configuration
**Boundary** — "Local" is not a guarantee; audit telemetry, model endpoints, plugins and backups
**Failure signal** — Answers cite a document name but cannot identify the supporting passage
**Builds on:** reuse the provenance rules from [document extraction](#document-extraction) and the cases and rubric from the [ten-run evaluation](#ten-run-evaluation) . Open implementations: [Docling](https://github.com/docling-project/docling) , [AnythingLLM](https://github.com/Mintplex-Labs/anything-llm) and [Obsidian](https://obsidian.md/help/) .
06
Draft customer-support replies
Add approved retrieval and escalation to the measured, draft-only workflow.
Half day
Intermediate
Draft only
**Deliverable:** a response draft, cited policy passage, confidence state and escalation recommendation.
### Build it
1. 01 **Assemble the approved context.** Use current policy, product documentation and 20 resolved examples. Exclude exceptions that should remain human-only.
2. 02 **Define routing before wording.** Create categories, mandatory escalations and deterministic tools for prices, account state or refunds.
3. 03 **Draft from retrieved policy.** Require the response to cite its governing passage internally and to ask for missing information rather than assume it.
4. 04 **Grade representative tickets.** Review correctness, policy fit, tone, data handling and escalation across at least 20 historical cases before live use.
**Draft rule**
`Do not invent policy, price, account state or an exception. Cite the internal source used. If the evidence conflicts or is incomplete, draft a clarification and mark the case for review.`
Acceptance test
**All factual commitments come from approved sources or deterministic tools, and ambiguous cases reach a person.**
**Keep** — Test cases, rubric, source version and reviewer corrections
**Boundary** — Redact unnecessary personal data and keep sending behind human approval
**Failure signal** — A fluent draft promises a refund, price or policy exception without authority
**Builds on:** reuse the cited retrieval and absence tests from [private documents](#private-documents) , then apply the [ten-run evaluation](#ten-run-evaluation) to historical tickets. Vendor reference pattern: [Anthropic support guide](https://platform.claude.com/docs/en/about-claude/use-case-guides/customer-support-chat) . Independent field evidence for a different support-assistant deployment: [NBER study](https://www.nber.org/papers/w31161) .
07
Turn an issue into a tested patch
Apply the same contract, baseline and independent-checking pattern to repository work.
60-90 min
Intermediate
Local or cloud
**Deliverable:** a reviewable diff, passing relevant checks and a concise implementation report.
### Build it
1. 01 **Convert the issue into acceptance criteria.** State the current behavior, desired behavior, affected surface and explicit non-goals.
2. 02 **Establish the baseline.** Use a clean branch, run the existing checks and record failures that already exist before the agent edits anything.
3. 03 **Constrain the agent loop.** Ask it to inspect, plan, edit and test. Require it to stop when evidence is missing or the task expands beyond the contract.
4. 04 **Review the artifact yourself.** Read the diff, rerun the claimed commands and inspect for unrelated edits, new dependencies and weakened tests.
**Starter specification**
`Implement only the stated acceptance criteria. Preserve public behavior outside scope. Run the narrowest relevant tests, then the repository check. Report files changed, commands run, results and residual risks.`
Acceptance test
**The new behavior is covered by a test, relevant checks pass and the diff contains no unexplained changes.**
**Keep** — Issue, spec, diff, command output and review notes
**Boundary** — No production secrets, direct deployment or automatic merge
**Failure signal** — The agent changes tests to accept the bug or reports checks it did not run
Tools and references: first learn the complete [reviewable coding loop](https://isaiuseful.com/thinking-with-ai.html.md#loop) , then compare [OpenCode](https://opencode.ai/docs/) , [OpenSpec](https://github.com/Fission-AI/openspec) and the [SWE‑bench family guide](https://isaiuseful.com/benchmarks.html.md#swe-bench) . If this becomes recurring work, use the [loop engineering guide](#loops) before scheduling it.
08
Compare two agent harnesses on one bounded task
Reuse the tested-patch fixture to measure what the surrounding tools, context and controls change.
60–90 min
Intermediate
Disposable workspace
**Deliverable:** a reproducible harness scorecard with run traces, artifacts and an adopt, revise or reject decision.
### Build it
1. 01 **Choose a task with a mechanical finish.** Use a small repository repair, document extraction or read-only investigation with a frozen input and a deterministic test. Record the clean starting state and expected artifact before either harness sees it.
2. 02 **Freeze everything except the harness.** Use the same model, model settings, task specification, starting files, tool permissions, network rule and time, token and attempt caps. Run each candidate in a fresh worktree, container or disposable folder without production credentials.
3. 03 **Run repeatedly and retain the trace.** Give each harness at least five independent attempts. Save elapsed time, input and output tokens, tool calls, retries, changed files, test results, approvals and the complete transcript; a single lucky pass is not the comparison.
4. 04 **Grade the artifact outside the agent.** Run the same deterministic checks on every result, inspect unrelated changes and record manual correction time. Compare pass rate, repeatability, total cost, wall-clock time and boundary violations—not polish or confidence.
5. 05 **Choose the smallest adequate scaffold.** Adopt only if one candidate improves the named workflow at an acceptable cost and authority level. Otherwise narrow the task, change the test or reject both; do not compensate for weak verification with more retries.
**Shared task contract**
`Complete only the supplied task in the disposable workspace. Do not change the acceptance test, access credentials, contact external services or expand scope. Run the named checks, preserve the trace and stop at the time, token or attempt cap. Report the artifact, evidence, failures and remaining risk.`
Acceptance test
**Another person can reset the fixture, rerun both harnesses under the same limits and reproduce the scorecard from saved traces and independently graded artifacts.**
**Keep** — Fixture, task contract, harness and model versions, settings, permissions, traces, artifacts, grader output, cost and decision
**Boundary** — A worktree separates changes, not processes, networks or credentials; test those controls independently
**Failure signal** — One harness gets more context, attempts or authority, the agent edits its grader or only the best run survives
**Builds on:** reuse the fixture and acceptance checks from the [tested-patch guide](#tested-patch) , then grade repeated runs with the [evaluation method](#ten-run-evaluation) . Start with the [browser harness demo](https://isaiuseful.com/benchmarks.html.md#harness-demo) , choose candidates from [software-agent tools](https://isaiuseful.com/tools.html.md#tools-build-software-with-agents) , then compare your scorecard with the narrow [repeated-run efficiency study](https://isaiuseful.com/tools.html.md#harness-efficiency) . For harder reproducible tasks, inspect the [Terminal‑Bench harness](https://isaiuseful.com/benchmarks.html.md#terminal-bench) . These routes demonstrate evaluation methods; they do not establish a universal harness ranking.
09
Run a bounded recurring automation
Turn one already measured workflow into a logged recurrence with explicit stops and approvals.
Half day
Intermediate
Approval gated
**Deliverable:** one repeatable workflow with least-privilege tools, durable state, logs and a human release gate.
### Build it
1. 01 **Choose a boring loop.** Use a repeated input, a known transformation and a draft output: classify an inbox, prepare a report or update a queue.
2. 02 **Constrain tools and authority.** Give the workflow the smallest read scope possible. Separate drafting from sending, deleting, buying or deploying.
3. 03 **Make state and stops explicit.** Record each input, tool call, output and approval. Define completion, timeout, retry and duplicate-prevention behavior.
4. 04 **Supervise before scheduling.** Run ten representative cases manually, convert recurring failures into rules or tests, then schedule with alerts.
**Operating contract**
`You may read the allowed inputs and prepare a draft. Stop on missing data, tool failure, conflicting instructions or any action outside scope. Never send, delete, purchase or deploy.`
Acceptance test
**The workflow can resume safely, never duplicates side effects and stops cleanly when evidence or authority is missing.**
**Keep** — Workflow contract, permissions, run log, failure set and approval record
**Boundary** — External side effects stay gated until the risk is understood and separately controlled
**Failure signal** — Retries send twice, state disappears between runs or the agent quietly expands scope
**Builds on:** schedule only a workflow that has passed the [ten-run evaluation](#ten-run-evaluation) ; use the [harness comparison](#harness-bake-off) when the runtime itself is still a choice. Design references: [building effective agents](https://www.anthropic.com/engineering/building-effective-agents) and [long-running harnesses](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) .
10
Run a local, permission-bounded assistant
Combine private retrieval and bounded automation, then control the local runtime safely from a phone.
Half day
Intermediate
Local GPU helpful
**Deliverable:** a locally hosted assistant with a documented data boundary, one useful skill, an action log and a kill switch.
### Build it
1. 01 **Draw the boundary first.** List allowed folders, network endpoints, users and retention. Check telemetry, plugins, backups and fallback providers.
2. 02 **Prove the model on the task.** Test a small local model on ten representative cases before buying hardware or adding tools.
3. 03 **Add one read source and one draft tool.** For example: search a notes folder and prepare a daily brief. Do not start with email sending, purchasing or shell access.
4. 04 **Add pocket control at the gateway.** Put an authenticated browser, companion app or messaging surface in front of the agent. Reach that surface through a private Tailnet or SSH tunnel; keep Ollama, LM Studio, vLLM and retrieval services on loopback or an internal network.
5. 05 **Log, review and stop.** Show whether the model runs on the phone, home host or cloud; record retrieval and tool calls, cap loop length and cost, and keep an obvious way to disable the service and revoke credentials.
**Permission contract**
`You may read only the allowlisted source and prepare a draft. Do not send, delete, rename, purchase, execute or contact external services. Stop after one failed tool call and report it.`
Acceptance test
**From cellular data, the phone reaches only the authenticated control surface; the assistant completes the workflow, makes no unapproved network call and cannot access a deliberately disallowed file.**
**Keep** — Boundary diagram, configuration, model version, ten-run result and tool log
**Boundary** — Local inference and remote transport are separate: include the tunnel, relay or messaging provider in the privacy review
**Failure signal** — A raw backend port is public, or a plugin, fallback model, relay or backup silently sends private context elsewhere
**Builds on:** combine the allowlisted corpus from [private documents](#private-documents) with the stops, logs and approval boundary from [bounded automation](#bounded-automation) . Remote-control references: [OpenClaw gateway access](https://docs.openclaw.ai/gateway/remote) , [Hermes messaging](https://hermes-agent.nousresearch.com/docs/user-guide/messaging/) and the [private operator pattern](https://isaiuseful.com/remote-spark.html.md#operators) . Other local-assistant options include [AnythingLLM](https://github.com/Mintplex-Labs/anything-llm) .
11
Audit and harden a UniFi network with AI
Move from screenshots to a GET-only API audit, then gate every firewall, zone and VLAN change.
Half day
Intermediate
Privileged infrastructure
**Deliverable:** a redacted network inventory, zone and policy matrix, prioritized findings, an ordered change plan and pre/post-change test evidence.
### Build it in three levels
1. 01 **Prepare the boundary and recovery path.** Use a local coding agent that can run a small API client. Confirm you can reach the UniFi console without the path you may change, take a current backup and record the Network version. Generate a dedicated API key in Site Manager or the console's Integrations area; prefer a one-month expiry, or set a dated reminder to revoke it.
2. 02 **Level 1 · Explain screenshots.** Start without API access. Redact public IPs, MAC addresses, SSIDs, client names and other identifiers from roaming logs, connectivity failures or RF charts, then ask for a plain-English diagnosis, competing explanations and the next measurement that would distinguish them.
3. 03 **Level 2 · Build a GET-only audit.** Keep the key in an ignored local `.env` file—never in chat, source control, command output or a generated dashboard. Put a mechanical wrapper in front of the agent that permits only `GET` , then inventory sites, devices, clients, networks, Wi-Fi broadcasts, firewall zones and ordered policies. Save a redacted baseline and build dashboards only from that bounded read path.
4. 04 **Turn findings into a testable design.** Have the agent map users, servers, IoT, cameras, guests, management and VPN access into named trust zones. For every proposed VLAN or policy, require source, destination, protocol or service, rule order, business reason, affected dependencies, expected result and rollback. Treat undocumented legacy-firewall objects, NAT, VPN and implicit rules as unresolved—not invitations to guess.
5. 05 **Level 3 · Approve one mutation at a time.** Review the exact API method, URL and redacted body before it runs. Prefer a new disabled policy when supported, keep management access untouched and never batch a firewall migration. After each approved write, re-read live state and test both traffic directions plus DNS, DHCP, admin, VPN and any device-discovery flows the change may affect.
6. 06 **Close the access.** Compare the final state with the plan, preserve the change log without secrets, verify rollback instructions and revoke the project key when the audit or migration ends. Renew a dashboard key intentionally rather than leaving forgotten access active.
**Staged operating contract**
`Start in screenshot or read-only mode. Load UNIFI_API_KEY only from the ignored local .env file; never print, paste, log or commit it. Permit GET requests only and save a redacted baseline. Return a zone/VLAN/firewall audit, evidence, an ordered proposal, dependency tests and rollback steps. Do not issue POST, PUT, PATCH or DELETE until I approve one exact request. After each approved change, re-read state, run the named tests and stop on any unexpected result.`
Acceptance test
**The audit is reproducible without exposing the key, every live change has one recorded approval and rollback, and the administrator can still reach the console after the network tests pass.**
**Needs** — A UniFi gateway or firewall, its Network API, network devices to inspect and a local tool-capable AI environment; a legacy firewall export is optional
**Keep** — Network and API versions, redacted inventory, policy order, proposals, approvals, responses, tests, rollback evidence and key revoke date—never the key
**Boundary** — AI can expose stale rules and prepare precise changes; it does not replace networking knowledge, an independent review or an out-of-band recovery path
**Failure signal** — A secret appears in chat or Git, the audit tool accepts write verbs, a rule lacks dependency tests or several live policies change before verification
**Builds on:** use the permission and kill-switch pattern from the [local assistant](#local-assistant) , the repeated checks from the [evaluation guide](#ten-run-evaluation) and the approval boundary from [bounded automation](#bounded-automation) . Official references: [UniFi API overview](https://help.ui.com/hc/en-us/articles/30076656117655-Getting-Started-with-the-Official-UniFi-API) , [current Network API](https://developer.ui.com/network) , [zone-based firewalling](https://help.ui.com/hc/en-us/articles/115003173168-Zone-Based-Firewalls-in-UniFi) and [virtual networks and VLAN assignment](https://help.ui.com/hc/en-us/articles/9761080275607-Creating-Virtual-Networks-VLANs) .
12
Build a linked second brain with a reviewable AI loop
Combine durable notes, derived retrieval, fact memory and the earlier review-only loop.
Half day
Intermediate
Local-first
**Deliverable:** a portable Markdown vault, a disposable search index, an optional correctable fact graph and one logged loop that prepares a review queue.
### Build it
1. 01 **Make the vault the source of truth.** Create `Inbox` , `Sources` , `Notes` , `Projects` and `AI Review` folders. Each durable note needs a clear title, source link, capture date and status. Use ordinary Markdown and links so the knowledge survives any one app or model.
2. 02 **Build a derived retrieval layer.** Parse only allowlisted notes, split Markdown by headings and meaning with Chonkie, preserve each file path and heading, then embed the chunks into a replaceable vector store. Rebuild the index from the vault; never make the index the only copy. [Follow the RAG build sequence](https://isaiuseful.com/rag.html.md#build) .
3. 03 **Separate documents from facts.** Use retrieval for quotations, arguments and changing source material. If you need stable relationships such as people, projects, systems and ownership, test FaultLine as a separate, correctable fact graph. Require a source note or explicit human confirmation before a fact becomes authoritative.
4. 04 **Add one loop-engineering job.** On a schedule, read new inbox notes, suggest links to existing notes, identify duplicates and draft one synthesis note in `AI Review` . Persist the last processed file and run result outside the conversation so the next run can resume without guessing.
5. 05 **Verify before promotion.** A separate check confirms every proposed link resolves, every quoted claim points to a source and no canonical note was changed. A person accepts, edits or rejects each draft before moving it into `Notes` .
**Loop contract**
`Read only allowlisted vault folders. Write only to AI Review. For each suggestion, cite the source note and heading. Do not delete, rename, retag or rewrite canonical notes. Stop on a broken link, missing provenance or conflicting fact and add it to the review queue.`
**What the graph means**
Obsidian's graph view visualizes links between notes. It can look neural, but it is not a neural network and dense connectivity is not proof of useful knowledge. Judge the system by whether you can recover a source, answer a real question and keep incorrect facts correctable.
Acceptance test
**Five known questions retrieve the right source passages, one absent answer is refused, every suggested link resolves and the scheduled run changes only the review folder.**
**Keep** — Markdown vault, attachments, source metadata, index recipe, fact corrections, loop state and review decisions
**Boundary** — Community plugins execute code; audit them and any sync, embedding or model endpoint before allowing private notes
**Failure signal** — The generated summary replaces the source, the graph rewards link spam or an unverified memory silently overrides a cited note
**Builds on:** extend the [private-document collection](#private-documents) with the review-only recurrence from [bounded automation](#bounded-automation) . Building blocks: [Obsidian](https://obsidian.md/help/) , [Chonkie](https://github.com/feyninc/chonkie) and [FaultLine](https://github.com/tkalevra/FaultLine) . Loop design references: the [Loop Engineering repository](https://github.com/cobusgreyling/loop-engineering) and [Addy Osmani's overview](https://addyosmani.com/blog/loop-engineering/) .
13
Confirm tomorrow's appointments and refill cancellations
Add voice, calendar state and reversible side effects to the bounded-automation pattern.
Pilot project
Intermediate
Local or cloud
**Deliverable:** a supervised appointment-confirmation loop with a locked waitlist offer, an auditable outcome for every call and a reconciliation report against the calendar.
### Build it
1. 01 **Snapshot tomorrow's eligible appointments.** At 17:00, let Logic Apps, n8n or another scheduler read the shared calendar and create one work item per appointment. Keep a stable appointment ID, start time, version and status; exclude entries without contact permission or enough data to match the attendee safely.
2. 02 **Make a deliberately small call.** Have a worker place the call through Azure Communication Services, Twilio or your own PBX and SIP operator. State who is calling and why, disclose automation where required, and offer fixed choices: confirm, cannot attend, repeat or speak to a person. Prefer keypad confirmation; map speech only to those states.
3. 03 **Confirm before releasing the slot.** Read the date and time back and require a second explicit answer before changing the booking. Silence, voicemail, low-confidence speech, identity uncertainty or a request outside the script becomes unresolved—never a cancellation.
4. 04 **Offer one locked slot at a time.** Query PostgreSQL for the next eligible, opted-in person using your existing priority rules. Atomically create an expiring offer for the open slot, call that person and continue only after decline or expiry. On acceptance, commit the booking and verify the calendar write before ending the call.
5. 05 **Reconcile, then continue.** Use idempotency keys for appointments, calls and offers so retries cannot cancel or book twice. After each outcome—and once at the end—compare the queue, database and calendar; alert a person about conflicts, failed writes and appointments that remain unconfirmed.
**Conversation contract**
`You may classify the caller's answer only as CONFIRM, CANNOT_ATTEND, REPEAT, HUMAN or UNCLEAR. Never invent, move or promise a time. A booking changes only after explicit confirmation and a successful scheduling-system response.`
True OSS route
### Keep the carrier contract. Own the application stack.
The phone network is still a paid service, but it does not require a communications cloud. Buy a SIP trunk and phone number from a local operator; keep scheduling, call control, speech and workflow state on infrastructure you operate.
> Visual: Open-source appointment calling stack
**Visual reading order:**
1. **01 · Calendar** **CalDAV** Nextcloud or Radicale; use a direct Microsoft Graph adapter only when the existing calendar is Microsoft 365.
2. **02 · Durable workflow** **Temporal + .NET worker** One workflow per appointment; timers, retries and signals survive restarts. Node-RED is the Apache-licensed visual alternative; n8n remains a self-hostable fair-code option.
3. **03 · Source of truth** **PostgreSQL** Appointments, consent, waitlist rank, call attempts and expiring offers live in relational transactions—not model memory.
4. **04 · Calls** **Asterisk or FreeSWITCH** Connect the PBX to the operator's SIP trunk. Use Asterisk ARI or FreeSWITCH ESL for originate, answer, DTMF, playback, hangup and call events.
5. **05 · Speech to text** **Speaches + faster-whisper** Stream narrow call audio to a local OpenAI-compatible STT endpoint. whisper.cpp is a lean CPU/edge alternative; always retain a DTMF path.
6. **06 · Text to speech** **Speaches or Kokoro-FastAPI** Generate local prompts through an OpenAI-compatible speech endpoint. openedai-speech documents the older Piper/XTTS pattern but is archived and should not be the new default.
7. **07 · Media loop** **ARI external media or Pipecat** Bridge RTP/WebSocket audio, voice activity, interruption and transcoding. A local LLM may classify free speech, but only into the five allowed states.
8. **08 · Operations** **OpenTelemetry + Grafana** Measure answer rate, STT latency, intent confidence, retries, slot locks and reconciliation failures without storing call audio by default.
### Wire the OSS path
1. A **Prove the telephone boundary first.** Get the SIP trunk, outbound caller ID and allowed calling regions in writing. From Asterisk or FreeSWITCH, place test calls, receive DTMF, play a WAV prompt and confirm hangup events before adding speech AI.
2. B **Run one durable state machine per appointment.** A Temporal schedule starts the nightly scan; a .NET worker reads CalDAV, writes the snapshot to PostgreSQL and starts child workflows with IDs such as `appointment:{calendar-id}:{version}` . Signals carry call outcomes back; activities perform all network and database I/O.
3. C **Keep media separate from booking authority.** Asterisk ARI originates the PJSIP channel and attaches an external-media channel. The media service converts the negotiated codec to the format expected by Speaches, applies VAD/barge-in and sends synthesized PCM back. It returns an intent event—not a calendar mutation.
4. D **Make the final choice deterministic.** For cancellation or offer acceptance, play a generated read-back and require DTMF or a second high-confidence answer. The .NET worker validates the current appointment version and performs one PostgreSQL transaction before writing the calendar.
5. E **Test crash and ambiguity paths.** Kill the worker mid-call, replay a webhook/event, drop STT, let an offer expire and change the calendar concurrently. The workflow must resume without a second call, release abandoned locks and route uncertainty to the human queue.
- [Temporal .NET SDK →](https://github.com/temporalio/sdk-dotnet)
- [Asterisk external media →](https://docs.asterisk.org/Development/Reference-Information/Asterisk-Framework-and-API-Examples/External-Media-and-ARI/)
- [Speaches →](https://github.com/speaches-ai/speaches)
- [Pipecat →](https://github.com/pipecat-ai/pipecat)
Acceptance test
**Every next-day appointment ends confirmed, explicitly released or visibly unresolved; each open slot has at most one live offer; and the final calendar contains no duplicate booking.**
**Keep** — Consent basis, appointment version, call and offer state, calendar response, retries and final reconciliation
**Boundary** — Apply local calling-hour, identification, recording, privacy and opt-out rules; disclose only the minimum appointment detail
**Failure signal** — Speech recognition releases a slot, two people receive a live offer, or the caller hears success before the calendar confirms it
**Builds on:** promote the state, idempotency and approval rules from [bounded automation](#bounded-automation) before adding outbound calls or calendar writes. Azure building blocks: [Logic Apps + Outlook calendar](https://learn.microsoft.com/en-us/azure/connectors/connectors-create-api-office365-outlook) and [ACS outbound call + choice recognition](https://learn.microsoft.com/en-us/azure/communication-services/quickstarts/call-automation/quickstart-make-an-outbound-call) . Alternative voice layer: [Twilio outbound calls](https://www.twilio.com/docs/voice/tutorials/how-to-make-outbound-phone-calls) and [speech/DTMF gathering](https://www.twilio.com/docs/voice/twiml/gather) .
OSS building blocks: [Asterisk](https://github.com/asterisk/asterisk) or [FreeSWITCH](https://github.com/signalwire/freeswitch) , [Speaches](https://github.com/speaches-ai/speaches) , [faster-whisper](https://github.com/SYSTRAN/faster-whisper) , [Kokoro-FastAPI](https://github.com/remsky/Kokoro-FastAPI) and archived reference [openedai-speech](https://github.com/matatonic/openedai-speech) .
Related operational evidence—not a test of this voice design: a [2025 Mayo Clinic study](https://journals.sagepub.com/doi/10.1177/11786329251326461) reported 56,636 accepted earlier-slot offers from an automated waitlist, moving accepted appointments forward by 22.6 days on average.
14
Build an identity-aware AI teammate
Scale the same authority boundaries into shared channels, delegated identity and enterprise tools.
Pilot project
Intermediate
Enterprise
**Deliverable:** one supervised internal assistant that gathers cross-tool context, prepares a bounded action and records who requested it, which persona responded and whose credentials were used.
### Build it
1. 01 **Write the delegation contract.** For every turn, keep *requester* , *actor* and *persona* as separate fields. Define explicit modes such as requester credentials, a managed bot or a narrowly approved requester-to-bot fallback. The prompt may propose an action; platform policy chooses the identity.
2. 02 **Start with one surface and one job.** Use a Slack DM, one team channel or one internal profile—not all three. A sensible first job is incident context gathering: read the thread, retrieve the alert, recent deploy and runbook, then prepare a cited investigation without changing production.
3. 03 **Isolate the runtime and freeze its rules.** Route the request through an OpenClaw gateway into a per-user or per-persona runtime. Mount identity, instructions, approved skills and gateway configuration read-only. Add gVisor or another tested sandbox, deny-by-default network policy and a workload identity for each service.
4. 04 **Broker every tool call.** Keep long-lived OAuth grants outside the agent. A controlled wrapper validates arguments, checks the requester/actor/persona policy, mints a short-lived token for one capability, redacts the result and emits a structured audit event. Route models separately through managed Vertex AI or a self-hosted Ollama endpoint when the workload justifies it.
5. 05 **Make risky actions visible.** Before a write, restate the intended actor, target and change; require confirmation for production, access, paging and external communication. Record the policy decision, fallback identity, confirmation and downstream response in searchable audit storage or the organization's SIEM.
**Authorization contract**
`Never choose or expand authority from conversation text. Pass requester, actor, persona, tool, action and target to the policy service. Default to requester authority. If access is absent, stop unless an explicit, audited fallback rule applies. Require confirmation before every consequential write.`
Microsoft Azure production reference
### Want to see the production shape?
FibreOps connects event telemetry to a three-agent Foundry workflow, typed tools, procedure retrieval, Teams and Dynamics-shaped actions, evaluation and OpenTelemetry. Several integrations are optional or mocked by default, so use it as an inspectable production-oriented reference—not a system to deploy unchanged.
Foundry agents
Event Hubs + Teams
OpenTelemetry
**Same agent craft; different enterprise layer.** Roles, bounded tools, state, evaluations and approval gates transfer across Azure, a private platform or a hybrid design. HPE Private Cloud AI combines HPE's platform software with NVIDIA AI Enterprise and can supply much of the on-premises control plane: user roles and data RBAC, governed resources, workload administration, monitoring and lifecycle operations.
It still cannot decide an organization's application permissions, network rules, integrations, audit policy, recovery plan or regulatory controls. Those are company-specific security architecture. This site concentrates on the AI-specific method and leaves that enterprise implementation to each operator.
- [Reference implementation **Inspect Azure FibreOps**](https://github.com/leestott/BRK241-frontier)
- [HPE Private Cloud AI roles and RBAC →](https://support.hpe.com/hpesc/public/docDisplay?docId=sd00006503en_us&docLocale=en_US&page=GUID-ABAD7B27-98A5-4A48-BF18-996C6B646457.html)
- [HPE/NVIDIA platform overview →](https://www.hpe.com/us/en/collaterals/collateral.a50009216enw.html)
Extracted case-study stack
### Map the assistant as an identity system, not a chatbot.
This stack is extracted from Damian F.'s reported Maestro implementation at Felix Pago. Product names describe that case; the controls are the reusable part of the pattern.
> Visual: Identity-aware OpenClaw teammate stack
**Visual reading order:**
1. **01 · Work surface** **Slack + internal web profiles** DMs, shared channels and published employee context are separate interaction modes with different actor policy.
2. **02 · Agent gateway** **OpenClaw** Routes channels, sessions, tools and model backends into isolated user or persona runtimes.
3. **03 · Runtime fleet** **GKE + Kubernetes + gVisor** Per-agent pods, lifecycle management and a stronger sandbox boundary around model-driven code and tools.
4. **04 · Service identity** **Istio + mTLS + SPIFFE** Workloads authenticate each other; network policy limits which runtimes can reach gateways, vaults and token services.
5. **05 · Model routing** **Vertex AI + Ollama** Managed models serve higher-stakes work while self-hosted GPU models cover suitable routine or high-volume tasks.
6. **06 · Delegated access** **OAuth vault + token minting** Durable grants stay outside the runtime; wrappers receive scoped, short-lived credentials for one invocation.
7. **07 · Work tools** **GitHub + Workspace + ClickUp + Notion** Constrained wrappers gather code, documents, tickets and knowledge under the selected actor's existing permissions.
8. **08 · Incident evidence** **PagerDuty + New Relic + SIEM** Alerts and telemetry become cited investigation context; session and tool events feed operational and security review.
- [OpenClaw →](https://docs.openclaw.ai/)
- [Slack platform →](https://docs.slack.dev/)
- [GKE →](https://cloud.google.com/kubernetes-engine/docs)
- [gVisor →](https://gvisor.dev/docs/)
- [Istio →](https://istio.io/latest/docs/)
- [SPIFFE →](https://spiffe.io/docs/latest/)
- [Vertex AI →](https://cloud.google.com/vertex-ai/generative-ai/docs/overview)
- [Ollama →](https://docs.ollama.com/)
Acceptance test
**A profile or shared persona can answer from published context, but an external lookup uses the requester's access; a denied request cannot borrow the profile owner's identity; and every write records requester, actor, persona, tool, target and confirmation.**
**Keep** — Identity tuple, policy decision, token scope and lifetime, tool request and response, confirmation, model route and audit event
**Boundary** — No long-lived credentials in the runtime; no prompt-selected identity; no shared persona or fallback bot without an owner and reviewed policy
**Failure signal** — A profile assistant reads as the profile owner, a shared channel hides the human requester or a tool receives a broad reusable token
**Builds on:** carry forward the least-privilege runtime from the [local assistant](#local-assistant) and the explicit side-effect gates from [bounded automation](#bounded-automation) . Case source: Damian F., [Building an Identity-Aware AI Teammate for Real Work Using OpenClaw](https://www.linkedin.com/pulse/building-identity-aware-ai-teammate-real-work-using-openclaw-finol-mpbrc/) (published June 26, 2026). The architecture and scale statements above are the author's report, not independently audited claims.
Loop engineering
## Design the system that prompts the agent.
A good prompt can finish one task. A loop repeatedly discovers work, carries state forward, verifies the result and knows when to stop or ask a person.
Stay one-shot
### Keep a person in the prompt.
Use a supervised agent session when the request is rare, the goal is ambiguous, the acceptance test is subjective or the action could affect production, money, identity or customer data.
Engineer a loop
### Automate only a boring recurrence.
Consider a loop when the same trigger recurs, inputs and permissions can be bounded, completion is independently testable and a failed attempt can stop without causing harm.
> Visual: Anatomy of a safe agent loop
**Visual reading order:**
1. **01** **Trigger** Run on a useful event or cadence. Exit cheaply when there is no work.
2. **02** **Skill** State one job, non-goals, watched scope and a structured output.
3. **03** **State** Read and update durable memory outside the conversation; prune stale items.
4. **04** **Isolation** Give each change its own branch or worktree and never expose secrets.
5. **05** **Maker + checker** Use separate implementation and verification contexts. The maker cannot approve itself.
6. **06** **Gate + log** Cap attempts and spend, record outcomes, then act only inside an allowlist or escalate.
Pattern picker
### Match the loop to the event, evidence and authority.
Choose the smallest control flow that fits the recurrence. “Shape” describes how work moves through the system—not how many agents to hire—and the final column is a starting ceiling, not a target for later autonomy.
| Situation | Trigger | Shape | Independent verification | Starting authority ceiling |
| --- | --- | --- | --- | --- |
| Daily report | Schedule | Chain Read → compare → report | Required schema, source links and a deterministic change diff | Report only |
| CI failure | Failed check | Router → merge-verify Classify → isolate → test | Named tests plus a fresh-context review of the complete diff | Prepare an isolated patch; a person merges |
| Issue implementation | Manual kickoff | Maker → checker Plan → edit → verify | Acceptance tests, denylisted-path check and unexplained-diff review | Propose a reviewable change; no deploy |
| Inbox triage | New item | Router Classify → queue → escalate | Sampled human review, false-alarm rate and missed-item audit | Draft or label only; do not send |
**Step 1 · Contract**
### Make “done” mechanical.
Choose one repository, branch or ticket queue. Write the goal, explicit non-goals, required tests, denylisted paths and the conditions that require human review.
**Step 2 · Memory**
### Persist facts, not chat.
Keep a small state file or board with item IDs, status, timestamps, attempt count and the last verified result. Read it first and clean it after every run.
**Step 3 · Control**
### Separate making from checking.
Let one context prepare the patch and a fresh context inspect the diff and rerun tests. Stop after a fixed number of failed attempts; do not retry a flaky test into submission.
**Step 4 · Operations**
### Budget the cadence.
Log runs, cost, actions and escalations. Check permissions mechanically, keep a kill switch and notify a person only when a decision or intervention is required.
**L0**
**Draft**
Purpose and boundaries documented.
**L1**
**Report**
Triage and update state; take no action.
**L2**
**Assist**
Prepare small, isolated changes for independent verification.
**L3**
**Unattended**
Run only after every safety, budget and observability gate is proven.
- [Video: 6 days autonomous coding](https://www.youtube.com/watch?v=VG2OWqBk0Gk)
Inspiration · Qwen vendor demonstration
### See what a sustained coding loop can aim for.
Qwen presents a six-day autonomous coding run as an ambitious view of what loop engineering can achieve. Use it to look for ideas about task selection, state, verification and recovery—but treat the video as inspiration, not evidence that your own unattended loop is ready. Reproduce the result on isolated, reversible work and measure it against the gates above.
A safer first loop
### Start with a daily, report-only triage.
Summarize actionable issues or failing checks into a state file. Run it supervised at least ten times. Measure missed items, false alarms, stale state, time saved and cost before allowing even a small patch.
**Starter loop contract**
`Read the allowlisted queue and prior state. Report only new, changed or blocked items in the required schema. Do not edit code, comment, merge, deploy or contact anyone. If evidence is missing, permissions fail or an item has reached the attempt cap, record the reason and escalate. Update the run log, then stop.`
**Stop the loop when**
the same item fails three times
state no longer matches the live system
the checker shares the maker's context
cost exceeds its daily cap
a change touches a denied path
Primary practitioner reference: [Loop Engineering repository](https://github.com/cobusgreyling/loop-engineering) . Go deeper with its [five-minute quickstart](https://github.com/cobusgreyling/loop-engineering/blob/main/docs/QUICKSTART.md) , [pattern picker](https://github.com/cobusgreyling/loop-engineering/blob/main/docs/pattern-picker.md) , [readiness checklist](https://github.com/cobusgreyling/loop-engineering/blob/main/docs/loop-design-checklist.md) and [safety guide](https://github.com/cobusgreyling/loop-engineering/blob/main/docs/safety.md) . Before scheduling anything, use the [hands-on harness comparison](#harness-bake-off) to test the surrounding scaffold on your own task. These are practitioner patterns, not evidence that unattended agents are reliable.
Implementation paths
## Pick an operating model before a product.
The right path follows data sensitivity, workflow stability, internal skills and the cost of operating another system.
| Path | Best when | Data boundary | Operating load | First proof |
| --- | --- | --- | --- | --- |
| Hosted assistant Fastest start | One person or team needs a general model now and approved data can leave the organization. | Vendor account, retention and training terms must match policy. | Low | Run ten redacted cases in an existing product; export outputs and corrections. |
| Local desktop Private baseline | One user has repeated private work that a smaller local model handles well. | Can stay on-device only after endpoints, telemetry, plugins and fallbacks are checked. | Medium | Use existing hardware, one runtime and one interface before buying a workstation. |
| Rented GPU job Burst compute | Fine-tuning, batch inference or evaluation needs more memory for hours or days—not a server purchase. | Region, disks, snapshots, logs, subprocessors, deletion and weight export must all meet policy. | Medium | Run one LoRA job on a redacted dataset, export the adapter and destroy the instance. |
| Self-hosted workspace Shared control | A team needs shared retrieval, identity, model routing and centralized logs. | Organization-controlled network, storage, auth, backups and model endpoints. | High | One collection, one group, read-only access and a named service owner. |
| Agent inside a work tool Artifact first | The deliverable lives in a repository, IDE, ticket system or other reviewable surface. | Scoped repository, account and tool permissions; secrets excluded. | Medium | Draft or patch only; require a diff, tests, citations or a preview. |
| Custom workflow or API Stable volume | Inputs, outputs and business rules are stable enough to justify engineering and evaluation. | Explicit per step; minimize context sent to each model and tool. | High | One queue, fixed schema, deterministic validations and human release. |
> Visual: Reference AI workflow architecture
**Visual reading order:**
1. **01** **Intake** Authenticate, classify and minimize the input.
2. **02** **Context** Retrieve allowlisted sources with provenance.
3. **03** **Route** Choose local, hosted or no model by policy.
4. **04** **Tools** Use deterministic code for facts and side effects.
5. **05** **Approve** Gate consequential actions and uncertainty.
6. **06** **Evaluate** Log outcome, correction, cost and failure.
Hardware decision
## Buy only after the workflow passes.
Model size alone does not determine usefulness. Measure quality, latency, concurrency, context and sustained throughput on your own cases.
Start here
### Existing computer
Use hosted models or test a small local model. Best for proving demand without a dedicated purchase.
One power user
### Consumer GPU or high-memory desktop
Good for private drafting, coding, extraction and transcription when a quantized model passes the test set.
Dedicated service
### AI workstation or server
Consider for larger models, long context or multiple users only after utilization and support ownership are clear.
Short training runs
### Rented European GPU
Rent the memory topology the job needs, export checkpoints promptly and verify the region and deletion path—not only the provider's headquarters.
Irregular frontier work
### Hosted model API
Fits bursty demand and advanced models without operating GPUs. Add budgets, data rules, caching and visible escalation.
Local and rented fine-tuning
## Tune behavior—not a changing knowledge base.
Fine-tuning is a measured training project, not a stronger prompt. Start only after a baseline shows a repeated, teachable failure.
**Prompt first**
If a clear instruction plus a few examples works, keep the base model. It is cheaper to change and easier to audit.
**Retrieve facts**
Use [RAG](https://isaiuseful.com/rag.html.md) for private, cited or frequently changing information. Weight updates are a poor document database.
**LoRA next**
Use an adapter for stable format, tone or task behavior that examples repeatedly improve. Keep the base and adapter versions separate.
**Full SFT last**
Consider full supervised fine-tuning only when an adapter is insufficient and the evaluation gain justifies much more compute and operational risk.
A defensible run
### Dataset → baseline → adapter → holdout → export
Remove secrets and unlicensed material; split train, validation and untouched test cases; record the base model, chat template, code, seed and hyperparameters; then compare the adapter against the unchanged baseline. Test general capability retention as well as the target task.
1. 01 **Write the target.** Name the behavior and a numeric or rubric-based acceptance threshold.
2. 02 **Inspect every source.** Record provenance, consent, license, PII treatment and deduplication.
3. 03 **Start small.** Run LoRA or QLoRA on a small model and a sample before scheduling expensive compute.
4. 04 **Reject regressions.** Use an untouched holdout and human review; an LLM judge is supplementary evidence.
5. 05 **Export before teardown.** Save adapter, tokenizer, config, metrics and a reproducible run manifest.
Cheap iteration
### Existing 24 GB-class GPU
Good for data-pipeline tests and parameter-efficient tuning of smaller models. Memory depends on model, sequence length, batch size, optimizer and quantization—do not size a job from parameter count alone.
Integrated lab
### DGX Spark
NVIDIA publishes Spark recipes for full SFT of a 3B model, LoRA for 8B and 70B models, and 70B QLoRA. Treat those as reproducible starting points, not throughput promises.
- [Use the Spark plan →](https://isaiuseful.com/remote-spark.html.md#spark-training)
Hours to days
### European GPU rental
Choose the exact GPU count, interconnect and region the run needs. Scaleway publishes H100/H100-SXM options and OVHcloud publishes one-to-four-H100 configurations.
- [Scaleway →](https://www.scaleway.com/en/choose-the-right-gpu/)
- [OVHcloud →](https://www.ovhcloud.com/en-gb/public-cloud/gpu/h100/)
Application route
### EuroHPC AI Factories
Eligible European startups and SMEs can apply for industrial access, including playground, fast-lane and larger allocations. This is allocation-based infrastructure—not an instant self-service rental.
- [Official access modes →](https://www.eurohpc-ju.europa.eu/ai-factories/ai-factories-access-modes_en)
European options
## European does not automatically mean sovereign.
Separate the model maker, legal entity, processing region, control plane, support access, subprocessors and key ownership before making a residency claim.
French model company
### Mistral Vibe
Mistral's current user product spans web, mobile, terminal and editor surfaces with chat, work and code modes. It is the current name for the product previously presented as Le Chat.
- [Platform overview →](https://docs.mistral.ai/getting-started/platform-overview)
API + operations
### Mistral Studio and Admin
Studio is the developer console and API surface for models, agents, evaluations and usage; Admin covers organizations, billing, SSO and access policy. Hosted use still requires a data-processing review.
- [Product documentation →](https://docs.mistral.ai/getting-started/platform-overview)
Open-weight route
### Mistral models on your hardware
Mistral documents local deployment through vLLM, TensorRT-LLM, TGI and other runtimes, from a desktop GPU to multi-node servers. Most open models use Apache 2.0, but some use modified terms—check the exact model card before training or redistribution.
- [Deployment guide →](https://docs.mistral.ai/models/deployment)
- [Licensing →](https://help.mistral.ai/en/articles/347393-under-which-license-are-mistral-s-open-models-available)
Public infrastructure
### EuroHPC AI Factories
A European compute and support route for research, startups, SMEs and industry. Access modes differ in eligibility, review and GPU-hour scale, so plan lead time and portability.
- [European Commission overview →](https://digital-strategy.ec.europa.eu/en/policies/ai-factories)
Commercial compute
### Telia, Scaleway and OVHcloud
These cover sales-led sovereign AI infrastructure and self-service European GPU instances. Compare actual region, GPU topology, identity, storage, network egress, snapshots and deletion—not the flag on the homepage.
- [Telia Helsinki →](https://www.telia.fi/en/telia-yrityksena/medialle/artikkeli/gcore-selects-telia-for-ai-gpu-cloud-newsroom)
- [Scaleway →](https://www.scaleway.com/en/choose-the-right-gpu/)
- [OVHcloud →](https://www.ovhcloud.com/en-gb/public-cloud/gpu/h100/)
Procurement test
### Ask seven concrete questions
Where are compute and backups? Who can administer them? Which subprocessors see data? Who holds keys? What is logged? How are disks and snapshots deleted? Can checkpoints and adapters leave without conversion or egress lock-in?
- [Apply it to a training run →](#finetuning)
Authority ladder
## Trust grows one permission at a time.
**01**
**Summarize**
Read supplied material
**02**
**Recommend**
Propose an action
**03**
**Prepare**
Create a draft or patch
**04**
**Act with approval**
Wait at the final gate
**05**
**Act narrowly**
Only after measured reliability
The smallest credible deployment
One workflow. One owner. One test set. One explicit approval boundary.
- [Choose a recipe](#chooser)
- [Check the economics](https://isaiuseful.com/index.html.md#roi)
- [Scale what works](https://isaiuseful.com/adoption.html.md#journey)
---
## https://isaiuseful.com/eu-ai-act (`/eu-ai-act.html.md`)
# The EU AI Act is more than labels.
Canonical source: [https://isaiuseful.com/eu-ai-act](https://isaiuseful.com/eu-ai-act)
- [Build](https://isaiuseful.com/guides.html.md)
· EU AI Act · Regulation plus enacted 2026 amendment
A support bot, applicant ranker, medical-device safety component and general-purpose model do not follow one checklist. Start with what the system does, where it reaches, which role you hold and when that route applies.
- [Find your Act routes](#checker)
- [See the whole framework](#risk-map)
- [Check enacted dates](#dates)
**Operational guide, not legal advice.** Legal position checked 2 August 2026. This page reads Regulation (EU) 2024/1689 together with Regulation (EU) 2026/1744, which entered into force on 27 July 2026.
- [▶ Watch the Commission explainer](https://youtu.be/0Xtd8vhytuo)
Official explainer · European Commission
**How You’ll Know When It’s AI**
Published 2 August 2026
**5**
routes can stack
PROHIBITED · HIGH-RISK · GPAI · TRANSPARENCY · OTHER
**6+**
roles can carry duties
PROVIDER ≠ DEPLOYER ≠ SUPPLY CHAIN
**2024→30**
phased legal timeline
ENTRY INTO FORCE ≠ APPLICATION DATE
01 · Situation first
## Start with the situation, not the tool.
The same model can sit in an ordinary drafting workflow, a prohibited practice, a high-risk decision system or several routes at once.
Internal assistant
### Staff use an off-the-shelf copilot for drafts
Usually no special risk tier on those facts, but the organisation still needs contextual AI-literacy measures and must check privacy, confidentiality, copyright and output quality.
Baseline duties apply
Employment
### Software ranks job applicants
Recruitment and candidate evaluation are Annex III use cases. Prepare the provider and deployer controls before the amended high-risk rules apply on 2 December 2027.
High-risk route
Regulated product
### AI performs a medical-device safety function
If the Section A product route and third-party health-or-safety conformity trigger are met, it is Annex I high-risk. The corresponding AI Act rules apply from 2 August 2028.
Product high-risk route
Human interaction
### A customer talks to a support bot
The interactive-system provider normally designs an AI notice into the first interaction. Article 50 applies from 2 August 2026.
Transparency route
Model release
### A company trains and releases a GPAI model
Model documentation, downstream information, copyright policy and a public training-content summary are separate from the rules for an app built on that model.
GPAI route
Harmful practice
### AI manipulates vulnerable people or creates social scores
Do not jump to disclosure. Screen the Article 5 prohibition first; the original prohibited-practice families have applied since 2 February 2025.
Stop and classify
Synthetic media
### A realistic CEO video or public-interest article is generated
A professional deployer may need a deepfake or text disclosure. Substantive review with an accountable publisher matters for the public-interest-text exception.
Article 50 route
Affected person
### AI materially influences a benefit, credit or other decision
The system may be high-risk, the deployer may owe notice, and the affected person may have complaint and explanation routes. Other equality and data-protection rights continue to apply.
Rights and high-risk routes
**Binding baseline** The Act is a set of stackable duties, not one risk score. Read [Regulation (EU) 2024/1689](https://eur-lex.europa.eu/eli/reg/2024/1689/oj) together with the enacted [Regulation (EU) 2026/1744](https://eur-lex.europa.eu/eli/reg/2026/1744/oj) .
02 · Act-wide pathfinder
## Find the routes worth opening.
Choose every situation that fits. The result is a linked reading path, not a verdict that the system is lawful or compliant.
Every route remains available without the interactive pathfinder. Start with [scope](#scope) , [roles](#actors) and the [prohibited-practice screen](#prohibited) .
03 · Scope
## First ask whether the Act reaches the system.
“AI” in marketing copy is not the test. The Act defines an AI system by how a machine-based system infers outputs that can influence physical or virtual environments.
**AI system:** a machine-based system designed to operate with varying autonomy, which may adapt after deployment and infers from inputs how to generate predictions, content, recommendations or decisions for explicit or implicit objectives.
EU reach
### Market, establishment or output
The Act can reach non-EU providers placing systems or GPAI models on the EU market, EU deployers, and third-country providers or deployers whose system output is used in the EU.
Narrow exclusions
### Purpose and phase matter
AI systems or models specifically developed and put into service solely for scientific R&D, and research, testing or development activity before market placement or putting into service, can be excluded. Other systems merely used in research remain covered; real-world testing is regulated, not generally excluded.
Personal use
### Only deployer duties drop out
A natural person’s purely personal, non-professional use is excluded from deployer obligations. It does not legalise fraud, harassment, privacy violations or unlawful content.
Open source
### Not a blanket exemption
The system-level exemption does not cover Article 5, Article 50 or high-risk systems. GPAI models have a separate, narrower open-source treatment.
Not an island
### Other law continues
GDPR, ePrivacy, equality, employment, copyright, consumer, product-safety, sector and national law continue alongside the AI Act.
Borderline software
### Document the inference test
Simple deterministic rules may fall outside the definition; sophisticated optimisation or learned prediction may not. Record the architecture, inputs, outputs, autonomy and inference method.
**Official classification help** Use Articles 2 and 3 of the [binding Act](https://eur-lex.europa.eu/eli/reg/2024/1689/oj) and the Commission’s [non-binding AI-system definition guidelines](https://digital-strategy.ec.europa.eu/en/library/commission-publishes-guidelines-ai-system-definition-facilitate-first-ai-acts-rules-application) .
04 · Actors
## Assign the role before the obligation.
One organisation can be provider, deployer and product manufacturer at the same time. Put the role in the system register instead of assuming the vendor owns every duty.
### Provider
Develops or commissions a system or GPAI model and places it on the market or puts the system into service under its name, whether paid or free.
### Deployer
Uses an AI system under its authority for professional or organisational activity. Employees acting under the organisation’s instructions are normally part of that deployment.
### Importer and distributor
An importer brings a third-country provider’s branded system to the EU market; a distributor makes a system available elsewhere in the supply chain.
### Product manufacturer
Places a product on the market or into service with an AI system under its own name. Product-law and AI Act routes can meet in one conformity process.
### Authorised representative
An EU-established representative accepts a written mandate for a non-EU provider and performs the specified documentation, cooperation and contact duties.
### Affected person
A person in the EU may receive notices and can use complaint or explanation routes. “Affected person” is not an operator role, but it changes the control design.
**Role-transfer trap:** an importer, distributor, deployer or other third party can become the provider by putting its name on the system, substantially modifying it, or changing the intended purpose so that it becomes high-risk. Contract for documentation and technical access before that happens.
05 · Horizontal duties
## “Minimal risk” does not mean “nothing applies.”
General duties, other law and voluntary controls sit underneath the special risk routes.
Article 4 · applies since 2 February 2025
### Support practical AI literacy
Tailor measures to people’s knowledge, experience, training, use context and affected groups. The amended rule does not require one certificate or guarantee a fixed level for every person.
- Map roles, systems and foreseeable mistakes.
- Teach limits, escalation, data handling and review controls.
- Keep the materials, audience, date and refresh trigger.
Article 4a · strict permission
### Bias work does not waive data law
High-risk providers may exceptionally process special-category personal data where strictly necessary for Article 10 bias detection and correction. Providers and deployers of other AI systems or models, and deployers of high-risk systems, may do so only for biases likely to affect health or safety, harm fundamental rights or cause discrimination prohibited by Union law, and only with every Article 4a safeguard. This creates no duty to conduct bias work.
- Show why anonymised, synthetic or other data cannot work.
- Restrict access, reuse, sharing and retention.
- Document necessity, security and deletion.
Voluntary layer
### Scale controls to the real risk
For systems outside high-risk rules, voluntary codes can still cover environmental performance, accessibility, inclusion, risk testing and governance. Other binding law may demand the same controls independently.
- Set a named owner and acceptable-use boundary.
- Measure quality, harm, cost and incident signals.
- Reclassify after purpose, model or audience changes.
**Amended literacy rule** The 2026 amendment changed “ensure a sufficient level” to taking measures that support development. The obligation remains. See the Commission’s [AI-literacy Q&A](https://digital-strategy.ec.europa.eu/en/faqs/ai-literacy-questions-answers) and Article 4 as amended by [Regulation (EU) 2026/1744](https://eur-lex.europa.eu/eli/reg/2026/1744/oj) .
06 · Route map
## Risk tiers are routes, not a single score.
Check from the top down. A high-risk system can also have Article 50 duties, and an application can inherit obligations from a GPAI model without becoming the model provider.
- [01 · Stop screen **Prohibited practices** Do not proceed unless a precise statutory exception applies.](#prohibited)
- [02 · Controlled route **High-risk systems** Annex I products and Annex III uses take different routes on the enacted 2027 and 2028 application dates.](#high-risk)
- [03 · Human-facing route **Article 50 transparency** Interaction, synthetic-output marking, biometrics, deepfakes and public-interest text.](#rules)
- [04 · Baseline route **Other or minimal risk** Literacy and other law still apply; voluntary controls remain useful.](#general-duties)
- [Separate model axis **GPAI models** Model-provider duties can stack with any system-level route downstream.](#gpai)
07 · Article 5
## Stop prohibited practices before designing controls.
Eight original practice families have applied since 2 February 2025. Two enacted additions apply from 2 December 2026; other law may already prohibit the same conduct.
1. **Harmful manipulation or deception** Subliminal, purposefully manipulative or deceptive techniques that materially distort an informed decision and cause, or are reasonably likely to cause, significant harm.
2. **Exploiting vulnerability** Using age, disability or a specific social or economic situation to materially distort behaviour in a significantly harmful way.
3. **Social scoring** Scoring people over time from social behaviour or personal characteristics where detrimental treatment is unrelated, unjustified or disproportionate.
4. **Individual criminal prediction** Assessing or predicting criminal-offence risk based solely on profiling or personality traits, subject to the narrow objective-facts and human-assessment boundary.
5. **Untargeted facial-image scraping** Creating or expanding facial-recognition databases by indiscriminately scraping the internet or CCTV footage.
6. **Emotion inference at work or school** Inferring emotions in workplaces or education institutions, except for a use intended for medical or safety reasons.
7. **Sensitive biometric categorisation** Using biometric data to infer race, political opinions, trade-union membership, religion, philosophical beliefs, sex life or sexual orientation, subject to narrow statutory boundaries.
8. **Real-time remote biometric identification** Law-enforcement use in publicly accessible spaces, except tightly limited victim, imminent-threat and serious-offence cases with legal, necessity, proportionality and authorisation safeguards.
Enacted · applies 2 December 2026
### Non-consensual intimate depictions
Realistic intimate or sexually explicit material of an identifiable person without the specified explicit consent. The amendment defines when provider and deployer conduct is caught and how safeguards matter.
Enacted · applies 2 December 2026
### Child sexual abuse material
AI systems used or designed for generation or manipulation of covered material, subject to the amendment’s purpose, foreseeability, safeguard and “without right” provisions.
**Do not reduce Article 5 to keywords** Purpose, effect, harm, context and exceptions are part of the legal test. Read [Article 5](https://eur-lex.europa.eu/eli/reg/2024/1689/oj) , the [2026 additions](https://eur-lex.europa.eu/eli/reg/2026/1744/oj) and the Commission’s [non-binding prohibition guidelines](https://ai-act-service-desk.ec.europa.eu/en/ai-act/resource/guidelines-prohibited-ai-practices) .
08 · Chapter III
## High-risk depends on product and intended use.
The enacted 2026 amendment moved Annex III duties to 2 December 2027 and Annex I product duties to 2 August 2028. Those are application dates, not permission to ignore applicable sector, data or equality law.
Article 6(1) · Annex I
### Regulated-product route
AI is a safety component of, or is itself, a product listed in Annex I and the product must undergo third-party conformity assessment for health or safety risks under that product law. Non-safety convenience or efficiency functions are excluded unless failure could endanger health or safety. Section A products take the direct Chapter III route, subject to any Article 2(13) delegated limitation; for Section B products, Article 2(2) routes most requirements through sector law.
Applies 2 August 2028
Article 6(2) · Annex III
### Listed-use route
The intended use falls within a specific Annex III case in biometrics, infrastructure, education, employment, essential services, law enforcement, migration, justice or democratic processes.
Applies 2 December 2027
Article 6(3)
### Narrow non-high-risk finding
An Annex III system may be found not high-risk only where it does not pose a significant risk of harm to health, safety or fundamental rights—including by not materially influencing a decision outcome—and performs at least one listed narrow procedural, preparatory, completed-work or pattern-detection task. Profiling of natural persons remains high-risk. Document and register the finding.
Document and register
Biometrics
Critical infrastructure
Education and training
Employment and workers
Essential services and benefits
Law enforcement
Migration and borders
Justice and democracy
Provider · where Chapter III duties apply
### Prove the system across its lifecycle
- Continuous risk management and pre-market testing.
- Data governance where training, validation or test data are used.
- Technical documentation, automatic logs and instructions.
- Human-oversight design, accuracy, robustness and cybersecurity.
- Quality management, conformity assessment, declaration, CE marking and registration.
- Post-market monitoring, corrective action and serious-incident reporting.
Deployer · where Chapter III duties apply
### Control the real use
- Follow instructions and assign competent, trained, authorised oversight.
- Keep controlled input data relevant and representative.
- Monitor, suspend and report risks or serious incidents.
- Retain controlled logs for at least six months unless other law says otherwise.
- Give workplace and affected-person notices where required.
- Complete registration, DPIA and fundamental-rights impact work where applicable.
**Do not flatten Annex I:** for Section B products, Article 2(2) directly retains only Article 6(1), Article 60a and Articles 102–112; Articles 57–59 follow only as sector law integrates the high-risk requirements. For Section A, Article 2(13) permits delegated limits on duplicated duties where product law provides equal or greater protection. Verify any applicable delegated act rather than assuming those duties are switched off.
**Supply-chain control:** importers and distributors verify provider, conformity, documentation, marking and storage or transport duties before making a high-risk system available. A rebrand, substantial modification or changed high-risk purpose can transfer provider responsibility.
**Legacy high-risk transition:** apart from specified large-scale Union IT systems, Article 111 generally brings a high-risk system placed on the market or put into service before its relevant Chapter III date into the Regulation only if its design is significantly changed from that date. High-risk systems intended for public-authority use must comply by 2 August 2030 in any case. Article 5 remains unaffected.
**Classification is still factual** Use Article 6, Annexes I and III and the Commission’s [high-risk overview](https://digital-strategy.ec.europa.eu/en/faqs/navigating-ai-act) . On 2 August 2026, its [detailed classification guidelines](https://digital-strategy.ec.europa.eu/en/library/draft-commission-guidelines-classification-high-risk-ai-systems) remained draft guidance, not binding law.
09 · Chapter V
## A GPAI model is a separate compliance layer.
Calling a model through an API does not make every customer its provider. Training, adapting, releasing and integrating a model can create different roles along the value chain.
All GPAI providers
### Document and enable downstream compliance
- Maintain model technical documentation.
- Give downstream system providers capability, limitation and integration information.
- Maintain a policy for EU copyright and rights reservations.
- Publish the required training-content summary using the AI Office template.
- Appoint an EU representative when required and cooperate with authorities.
Systemic-risk GPAI
### Add evaluation, mitigation and security
A model is presumed to have high-impact capabilities above 10 25 training FLOPs, or can be designated on equivalent capability or impact. Notify the Commission within two weeks after the threshold is met or known, subject to the rebuttal process.
- State-of-the-art evaluations and adversarial testing.
- Union-level systemic-risk assessment and mitigation.
- Serious-incident tracking and reporting.
- Model and physical-infrastructure cybersecurity.
Open source
### The exception is limited
Qualifying free and open-source GPAI models can be exempt from some technical and downstream documentation and representative duties. Copyright policy and training-summary duties remain, and systemic-risk models do not receive the same exemption.
Since 2 August 2025
### New GPAI models
Chapter V obligations apply to providers placing covered models on the EU market from that date.
From 2 August 2026
### Commission enforcement
The AI Office can enforce Chapter V and the GPAI fine regime. Code adherence is voluntary, not immunity.
By 2 August 2027
### Legacy GPAI models
Providers of models placed on the market before 2 August 2025 must bring them into compliance.
**Implementation route** Use the Commission’s [GPAI provider guidelines](https://digital-strategy.ec.europa.eu/en/policies/guidelines-gpai-providers) , [obligation summary](https://digital-strategy.ec.europa.eu/en/factpages/general-purpose-ai-obligations-under-ai-act) , mandatory training-content template and the voluntary GPAI Code of Practice.
10 · Article 50 · applies from 2 August 2026
## Transparency is one chapter, not the whole Act.
Each card names a different actor and trigger. A disclosure never makes a prohibited, unsafe or otherwise unlawful system lawful.
Article 50(1)
### Interactive AI systems
Providers design systems for direct human interaction so people are told they are interacting with AI by the first interaction, unless that fact is genuinely obvious in context.
**Actor** — **Provider**
**Keep** — First-screen, message or audio capture; wording; context; version; accessibility test.
Article 50(2)
### Machine-readable synthetic-output marking
Providers of systems generating synthetic text, audio, images or video make outputs marked and detectable using effective, interoperable, robust and reliable technical solutions as far as feasible.
**Actor** — **Provider** , including covered GPAI systems
**Boundary** — Standard editing, non-substantial alteration and narrow law-enforcement cases.
Article 50(3)
### Emotion and biometric exposure
Deployers inform people exposed to emotion-recognition or biometric-categorisation systems and comply with applicable personal-data law. Check Article 5 first because some uses are prohibited.
**Actor** — **Deployer**
**Keep** — Exposure map, notice, purpose, legal assessment, data-protection record and complaint route.
Article 50(4)
### Deepfake image, audio and video
Deployers disclose AI-generated or manipulated media that resembles existing people, objects, places, entities or events and could falsely appear authentic or truthful. Evident creative works receive a less disruptive disclosure accommodation, not silence.
**Actor** — **Professional deployer**
**Keep** — Authenticity assessment, expected audience, labelled final asset and placement.
Article 50(4)
### Public-interest text
Deployers disclose AI-generated or manipulated text published to inform the public on matters of public interest, unless it receives qualifying human review or editorial control and a person or entity holds editorial responsibility.
**Actor** — **Professional deployer**
**Keep** — Public-interest decision, sources, substantive changes, final approval and accountable publisher.
Human review
### “Someone looked at it” is not enough.
For the public-interest-text exception, a competent person reviews meaning, facts and completeness, can change or reject the text, approves the final published version and sits inside a process with identifiable editorial responsibility.
**Examine substance**
Meaning, accuracy and completeness—not only spelling or style.
**Check sources**
Verify factual claims and correct unsupported content.
**Exercise authority**
Approve, change or reject the substance.
**Lock the final**
Prevent unreviewed AI changes after approval.
Clear wording
### Say what happened at first contact.
The Act does not mandate one sentence. Make the information clear, distinguishable, accessible and available no later than the first interaction or exposure.
Chat
You are chatting with an AI assistant.
Voice
This call is handled by an AI voice assistant.
Deepfake
This video contains AI-generated or AI-manipulated people or events.
Public-interest text
This article was generated or altered with AI and has not received substantive human editorial review.
Official EU labels · optional
### Use the symbols as labels—not as a compliance stamp.
The Commission makes three human-visible labels freely reusable without attribution. They support Article 50(4) disclosure; they are not the provider-side machine-readable marking required by Article 50(2).

Basic icon
#### AI involved
Use with a nearby plain-language label or accessible second layer that explains what AI did.
- [Download light-on-dark SVG →](https://isaiuseful.com/assets/images/eu-ai-label-basic.svg)

Fully AI-generated
#### AI generated
Use when the entire in-scope item was generated by AI, apart from prompting, with no human-created elements or qualifying editorial control.
- [Download light-on-dark SVG →](https://isaiuseful.com/assets/images/eu-ai-label-generated.svg)

Partially AI-modified
#### AI modified
Use when pre-existing human-made content was changed with AI into an in-scope deepfake or public-interest text.
- [Download light-on-dark SVG →](https://isaiuseful.com/assets/images/eu-ai-label-modified.svg)
**Visible label ≠ machine-readable watermark.** For Code signatories, provider-side marking uses signed metadata and, where applicable, an imperceptible watermark; there is no single universal Article 50(2) watermark file. These symbols are the human-perceptible disclosure layer. Non-signatories may use them, but must not imply that they signed the Code.
- [Read the official icon and placement guidance →](https://digital-strategy.ec.europa.eu/en/policies/eu-icons-labelling-ai-generated-content)
- [Download the complete official SVG pack →](https://ec.europa.eu/newsroom/dae/redirection/document/129546)
- [Download the complete official PNG pack →](https://ec.europa.eu/newsroom/dae/redirection/document/129547)
- [Read the voluntary transparency Code →](https://ec.europa.eu/newsroom/dae/redirection/document/129555)
**Final non-binding guidance on binding law** The [Commission’s final Article 50 guidelines](https://digital-strategy.ec.europa.eu/en/library/guidelines-transparency-obligations-providers-and-deployers-ai-systems) explain the binding rule. The [transparency Code](https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content) is voluntary, and use of the [official icon set](https://digital-strategy.ec.europa.eu/en/policies/eu-icons-labelling-ai-generated-content) alone is neither proof of compliance nor proof of Code membership.
11 · Rights and remedies
## Design the route back to a responsible human.
The AI Act does not create one universal right to a model explanation or compensation. It creates specific notices and remedies while other EU and national rights continue.
Article 26(11) · from 2 December 2027
### Notice of high-risk decision support
Where an Annex III high-risk system makes or assists decisions about people, the deployer informs them that they are subject to its use, subject to the Act’s scope and dates.
Article 86 · applies from 2 August 2026
### Meaningful explanation
A person subject to a deployer decision based on output from an Annex III high-risk system—excluding point 2 critical-infrastructure systems—can obtain a clear, meaningful explanation of the system’s role and the main decision elements where the decision has legal or similarly significant effects which they consider adverse to health, safety or fundamental rights. Statutory and parallel-Union-law limits apply.
Article 85 · applies from 2 August 2026
### Complaint to the authority
Any natural or legal person with grounds to suspect an infringement can complain to the relevant market-surveillance authority.
Article 87 · applies from 2 August 2026
### Protected reporting
EU whistleblower protections apply to reporting AI Act infringements. Downstream GPAI providers also have a specific complaint route to the AI Office.
Parallel rights
### GDPR and equality law remain
Access, information, objection, human intervention, non-discrimination and judicial remedies may arise under other law even where the AI Act route is unavailable.
Practical control
### Make contact findable
Give the affected person a decision owner, channel, response process, evidence-preservation rule and escalation path—not only a generic privacy inbox.
**Remedy boundary** Articles 85–87 create complaint, explanation and reporting routes. The AI Act itself does not turn every breach into a criminal offence or guarantee compensation; other Union and national remedies may apply.
12 · Innovation and governance
## Test under supervision, then keep owning the risk.
Sandboxes can improve legal certainty and evidence. They do not waive Article 5, data law, safeguards or liability for harm.
Regulatory sandboxes
### Controlled, time-limited development
Each Member State must have at least one national or jointly equivalent sandbox operational by 2 August 2027. The AI Office may establish a Union-level sandbox within its competence; the EDPS may do so for Union bodies.
- Agree a sandbox plan, scope, safeguards and supervision.
- Use written proof and exit reports as compliance evidence—not certification.
- Mitigate significant risks or expect testing to be suspended.
Real-world testing
### People are not an ungoverned test set
Covered high-risk testing outside a sandbox needs a plan, authority route, registration and oversight. Informed consent, reversibility, incident handling, time limits and data safeguards apply subject to precise exceptions.
- Separate laboratory validation from regulated real-world testing.
- Preserve consent, approvals, versions, incidents and outcomes.
- Keep a prompt recall, suspension and deletion process.
### National authorities
Market-surveillance authorities enforce most system duties; notifying authorities oversee conformity-assessment bodies. Member States provide a single contact point.
### European AI Office
The Commission, acting through the AI Office, exclusively enforces Chapter V and the specified Article 75(1) system categories, including same-undertaking GPAI-based systems and VLOP/VLOSE systems, subject to statutory sector and public-authority exceptions. This system competence reaches deployers only where they are also the provider or belong to the same undertaking; other deployers remain under national supervision.
### AI Board and expert bodies
The European AI Board coordinates national application. The Scientific Panel and Advisory Forum add technical and stakeholder expertise.
### EDPS and notified bodies
The EDPS supervises Union institutions. Designated notified bodies perform the third-party conformity work required for covered high-risk systems.
**Sandbox safe harbour has limits:** participants remain liable under applicable Union and national law for damage to third parties. Prospective providers who follow the sandbox plan and terms and act in good faith on competent-authority guidance are protected from AI Act administrative fines for sandbox infringements, but supervisory and corrective powers remain.
**Small-company support** Member States give qualifying EU-established SMEs priority access to national sandboxes; an AI Office sandbox gives priority to SMEs and small mid-caps. Article 11 provides simplified technical documentation for SMEs and small mid-caps, while Article 63’s simplified quality-management route is limited to qualifying SMEs without partner or linked enterprises. Substantive protection duties remain.
13 · Dates
## Entry into force and application are different dates.
The Act entered into force on 1 August 2024. The timeline below reflects the enacted July 2026 amendment, not the earlier proposal.
Position from 2 August 2026
General provisions and Article 50 apply; GPAI enforcement is active. Annex III and Annex I high-risk lifecycle duties remain on the later dates enacted below.
1 August 2024
### Act enters into force
Regulation (EU) 2024/1689 enters into force, with phased application.
2 February 2025
### Definitions, literacy and original bans
Chapters I and II apply, including Article 4 and the eight original Article 5 practice families.
2 August 2025
### GPAI, governance and penalties
Chapter III Section 4, Chapter V, Chapter VII, Chapter XII and Article 78 apply, except Article 101; providers of GPAI models placed before this date retain the 2 August 2027 transition.
27 July 2026
### AI Omnibus enters into force
Regulation (EU) 2026/1744 enacts the revised dates and targeted changes; amended Articles 102–110 also apply from this date.
2 August 2026
### General date and Article 50
Most remaining provisions, Article 50 duties and Commission GPAI enforcement apply.
2 December 2026
### New bans and marking transition
The intimate-depiction and CSAM prohibitions apply. Providers of synthetic audio, image, video or text systems, including GPAI systems, placed on the market before 2 August 2026 must comply with Article 50(2)’s machine-readable marking duty by this date; the other Article 50 duties already apply from 2 August 2026.
2 August 2027
### Sandboxes and legacy GPAI
National sandboxes must be operational; pre-2 August 2025 GPAI models must comply.
2 December 2027
### Annex III high-risk duties
Chapter III Sections 1–3, except Article 6(5), apply to Article 6(2) listed-use systems, subject to Article 111’s legacy-system transition.
2 August 2028
### Annex I product high-risk duties
Chapter III Sections 1–3, except Article 6(5), apply to Article 6(1) systems, subject to Section B sector-law treatment, Article 2(13) delegated limitations and Article 111’s legacy-system transition.
2 August 2030
### Public-authority legacy systems
Providers and deployers of pre-existing high-risk systems intended for use by public authorities must comply by this date. AI components of Annex X large-scale IT systems placed on the market or put into service before 2 August 2027 must comply by 31 December 2030.
Prohibited practice
### Up to €35m or 7%
Whichever ceiling is higher for an undertaking, subject to proportionality and the special SME treatment.
Listed operator and Article 50 duties
### Up to €15m or 3%
Covers the specified duties in Articles 16, 22–26, 31, 33, 34 and 50, including the amended Article 25(2) and (4) value-chain cooperation duties.
Misleading information
### Up to €7.5m or 1%
For incorrect, incomplete or misleading information supplied to a notified body or competent authority in reply to a request.
GPAI providers
### Up to €15m or 3%
A separate Article 101 ceiling for intentional or negligent Chapter V and enforcement failures: whichever is higher for an undertaking.
**AI Office operator enforcement:** for Article 75(1) operators, Article 75c can apply the €15m/3% band to any applicable AI Act infringement, even one not enumerated in Article 99(4), and to failures to comply with enforcement decisions, measures or binding commitments. Misleading replies use the €7.5m/1% band. Periodic payments can reach 5% of average daily income or worldwide annual turnover in the preceding financial year per day.
**Union institutions have separate ceilings:** Article 100 allows the EDPS to impose up to €1.5m for Article 5 infringements and up to €750,000 for other infringements.
**Ceilings are not tariffs** Under Article 99(3)–(5), SMEs receive the lower fixed or percentage ceiling. Small mid-caps receive that lower treatment only for the €15m/3% and €7.5m/1% bands. Article 101’s GPAI ceiling remains the higher amount. National regimes can also use warnings and non-monetary measures, and not every duty—Article 4 is one example—sits in a fixed €15m/3% band. Actual enforcement must remain proportionate; Member States decide how administrative fines apply to their public authorities.
14 · Operating contract
## One register, then route-specific evidence.
Do not build separate spreadsheets for labels, literacy, GPAI and high-risk projects. Keep one source of truth and attach the evidence each route requires.
| Register field | Record | Acceptance check |
| --- | --- | --- |
| System and purpose | Owner, version, model, inputs, outputs, users, affected groups and intended purpose | Describes the actual workflow, not “AI-powered” |
| Scope and EU link | AI-system or GPAI rationale, market, establishment, output use and any exclusion | Every exclusion cites facts and a legal basis |
| Roles and supply chain | Provider, deployer, importer, distributor, manufacturer, representative and affected people | Contracts provide required documents and technical access |
| Route classification | Article 5 screen, Annex I/III test, GPAI layer, Article 50 triggers and other law | Concurrent routes remain visible |
| Date and owner | Applicable date, transition, accountable role and decision authority | No bare “2026 deadline” or other relative-date language |
| Controls and evidence | Literacy, data, testing, oversight, notices, conformity, review, logs and approvals | Each control has an artefact and a pass condition |
| Monitoring and change | Incidents, complaints, performance, model or purpose changes and review date | A trigger reopens classification before release |
Inventory
### Know every system
Include shadow use, vendor features, embedded product AI and retired versions still affecting people.
Decision gate
### Stop before deployment
No release until scope, role, prohibition, route, date and control owner are recorded.
Change control
### Reclassify material changes
New purpose, model, audience, data, autonomy, integration or brand can change risk and role.
Official sources · checked 2 August 2026
## Use the law for the final decision.
Guidance and codes explain implementation but do not replace the Regulation. Check EUR-Lex for a consolidated text; where it does not yet include the amendment, read both Regulations together.
- [Regulation (EU) 2024/1689 — original binding AI Act →](https://eur-lex.europa.eu/eli/reg/2024/1689/oj)
- [Regulation (EU) 2026/1744 — enacted AI Omnibus amendment →](https://eur-lex.europa.eu/eli/reg/2026/1744/oj)
- [AI Act Service Desk — official implementation timeline →](https://ai-act-service-desk.ec.europa.eu/en/ai-act/timeline/timeline-implementation-eu-ai-act)
- [AI Act Service Desk — official detailed compliance checker →](https://ai-act-service-desk.ec.europa.eu/en/eu-ai-act-compliance-checker)
- [Commission — scope, high-risk, GPAI and governance Q&A →](https://digital-strategy.ec.europa.eu/en/faqs/navigating-ai-act)
- [Commission — AI-system definition guidelines →](https://digital-strategy.ec.europa.eu/en/library/commission-publishes-guidelines-ai-system-definition-facilitate-first-ai-acts-rules-application)
- [Commission — draft high-risk classification guidelines as of 2 August 2026 →](https://digital-strategy.ec.europa.eu/en/library/draft-commission-guidelines-classification-high-risk-ai-systems)
- [Commission — Article 4 AI-literacy Q&A →](https://digital-strategy.ec.europa.eu/en/faqs/ai-literacy-questions-answers)
- [Commission — prohibited-practice guidelines →](https://ai-act-service-desk.ec.europa.eu/en/ai-act/resource/guidelines-prohibited-ai-practices)
- [Commission — GPAI provider guidelines →](https://digital-strategy.ec.europa.eu/en/policies/guidelines-gpai-providers)
- [Commission — final Article 50 transparency guidelines →](https://digital-strategy.ec.europa.eu/en/library/guidelines-transparency-obligations-providers-and-deployers-ai-systems)
- [Commission — governance and enforcement →](https://digital-strategy.ec.europa.eu/en/policies/ai-act-governance-and-enforcement)
The practical default
Classify the situation. Follow every route it triggers.
- [Build a route map](#checker)
- [Run the stop screen](#prohibited)
- [Build the register](#record)
---
## https://isaiuseful.com/tools (`/tools.html.md`)
# Which AI Tool Fits Your Workflow and Data Boundary?
Canonical source: [https://isaiuseful.com/tools](https://isaiuseful.com/tools)
A replaceable stack, not a shopping list
Search practical local, hosted and agent tools by what they do, where they run, how they are licensed and the caveat most likely to change your decision—including whether a phone can be the controller.
- [Search the catalogue](#catalog)
- [Use the selection test](#how-to-choose)
- [Start with a workflow](https://isaiuseful.com/guides.html.md)
**179**
searchable entries
PRIMARY LINKS + CAVEATS
**22**
practical topics
JOB BEFORE BRAND
**0**
automatic endorsements
VERIFY THE EXACT VERSION
Vetted tool catalogue
## How do you choose the right AI tool?
These are replaceable examples, not endorsements or a required stack. Each entry names the practical fit and the caveat that most changes the decision.
**Control surface**
The browser, native app or messaging client you actually use from the phone.
**Gateway**
Authenticates the person and device, owns sessions, constrains tools and routes requests.
**Backend**
Loads the model or provides retrieval and automation. It is not a mobile interface by itself.
**Execution location**
Keep phone, home and cloud visible separately; “local model” does not describe the transport.
**Pocket-control rule:** expose the authenticated interface or gateway through a private route—not a raw inference, vector-database or agent-control port. Search **phone** or **mobile** below for current entries that describe a direct control path.
Search names, jobs, platforms, licenses and caveats. Search combines with the topic filter below.
**Filter by topic**
### Run models locally
heterogeneous LLM inference + SFT
### KTransformers
An Apache-2.0 research framework for running and fine-tuning large mixture-of-experts models across CPU and GPU memory. Its current path exposes hardware-specific kernels and LLaMA-Factory integration; follow the exact model and instruction-set tutorial, and treat project throughput figures as configuration-specific rather than a desktop promise.
CPU + GPU
MoE models
Research project
- [KTransformers repository →](https://github.com/kvcache-ai/ktransformers)
- [CUDA libraries reference →](https://docs.nvidia.com/cuda-libraries/index.html)
runtime + interface
### Ollama + Open WebUI
A straightforward local API and browser interface for trying multiple models. **Mobile role:** open the authenticated WebUI from the phone over a private route while Ollama and the model stay on the host. Audit exposed network interfaces, user access and any cloud fallback. Current Open WebUI releases use a branding-restricted license that is not OSI-approved; call it self-hostable, not strict OSS.
Phone browser
Home inference
License caveat
- [Ollama API →](https://docs.ollama.com/api)
- [Open WebUI →](https://docs.openwebui.com/)
- [License explanation →](https://docs.openwebui.com/license/)
self-hosted team interface
### LibreChat
A multi-user web interface for local and hosted models, custom OpenAI-compatible endpoints, agents, files and MCP tools. It fits in front of vLLM or several providers; treat authentication, MCP credentials and code execution as production services rather than desktop conveniences.
Self-hosted
Multi-provider
Agents + MCP
- [Documentation →](https://www.librechat.ai/docs)
- [Compatibility matrix →](https://www.librechat.ai/docs/compatibility)
power-user local LLM frontend
### SillyTavern
An AGPL-3.0, locally installed interface for switching among local and hosted text models, image generators and speech engines, with detailed prompt, character, lorebook and extension controls. It is deliberately flexible rather than beginner-simple. Keep it on a trusted network: the project warns against exposing the server directly, stores multi-user data and API keys in plain text on the server, and gives UI extensions broad access while server plugins are unsandboxed.
Local interface
Many backends
Extensions need trust
- [SillyTavern repository →](https://github.com/SillyTavern/SillyTavern)
- [Documentation →](https://docs.sillytavern.app/)
- [Plugin security boundary →](https://docs.sillytavern.app/for-contributors/server-plugins/)
runtime
### llama.cpp
A portable, low-level local inference foundation with broad quantization and hardware support. Best when you want control and can own the configuration.
Local
Technical
Portable
- [Project repository →](https://github.com/ggml-org/llama.cpp)
desktop
### LM Studio
A desktop route for downloading, comparing and serving local models without assembling a full command-line stack. **Mobile role:** backend only—enable API-token authentication, keep the server private and put a phone-friendly interface or gateway in front.
Home backend
API token available
Private route
- [Official site →](https://lmstudio.ai/)
- [API authentication →](https://lmstudio.ai/docs/developer/core/authentication)
GPU serving
### vLLM
A high-throughput, OpenAI-compatible serving engine for shared GPU inference. **Mobile role:** remote-ready backend, not the public phone endpoint; keep it behind a private network and authenticated gateway. Use it when concurrency and batching matter; for DGX Spark follow NVIDIA's ARM64/Blackwell recipe and current model matrix instead of generic x86 installation instructions.
Home/server backend
OpenAI-compatible API
Gateway required
- [vLLM docs →](https://docs.vllm.ai/en/stable/)
- [Spark recipe →](https://build.nvidia.com/spark/vllm/instructions)
distributed inference + RL rollouts
### SGLang
An Apache-2.0 serving framework for language and multimodal models, from one GPU to distributed clusters. Its runtime covers prefix caching, continuous batching, prefill-decode disaggregation, speculative decoding, structured outputs and multiple parallelism strategies, with integrations for reinforcement-learning rollout generation. Treat project performance claims as configuration-specific: pin the model, engine, kernels, hardware and traffic shape, then measure latency, throughput, correctness and recovery on your workload.
GPU service
Distributed
RL integration
- [SGLang documentation →](https://docs.sglang.io/)
- [Repository and Apache-2.0 source →](https://github.com/sgl-project/sglang)
- [Benchmarking guide →](https://docs.sglang.io/docs/developer_guide/bench_serving)
agentic GPU serving
### TokenSpeed
An MIT-licensed, OpenAI-compatible inference engine from the LightSeek Foundation, built around a compiler-backed model layer, C++ scheduler and pluggable kernels for agentic and large-MoE workloads. Current recipes cover Kimi K3, DeepSeek V4, GLM‑5.2, GPT‑OSS and other model families across selected NVIDIA and AMD systems. This is an operator-grade, fast-moving stack—not a generic desktop runtime: pin the checkpoint, container, kernel bundle and topology, then validate correctness, context capacity and throughput together.
GPU service
NVIDIA + AMD recipes
Fast-moving
- [TokenSpeed documentation →](https://lightseek.org/tokenspeed/)
- [Repository and MIT-licensed source →](https://github.com/lightseekorg/tokenspeed)
- [Model and hardware recipes →](https://lightseek.org/tokenspeed/recipes/models)
multi-model inference server + cluster
### Superlinked SIE
An Apache-2.0 self-hosted server and Kubernetes stack that puts embedding, reranking, document conversion, structured output, content screening and text generation behind one OpenAI-compatible API. It loads models on demand through task-specific CPU and NVIDIA bundles, with a local Apple-silicon route. Pin the model weights and container or chart version, validate each model’s licence and hardware fit, and account for first-call downloads and opt-out usage telemetry before production.
Apache-2.0
Local + Kubernetes
Many model tasks
- [SIE repository →](https://github.com/superlinked/sie)
- [SIE documentation →](https://superlinked.com/docs/)
model gateway
### LiteLLM Proxy
A central OpenAI-compatible gateway for multiple local and hosted endpoints, with authentication hooks, budgets, rate limits and spend tracking. **Mobile role:** it can keep one endpoint stable while a phone session routes between home and cloud models, but a public app still needs end-user identity and session handling in front. It becomes a security boundary: keep it patched, authenticated and private.
Hybrid routing
Budgets
Shared service
- [Official documentation →](https://docs.litellm.ai/)
European models + API
### Mistral
Use Vibe for an end-user assistant, Studio/API for applications, or supported open-weight models through a local runtime. Availability and licenses differ by model; “Mistral” is not one deployment or one data boundary.
Hosted + local
Open-weight options
EU company
- [Compare Mistral routes →](https://isaiuseful.com/guides.html.md#europe)
on-device app runtime + SDK
### Foundry Local
Microsoft's GA runtime embeds curated ONNX Runtime-optimized chat and speech models into Windows, Apple-silicon macOS and Linux applications. C#, JavaScript, Python and Rust SDKs manage model download, hardware-specific variants and in-process inference; an optional OpenAI-compatible server supports local integrations. It needs no Azure subscription or API key, and it is designed for single-user on-device apps—not hosted Microsoft Foundry or multi-user serving like vLLM. The SDK is MIT-licensed, the CLI uses Microsoft terms and each model keeps its own license.
On-device
C# + JS + Python + Rust
Curated models
- [Foundry Local documentation →](https://learn.microsoft.com/en-us/azure/foundry-local/get-started)
- [SDK, samples and licenses →](https://github.com/microsoft/Foundry-Local)
OpenAI-compatible extension service
### Open WebUI Pipelines
Run Python filters, provider adapters and compute-heavy pre/post-processing behind an OpenAI-compatible endpoint. The project itself recommends built-in Functions for simple filters. Pipelines execute arbitrary code and are neither a durable scheduler nor a permission boundary.
Python
Filters + adapters
Not orchestration
- [Pipelines repository →](https://github.com/open-webui/pipelines)
multi-provider LLM CLI + Python
### LLM CLI
Simon Willison's Apache-2.0 command-line tool and Python library can call hosted, OpenAI-compatible and locally installed models, extract structured data, create embeddings and invoke tools through plugins. It logs prompts and responses to a local SQLite database by default; turn logging off or set retention and access controls before processing sensitive material. Plugins are executable Python packages, so pin and review them rather than treating the directory as a trust list.
Apache-2.0
Hosted + local models
Local history by default
- [LLM documentation →](https://llm.datasette.io/en/stable/)
- [Source and releases →](https://github.com/simonw/llm)
- [Logging controls →](https://llm.datasette.io/en/stable/setup.html#turning-sqlite-logging-on-and-off)
### Operate AI infrastructure
NVIDIA application development
### NVIDIA NIM + NeMo Framework + RAPIDS
NIM packages supported model-serving APIs, NeMo supplies model-development and customization workflows, and RAPIDS accelerates GPU data science. They are complementary products rather than one required bundle; confirm the supported model, container, hardware and commercial entitlement for the exact deployment.
Model APIs
Development
GPU analytics
- [NIM →](https://docs.nvidia.com/nim/)
- [NeMo Framework →](https://docs.nvidia.com/nemo-framework/user-guide/latest/overview.html)
- [RAPIDS →](https://docs.rapids.ai/)
NVIDIA inference serving
### TensorRT-LLM + Triton Inference Server + NVIDIA Dynamo
TensorRT-LLM optimizes supported large-model inference, Triton exposes production model-server endpoints and Dynamo coordinates distributed inference. Select only the components the service needs, then pin the model, kernels, driver, container and topology used in acceptance testing.
Optimization
Serving
Distributed inference
- [TensorRT-LLM →](https://nvidia.github.io/TensorRT-LLM/)
- [Triton →](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/)
- [Dynamo →](https://docs.nvidia.com/dynamo/)
commercial GPU operations
### NVIDIA Run:ai + Mission Control + Base Command Manager
Run:ai schedules and governs shared GPU workloads, Mission Control manages supported AI-factory configurations, and Base Command Manager provisions and operates clusters. Licensing, platform support and control-plane requirements differ; verify the exact product release and order rather than assuming hardware includes them.
Scheduling
Fleet operations
Entitlement-sensitive
- [Run:ai →](https://docs.nvidia.com/run-ai/index.html)
- [Mission Control →](https://docs.nvidia.com/nvidia-mission-control/)
- [Base Command Manager →](https://docs.nvidia.com/base-command-manager/index.html)
free NVIDIA development components
### NVIDIA Omniverse + AI Workbench
As of May 2026, Omniverse is free for development, production and redistribution without NVIDIA AI Enterprise; free use has community support, while enterprise support needs the applicable subscription. NVIDIA AI Workbench is also free for development workflows. Neither is a free-forever entitlement to the complete supported NVIDIA AI Enterprise production suite.
Development
Community support
Not the full suite
- [Omniverse license and support route →](https://docs.omniverse.nvidia.com/dev-guide/latest/common/NVIDIA_Omniverse_License_Agreement.html)
- [AI Workbench introduction →](https://docs.nvidia.com/ai-workbench/user-guide/latest/overview/introduction.html)
NVIDIA fabric operations
### NVIDIA UFM + NetQ
UFM manages supported InfiniBand fabrics while NetQ observes and troubleshoots supported Ethernet and network environments. They cover different fabrics and support matrices; confirm topology, device software, licensing and telemetry retention before choosing either.
InfiniBand
Network telemetry
Support matrix
- [UFM →](https://networking-docs.nvidia.com/ufmenterpriseum/6221)
- [NetQ →](https://docs.nvidia.com/networking-ethernet-software/cumulus-netq/)
NVIDIA system foundation
### DGX OS + DCGM + CUDA + NVIDIA Container Toolkit
DGX OS supplies the supported system baseline, DCGM exposes GPU health and telemetry, CUDA provides the GPU programming platform, and Container Toolkit gives containers controlled GPU access. Pin their compatibility chain; installing these components does not by itself prove an NVIDIA AI Enterprise entitlement.
System software
GPU telemetry
Containers
- [DGX OS →](https://docs.nvidia.com/dgx/dgx-os-7-user-guide/)
- [DCGM →](https://docs.nvidia.com/datacenter/dcgm/latest/index.html)
- [CUDA →](https://docs.nvidia.com/cuda/)
- [Container Toolkit →](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/)
open workload control
### Kubernetes + Kueue + Volcano + Slurm
Kubernetes is the container-orchestration baseline, Kueue adds quota-aware batch admission, Volcano adds Kubernetes batch scheduling, and Slurm supplies an established HPC workload manager. Choose one primary queue and scheduling authority; overlapping control loops make priority, recovery and capacity ownership harder to reason about.
GPU scheduling
Batch queues
Operator-owned
- [Kubernetes →](https://kubernetes.io/docs/tasks/manage-gpus/scheduling-gpus/)
- [Kueue →](https://kueue.sigs.k8s.io/docs/)
- [Volcano →](https://volcano.sh/docs/home/introduction/)
- [Slurm →](https://slurm.schedmd.com/gres.html)
training + model lifecycle
### PyTorch + Transformers + PEFT + TRL + MLflow
PyTorch supplies the training framework; Transformers provides model implementations; PEFT supplies parameter-efficient adapters; TRL provides supervised and preference-training loops; and MLflow records runs and artifacts. Pin the model, dataset, evaluation and environment together: this is a composable lifecycle stack, not a pre-integrated production platform.
Training
Adapters + preference
Experiment evidence
- [PyTorch →](https://pytorch.org/docs/stable/index.html)
- [Transformers →](https://huggingface.co/docs/transformers/index)
- [PEFT →](https://huggingface.co/docs/peft/index)
- [TRL →](https://huggingface.co/docs/trl/index)
- [MLflow →](https://mlflow.org/docs/latest/)
GPU enablement + AI scheduling
### NVIDIA GPU Operator + KAI Scheduler
NVIDIA GPU Operator manages the Kubernetes software components needed to run NVIDIA GPUs; KAI Scheduler adds AI-oriented scheduling, including gang scheduling, to the workload-control layer. They do not replace a clear queue owner, compatibility testing or the hardware-vendor support boundary.
Kubernetes GPUs
Gang scheduling
Operator-owned
- [NVIDIA GPU Operator →](https://docs.nvidia.com/gpu-operator/latest/)
- [KAI Scheduler →](https://github.com/kai-scheduler/KAI-Scheduler)
delivery + artifact supply chain
### Argo Workflows + Harbor
Argo Workflows runs Kubernetes-native DAG and step workflows; Harbor provides a private artifact registry. Pair them only with explicit access, signing, retention and promotion rules—an image in a registry is not a safe or approved deployment by itself.
Workflow automation
Private registry
Policy required
- [Argo Workflows →](https://argoproj.github.io/argo-workflows/)
- [Harbor →](https://goharbor.io/docs/)
bare-metal lifecycle + configuration
### MAAS + Foreman + Ansible
MAAS and Foreman are alternative routes for provisioning and host lifecycle management; Ansible applies repeatable configuration and operational automation. None removes the need for tested backups, firmware coordination, break-glass access and a documented rebuild path.
Bare metal
Configuration
Recovery ownership
- [MAAS →](https://maas.io/docs)
- [Foreman →](https://docs.theforeman.org/)
- [Ansible →](https://docs.ansible.com/)
storage + network policy + secrets
### Ceph + MinIO + Cilium + Vault
Ceph supplies distributed object, block and file storage; MinIO's current vendor route is AIStor while the former community repository is archived; Cilium provides eBPF networking and policy; and Vault manages secrets under HashiCorp's current terms. Treat this as four separately operated services, not a pre-integrated platform.
Storage
Network policy
Secrets
- [Ceph →](https://docs.ceph.com/en/latest/)
- [MinIO / AIStor →](https://docs.min.io/aistor/)
- [Cilium →](https://docs.cilium.io/en/stable/)
- [Vault →](https://developer.hashicorp.com/vault/docs)
hardware and vendor boundary
### Redfish + NVIDIA Driver + NCCL + platform firmware
Redfish standardizes server-management APIs, the NVIDIA driver connects the operating system to the accelerator, NCCL coordinates collective communication and platform firmware remains OEM-specific. CUDA is catalogued with the NVIDIA system foundation above. Validate the whole compatibility chain and support route before an upgrade.
Hardware API
GPU runtime
OEM boundary
- [Redfish →](https://www.dmtf.org/standards/redfish)
- [NVIDIA Driver →](https://docs.nvidia.com/datacenter/tesla/driver-installation-guide/)
- [NCCL →](https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/)
- [Platform firmware example →](https://docs.nvidia.com/dgx/dgxb300-fw-update-guide/)
### Choose model and embedding services
hosted models + embeddings
### Cohere + Cohere Embed
Cohere provides hosted chat, embedding and reranking APIs. Treat each endpoint as a separate design choice: record the model version, input type and data boundary, and evaluate retrieval with your own corpus before committing to an index.
Hosted API
Embed + rerank
Version explicitly
- [Cohere documentation →](https://docs.cohere.com/)
- [Cohere Embed →](https://docs.cohere.com/docs/embeddings)
open-weight local embeddings
### Nomic Embed + Sentence Transformers
Nomic Embed v2 is an Apache-2.0 multilingual embedding model that runs locally through Sentence Transformers or Transformers. Benchmark retrieval on your own languages, use the required query and document prefixes, and pin the model revision and output dimension—changing either means re-embedding the corpus.
Local weights
Multilingual retrieval
Pin the index contract
- [Nomic Embed v2 model card →](https://huggingface.co/nomic-ai/nomic-embed-text-v2-moe)
- [Sentence Transformers →](https://sbert.net/)
community model + provider metadata
### Models.dev
An MIT-licensed, community-maintained catalogue and JSON API for model identifiers, provider routes, context limits, capabilities and prices, with a type-safe SDK and offline snapshot. Use it to seed selectors, routing tests and cost comparisons—not as the final procurement source. Pin a snapshot and reconcile consequential limits, regional availability, terms and prices with each provider's current documentation.
MIT
JSON API + SDK
Verify with providers
- [Browse Models.dev →](https://models.dev/)
- [Data source and licence →](https://github.com/anomalyco/models.dev)
### Build speech and calling
open speech models
### Microsoft VibeVoice
An MIT-licensed codebase for long-form speech recognition and text to speech. The project publishes an ASR model with speaker, timestamp and transcript structure plus a real-time TTS model; model weights keep their own terms. Test language, diarization and timestamp accuracy on your audio, and require consent for cloned or synthetic voices.
ASR + TTS
Local models
Consent required
- [VibeVoice repository →](https://github.com/microsoft/VibeVoice)
local voice studio
### Voicebox
An MIT-licensed desktop stack for voice cloning, speech generation, dictation and MCP voice output across several model engines. It can keep captures and models local, but GPU support, weight licenses and quality vary by engine. Clone only voices you own or have permission to use, and keep disclosure and anti-impersonation controls around generated audio.
Local-first
STT + TTS
Voice cloning
- [Voicebox repository →](https://github.com/jamiepine/voicebox)
- [Responsible-use policy →](https://github.com/jamiepine/voicebox/blob/main/RESPONSIBLE_USE.md)
open-source desktop dictation
### OpenFlow
A free MIT-licensed voice-to-text app for macOS and Windows that can type into the active application. Speech recognition and optional AI cleanup can each run locally, through a self-hosted endpoint or through a remote provider, so confirm both settings before calling a workflow private. Current downloads are unsigned and trigger an operating-system warning; verify the release or build from source rather than treating the bypass as routine.
Local-first
macOS + Windows
Unsigned builds
- [OpenFlow →](https://openflow.computer)
- [Source and license →](https://github.com/avijeett007/openflow)
desktop + mobile AI dictation
### Superwhisper
A proprietary dictation app for macOS, Windows and iOS with custom vocabulary, task-specific modes, file transcription and meeting notes. It offers local and cloud speech and language models and can work offline, but privacy and connectivity depend on the models and mode you select; the vendor notes that offline models perform best on Apple-silicon Macs.
Local + cloud
Desktop + iOS
Free + paid tiers
- [Superwhisper →](https://superwhisper.com)
- [Product documentation →](https://superwhisper.com/docs/get-started/introduction)
OSS PBX + direct SIP
### Asterisk + FreeSWITCH
Connect a local operator's SIP trunk to an open-source PBX, then control outbound calls, DTMF, media and hangup events through Asterisk ARI or FreeSWITCH ESL. The carrier, number rights, emergency restrictions and call rates remain contractual services.
Self-operated
SIP/RTP
Call control
- [Asterisk →](https://github.com/asterisk/asterisk)
- [ARI →](https://docs.asterisk.org/Configuration/Interfaces/Asterisk-REST-Interface-ARI/)
- [FreeSWITCH ESL →](https://developer.signalwire.com/freeswitch/integration/event-socket/)
local STT + TTS service
### Speaches
An MIT-licensed, OpenAI-compatible server for streaming transcription, translation and speech generation. It combines faster-whisper for STT with Kokoro or Piper-family TTS paths and supports CPU or GPU deployment.
MIT
OpenAI-compatible
Streaming
- [Speaches →](https://github.com/speaches-ai/speaches)
- [faster-whisper →](https://github.com/SYSTRAN/faster-whisper)
- [whisper.cpp →](https://github.com/ggml-org/whisper.cpp)
local text to speech
### Kokoro-FastAPI + openedai-speech
Kokoro-FastAPI is the active Apache-licensed, OpenAI-compatible TTS wrapper to test first. openedai-speech remains a useful AGPL reference for Piper/XTTS and voice mapping, but its maintainer archived it in January 2026 and calls it mostly obsolete.
Local voices
OpenAI speech API
Archive noted
- [Kokoro-FastAPI →](https://github.com/remsky/Kokoro-FastAPI)
- [openedai-speech archive →](https://github.com/matatonic/openedai-speech)
real-time voice pipeline
### Pipecat + LiveKit Agents
Open-source frameworks for streaming transports, VAD, interruptions, STT, LLM and TTS stages. Use them for the media conversation loop; keep appointment locks and business state in PostgreSQL and a durable workflow engine.
Real-time media
Pluggable speech
Barge-in
- [Pipecat →](https://github.com/pipecat-ai/pipecat)
- [LiveKit Agents →](https://github.com/livekit/agents)
### Work with documents and media
research OCR model
### DeepSeek-OCR
A research release for turning document images into compact textual context through vLLM or Transformers. The repository is MIT-licensed and its reference environment pins CUDA 11.8, PyTorch 2.6 and FlashAttention. Treat “optical compression” as a model technique—not lossless storage—and measure transcription, reading order, tables and hallucinations against page images.
OCR
GPU-oriented
Verify against source
- [DeepSeek-OCR repository →](https://github.com/deepseek-ai/DeepSeek-OCR)
- [CUDA libraries reference →](https://docs.nvidia.com/cuda-libraries/index.html)
agent-native Office file CLI
### OfficeCLI
An Apache-2.0 command-line tool for agents to read and edit Word, Excel and PowerPoint files without an Office installation. It offers a local preview and single-binary releases. Work on copies, render the result and compare formulas, charts, fonts and pagination in a real Office-compatible viewer; ZIP-format access does not guarantee perfect application fidelity.
DOCX + XLSX + PPTX
No Office install
Render and verify
- [OfficeCLI repository →](https://github.com/iOfficeAI/OfficeCLI)
local generative-media workflows + image repair
### ComfyUI + IOPaint
ComfyUI is a GPL-3.0 node-graph engine for repeatable local image, video, audio and 3D workflows; IOPaint is an Apache-2.0, self-hosted route for erasing, replacing and extending images on CPU, GPU or Apple silicon. Model weights and outputs keep their own licences. Pin stable releases and known workflows, keep ComfyUI's optional API nodes disabled for a fully offline path, and treat third-party models and custom nodes as untrusted code and artifacts.
Local + self-hosted
Visual workflows
Audit nodes + weights
- [ComfyUI documentation →](https://docs.comfy.org/)
- [ComfyUI source and releases →](https://github.com/Comfy-Org/ComfyUI)
- [IOPaint source →](https://github.com/Sanster/IOPaint)
local video creation
### FramePack Studio
An Apache-2.0 local application based on FramePack for image-to-video and text-to-video work, with queueing, prompt timelines and blending, LoRA, upscale and post-processing paths. It requires a CUDA-compatible GPU: 8 GB VRAM is the stated minimum, while 16 GB or more and 80 GB or more of storage are recommended.
Local video
CUDA GPU
8 GB VRAM minimum
- [FramePack Studio source →](https://github.com/FP-Studio/framepack-studio)
- [FramePack Studio documentation →](https://docs.framepackstudio.com/)
local media studio
### Amuse
A local Windows studio for image, video, audio and text pipelines plus editing, upscale and interpolation. Its current reference route uses CUDA 13 and recommends RTX hardware. The software licence is personal and non-commercial; commercial use needs a separate licence, and model weights keep their own terms.
Windows
CUDA 13 + RTX
Commercial licence needed
- [AmuseAI source →](https://github.com/saddam213/AmuseAI)
- [Amuse licence →](https://github.com/saddam213/AmuseAI/blob/master/LICENSE)
local AI-app launcher
### Pinokio
An MIT-licensed local launcher and runtime for installing and running AI apps. Its scripts can execute commands and download code with your authority: review source, prefer verified or frozen scripts, and isolate credentials before running an app.
Local launcher
MIT
Scripts need review
- [Pinokio desktop →](https://desktop.pinokio.co/)
- [Pinokio documentation →](https://desktop.pinokio.co/docs/)
- [Pinokio source →](https://github.com/pinokiocomputer/pinokio)
text and image to 3D research
### threestudio
An Apache-2.0 local research framework for text, image and few-shot 3D generation across multiple methods and extensions. It is NVIDIA/CUDA-oriented, Ubuntu-tested and states a 6 GB VRAM minimum. Dependencies and model licences vary, and this is a research stack rather than a beginner desktop default.
NVIDIA + CUDA
6 GB VRAM minimum
Research stack
- [threestudio source →](https://github.com/threestudio-project/threestudio)
### Build AI applications and agents
general-purpose agent research
### OpenManus
An MIT-licensed implementation of a general tool-using agent with browser, code, retrieval, sandbox and multi-agent paths. The maintainers describe it as a simple prototype. Model endpoints, browser access and executable tools create the real data and permission boundary, so start in a disposable environment with narrow credentials and an external acceptance test.
Python
General agent
Prototype
- [OpenManus repository →](https://github.com/FoundationAgents/OpenManus)
multi-agent orchestration
### CrewAI + CAMEL
CrewAI packages agents, tasks, crews and stateful flows; CAMEL focuses on multi-agent research and composable agent societies. Use either only when separate roles improve a measured result—more agents add coordination cost, failure modes and permissions.
Python
Multi-agent
Role-based
- [CrewAI →](https://docs.crewai.com/)
- [CAMEL →](https://docs.camel-ai.org/)
experimental persona simulation
### TinyTroupe
Microsoft's MIT-licensed Python library simulates configurable personas and focus groups for research and product exploration. It runs on your machine, but OpenAI and Azure OpenAI remain the primary model routes and Ollama support is experimental. Treat synthetic responses as hypotheses, validate them against real people and data, and read the project's legal disclaimer before use.
Local library
Persona simulation
Research only
- [TinyTroupe repository →](https://github.com/microsoft/TinyTroupe)
event workflows + RAG pipelines
### LlamaIndex Workflows + Haystack
LlamaIndex Workflows provides event-driven, step-based control for retrieval and agents. Haystack composes model, retrieval and processing components into explicit pipelines. Both suit knowledge-heavy applications; keep source provenance and evaluation independent of the framework. [Plan the retrieval architecture](https://isaiuseful.com/rag.html.md#architecture) .
Python
Retrieval
Explicit pipelines
- [LlamaIndex Workflows →](https://developers.llamaindex.ai/python/llamaagents/workflows/)
- [Haystack →](https://docs.haystack.deepset.ai/docs/intro)
typed agent frameworks
### Agno + Pydantic AI
Agno supplies a broad agent stack with teams, workflows and knowledge integrations. Pydantic AI emphasizes typed dependencies, validated outputs, tools and testable Python application patterns. Choose the smaller abstraction that meets the acceptance test and keep business authorization outside the model loop.
Python
Structured output
Tools + workflows
- [Agno →](https://docs.agno.com/)
- [Pydantic AI →](https://pydantic.dev/docs/ai/overview/)
program and optimize model behavior
### DSPy
Define model programs from typed modules and signatures, then optimize prompts or weights against examples and metrics. DSPy is most useful when you have a representative dataset and a meaningful objective; optimization against a weak metric simply automates overfitting.
Code-first
Optimization
Evaluation required
- [DSPy documentation →](https://dspy.ai/)
Model Context Protocol development
### FastMCP + Official MCP Registry
FastMCP helps build MCP clients and servers; the official registry publishes discoverable server metadata. A registry listing is not a security review: pin packages, inspect requested capabilities, isolate credentials and test every tool's authorization boundary.
MCP
Python
Discovery, not trust
- [FastMCP →](https://gofastmcp.com/)
- [Official MCP Registry →](https://registry.modelcontextprotocol.io/)
local browser automation + MCP
### agent-browser
An Apache-2.0 native Rust CLI that automates a browser locally and can start a local MCP stdio server for an agent. It can inspect the accessibility tree, navigate, click, fill forms, upload files, and read cookies or network requests. The browser still opens real services under a real signed-in session: isolate the profile and credentials, allow only authorized domains and accounts, require confirmation for uploads or consequential actions, and do not mistake local execution for a local data boundary. First setup downloads Chrome for Testing unless a compatible browser is already available.
Apache-2.0
CLI + local MCP
Live-session authority
- [Source, releases and installation →](https://github.com/vercel-labs/agent-browser)
- [Apache-2.0 license →](https://github.com/vercel-labs/agent-browser/blob/main/LICENSE)
### Use agent skills and playbooks
open skill format
### Agent Skills specification
A lightweight open format in which a folder contains a required `SKILL.md` plus optional scripts, references and assets. It provides a portable packaging convention, not a trust or permission boundary: read the instructions and executable files, pin the source and test the skill with least privilege before installing it broadly.
Open format
Portable instructions
Inspect before use
- [Agent Skills overview →](https://agentskills.io/home)
skill optimization research
### Microsoft SkillOpt
An MIT-licensed optimizer that revises natural-language skills from agent trajectories and promotes an artifact only through validation gates. It changes prompts and procedures rather than model weights. Keep an untouched holdout, version every candidate and include model/API cost and data exposure in the experiment; repeated optimization can overfit a weak evaluator.
Python
Skills + evals
Holdout required
- [SkillOpt repository →](https://github.com/microsoft/SkillOpt)
Apple development skill library
### Claude Code Apple Skills
An MIT-licensed collection covering iOS, macOS, product work, testing, App Store preparation and code generation. Marketplace installs track `main` ; use the repository's era tags when reproducibility matters. Treat generated patterns and legal copy as starting points, then verify them against the target Xcode, SDK, platform guidance and counsel where needed.
Apple platforms
Prompt library
Pin for stability
- [Apple Skills repository →](https://github.com/rshankras/claude-code-apple-skills)
software-development workflow skills
### Superpowers
An MIT-licensed, cross-agent skill set that enforces discovery, specification, planning, test-driven implementation and review stages. It is an opinionated engineering method rather than evidence that autonomous coding is reliable. Adopt the parts that improve your acceptance tests and remove roles or ceremony that do not improve measured outcomes.
Planning
TDD
Cross-agent
- [Superpowers repository →](https://github.com/obra/superpowers)
specialist persona library
### Agency Agents
An MIT-licensed catalogue of role prompts spanning engineering, design, marketing, finance, healthcare and other functions. These are reusable briefs and personalities—not verified professionals or validated operating procedures. Test deliverables against primary sources and named reviewers, especially in regulated, financial, legal or health work.
Prompt library
Many domains
No credential guarantee
- [Agency Agents repository →](https://github.com/msitarzewski/agency-agents)
legal workflow references
### Claude for Legal
Anthropic's Apache-2.0 reference agents, skills and connectors for commercial, privacy, product, corporate, employment, litigation, regulatory and learning workflows. They can run as Claude plugins or behind the Managed Agents API. They are not legal advice: preserve privilege, apply the correct jurisdiction, minimize connector access and require qualified counsel to release consequential work.
Claude ecosystem
Legal workflows
Human counsel
- [Claude for Legal repository →](https://github.com/anthropics/claude-for-legal)
SEO skill + specialist agents
### Claude SEO
An MIT-licensed Claude Code plugin with technical, content, schema, local, international and AI-search audit workflows. The repository says its recommendations are grounded in primary Google guidance, while optional providers add external data collection. Recheck every recommendation, respect robots and service terms, and treat rankings or traffic as measured outcomes—not promises.
SEO audits
Claude Code
Optional data APIs
- [Claude SEO repository →](https://github.com/AgriciDaniel/claude-seo)
opinionated Claude Code workflow
### gstack
An MIT-licensed set of Claude Code skills for product, design, engineering, QA, security review and release work. Its productivity examples are the author's own measurements, not an independent benchmark. Some workflows can browse, change code or deploy, so inspect each skill, narrow credentials and keep repository review, tests and release approval outside the persona.
Claude Code
Product to release
Self-reported results
- [gstack repository →](https://github.com/garrytan/gstack)
### Build knowledge and workflows
consent-gated data-broker removal skill
### Hermes unbroker
An MIT-licensed Hermes skill that finds a consenting person's exposure on people-search sites, submits supported opt-outs and records human-only steps. It is US-first and may use browser sessions, email credentials and a dossier containing the exact personal data being removed. Encrypt state, minimize disclosure, verify consent and removals, and recheck legal and site requirements.
Privacy workflow
Hermes skill
Sensitive dossier
- [unbroker skill →](https://github.com/NousResearch/hermes-agent/tree/main/optional-skills/security/unbroker)
agent-maintained Markdown knowledge
### OpenWiki
An MIT-licensed CLI that creates and updates codebase documentation or a personal local wiki in a portable Markdown format. The configured model can receive repository or note context, and scheduled runs send anonymous reliability telemetry unless disabled. Review generated changes like code: pin the provider and prompt, require diffs and keep primary documentation canonical.
Code docs + personal wiki
Markdown
Review updates
- [OpenWiki repository →](https://github.com/langchain-ai/openwiki)
self-hosted receipt + invoice extraction
### TaxHacker
An MIT-licensed application for extracting fields from receipts, invoices and transactions into a structured accounting view. Self-hosting the app does not make a remote model local, and the documents contain financial and personal data. Keep originals and your accounting system canonical, configure the model boundary explicitly and require human review for categories, tax and filings.
Receipts + invoices
Self-hosted app
Not accounting advice
- [TaxHacker repository →](https://github.com/vas3k/TaxHacker)
self-hosted web acquisition
### Firecrawl + Scrapy
Firecrawl packages search, scrape and crawl operations behind a self-hostable API; its core is AGPL-3.0 and the hosted service includes additional features. Scrapy is a BSD-licensed Python framework for explicit spiders and extraction pipelines. Whichever route you choose, respect robots directives, site terms, privacy and rate limits, and store source URLs and retrieval times with extracted content.
Self-hosted
Web crawling
Policy required
- [Firecrawl and self-hosting guide →](https://github.com/firecrawl/firecrawl)
- [Scrapy repository →](https://github.com/scrapy/scrapy)
local knowledge workspace
### Obsidian
A free-to-use, proprietary app that stores the canonical vault locally as plain-text Markdown. Backlinks, properties and graph view help navigate relationships; they do not make the vault a neural network. Treat community plugins as executable code and audit optional Sync, Publish and AI endpoints separately.
Local files
Portable Markdown
Plugin boundary
- [Obsidian Help →](https://obsidian.md/help/)
- [License overview →](https://obsidian.md/license)
- [Second-brain guide →](https://isaiuseful.com/guides.html.md#second-brain)
continuous local activity memory
### screenpipe
A source-available desktop recorder that indexes screen text, screenshots, app and browser context, audio transcripts and user input locally, then exposes search, an API, MCP and scheduled agent workflows. That breadth creates a high-risk monitoring dataset even when storage stays on-device. For workplace use, define purpose and legal basis, consult workers where required, complete a DPIA where systematic monitoring triggers it, minimize captured apps and fields, set short retention and deletion paths, encrypt storage, disable default analytics and keep cloud transcription, sync and hosted models off unless their processors and transfers are approved.
Local by default
Source-available
DPIA likely at work
- [screenpipe documentation →](https://docs.screenpipe.com/home)
- [Source, licence and data-flow notes →](https://github.com/screenpipe/screenpipe)
RAG ingestion + chunking
### Chonkie
An MIT-licensed Python library for turning text, Markdown, tables and code into retrieval chunks. It offers token, sentence, recursive, semantic, neural and code-aware chunkers plus optional embedding and vector-store integrations. Preserve file and heading provenance, and evaluate retrieval on your documents instead of trusting a chunking benchmark. [Follow the RAG build sequence](https://isaiuseful.com/rag.html.md#build) .
MIT
Local-capable
Replaceable pipeline
- [Repository →](https://github.com/feyninc/chonkie)
- [Documentation →](https://docs.chonkie.ai/)
validated personal memory
### FaultLine
An early-stage AGPL service that exposes correctable, structured facts through MCP and OpenWebUI tools. Its repository describes a validation gate, confidence strengthening and archived corrections; treat those as project claims to test. Keep source notes canonical, secure the MCP endpoint and verify user isolation, backup and retraction behavior.
AGPL-3.0
Local service
MCP fact graph
- [Repository →](https://github.com/tkalevra/FaultLine)
- [Project limitations →](https://github.com/tkalevra/FaultLine/blob/main/HONESTY.md)
documents
### Docling + AnythingLLM
Parse complex files with provenance, then retrieve over a bounded collection with a model endpoint you can replace. [See how the stages fit together](https://isaiuseful.com/rag.html.md#model) .
Local-capable
RAG
Page provenance
- [Docling →](https://github.com/docling-project/docling)
- [AnythingLLM →](https://github.com/Mintplex-Labs/anything-llm)
end-to-end RAG applications
### RAGFlow + Embedchain
RAGFlow is a self-hostable RAG platform with document processing, retrieval and agent workflows; Embedchain is a lighter Python framework for adding data and querying an application. Use them to accelerate a prototype, but retain the original files, chunk metadata and an exportable evaluation set. [Define the acceptance test](https://isaiuseful.com/rag.html.md#evaluate) .
RAG
Application layer
Keep provenance
- [RAGFlow →](https://ragflow.io/docs/dev/)
- [Embedchain →](https://docs.embedchain.ai/)
document preparation + graph retrieval
### Unstructured + Microsoft GraphRAG
Unstructured partitions and stages documents for downstream retrieval. Microsoft GraphRAG builds graph-based indexes and query paths for relationship-heavy corpora. Neither removes the need to inspect extraction quality, preserve page-level provenance and compare the result with simpler retrieval. [Compare RAG routes](https://isaiuseful.com/rag.html.md#choose) .
Ingestion
GraphRAG
Evaluate complexity
- [Unstructured →](https://docs.unstructured.io/open-source/introduction/overview)
- [Microsoft GraphRAG →](https://microsoft.github.io/graphrag/)
vector retrieval
### Qdrant
Add Qdrant when retrieval needs explicit collections, metadata filters, hybrid search, persistence or multiple applications. A small single-user prototype may not need another service. Self-hosted Qdrant starts without authentication or encryption by default, so never expose its ports unchanged. [Apply the retrieval security checklist](https://isaiuseful.com/rag.html.md#operate) .
Vector search
Local or cloud
Secure explicitly
- [Documentation →](https://qdrant.tech/documentation/)
- [Security guide →](https://qdrant.tech/documentation/operations/security/)
local + self-hosted vector retrieval
### Chroma
Chroma offers an approachable local or self-hosted retrieval store, with a managed route available separately. Define metadata filters, tenancy, authentication, backups and an export path before relying on it—the vector index is rebuildable infrastructure, not the source of truth. [See the reference architecture](https://isaiuseful.com/rag.html.md#architecture) .
Local or self-hosted
Vector search
Export plan
- [Chroma documentation →](https://docs.trychroma.com/)
vector databases at service scale
### Weaviate + Milvus
Both support vector search, metadata filtering and production deployments across self-managed and hosted routes. They introduce a real database service: size indexes, define tenancy, secure every endpoint and test restore and re-embedding before relying on either. [Follow the production build gates](https://isaiuseful.com/rag.html.md#build) .
Vector database
Self-hosted options
Operate deliberately
- [Weaviate →](https://docs.weaviate.io/weaviate)
- [Milvus →](https://milvus.io/docs)
graph retrieval · optional
### Neo4j + FalkorDB
Add a graph only when relationships and multi-hop traversals are part of the acceptance test; PostgreSQL is enough for an appointment queue. Neo4j Community is GPL-3.0. FalkorDB targets GraphRAG but uses SSPL, so it is source-available rather than OSI open source. GraphAware is a commercial graph-services layer, not a required database component. [Use the GraphRAG decision gate](https://isaiuseful.com/rag.html.md#graph) .
Knowledge graphs
GraphRAG
License check
- [Neo4j Community →](https://github.com/neo4j/neo4j)
- [FalkorDB →](https://github.com/FalkorDB/FalkorDB)
- [GraphAware →](https://graphaware.com/)
self-hosted agent memory
### Mem0 + Graphiti
Mem0 provides an open-source memory layer; Graphiti is the self-hosted temporal context-graph engine behind hosted Zep. Neither makes extracted memories authoritative: preserve raw records and consent rules, namespace every user and agent, and test correction, deletion, provenance and stale-memory behavior before using recalled facts in decisions.
Long-term memory
Self-hosted
Retraction test
- [Mem0 →](https://docs.mem0.ai/)
- [Graphiti →](https://github.com/getzep/graphiti)
knowledge-graph memory
### Cognee
Cognee is a self-hosted memory framework that ingests supplied data into graph and vector structures and exposes add, cognify, search and memory workflows. It can use local databases and OpenAI-compatible local model endpoints; configure both generation and embeddings explicitly, keep source records outside the derived graph, and test deletion and re-indexing before production use.
Self-hosted
Graph + vector
Source records required
- [Cognee repository →](https://github.com/topoteretes/cognee)
stateful agents + fast state
### Letta + Redis
Letta builds stateful agents around persistent memory and tools. Redis can hold fast operational state and provide search primitives, but it does not define memory policy for you. Separate conversation state, durable facts and source records; give each a retention and deletion path.
Agent memory
Structured state
Retention required
- [Letta →](https://docs.letta.com/)
- [Redis for AI and search →](https://redis.io/docs/latest/develop/ai/)
visual workflow automation
### Node-RED + n8n
Both provide visual triggers, API calls and approvals. Node-RED is Apache-2.0 and fits a strict OSS stack. n8n has the larger packaged integration surface but uses a Sustainable Use fair-code license; keep that distinction visible in procurement.
Visual
Connectors
Different licenses
- [Node-RED →](https://github.com/node-red/node-red)
- [n8n →](https://github.com/n8n-io/n8n)
managed integration automation
### Zapier + Make + Pipedream
Zapier and Make emphasize visual SaaS automation; Pipedream combines hosted connectors with code steps. They are useful for bounded integrations and approvals, but credentials and payloads cross a managed control plane—review data handling, retries, run history and exportability.
Managed
Connectors
Visual + code
- [Zapier →](https://zapier.com/ai/)
- [Make →](https://www.make.com/en)
- [Pipedream →](https://pipedream.com/docs)
durable business workflows
### Temporal
Use Temporal when timers, callbacks, retries and state must survive process restarts for hours or days. Its MIT-licensed server and .NET SDK fit the appointment workflow: deterministic workflow code owns state while activities perform calendar, PBX and database I/O.
MIT
Durable state
.NET SDK
- [Temporal server →](https://github.com/temporalio/temporal)
- [.NET SDK →](https://github.com/temporalio/sdk-dotnet)
data + batch orchestration
### Prefect + Apache Airflow + Kestra
Use these for scheduled data flows, backfills and observable batch dependencies. Prefect and Airflow are Python-centered; Kestra defines workflows declaratively and spans more runtimes. They can launch preparation or reporting, but long-lived business conversations may be clearer in Temporal or application code.
Schedules
Backfills
Observable runs
- [Prefect →](https://docs.prefect.io/)
- [Apache Airflow →](https://github.com/apache/airflow)
- [Kestra →](https://kestra.io/docs)
AI workflow builders
### Dify + Flowise
Visual options for retrieval, model routing and tool-using flows. Useful for prototypes when evaluation and export remain part of the design.
Visual
Self-hostable
Prototype
- [Dify →](https://dify.ai/)
- [Flowise →](https://flowiseai.com/)
### Move, transform and serve data
OSS ingestion libraries
### Meltano + dlt
Code-first, open-source routes for moving API and database data into an analytical store without adopting a managed connector platform. Use them when the connector set is modest and versioned pipelines matter more than a large visual catalog.
OSS
ELT
Code-first
- [Meltano →](https://github.com/meltano/meltano)
- [dlt →](https://github.com/dlt-hub/dlt)
connector platforms
### Airbyte + Fivetran
Airbyte offers a broad self-hosted connector catalog but its main platform uses Elastic License 2.0, not an OSI license. Fivetran is a managed proprietary service. Compare them for connector coverage and operations—not as equivalents to a true OSS stack.
Many connectors
ELT
License boundary
- [Airbyte →](https://github.com/airbytehq/airbyte)
- [Fivetran →](https://www.fivetran.com/)
analytics transformation
### dbt Core
Version SQL transformations, tests and documentation after data lands in an analytical database. dbt is not an ingestion tool, operational scheduler or appointment transaction engine; use it for trusted reporting models.
SQL
Tests
Lineage
- [dbt Core →](https://github.com/dbt-labs/dbt-core)
lakehouse table formats
### Apache Iceberg + Delta Lake
Transactional table formats for large analytical data lakes, schema evolution and multiple compute engines. They become relevant when call/event history reaches lake scale; they do not replace PostgreSQL for live slot locks.
Apache-2.0
Data lake
Large scale
- [Apache Iceberg →](https://github.com/apache/iceberg)
- [Delta Lake →](https://github.com/delta-io/delta)
dashboards + internal apps
### Apache Superset + Dashjoin
Superset is an Apache-licensed BI and data-exploration layer for outcome dashboards. Dashjoin is an AGPL low-code platform that can put forms and actions over existing sources. Neither should become the authoritative booking state.
BI
Low-code UI
Read-mostly
- [Apache Superset →](https://github.com/apache/superset)
- [Dashjoin →](https://github.com/Dashjoin/platform)
self-hosted business intelligence
### Metabase
An AGPL-licensed, self-hostable service for queries, dashboards and self-service exploration when a simpler BI interface fits better than Superset. Run the Open Source Edition as a Java JAR or Docker container with a production application database. Self-hosting keeps Metabase the company out of your data unless you opt in to anonymized usage statistics, but it does not make connected data sources local or safe by default. Use a dedicated least-privileged, read-only database user, keep results and dashboard access behind authenticated groups, and keep transactional state in its application service. Fine-grained row and column controls require a self-hosted Pro or Enterprise plan.
AGPL-3.0
JAR or Docker
Least-privilege data access
- [Self-hosting documentation →](https://www.metabase.com/docs/latest/installation-and-operation/installing-metabase)
- [Open Source and commercial license terms →](https://www.metabase.com/license)
- [Database roles and privileges →](https://www.metabase.com/docs/latest/databases/users-roles-privileges)
### Build operational intelligence systems
property graph + traversal
### JanusGraph + TinkerPop + Cassandra
JanusGraph supplies a distributed property-graph layer, TinkerPop defines the Gremlin traversal API and Cassandra can provide the storage backend. Use this combination only when relationship-heavy queries and scale justify three moving parts; keep raw and curated records outside the graph so it remains a rebuildable serving model.
Open source
Distributed graph
Operational burden
- [JanusGraph architecture →](https://docs.janusgraph.org/getting-started/architecture/)
- [TinkerPop reference →](https://tinkerpop.apache.org/docs/current/reference/)
- [Cassandra documentation →](https://cassandra.apache.org/doc/latest/)
events + visible ingestion
### Apache Kafka + Apache NiFi
Kafka fits durable event streams and replay; NiFi fits visual routing, provenance and source-to-destination controls. A nightly file does not need either. Name which service owns retry and deduplication, retain raw inputs and test what happens when the source sends late, duplicated or malformed records.
Apache-2.0
Streaming
Provenance
- [Kafka documentation →](https://kafka.apache.org/documentation/)
- [NiFi documentation →](https://nifi.apache.org/docs.html)
versioned lakehouse + owned object storage
### Apache Iceberg + S3-compatible storage
Iceberg adds schema evolution, partition management and snapshots over object storage. Default new self-hosted builds to SeaweedFS; consider Garage for lightweight multi-site replication or Ceph RGW when Ceph already has an operations team. MinIO Community is no longer a default: its upstream repository is archived and its distribution is source-only. Pin the catalog, engine and store versions, then prove compatibility and a consistent restore.
SeaweedFS default
Garage + Ceph alternatives
MinIO migration only
- [Iceberg documentation →](https://iceberg.apache.org/docs/latest/)
- [SeaweedFS →](https://github.com/seaweedfs/seaweedfs)
- [Garage →](https://garagehq.deuxfleurs.fr/)
- [Ceph RGW →](https://docs.ceph.com/en/latest/radosgw/)
- [MinIO status →](https://github.com/minio/minio)
transform + federated query
### Apache Spark + Trino + DuckDB
Spark handles distributed transforms, Trino provides interactive SQL across services and DuckDB is often enough for the first single-machine proof. Start with DuckDB, preserve standard table formats and add distributed engines only when measured data size, concurrency or latency demands them.
SQL + compute
Scale later
Open formats
- [Spark documentation →](https://spark.apache.org/docs/latest/)
- [Trino documentation →](https://trino.io/docs/current/)
- [DuckDB documentation →](https://duckdb.org/docs/stable/)
text + faceted + geospatial index
### Apache Solr
Solr provides text search, faceting and geospatial queries and can act as a JanusGraph mixed index. Treat the index as derived and rebuildable, apply document-level authorization before results leave the service and test stale-index behavior rather than letting search become an accidental system of record.
Apache-2.0
Search
Derived index
- [Solr reference guide →](https://solr.apache.org/guide/solr/latest/)
raster processing + map clients
### GeoTrellis + CesiumJS + MapLibre
GeoTrellis handles tiled raster and terrain work on Scala/Spark; CesiumJS renders a 3D globe; MapLibre renders conventional 2D and 2.5D maps. These are processing and presentation components, not authoritative data. Keep source time, accuracy and attribution visible in the map and request only the viewport the operator may see.
Open source
2D + 3D
Geospatial
- [GeoTrellis →](https://geotrellis.io/)
- [CesiumJS →](https://cesium.com/platform/cesiumjs/)
- [MapLibre →](https://maplibre.org/)
operational UI + analytical dashboard
### LocStat + Apache Superset
LocStat is useful as a Palantir-style interface reference; verify code availability, support and licence before making it a dependency. Superset is an Apache-licensed analytical dashboard, not a transactional workflow engine. Borrow interaction patterns, but keep approvals and case state in a maintained application service.
Interface reference
BI
Verify licence
- [LocStat →](https://www.locstat.co.za/)
- [Superset documentation →](https://superset.apache.org/docs/intro)
operational data pipelines
### Apache Airflow + Prefect
Either can schedule, retry and observe batch ingestion and ontology projection. Choose one owner for each retry, make tasks idempotent and keep long-lived human cases in application state or a durable workflow engine. A scheduler succeeding does not prove that the resulting data is correct.
Python
Schedules
One retry owner
- [Airflow documentation →](https://airflow.apache.org/docs/)
- [Prefect documentation →](https://docs.prefect.io/)
data policy + application policy
### Apache Ranger + Open Policy Agent
Ranger centralizes policies and audit for supported data services; OPA evaluates policy as code in applications and infrastructure. Neither creates identities, secrets or correct object-level permissions automatically. Test denied rows, fields, actions and exports with the exact service versions in production.
Authorization
Audit
Test denies
- [Apache Ranger →](https://ranger.apache.org/)
- [OPA documentation →](https://www.openpolicyagent.org/docs/latest/)
private datacentre control plane
### vSphere Supervisor + Cloud Director CSE + vSphere CSI
Use ordinary vSphere VMs for a small stateful deployment unless the datacentre already operates Kubernetes. Supervisor and Cloud Director CSE add cluster or tenant control planes; the CSI driver maps persistent volumes to vSphere storage policies. Tenant boundaries do not replace application authorization, and VM snapshots do not replace application-consistent backup.
VMware
Private deployment
Recovery test
- [vSphere Supervisor →](https://techdocs.broadcom.com/us/en/vmware-cis/vsphere/vsphere-supervisor/8-0.html)
- [Cloud Director CSE →](https://github.com/vmware/container-service-extension)
- [vSphere CSI →](https://github.com/kubernetes-sigs/vsphere-csi-driver)
Foundry + Gotham UI references
### OpenFoundry + Osiris interface references
These repositories can make ontology, lineage, map, layer and event-card choices tangible. Treat them as workshop material until their licences, maintenance, identity model and data boundaries pass review. Reimplement only the operator journey your own services and acceptance tests require.
Prototype
Interface study
Not a platform proof
- [OpenFoundry →](https://github.com/cherishwins/OpenFoundry)
- [Osiris →](https://github.com/simplifaisoul/osiris)
- [Use the interface guide →](https://isaiuseful.com/diy-palantir.html.md#clones)
routing + constraint optimisation
### Google OR-Tools
Use established solvers for vehicle routing, assignment, scheduling and other constrained decisions before asking a language model to improvise. Define hard constraints separately from preferences, preserve the solver inputs and compare feasible solutions against the current operational baseline.
Apache-2.0
Optimisation
Deterministic baseline
- [OR-Tools documentation →](https://developers.google.com/optimization)
crowdsourced aircraft state vectors
### OpenSky Network
A research-oriented REST API for live aircraft state vectors, tracks and limited flight history. Coverage and identity are receiver-dependent, quotas apply and OpenSky now requires a written agreement for operational or commercial REST API use. Use licensed snapshots as realistic fixtures; use an approved aviation feed or your own authorized receivers when decisions depend on completeness.
Live + historical
Written licence for operations
Not authoritative
- [REST API →](https://openskynetwork.github.io/opensky-api/rest.html)
- [Terms and data licence →](https://opensky-network.org/about/terms-of-use)
crowdsourced aircraft + military filter
### ADS-B Exchange
The commercial API exposes changing aircraft positions and a military-filtered endpoint. It is useful as a supplementary or development feed, not proof of identity, intent or complete coverage. Budget for subscription and rate limits, preserve source timestamps and reconcile against the authority responsible for the air picture.
Commercial API
Live aircraft
Supplement only
- [API documentation →](https://gateway.adsbexchange.com/api/aircraft/v2/docs/index.html?url=%2Fapi%2Faircraft%2Fv2%2Fdocs%2Fopenapi.json)
- [Data products →](https://www.adsbexchange.com/data-products/)
satellite orbital elements
### CelesTrak GP data
CelesTrak publishes current general-perturbations element sets in TLE, OMM JSON, XML, KVN and CSV formats. Rendered paths are propagated estimates from an element epoch—not continuous live sensor positions. Pin dated fixtures for development and surface element age, propagation time and uncertainty in operational views.
Public data
Orbit propagation
Show data age
- [GP formats and queries →](https://celestrak.org/NORAD/documentation/gp-data-formats.php)
- [Current element sets →](https://celestrak.org/NORAD/elements/)
open road + feature data
### OpenStreetMap data
OpenStreetMap can supply roads and mapped features for routing, simulation and basemaps; it does not supply live vehicle flow. Attribute contributors and respect the ODbL. The community tile service is best-effort and forbids bulk or offline use, so production systems should self-host approved extracts and tiles or use a provider with a contract.
Open data
Attribution required
Self-host production tiles
- [Attribution and licence →](https://www.openstreetmap.org/copyright/attribution-guide/)
- [Community tile policy →](https://operations.osmfoundation.org/policies/tiles/)
public camera locations + snapshots
### City of Austin Traffic Cameras
A public-domain dataset with camera identifiers, approximate locations, status and—where publication is allowed—a link to the latest screenshot on a five-minute cadence. It is useful for a public-feed prototype and outage fixtures. The city says footage is not retained and the locations are not suitable for legal, engineering or surveying use.
Public domain
Five-minute screenshots
Approximate locations
- [Official dataset →](https://data.austintexas.gov/Transportation-and-Mobility/Traffic-Cameras/b4k4-adkb)
local NVR + object detection
### Frigate
An MIT-licensed, self-hosted NVR that performs real-time object detection locally for IP cameras, retains recordings according to detected objects and integrates with Home Assistant and MQTT. Use a supported GPU or AI accelerator as the project recommends, and keep its unauthenticated internal API off exposed networks; use the authenticated endpoint or a carefully configured reverse proxy. “Local” still leaves sensitive footage, camera credentials and retention policy for you to secure.
MIT
Local object detection
Accelerator recommended
- [Frigate NVR →](https://frigate.video/)
- [Source, releases and licence →](https://github.com/blakeblackshear/frigate)
- [Hardware guidance →](https://docs.frigate.video/frigate/hardware)
- [Authentication boundary →](https://docs.frigate.video/configuration/authentication/)
hosted photorealistic city mesh
### Google Photorealistic 3D Tiles
A high-resolution textured 3D basemap that can be rendered with CesiumJS. It requires billing, an API key and visible attribution. Google's policies restrict prefetching, storage, offline use, extraction and machine analysis, so keep every operational layer independent and retain a fallback to government terrain, imagery or 3D city models.
Hosted API
Commercial terms
Basemap only
- [3D Tiles documentation →](https://developers.google.com/maps/documentation/tile/3d-tiles)
- [Map Tiles policies →](https://developers.google.com/maps/documentation/tile/policies)
owned government-grade GIS backbone
### PostGIS + GeoServer + QGIS
PostGIS stores governed spatial records in PostgreSQL, GeoServer publishes standards-based services and QGIS supports desktop editing, analysis and review. This is the airtight fallback for an organisation that already owns authoritative GIS data: keep identity and workflow state in maintained services, issue read-only layers to the browser and test disconnected operation and restore.
Self-hosted
OGC services
Authoritative data
- [PostGIS documentation →](https://postgis.net/documentation/)
- [GeoServer documentation →](https://docs.geoserver.org/)
- [QGIS documentation →](https://docs.qgis.org/latest/en/docs/)
### Build software with agents
deterministic CLI output compression
### RTK
An Apache-2.0 Rust proxy that rewrites supported shell commands and filters, groups, deduplicates or truncates their output before it reaches a coding agent. The project estimates 60–90% token savings on common commands; measure that on your workload. Retain a route to the raw command output because compression can hide the line that matters during diagnosis.
Rust binary
Agent hooks
Keep raw output
- [RTK repository →](https://github.com/rtk-ai/rtk)
PRD-driven coding loop
### Ralph
An MIT-licensed shell loop that turns a PRD into prioritized stories, launches a fresh supported coding-agent instance per story, runs checks and commits passing work until a limit is reached. It is a useful scaffold, not proof that unattended agents finish correctly, and the selected agent keeps its own model and data boundary. Use a disposable branch, cap iterations and review every commit.
External coding agent
Fresh contexts
Bound the loop
- [Ralph repository →](https://github.com/snarktank/ralph)
response-style compression skill
### Caveman
An MIT-licensed skill that asks coding agents to answer in terse “caveman” language while preserving code and commands. Its 65% output-token claim comes from the project and does not reduce input context. Test whether terseness drops rationale, warnings or accessibility for your team; avoid it where a complete handoff matters more than output cost.
Many agents
Output only
Vendor benchmark
- [Caveman repository →](https://github.com/juliusbrussee/caveman)
open-source coding CLI + local providers
### Codex CLI
OpenAI's open-source terminal coding agent can run against Ollama or LM Studio with `--oss` ; `--local-provider` selects the provider and `oss_provider` sets a default. Local inference is not a containment boundary: the agent can still read, edit and run commands. Keep sandboxing and approvals enabled, scope files and credentials, choose a model that handles tools reliably and verify the result with tests and review.
Open source
Ollama + LM Studio
Sandbox + approvals
- [Learn the coding-agent operating method →](https://isaiuseful.com/thinking-with-ai.html.md#tools)
- [Codex CLI repository →](https://github.com/openai/codex)
- [Local-provider configuration →](https://learn.chatgpt.com/docs/config-file/config-advanced#oss-mode-local-providers)
community external-model router for Codex
### Codex Router
An independent MIT-licensed beta that adapts selected external providers to the Responses API and merges their models into the Codex App and CLI picker through a loopback service. It writes marked Codex configuration blocks, installs a per-user background service and stores provider credentials in protected local files. Review a tagged installer instead of piping the moving default branch, keep its config backup and rollback path, run the included doctor and re-test after Codex updates. Provider billing, data handling and terms still apply, and routed models do not become OpenAI-supported Codex models.
MIT
Beta
Config + credential boundary
- [Codex Router repository →](https://github.com/duolahypercho/codex-router)
- [Security model →](https://github.com/duolahypercho/codex-router/blob/main/SECURITY.md)
- [Pinned releases and checksums →](https://github.com/duolahypercho/codex-router/releases)
multi-provider coding agent + desktop
### Code Buddy
An MIT-licensed coding agent with a terminal UI, Electron-based Cowork desktop app, HTTP/WebSocket server, local Ollama support, hosted model providers, MCP and a broad tool surface. Local inference does not contain file, shell, browser or connector access: begin with narrow permissions, isolate credentials and repositories, keep optional background and self-improvement features off until tested, and review generated changes and external actions.
MIT
CLI + desktop
Local + hosted models
- [Code Buddy repository →](https://github.com/phuetz/code-buddy)
Rust coding harness + agent memory
### JCode
An MIT-licensed terminal coding agent with hosted and local model routes, MCP, semantic memory, persistent background work and coordinated swarms. The project publishes its own optimization benchmark plus Linux PSS and startup comparisons; treat those as project-run evidence and reproduce them with pinned versions. Scope files, shell, network and credentials, and set retention rules so recalled memory does not preserve stale decisions or sensitive material.
Rust
Local + hosted models
Memory + swarms
- [JCode repository →](https://github.com/1jehuang/jcode)
- [JCode Bench design →](https://jcode.sh/bench)
- [Compare published memory usage →](#harness-memory)
AI commit messages
### OpenCommit
An MIT-licensed CLI that generates commit messages from staged changes through a hosted model or local Ollama. A remote provider may receive the diff, and the default flow can stage files for you. Stage deliberately, exclude secrets, inspect the final diff and edit the message so it describes the actual change rather than the model's guess.
Git CLI
Hosted or local model
Review before commit
- [OpenCommit repository →](https://github.com/di-sukharev/opencommit)
agent-native version control
### GitButler
A Git-based desktop and CLI workflow for organizing coding-agent changes into parallel or stacked branches, selected hunks and reviewable commits in one working directory. Its agent setup can install version-control instructions, but those instructions are not access controls; keep repository permissions and release approval separate. Workspace mode also changes branch semantics, so follow GitButler's commands instead of mixing in Git index, checkout or branch operations. The source uses a Fair Source licence that converts to MIT after two years and restricts competing products during that period.
Desktop + CLI
Parallel branches
Fair Source
- [GitButler →](https://gitbutler.com)
- [Agent workflow documentation →](https://docs.gitbutler.com/ai-agents/overview)
- [Source and licence →](https://github.com/gitbutlerapp/gitbutler)
human + agent collaboration workspace
### Buzz
Block's Apache-2.0 workspace puts people and model-agnostic agents—including Codex, Claude Code and Goose harnesses—into shared channels around conversation, workflows and code. Events live on a Nostr relay that can be hosted by Block or self-hosted. Treat it as a collaboration and audit layer, not a sandbox: each underlying agent still needs narrowly scoped repositories, tools, network access and credentials. Buzz is pre-1.0, so check the current release and the project's works-now / in-progress table before standardizing a team workflow.
Apache-2.0
Hosted + self-hosted
Human + multi-agent
- [Try Buzz →](https://buzz.xyz)
- [Source, releases and current status →](https://github.com/block/buzz)
- [Block's introduction →](https://block.xyz/inside/introducing-buzz-where-humans-and-agents-work-together)
multi-agent terminal manager
### Claude Squad
An AGPL-3.0 terminal interface for running Claude Code, Codex, Gemini, Aider and other agents in separate Git worktrees. Worktrees reduce merge collisions but are not process, network or credential sandboxes; its auto-accept mode further expands authority. Review changes before applying them and isolate secrets and external accounts separately.
TUI
Git worktrees
AGPL-3.0
- [Claude Squad repository →](https://github.com/smtg-ai/claude-squad)
graphical worktree launcher
### ParallelCode
A free desktop Git worktree manager for opening Cursor, Claude Code, Copilot and other local-repository agents on separate branches. It is the focused graphical choice when parallel worktrees—not a full task platform—are the requirement. Worktrees separate files and branches, not processes, ports, networks or credentials, so coordinate shared services and review every merge.
Desktop GUI
Git worktrees
Free
- [ParallelCode overview and downloads →](https://parallelcode.dev/)
- [Worktree workflow documentation →](https://parallelcode.dev/docs)
Windows terminal workspace
### Windows Terminal
Microsoft's MIT-licensed terminal host puts PowerShell, Command Prompt, WSL and SSH profiles into configurable tabs and split panes. It is a clean baseline for watching several Windows-side agent CLIs, but it does not add agent status, task routing, worktrees or review workflow. Use scripted `wt` layouts for repeatability and pair it with Git or a session manager when branches and resumable state matter.
Windows
Tabs + panes
MIT
- [Windows Terminal guide →](https://learn.microsoft.com/en-us/windows/terminal/)
- [Source and releases →](https://github.com/microsoft/terminal)
visual multi-agent workspace
### Nimbalyst
A free, MIT-licensed desktop workspace for parallel Codex and Claude Code sessions, Kanban task management, worktrees, terminals and visual editing of documents, diagrams, mockups and code. It is the strongest fit here when non-terminal review and a mobile companion matter. Core content stays in local files, but anonymous analytics are enabled unless you opt out and mobile collaboration uses a separately operated sync service; verify both boundaries before using sensitive repositories.
Visual workspace
Codex + Claude Code
MIT
- [Parallel-session tour →](https://nimbalyst.com/features/session-management/)
- [Source, license and privacy notes →](https://github.com/nimbalyst/nimbalyst)
cross-platform agent development environment
### Orca
An MIT-licensed desktop workspace for running Codex, Claude Code, OpenCode and other CLI agents side by side in Git worktrees, with terminals, editing, per-worktree browsers, diff review and SSH. **Mobile role:** its beta iPhone and Android companion is a phone remote control for desktop-owned sessions: watch status and terminal output, browse files, reply or dictate, review changes and reconnect over LAN or Tailscale. Orca's current docs say supported agents launch with their full-autonomy flags pre-applied by default. Worktrees separate branches, not process, network or credential authority, so add real containment or change that launch policy before treating a pocket approval as a safety boundary. Packaged builds also send opt-out anonymous usage telemetry to PostHog's US region.
Phone companion beta
Desktop source of truth
MIT
- [Mobile companion and Tailscale path →](https://www.onorca.dev/docs/mobile)
- [Agent launch flags and authority →](https://www.onorca.dev/docs/agents/supported)
- [Source, releases and licence →](https://github.com/stablyai/orca)
- [Privacy and telemetry boundary →](https://www.onorca.dev/docs/telemetry)
multi-provider agent workbench
### Kandev
An AGPL-3.0, self-hostable control plane with Kanban and pipeline views, worktrees, an integrated editor, terminal and review flow. Its broad agent registry includes Claude Code, Codex, Copilot, Gemini, OpenCode, Kimi and Grok, while executors can run locally, in Docker, over SSH or in a cloud environment. That breadth brings more operational surface: pin agent adapters, scope executor secrets and treat container or remote-host policy as a separate security boundary.
Many agent vendors
Local + remote execution
AGPL-3.0
- [Supported agents and profiles →](https://kandev.ai/docs/agents-and-profiles)
- [Architecture, executors and license →](https://github.com/kdlbs/kandev)
terminal-native agent dashboard
### Agent Deck
An MIT-licensed TUI that adds agent-aware running, waiting and completed states to tmux, plus groups, search, session forking, Git worktrees, cost tracking and local or remote sessions. It is the lightweight terminal-native choice and installs through Homebrew, Go or a release script; Windows use runs through WSL. Session organization is not command isolation, so keep each underlying agent's approvals, sandbox and credential scope intact.
TUI + tmux
Homebrew
MIT
- [Install Agent Deck →](https://github.com/asheshgoplani/agent-deck#installation)
- [Features, platforms and caveats →](https://github.com/asheshgoplani/agent-deck#faq)
multi-vendor terminal session manager
### CCManager
An MIT-licensed terminal manager for parallel Claude Code, Codex, Gemini, Cursor Agent, Copilot, Cline, OpenCode and Kimi CLI sessions across projects and Git worktrees. Busy, waiting and idle indicators make it a straightforward choice when status visibility and broad CLI support matter more than a graphical workspace. Its worktree merge and delete actions change Git state, and experimental auto-approval expands agent authority; keep both deliberate and recoverable.
TUI
Kimi + major CLIs
MIT
- [CCManager repository and setup →](https://github.com/kbwo/ccmanager)
- [Supported agent commands →](https://github.com/kbwo/ccmanager#supported-ai-assistants)
persistent coding-agent terminals
### herdr
An Apache-2.0 Rust terminal runtime that keeps coding-agent panes alive in a background server, marks them working, blocked or idle and supports detach, reattach, SSH access and an agent-facing socket API. It runs existing tools such as Codex, Claude Code, Cursor and OpenCode rather than replacing them. Persistence and status tracking are not containment: each pane still inherits the filesystem, network and credentials available to its process, so keep agent approvals and sandboxing in place.
Rust binary
Persistent sessions
Apache-2.0
- [herdr repository →](https://github.com/herdrdev/herdr)
agent context compression layer
### Headroom
An Apache-2.0 local-first library, proxy and MCP server that compresses tool output, files, logs and retrieval chunks before model use. The project reports 15–20% savings for coding agents and 60–95% for JSON, with reversible paths. Benchmark answer quality as well as token count, keep raw artifacts and bypass compression for forensic or exact-text work.
Library + proxy + MCP
Local-first
Keep originals
- [Headroom repository →](https://github.com/headroomlabs-ai/headroom)
self-hosted coding-agent control center
### OpenHands Agent Canvas
A beta, MIT-licensed interface for running OpenHands, Claude Code, Codex, Gemini and ACP-compatible agents across local, remote or cloud backends. Local UI does not mean local inference or isolated execution: inspect the selected backend, mounts, network, provider credentials and automation triggers before leaving it always on.
Beta
Many agent backends
Local + remote
- [Agent Canvas repository →](https://github.com/OpenHands/agent-canvas)
mainstream IDE + BYOK
### Visual Studio Code
VS Code's built-in language-model picker can connect chat and local-agent workflows to Ollama or a custom OpenAI-compatible endpoint without a Copilot plan or GitHub sign-in. Native BYOK does not provide standard inline completions, semantic search or embedding-backed features. On DGX Spark, use NVIDIA's guide for a direct ARM64 installation or remote access through NVIDIA Sync.
Local or remote
BYOK
Chat + agents
- [VS Code BYOK guide →](https://code.visualstudio.com/docs/agent-customization/language-models)
- [VS Code on DGX Spark →](https://build.nvidia.com/spark/vscode/overview)
VS Code + JetBrains extensions
### Continue
An open-source assistant for chat, edits, agents and role-specific autocomplete. It can discover Ollama models locally or use a remote Ollama address, which fits a Spark reached through a private network or SSH tunnel. Pick models by role: a good chat model is not automatically a good fill-in-the-middle completion model.
Local-capable
IDE extension
Role-specific models
- [Ollama setup →](https://docs.continue.dev/guides/ollama-guide)
- [Autocomplete model guide →](https://docs.continue.dev/customize/model-roles/autocomplete)
VS Code coding agent
### Roo Code
A model-flexible coding agent that can use Ollama, LM Studio or another OpenAI-compatible endpoint. Local operation still depends on a model with reliable tool calling and enough context; start with approvals enabled and test file, terminal and browser boundaries before increasing autonomy.
Local-capable
Agentic
Approval controls
- [Ollama integration →](https://docs.ollama.com/integrations/roo-code)
- [OpenAI-compatible setup →](https://roocodeinc.github.io/Roo-Code/providers/openai-compatible/)
JetBrains IDEs + BYOK
### JetBrains AI Assistant
IntelliJ IDEA, Rider and the other supported JetBrains IDEs can connect AI Assistant to Ollama, LM Studio or a private OpenAI-compatible endpoint. Local models can cover chat and assigned editor tasks, but local-model MCP tools, next-edit suggestions and some proprietary completion features remain unavailable.
Local or remote
IDE-native
Feature caveats
- [Custom and local models →](https://www.jetbrains.com/help/ai-assistant/use-custom-models.html)
local-first editor
### Zed
Zed can connect its Agent and Inline Assistant directly to Ollama, LM Studio, llama.cpp or a local OpenAI-compatible server. Remote endpoint URLs are supported, while local or self-hosted edit prediction has a separate configuration path.
Native editor
Local endpoints
Agent + inline
- [Local-model guide →](https://zed.dev/docs/ai/use-a-local-model)
coding + specification
### OpenCode + OpenSpec
A provider-flexible coding agent paired with a durable, reviewable agreement about what to build and what stays out of scope.
Terminal
Provider-flexible
Spec-driven
- [OpenCode →](https://opencode.ai/docs/)
- [OpenSpec →](https://github.com/Fission-AI/openspec)
coding
### Aider
A terminal coding assistant built around repository edits and version control. Use a clean branch and keep tests as the acceptance boundary.
Terminal
Git-aware
Multiple providers
- [Official documentation →](https://aider.chat/)
IDE + CLI agent
### Cline
An open-source coding agent for IDEs and the terminal with reviewable diffs and human-in-the-loop approval. Keep its repository scope narrow and review every command or enablement choice.
IDE + CLI
Model-flexible
Reviewable diffs
- [Project repository →](https://github.com/cline/cline)
agent harness + coding CLI
### Pi Agent Harness
The MIT-licensed `earendil-works/pi` project packages a multi-provider LLM API, tool-calling agent core, terminal UI and extensible coding-agent CLI. Pi explicitly does not provide a built-in permission system, so run it with only the filesystem, process, network and credentials it should have—prefer a tested container, micro-VM or policy sandbox for consequential work.
TypeScript
Multi-provider
External sandbox required
- [Pi repository →](https://github.com/earendil-works/pi)
- [Pi documentation →](https://pi.dev/)
persistent self-improving coding agent
### Prime Agent
An MIT-licensed coding and research agent built around a persistent Python environment, programmatic subagents and durable harness state that can retain reviewed memories, prompts, skills and subagent specifications. Background sessions, goals, schedules and bounded autonomous runs suit long work. It executes model-generated Python and project commands with the user’s permissions—not inside a security sandbox—so use a disposable worktree or external containment and inspect every retained refinement.
MIT
Terminal + background daemon
External sandbox required
- [Prime Agent repository →](https://github.com/PrimeIntellect-ai/prime-agent)
community local-first coding agent
### Nanocoder
An MIT-licensed TypeScript terminal agent with Ollama and other local-server routes, cloud providers, MCP, skills, subagents, checkpoints and scheduled or event-driven runs. “Local-first” describes where the selected model can run, not a permission boundary: file and shell tools still act with the process’s authority. Begin in normal or plan mode, isolate the repository and credentials, and measure its exploration cost on your own tasks.
Terminal + VS Code
Local + cloud models
Scope tool authority
- [Nanocoder repository →](https://github.com/Nano-Collective/nanocoder)
- [Nanocoder documentation →](https://docs.nanocollective.org/nanocoder/)
- [See the repeated-run harness benchmark →](https://isaiuseful.com/benchmarks.html.md#harness-efficiency)
native local editor
### LibreCode
A .NET 10 and Avalonia code editor with Ollama integration, terminal and model browsing for Windows, Linux and macOS. Its repository is public and local-first, but the custom license restricts redistribution, forks and SaaS use—treat it as source-available, not permissively open source.
.NET + Avalonia
Ollama
Custom license
- [Repository →](https://github.com/re4/LibreCode)
- [License →](https://github.com/re4/LibreCode/blob/main/LICENSE.md)
versioned code documentation
### Context7
Feeds current, version-specific library documentation into coding agents through a CLI skill or MCP server. Pin the library ID and version, and still open the primary documentation for consequential changes; a retrieved snippet is context, not a compatibility test. Hosted use sends the documentation query to Context7.
CLI or MCP
Current docs
Coding agents
- [Documentation →](https://context7.com/docs)
- [Repository →](https://github.com/upstash/context7)
version-matched package source
### opensrc
An Apache-2.0 Rust CLI that gives coding agents searchable source for npm, PyPI and crates.io packages plus GitHub, GitLab and Bitbucket repositories. It resolves an installed or requested version, shallow-clones the matching source and caches it locally so ordinary tools such as `rg` can inspect implementation, tests and examples. Pin the dependency version and verify the resolved tag: code is the strongest evidence of implementation behavior, while official release, support and security guidance still define the supported contract.
Apache-2.0
Rust CLI
Package + repository source
- [opensrc documentation →](https://opensrc.sh/)
- [Repository and source →](https://github.com/vercel-labs/opensrc)
### Fine-tune, evaluate and reproduce
model-file architecture inspector
### Netron
An MIT-licensed desktop, browser and Python viewer for ONNX, TensorFlow, PyTorch, Core ML, OpenVINO, Safetensors and many other model formats. Use it to inspect graph structure, operators, tensor shapes and metadata before conversion or deployment. A readable graph is not proof that a checkpoint is safe, correctly licensed or numerically equivalent after export; keep provenance, hash files and run format-specific validation.
MIT
Desktop + browser + Python
Inspect, then validate
- [Netron source and downloads →](https://github.com/lutzroeder/netron)
- [Open the browser viewer →](https://netron.app/)
agent diagnostic suite
### iFixAI
An Apache-2.0 CLI and agent skill that runs 45 inspections, with 32 core checks contributing to an A–F score and optional independent model judges. The project explicitly says the result is not certification or a safety guarantee. Version fixtures and judges, inspect false positives and unsupported capabilities, budget judge calls and disable disclosed telemetry if policy requires it.
Agent evals
CI-friendly
Not certification
- [iFixAI repository →](https://github.com/ifixai-ai/iFixAI)
training libraries
### Hugging Face PEFT + TRL
PEFT provides adapters such as LoRA; TRL provides supervised and preference-training loops. This is the flexible foundation when your team wants code-level control over datasets, trainers and evaluation.
LoRA
SFT + DPO
Code-first
- [PEFT →](https://huggingface.co/docs/peft/index)
- [TRL →](https://huggingface.co/docs/trl/index)
configuration-driven training recipes
### Axolotl
Packages full, LoRA, QLoRA, preference and reinforcement-learning runs into reusable YAML configuration, with distributed and multi-GPU routes. Pin the repository revision, base model, dataset, chat template and dependency environment; a portable recipe is not a guarantee that another GPU topology will reproduce the same speed or result.
Config-driven
Distributed training
Adapters + full tuning
- [Axolotl documentation →](https://docs.axolotl.ai/)
beta desktop AI studio + training library
### Unsloth
Unsloth Desktop is a beta, cross-platform app for running and training local text, audio, image and video models, with RAG, MCP, code execution, export and OpenAI- and Anthropic-compatible endpoints that can connect Codex or Claude Code. Its Core package is Apache-2.0; the Studio UI is AGPL-3.0. Unsloth reports up to 2× faster supported LLM training with up to 70% less VRAM, but treat that as vendor evidence and reproduce it on the exact model and hardware. Support differs across CPU, NVIDIA, AMD, Intel and Apple backends. Keep it on loopback first: web search, cloud providers, MCP servers and Cloudflare remote access each change the data or authority boundary, and anyone with the tunnel link and API key can run code.
Desktop + Core
Local multimodal + training
Beta · dual license
- [Desktop downloads →](https://unsloth.ai/)
- [Desktop launch and claims →](https://github.com/unslothai/unsloth/releases/tag/v0.1.70-beta)
- [Source, requirements and licenses →](https://github.com/unslothai/unsloth)
agent optimization + training
### Microsoft Agent Lightning
Agent Lightning is an MIT-licensed training layer that turns agent trajectories into inputs for reinforcement learning, supervised fine-tuning or prompt optimization without binding the agent to one framework. The trainer, inference engine and compute remain your responsibility: version rewards and learned artifacts, retain untouched holdouts and reject gains that disappear under repeat runs.
Self-deployed
Framework-agnostic
RL + SFT
- [Agent Lightning repository →](https://github.com/microsoft/agent-lightning)
experiment evidence
### MLflow
Track run configuration, artifacts and metrics, and compare an adapter with its baseline and holdout. Keep human task judgments alongside automated scores; a model judging another model is not independent proof.
Run tracking
Artifacts
Evaluation
- [GenAI evaluation docs →](https://mlflow.org/docs/latest/genai/eval-monitor/)
open-source tracing + evaluation
### Arize Phoenix + W&B Weave
Phoenix and Weave trace model and agent calls, organize datasets and run evaluations. Instrumentation can capture prompts, retrieved documents and tool arguments, so define redaction, sampling, access and retention before sending production traces anywhere.
Tracing
Datasets
Evals
- [Arize Phoenix →](https://arize.com/phoenix/)
- [Weights & Biases Weave →](https://docs.wandb.ai/weave)
self-hosted tracing + local test suites
### Langfuse + DeepEval
Langfuse self-hosts tracing, prompt management, datasets and evaluations, but production deployment brings Postgres, ClickHouse, Redis or Valkey and object storage. DeepEval is an Apache-2.0, pytest-style evaluation framework that can run in local development and CI. Local tooling does not guarantee local judging: configure the evaluator model deliberately and redact sensitive trace fields before capture.
Self-hosted
Tracing + evals
CI tests
- [Langfuse self-hosting →](https://langfuse.com/self-hosting)
- [DeepEval repository →](https://github.com/confident-ai/deepeval)
RAG and agent evaluation
### TruLens + Ragas
TruLens combines tracing with feedback functions for agent applications; Ragas provides evaluation workflows and metrics for RAG and other LLM systems. Calibrate automated scores against human judgments and keep a stable, versioned test set—metric names are not evidence by themselves. [Use the RAG evaluation plan](https://isaiuseful.com/rag.html.md#evaluate) .
RAG evals
Agent evals
Human calibration
- [TruLens →](https://www.trulens.org/)
- [Ragas →](https://docs.ragas.io/)
test suites + request observability
### Promptfoo + Helicone
Promptfoo runs repeatable model comparisons, assertions and adversarial tests; Helicone captures and analyzes LLM requests through an observability layer. Keep golden cases in version control, block releases on meaningful thresholds and avoid treating traffic dashboards as correctness tests.
Regression tests
Red teaming
Request traces
- [Promptfoo →](https://www.promptfoo.dev/)
- [Helicone →](https://docs.helicone.ai/)
authorized generative-AI red teaming
### garak + PyRIT
Apache-2.0 garak probes models for prompt injection, jailbreaks, data leakage, misinformation, toxicity and other failure modes; MIT-licensed PyRIT orchestrates multi-turn attacks, converters, scorers and result memory for structured security exercises. Run only against systems you own or are authorized to test, cap cost and concurrency, isolate credentials, and handle prompts, responses and JSONL logs as potentially harmful security fixtures. Community attack corpora such as L1B3RT4S can broaden coverage, but they are uncurated inputs—not a safety standard or a pass/fail oracle.
Attack automation
Many model routes
Authorized targets only
- [garak repository →](https://github.com/NVIDIA/garak)
- [PyRIT repository →](https://github.com/microsoft/PyRIT)
- [L1B3RT4S attack corpus →](https://github.com/elder-plinius/L1B3RT4S)
- [Choose a jailbreak benchmark →](https://isaiuseful.com/benchmarks.html.md#jailbreakbench-harmbench)
open-weight refusal-direction research
### Heretic
An AGPL-3.0 command-line tool that automatically finds and suppresses refusal directions in supported open-weight transformer models, then offers comparison and evaluation routes. That makes it useful for interpretability and controlled robustness research—and capable of deliberately removing safety behavior. Use only in an isolated, authorized lab; preserve the original checkpoint, block deployment and sharing by default, scan generated artifacts, and compare helpfulness, over-refusal and harmful compliance rather than celebrating a lower refusal count.
Directional ablation
Dual-use
Lab isolation
- [Heretic repository →](https://github.com/p-e-w/heretic)
- [Measure attacks and defenses →](https://isaiuseful.com/benchmarks.html.md#jailbreakbench-harmbench)
prompt-injection attack + defense labs
### Tensor Trust + Gandalf
Tensor Trust is an open research game where participants write defenses and attack other prompt-protected accounts; Gandalf is a guided prompt-injection challenge that becomes harder across levels. They make failure modes tangible for awareness training, but neither simulates tool permissions, retrieval poisoning or a production incident. Tensor Trust says submitted text will be released for research, so never enter secrets, personal data or proprietary prompts.
Hands-on learning
Attack + defense
Public submissions
- [Play Tensor Trust →](https://tensortrust.ai/)
- [Research data and method →](https://tensortrust.ai/paper/)
- [Try Gandalf →](https://gandalf.lakera.ai/)
### Run durable assistants and custom systems
company agent workspace + app sandbox
### Cloudflare OS
An Apache-2.0 early-access workspace for agent chat and AI-built, shareable “Gadgets” on Cloudflare Workers. Capability-scoped Gatekeepers mediate external resources, log actions and queue side-effect approvals. The quick local route is explicitly non-production, while documented deployment on your own `workerd` server is still forthcoming; audit each OAuth connector, sandbox boundary, simulated approval result and data-residency path before organizational use.
Apache-2.0
Workers + workerd
Early access
- [Cloudflare OS repository →](https://github.com/cloudflare/cloudflare-os)
- [Deploy to Cloudflare →](https://os.cloudflare.app/deploy)
embedded agent isolation + orchestration
### Rivet agentOS
An Apache-2.0 runtime that executes supported coding agents inside lightweight isolated Linux environments with deny-by-default filesystem, network and process permissions. Its cold-start and cost comparisons are project benchmarks. Test required syscalls, process behavior and deny paths on your workload; a lightweight VM boundary still needs patched hosts, secrets brokering and resource limits.
Linux environment
Deny by default
Verify compatibility
- [agentOS repository →](https://github.com/rivet-dev/agentos)
local agent memory
### MEMANTO
An MIT-licensed local memory service with remember, recall and answer operations for several coding-agent clients, without a hosted backend or vector database. Persistent memory can preserve mistakes, secrets and stale decisions as easily as useful context. Scope what is ingested, test retrieval on a known set and keep export, deletion, backup and retention controls visible.
Local
Persistent memory
No vector database
- [MEMANTO repository →](https://github.com/moorcheh-ai/memanto)
desktop + multi-device automation
### Microsoft UFO
UFO is an MIT-licensed Microsoft research project for Windows desktop automation and cross-device orchestration. The code runs on infrastructure you control, while model endpoints are configured separately. GUI control carries broad authority: begin with non-critical accounts, isolate credentials and files, require approval for consequential actions and measure recovery from focus changes and partial execution.
Self-deployed
Windows + devices
Research project
- [Microsoft UFO repository →](https://github.com/microsoft/UFO)
general-purpose local agent
### Goose
An Apache-2.0 desktop app, CLI and API for coding, research and automation across many hosted and local model providers, with MCP extensions, reusable recipes and ACP support. It includes tool permissions, prompt-injection checks and optional adversary review, but those controls do not make arbitrary extensions or model output trustworthy. Begin with a narrow directory, an extension allowlist and approval-required tools; inspect provider, subscription and telemetry paths separately.
Desktop + CLI + API
MCP + ACP
Permission controls
- [Goose documentation →](https://goose-docs.ai/)
- [Source and Apache-2.0 license →](https://github.com/aaif-goose/goose)
- [Security guidance →](https://goose-docs.ai/docs/guides/security/)
general agent
### Hermes Agent
A persistent assistant that can learn reusable skills and run on Linux, macOS, Windows or a remote server. **Mobile role:** talk to the same home- or cloud-hosted agent through Telegram, Discord, Slack, WhatsApp or another configured channel. The messaging provider remains part of the communications path.
Phone messaging
Remote-first
Skills
- [Messaging gateway →](https://hermes-agent.nousresearch.com/docs/user-guide/messaging/)
- [Official repository →](https://github.com/NousResearch/hermes-agent)
connected assistant
### OpenClaw + companion apps
OpenClaw keeps one gateway and state store on the always-on host, while Windows, macOS, Linux, mobile and messaging clients act as operators or narrowly scoped nodes. **Mobile role:** its iPhone and Android companions connect to that gateway; the phone does not host the gateway or silently reduce its tool authority.
iPhone + Android
Always-on host
Remote gateway
- [Remote gateway →](https://docs.openclaw.ai/gateway/remote)
- [iPhone app →](https://docs.openclaw.ai/platforms/ios)
- [Android app →](https://docs.openclaw.ai/platforms/android)
local multi-agent manager
### TripleBits Apprentice
A desktop route for creating scheduled agents with chosen models, memory, budgets, channels and tools. TripleBits says Apprentice runs agents in isolated local Docker containers, offers command and website controls and keeps run logs locally; treat those as vendor claims and test filesystem mounts, network denial, credential brokering and budget stops before sensitive work.
Local-first
Schedules + channels
Verify controls
- [TripleBits Apprentice →](https://triplebits.com/)
agent containment
### NVIDIA OpenShell
A policy and sandbox runtime for autonomous agents on Linux, macOS and Windows through WSL 2. Use it to bound files, networking and execution around an agent or tool runner; still verify the selected backend and test the deny path instead of treating “sandboxed” as a blanket guarantee.
Policy
Sandbox
Cross-platform
- [Official repository →](https://github.com/NVIDIA/OpenShell)
Python + TypeScript orchestration
### LangChain + LangGraph
LangChain supplies model and tool integrations; LangGraph adds explicit stateful agent graphs, checkpointing and human-in-the-loop control. They orchestrate agent reasoning—not telecom media, distributed transactions or durable business state by themselves.
Code-first
Agent state
Python + TypeScript
- [LangChain →](https://github.com/langchain-ai/langchain)
- [LangGraph →](https://github.com/langchain-ai/langgraph)
Microsoft agent orchestration + governance
### Microsoft Agent Framework + Governance Toolkit
Use Microsoft Agent Framework for new .NET or Python agents and graph-based workflows; Microsoft describes it as the successor to AutoGen and Semantic Kernel, and its current orchestration patterns cover the Magentic route without a separate legacy Magentic-One deployment. Pair it with the framework-agnostic Agent Governance Toolkit for policy enforcement, agent identity, sandboxing and tamper-evident audit records. Pin versions and test deny paths before production.
.NET + Python
Workflows
Governance controls
- [Agent Framework →](https://github.com/microsoft/agent-framework)
- [AutoGen migration →](https://learn.microsoft.com/en-us/agent-framework/migration-guide/from-autogen/)
- [Semantic Kernel migration →](https://learn.microsoft.com/en-us/agent-framework/migration-guide/from-semantic-kernel/)
- [Governance Toolkit →](https://github.com/microsoft/agent-governance-toolkit)
### Operate the stack
telemetry + dashboards + logs
### OpenTelemetry + Prometheus + Grafana + Loki
Trace one workflow across service boundaries, record bounded metrics, build dashboards and aggregate logs. Avoid putting transcripts, phone numbers, prompts or appointment details into metric labels and unbounded log streams.
Tracing
Metrics
Logs + alerting
- [OpenTelemetry →](https://opentelemetry.io/docs/)
- [Prometheus →](https://prometheus.io/docs/introduction/overview/)
- [Grafana →](https://grafana.com/docs/grafana/latest/)
- [Loki →](https://grafana.com/docs/loki/latest/)
identity + secrets
### Keycloak + OpenBao
Use Keycloak for operator identity and roles; use OpenBao for SIP, calendar and database credentials with short leases or rotation where supported. Neither replaces host hardening or application-level authorization.
OIDC
Secrets
Rotation
- [Keycloak →](https://www.keycloak.org/documentation)
- [OpenBao →](https://openbao.org/docs/)
data governance + compliance
### Microsoft Purview
Microsoft's portfolio spans data cataloging and governance, information protection, data loss prevention, audit, lifecycle, eDiscovery and related compliance controls. In an AI teammate, use Purview to classify and govern data around the workflow; it does not replace per-tool authorization, short-lived credentials or runtime containment.
Microsoft ecosystem
Data controls
Audit + compliance
- [Purview overview →](https://learn.microsoft.com/en-us/purview/)
- [Purview developer platform →](https://learn.microsoft.com/en-us/purview/developer/)
programmable input + output rails
### NVIDIA NeMo Guardrails + Guardrails AI
NeMo Guardrails adds configurable input, dialog, retrieval, execution and output rails around LLM applications. Guardrails AI focuses on validators and structured, checked outputs. Use either as explicit policy code with tested allow and deny cases—the framework name is not a safety guarantee.
Policy checks
Validation
Test deny paths
- [NVIDIA NeMo Guardrails →](https://docs.nvidia.com/nemo/guardrails/about-nemo-guardrails-library/overview)
- [Guardrails AI →](https://guardrailsai.com/guardrails/docs)
PII detection + de-identification
### Microsoft Presidio
Presidio detects and anonymizes personally identifiable information across text, images and structured data. Detection is probabilistic and domain-dependent: add organization-specific recognizers, measure false negatives and keep access control and minimization in place even after redaction.
PII
Self-hostable
Custom recognizers
- [Microsoft Presidio →](https://data-privacy-stack.github.io/presidio/getting_started/)
AI application + supply-chain security
### Prompt Security + Protect AI
Prompt Security focuses on discovering and protecting generative-AI use and applications. Protect AI covers AI security across models, artifacts and the software supply chain. Evaluate the exact product or open-source scanner you intend to deploy; neither replaces least privilege, sandboxing or incident response.
Application security
Model supply chain
Commercial platforms
- [Prompt Security →](https://prompt.security/)
- [Protect AI →](https://protectai.com/)
prompt-injection screening
### SafePrompt + Lakera Guard + LLM Guard
SafePrompt is the low-friction hosted choice for most developers: one provider-neutral HTTP call and a free starting tier. Its own April 2026 comparison reports most checks under 100 ms and accuracy above 95%; treat those as vendor claims and validate them against your traffic and attack set. Choose Lakera Guard when current SOC 2 evidence and enterprise support are procurement gates, or MIT-licensed LLM Guard when prompts must stay inside your infrastructure and you can operate and tune Python scanners. Detection is one layer—not a replacement for least-privilege tools, policy enforcement, output validation and incident logs.
Hosted default
Enterprise option
Self-hosted option
- [Vendor comparison and methodology →](https://safeprompt.dev/blog/best-prompt-injection-detection-tools)
- [SafePrompt quick start →](https://docs.safeprompt.dev/quick-start)
- [Lakera platform and compliance →](https://docs.lakera.ai/docs/platform)
- [LLM Guard repository →](https://github.com/protectai/llm-guard)
event transport · optional
### NATS + RabbitMQ
Add a broker when independently deployed services need backpressure, fan-out or replayable delivery. A single Temporal deployment plus PostgreSQL may not need one; do not introduce a second retry system until ownership and deduplication are explicit.
Messaging
Backpressure
Optional
- [NATS →](https://docs.nats.io/)
- [RabbitMQ →](https://www.rabbitmq.com/docs)
No tool matches this topic and search. Try a broader term or choose all.
Selection test
## Trial the boundary, not only the happy path.
A useful tool survives your real files, denial cases, export needs and maintenance budget.
**Fit**
Test ten representative cases against a manual baseline and a named acceptance threshold.
**Boundary**
Write down what leaves the device, which identities and tools it can use, and where traces remain.
**Failure**
Exercise denied files, missing credentials, hostile content, bad output and recovery after interruption.
**Exit**
Export data and configuration, reproduce one result and estimate the work to replace the tool.
**Fast-moving catalogue:** the catalogue was last updated from primary project or vendor sources on 15 August 2026. Recheck the exact release, model weights, license, pricing and deployment documentation before adoption.
Coding-agent overhead
## Efficiency: wall-clock and memory
Two narrow studies expose costs hidden by capability scores. They use different methods and workloads, so read each on its own terms—not as one combined ranking.
> Visual: Independent practitioner benchmark · updated 27 July 2026. Hold DeepSeek V4 Flash fixed; measure the wrapper. The study ran the same eight focused repository bug fixes with the same local model and grading method. Pi, OpenCode and Claude Code have 24 timed runs each; Nanocoder has 54 timed runs in the published wall-clock series. Dots are individual runs and diamonds are task-weighted averages.
Independent practitioner benchmark · updated 27 July 2026
### Hold DeepSeek V4 Flash fixed; measure the wrapper.
The study ran the same eight focused repository bug fixes with the same local model and grading method. Pi, OpenCode and Claude Code have 24 timed runs each; Nanocoder has 54 timed runs in the published wall-clock series. Dots are individual runs and diamonds are task-weighted averages.
Axis range
Zoomed to ten minutes. Choose Full to reveal every run at its true position.
**Visual entries (display order):**
- **Pi** 2.1m average 24 runs · 4 tools
- **OpenCode** 3.1m average 24 runs · 10 tools
- **Claude Code** 8.0m average 24 runs · 27 tools
- **Nanocoder** 5.2m average 54 timed runs · 15 tools
**Read this narrowly.** The result describes one model, codebase and task distribution. It does not establish a universal speed or quality ranking; it does show why harness, output tokens, tool calls, timeouts and repeated runs belong in a coding-agent evaluation.
- [Read the study and methodology →](https://nqawhc.github.io/articles/harness-efficiency-not-quality/)
- [Open the benchmark database entry →](https://isaiuseful.com/benchmarks.html.md#harness-efficiency)
> Visual: JCode project benchmark · checked 28 July 2026. Add another session; measure proportional RAM. JCode's corrected Linux rerun measures the slope from one to ten active clients. Bars show approximate extra proportional set size (PSS) per added session, which accounts for shared memory proportionally. Lower is better.
JCode project benchmark · checked 28 July 2026
### Add another session; measure proportional RAM.
JCode's corrected Linux rerun measures the slope from one to ten active clients. Bars show approximate extra proportional set size (PSS) per added session, which accounts for shared memory proportionally. Lower is better.
**Visual entries (display order):**
- **JCode · embeddings off** ~9.9 MB · baseline
- **JCode** ~10.4 MB · 1.1×
- **Codex CLI** ~21.6 MB · 2.2×
- **Pi** ~76.5 MB · 7.7×
- **Antigravity CLI** ~86.4 MB · 8.7×
- **Cursor Agent** ~157.5 MB · 15.9×
- **GitHub Copilot CLI** ~158.1 MB · 16.0×
- **Claude Code** ~212.7 MB · 21.5×
- **OpenCode** ~318.4 MB · 32.2×
**Read this as project evidence.** The measurements use one Linux machine and pinned client versions. PSS varies with operating system, version, authentication state, extensions, MCP servers and workload; it does not measure answer quality or model-serving memory. Re-run the comparison before capacity planning.
- [See the live comparison and method →](https://jcode.sh/#performance-and-resource-efficiency)
- [Inspect JCode source and tested versions →](https://github.com/1jehuang/jcode)
Next step
## Choose a workflow before assembling a stack.
The smallest tool set that passes a real acceptance test is usually the easiest one to secure, explain and replace.
- [Choose a guide](https://isaiuseful.com/guides.html.md#chooser)
- [Choose a local model](https://isaiuseful.com/local-models.html.md)
---
## https://isaiuseful.com/rag (`/rag.html.md`)
# What Is RAG? A Practical Guide to RAG, GraphRAG and Hybrid Retrieval
Canonical source: [https://isaiuseful.com/rag](https://isaiuseful.com/rag)
- [Build](https://isaiuseful.com/guides.html.md)
· Retrieval systems · checked 28 July 2026
Start with a small, cited retrieval baseline. Add graph structure only when the questions genuinely depend on entities, connections or whole-corpus themes—and keep both indexes rebuildable from governed source data.
- [Choose a retrieval route](#choose)
- [Build the pipeline](#build)
- [Define the test](#evaluate)
**3**
retrieval routes
VECTOR · GRAPH · HYBRID
**2**
systems to test separately
RETRIEVER + GENERATOR
**1**
governed source of truth
INDEXES ARE REBUILDABLE
01 · Mental model
## How does retrieval-augmented generation work?
Retrieval-Augmented Generation joins a generator to external, non-parametric memory. The documents remain outside the model weights.
### Prepare
Parse governed sources, preserve document and section identity, split them into useful retrieval units and create searchable representations.
### Retrieve
Use the question to select candidate passages, records or graph neighborhoods, then filter and rerank within the user's access boundary.
### Generate
Give the model the selected evidence, require an answer bounded by it and return citations that resolve to the original source—not merely to an index row.
**Evidence boundary:** the [original RAG paper](https://papers.neurips.cc/paper_files/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html) established generation with parametric and retrieved non-parametric memory on specific knowledge-intensive tasks. It did not establish that every modern document-chat system is accurate, secure or citation-faithful.
02 · Architecture decision
## Use the lightest route that can pass the real questions.
GraphRAG is an additional data product and query path, not an automatic upgrade.
**Vector RAG**
Start here when answers usually live in one or a few text passages: policies, manuals, case notes, contracts or knowledge-base articles.
**Search + rerank**
Add keyword retrieval, metadata filters or a reranker when dense similarity misses exact terms, identifiers or authoritative sources.
**GraphRAG**
Test graph retrieval when questions depend on connected entities, multi-document paths, controlled relationship types or themes across a corpus.
**Hybrid**
Combine text and graph context only when the evaluation set shows complementary failures and the improvement justifies two synchronized indexes.
| Route | Best first test | What it preserves | Primary failure | Operating burden |
| --- | --- | --- | --- | --- |
| Vector RAG | “What does this document say?” | Semantic similarity and source chunks | Related wording can outrank the needed fact; evidence may be split across chunks | Lowest of the three; still requires parsing, updates, ACLs and evaluation |
| GraphRAG | “How are these people, events or systems connected?” | Entities, typed relationships, neighborhoods and graph communities | Extraction and entity-resolution errors create missing or false paths | Higher; graph construction, resolution, provenance and query tuning |
| Hybrid RAG | A query needing both an exact source passage and a cross-source relationship | Broad textual evidence plus explicit structure | More context can add noise; ranking and synchronization become harder | Highest; two retrieval paths, merging, budgets and regression tests |
Research context: a [2025 systematic evaluation](https://arxiv.org/abs/2502.11371) reports distinct strengths for RAG and GraphRAG across question answering and query-focused summarization; [HybridRAG](https://arxiv.org/abs/2408.04948) reports gains from combined vector and graph retrieval on a financial-transcript experiment. These results motivate testing routes—they are not universal production guarantees.
03 · Reference architecture
## See exactly what the graph changes.
Both systems retrieve evidence for a generator. Conventional RAG retrieves passages; GraphRAG first builds explicit entities, relationships and corpus-level structure.
Conventional RAG
**Retrieve the best passages.**
Lowest useful complexity for document-grounded answers
**Index time**
Governed source
**Documents + records**
Version · ACL · owner
Prepare
**Parse + chunk**
Page and heading IDs survive
Derived memory
**Text + vector index**
Keywords · embeddings · metadata
**Question time**
Input
**User question**
Identity + allowed scope
Retrieve
**Search + filter + rerank**
Select top-k passages
Context
**Source chunks**
Passages + resolvable citations
Generate
**LLM answer**
Answer · abstention · citations
**Best fit**
Policies, manuals, contracts, case notes and answers that live in a few passages.
GraphRAG
**Retrieve connected evidence.**
Additional structure for relational and whole-corpus questions
**Index time**
Governed source
**Documents + records**
Version · ACL · owner
Extract + resolve
**Entities + relations**
Aliases · claims · timestamps
Derived memory
**Knowledge graph**
Edges · communities · source links
**Question time**
Input
**User question**
Entities + question class
Route + traverse
**Local, global or DRIFT**
Neighborhood · paths · communities
Context
**Subgraph + source chunks**
Nodes · edges · supporting text
Generate
**LLM answer**
Answer · path · citations
**Best fit**
Dependencies, ownership, lineage, investigations, multi-document paths and corpus-wide themes.
**The architectural difference is upstream of the model.** GraphRAG adds extraction, entity resolution, graph maintenance and graph-aware retrieval. A hybrid system keeps the passage path and adds the graph path only where the evaluation set proves it helps.
> Visual: Reference hybrid RAG pipeline
**Visual reading order:**
1. **01** **Govern sources** Rights, classification, owner, version, retention and access policy.
2. **02** **Parse + identify** Document, page, heading, record and stable source identifiers.
3. **03** **Build indexes** Text chunks and embeddings; optional entities, edges and community summaries.
4. **04** **Route + retrieve** Metadata filters, vector or keyword candidates and bounded graph traversal.
5. **05** **Merge + answer** Rerank, enforce a context budget and cite the original evidence.
6. **06** **Measure + refresh** Log versions, misses, path validity, latency, cost and user corrections.
**GraphRAG is more than a graph database.** Microsoft's current implementation extracts entities, relationships and claims, detects communities, produces summaries and embeds text. Its query engine separates entity-focused local search, whole-dataset global search, DRIFT and basic vector search. See the official [indexing overview](https://microsoft.github.io/graphrag/index/overview/) and [query overview](https://microsoft.github.io/graphrag/query/overview/) .
04 · Graph gate
## Earn the graph with questions that require one.
A graph pays for itself only when explicit structure improves an outcome that simpler retrieval cannot reach reliably.
**Strong signal**
### Relationships are the answer.
Ownership, dependency, lineage, supply chains, citations, organizational paths or event sequences must be traversed and explained.
**Strong signal**
### Questions span the corpus.
Readers need themes, clusters or connected evidence across many documents rather than the nearest matching passage.
**Conditional**
### A useful schema exists.
Stable identifiers, entity types and relationship rules already exist—or the workflow value can fund their creation and maintenance.
**Weak signal**
### “Graphs sound smarter.”
A product label, demo or vendor benchmark does not justify graph extraction when ordinary search already passes the acceptance set.
### A GraphRAG index is a governed data pipeline, not an LLM side effect.
Every stage creates an artifact that can be inspected independently. If the final answer is wrong, the trace should reveal whether the source was missing, the entity was split, the relationship was invented, the wrong neighborhood was traversed or the generator ignored valid context.
> Visual: Six stages in a GraphRAG indexing pipeline
**Visual reading order:**
1. **01 · Segment** **Text units** Preserve document, page, heading, version and ACL on every unit.
2. **02 · Extract** **Entities + claims** Find people, systems, events, concepts and candidate relationships.
3. **03 · Resolve** **Canonical identity** Merge aliases; keep homonyms, subsidiaries and versions separate.
4. **04 · Relate** **Typed edges** Direction, predicate, validity time, confidence and supporting source.
5. **05 · Organize** **Communities + summaries** Cluster the graph for broader questions without discarding raw evidence.
6. **06 · Publish** **Queryable index** Version the graph, embeddings, prompts and extraction configuration together.
### Choose the query mode by the shape of the question.
| Query mode | Question shape | Context assembled | What to test | Cost / failure boundary |
| --- | --- | --- | --- | --- |
| Basic text / vector | “What does policy 7.2 say about retention?” | Top matching passages | Exact source appears in top-k; citation resolves | Cheapest baseline; can miss distributed or relational evidence |
| Local graph search | “Who owns service A, and which incidents involved it?” | Seed entities, neighbors, relationships, community context and linked text | Entity mapping, edge direction, path validity and source support | Sensitive to duplicate entities and missing edges |
| Global graph search | “What recurring risks appear across the full incident archive?” | Community reports evaluated and reduced across the corpus | Theme coverage, minority evidence, aggregation bias and token budget | Resource-intensive; summaries can flatten exceptions |
| DRIFT / exploratory | “How might these local failures connect to broader operating patterns?” | Community-informed starting point plus detailed follow-up retrieval | Breadth gained, irrelevant branches, reproducibility and stop conditions | Broader search can add latency and plausible noise |
| Explicit graph query | “List approved suppliers two hops from programme X as of 30 June.” | Schema-bound traversal with filters and validity time | Exact path, filter semantics, authorization and empty-result behavior | Precise only when schema and graph data are precise |
### Graph quality is won or lost at four seams.
**Identity**
“ACME,” “ACME Ltd” and a product called “Acme” cannot be merged because an embedding says they look alike. Use stable keys where available, alias rules where necessary and a review queue for uncertain merges.
**Relationship**
“Uses,” “owns,” “approved by” and “mentioned with” are not interchangeable. Define direction and allowed entity types; reject an edge that cannot name its predicate and source.
**Time**
A graph without valid-from, valid-to and observed-at fields can answer with a relationship that was once true. Preserve event time separately from ingestion time and expose staleness.
**Provenance**
An extracted edge is navigation, not proof. Store the document and text-unit IDs behind it, and make every answer path resolve to original evidence a reviewer can inspect.
### Trace one relational question end to end.
> Visual: Worked GraphRAG query trace
**Visual reading order:**
1. **Question** **Which change caused the outage, and who approved it?** Requires an event, deployment, service, incident and person to connect.
2. **Seed** **Map outage + service** Resolve the incident ID and service alias before expanding.
3. **Traverse** **Incident → deployment → change** Follow only allowed, time-valid edge types within the incident window.
4. **Join** **Change → approval → person** Recover the approval record and responsible identity.
5. **Retrieve** **Open supporting passages** Deployment log, incident timeline and approval record enter context.
6. **Answer** **State path + uncertainty** Cite each hop; abstain if any required edge lacks evidence.
### The graph earns production only if it passes a separate acceptance set.
| Test family | Fixture | Pass rule | Failure it catches | Release action |
| --- | --- | --- | --- | --- |
| Entity resolution | Aliases, homonyms, mergers, renamed systems and versioned products | Known same entities merge; known different entities remain separate | False joins and broken neighborhoods | Block new resolver; review uncertain identity queue |
| Relationship extraction | Positive, negative, hypothetical and historical statements | Predicate, direction, time and source match the reference | Invented, reversed or timeless edges | Quarantine edge type or fall back to text retrieval |
| Path retrieval | Known one-, two- and three-hop questions plus impossible paths | Required path is returned; impossible path produces no fabricated bridge | Traversal gaps and graph completion by guessing | Tune seeding and hop limits; keep abstention |
| Global themes | Dominant, minority and contradictory themes across a frozen corpus | Material themes survive aggregation with traceable evidence | Summary flattening and majority bias | Change community level or return scoped results |
| Authorization | Users with overlapping but different document and entity rights | No node, edge, summary or citation crosses the caller's boundary | Relationship leakage across collections | Stop serving graph results until fixed |
Implementation references: Microsoft's [standard and fast indexing methods](https://microsoft.github.io/graphrag/index/methods/) , [query-mode overview](https://microsoft.github.io/graphrag/query/overview/) and the [systematic RAG versus GraphRAG evaluation](https://arxiv.org/abs/2502.11371) . Microsoft's method notes explicitly trade richer extraction against cost and noisier fast graphs; treat that as an engineering choice to benchmark on your corpus.
**From agent loop to shared graph memory.** The downloadable 11-page study note connects Karpathy's autoresearch loop and AgentHub commit DAG with multi-agent workflows, knowledge-graph provenance and a staged implementation path.
- [Download Graph Engineering (PDF) →](https://isaiuseful.com/downloads/Karpathy-Graph-Engineering-Systems.pdf)
**Graph data can be confidently wrong**
Keep source identifiers and confidence on every extracted claim. Test duplicate entities, aliases, missing edges, contradictory timestamps and unauthorized relationship leakage. If a path cannot resolve back to evidence, do not present it as a citation.
05 · Build sequence
## Build one measured slice before a platform.
The first deliverable is not “chat with everything.” It is a small set of real questions answered within a defined data and authority boundary.
| Stage | Minimum deliverable | Acceptance check | Do not hide | Scale trigger |
| --- | --- | --- | --- | --- |
| 1 · Contract | One audience, corpus, task, owner and answer policy | Known answerable, unanswerable and forbidden questions | Rights, personal data, stale sources and access boundaries | The workflow owner accepts the question set |
| 2 · Baseline | Keyword and vector retrieval over a small representative corpus | Correct evidence appears in top-k for held-out questions | Parser failures, empty pages, tables and duplicate versions | Retrieval misses are understood by category |
| 3 · Grounded answer | Answer, refusal and resolvable source citations | Claims match cited evidence; unsupported questions abstain | Prompt and model version, context used and truncation | A stable regression set passes repeatedly |
| 4 · Optional graph | Only the entity and relationship types needed by failed questions | Valid paths improve the named failures without harming simple QA | Resolution confidence, provenance and graph build cost | Measured gain exceeds added latency and upkeep |
| 5 · Operate | Incremental refresh, deletion, ACL enforcement, monitoring and rollback | Source changes appear on time; revoked data disappears everywhere | Index age, last successful build and partial failures | Restore and re-index drills work |
06 · Evaluation
## Score the retriever before blaming the model.
A fluent wrong answer can begin with a retrieval miss, a ranking mistake, an incomplete graph path or unsupported generation. Preserve the stage boundary.
Retriever
### Did the right evidence arrive?
Measure top-k hit rate, context precision and recall, metadata-filter accuracy and—when graph retrieval is used—entity and path coverage.
Generator
### Did the answer stay inside it?
Measure answer correctness, citation entailment, refusal behavior and faithfulness to the retrieved context. Review consequential failures manually.
System
### Was it useful repeatedly?
Track end-to-end latency, cost, freshness, access-control failures, user corrections and pass rate across repeated runs—not only a best attempt.
**Metric caution:** Ragas documents context precision, context recall, response relevance and faithfulness as separable RAG measures. Some of these use model-based scoring. Calibrate them against references and human review rather than treating one automated score as independent proof. See the [current metric catalogue](https://docs.ragas.io/en/latest/concepts/metrics/available_metrics/) and this site's [evidence policy](https://isaiuseful.com/evidence.html.md) .
**Ten real trials before more authority**
Freeze the corpus snapshot, questions, expected evidence and pass rules. Run vector, graph and hybrid variants on the same set. Save every retrieved context and answer. Expand only the route that improves the named workflow without unacceptable security, latency or maintenance regressions.
07 · Production boundaries
## Retrieval creates a new route to governed data.
Embeddings, chunks, graph edges, traces and caches can reveal sensitive material even when the original repository is protected.
**Authorize first**
Filter candidates by the user's entitlement before generation. Do not retrieve broadly and ask the model to redact afterward.
**Preserve provenance**
Carry source ID, section or page, observed time, version and classification through every index and response.
**Delete everywhere**
Define how revocation reaches chunks, vectors, graph facts, summaries, caches, traces, backups and derived evaluations.
**Expose freshness**
Show the source date and last successful index build. Stop or warn when an update fails instead of silently serving stale evidence.
### Map policy to every derived surface.
A source repository's controls do not automatically follow copied text, embeddings, community summaries or traces. Each derived surface needs an explicit owner, authorization check, retention rule and deletion path.
| Surface | What it can reveal | Minimum control | Deletion / correction path | Production signal |
| --- | --- | --- | --- | --- |
| Parsed text + chunks | Full passages, hidden fields, OCR mistakes and prior versions | Encrypted storage, document ACL, parser allowlist and quarantine for unreadable files | Remove by stable source/version ID, then rebuild affected indexes | Parse success, skipped content and last valid version |
| Embeddings + search index | Similarity, membership and retrievable sensitive concepts | Tenant/collection isolation, pre-retrieval filters, private endpoints and tested backups | Delete vector and metadata by source ID; verify it cannot be retrieved | Index age, document count, filter result and restore test |
| Graph nodes + edges | Sensitive relationships that no single document states plainly | Node and edge authorization, provenance, time validity and restricted traversal | Retract affected claims, summaries and neighborhoods; re-run resolution | Orphan rate, unresolved identities, invalid edges and ACL denials |
| Community summaries + caches | Cross-document conclusions and stale aggregate facts | Scope summaries to compatible ACL domains; attach input version set and expiry | Invalidate every summary whose dependency set changed | Dependency version, expiry and cache-hit freshness |
| Prompts, traces + eval sets | Questions, retrieved evidence, model answers, corrections and user intent | Redaction, access-limited observability, retention limits and separation from product analytics | Purge by run, user and source ID without destroying required audit records | Sampling rate, redaction failures, retention age and reviewer access |
### Treat refresh as a versioned release.
> Visual: Safe retrieval index refresh lifecycle
**Visual reading order:**
1. **01 · Detect** **Source changed** Create, update, revoke or classification change enters a durable queue.
2. **02 · Rebuild** **Derived artifacts** Reparse affected units; re-embed; re-extract graph claims and summaries.
3. **03 · Validate** **Quality + policy** Parser, retrieval, graph, ACL, deletion and freshness checks run before publish.
4. **04 · Publish** **Atomic version** Promote compatible index, graph, prompt and configuration versions together.
5. **05 · Observe** **Canary questions** Compare hit rate, groundedness, latency, empty results and authorization denials.
6. **06 · Roll back** **Last valid release** Keep a known-good manifest; never serve a half-updated vector/graph pair.
### Define the failure response before the first incident.
| Failure | How it appears | Automatic response | Owner decision | Evidence to retain |
| --- | --- | --- | --- | --- |
| Source or parser failure | Missing pages, zero-length chunks, broken tables or unexpected document-count drop | Quarantine source; keep last valid version with a visible stale flag | Repair parser, accept exclusion or stop the collection | File hash, parser version, errors and skipped ranges |
| Retrieval regression | Known evidence falls out of top-k or irrelevant context dominates | Fail canary; hold index promotion; fall back to last valid release | Change chunking, embeddings, filters or reranker | Question, expected source, candidates, scores and versions |
| Graph corruption | Entity explosion, suspicious new hubs, reversed edges or impossible paths | Disable affected edge types or graph route; retain text RAG | Re-run resolution, correct schema or rebuild graph | Extraction prompt, claim source, merge history and graph diff |
| Grounding failure | Answer claim is not supported by the supplied passage or path | Abstain or return evidence without synthesized claim | Adjust prompt, context assembly, model or answer policy | Full context, answer, citations, grader and human adjudication |
| Authorization leak | Restricted chunk, entity, relationship or summary reaches an unauthorized run | Stop affected route, revoke caches and preserve incident evidence | Breach process, user notification and safe re-enable criteria | Caller identity, policy decision, retrieved IDs and access logs |
| Refresh drift | Vector and graph versions disagree or revoked facts remain in summaries | Mark release unhealthy and route to a coherent known-good version | Complete rebuild or targeted dependency invalidation | Release manifest, dependency graph and deletion verification |
### Make ownership as explicit as the architecture.
Data owner
### Controls the source boundary.
Approves corpus, rights, classification, retention, authoritative versions and who may see which documents and relationships.
Retrieval owner
### Controls the derived memory.
Owns parsers, chunking, indexes, entity resolution, refresh, deletion, restore, latency and retrieval regression tests.
Workflow owner
### Controls the useful answer.
Defines questions, acceptance rules, abstention, human escalation, outcome measurement and the authority the answer may trigger.
**Cost lever, not a RAG property:** vector quantization can reduce index memory and storage, but it is lossy and system-specific. Azure AI Search currently documents up to 28× index-size reduction for binary quantization and recommends oversampling and rescoring to offset information loss. Treat compression as an experiment with the same recall set. [Read the official guide →](https://learn.microsoft.com/en-us/azure/search/vector-search-how-to-quantization)
08 · Replaceable components
## Choose interfaces before brands.
Keep the raw corpus, parsing output, retrieval evaluation and source identifiers portable. The index is derived infrastructure.
### Prepare + retrieve
Document parsers, chunkers, keyword search, vector stores, metadata filters and rerankers. Start with the fewest services that meet the data boundary.
- [Compare retrieval tools →](https://isaiuseful.com/tools.html.md#tools-build-knowledge-and-workflows)
- [See a bounded local knowledge recipe →](https://isaiuseful.com/guides.html.md#second-brain)
### Relate + traverse
Entity extraction, resolution, graph storage and query generation. A graph store is optional; source provenance and repeatable graph builds are not.
- [Compare graph options →](https://isaiuseful.com/tools.html.md#tools-build-knowledge-and-workflows)
- [See an operational graph stack →](https://isaiuseful.com/diy-palantir.html.md#stack)
### Trace + evaluate
Capture the query, route, retrieved evidence, graph path, prompt, answer, model and versioned grader result so failures can be reproduced.
- [Compare evaluation tools →](https://isaiuseful.com/tools.html.md#tools-fine-tune-evaluate-and-reproduce)
- [Choose an evaluation pattern →](https://isaiuseful.com/benchmarks.html.md#database)
The practical default
Start with cited vector retrieval. Let failed questions earn more structure.
- [Choose the route](#choose)
- [Build one slice](#build)
- [Compare retrieval with training](https://isaiuseful.com/training-models.html.md#chooser)
---
## https://isaiuseful.com/local-models (`/local-models.html.md`)
# Which Local AI Model Should You Run?
Canonical source: [https://isaiuseful.com/local-models](https://isaiuseful.com/local-models)
Agent-first model guide · checked 15 August 2026
A model package loading is only the first gate. A useful local agent also needs memory and tokens for **instructions, tool schemas, repository files, tool results, reasoning and the answer** . Choose the job first, then balance capability, context and memory.
- [Choose an agent role](#roles)
- [Route local + frontier work](#routing)
- [Use a phone as the controller](#mobile-control)
- [Budget the context](#context)
- [Match your memory](#memory)
**4**
agent roles
WORKER · CODER · RESEARCHER · PLANNER
**3**
context numbers
ADVERTISED · CONFIGURED · USABLE
**20–25%**
minimum memory headroom
OS · HARNESS · KV CACHE
Role selector
## Which local AI model fits each kind of work?
Size is a rough capability band, not an IQ score. Architecture and post-training can move a model up or down; your quant, runtime and tool grammar can move it again.
Practical band
**20–35B · Q4/Q6**
16–32K minimum · 32–64K preferred
### The everyday local-agent band.
Suitable for bounded multi-file edits, debugging with a reproduction, test-driven implementation and a small set of reliable tools. Keep a human checkpoint before migrations, security changes or broad refactors.
### Escalate the plan
For an ambiguous feature, let a 70–120B or strong hosted model produce the plan and review criteria. Save that plan as a file, compact or start a clean task, then let the cheaper 20–35B worker implement it.
Model band
Good default assignments
Expect trouble with
**2–4B**
Routing, tagging, format conversion, autocomplete and tightly constrained extraction.
Reliable planning, fuzzy tool choice, long chains and precise multilingual nuance.
**7–12B**
Grounded Q&A, drafting, single-file code, one or two simple tools and supervised workers.
Autonomous repo-wide work, conflicting requirements and recovery after several failed steps.
**20–35B**
Everyday coding agents, multi-document synthesis, structured research and several focused tools.
High-stakes judgment, novel architecture and long autonomous runs without tests or checkpoints.
**70–120B**
Planning, review, ambiguous debugging, cross-file refactors and coordinating specialist workers.
Guaranteed frontier parity, perfect long-context recall or unattended production changes.
**400B+**
Hosted or multi-GPU frontier open-weight work when quality justifies infrastructure. [Plan the provider tier →](https://isaiuseful.com/cloud-models.html.md)
Single 128 GB machines; active MoE parameters do not remove the need to store and load the full checkpoint.
High fidelity
### Q6 / Q8
Use when exact phrasing, code reliability or multilingual nuance matters and a smaller high-precision model already meets the task.
Start here
### Q4_K_M / official QAT
The normal quality/size starting point. Four-bit compression usually preserves broad ability, but agentic application accuracy can degrade more than perplexity suggests.
Trade carefully
### Q3
Useful when it unlocks a meaningfully larger model. Re-run tool use, structured output and language tests; subtle failures appear before chat becomes obviously bad.
Last-resort fit
### Q2 / IQ2
Use only after testing the exact build on your tasks. Very-low-bit quants can make a larger model load, but tool calls, structured output, code edits and multilingual nuance may degrade unevenly.
**Heuristic, not a benchmark grade.** The role bands synthesize model-card capability curves and practical agent constraints. [ACBench](https://isaiuseful.com/benchmarks.html.md#acbench) found that compression effects differ by model, task and quantization method; validate the exact build through the harness and tasks you will use.
- [Understand the ACBench signal →](https://isaiuseful.com/benchmarks.html.md#acbench)
- [Read the ACBench paper →](https://arxiv.org/abs/2505.19433)
- [See llama.cpp quant sizes →](https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md)
Hybrid agent architecture
## Route each step—not the whole application—to the right model.
A local model does not need to replace the frontier model. Let the strongest model resolve ambiguity, move routine execution to a cheaper specialist, and escalate when the trace shows difficulty.
> Visual: Hybrid model-routing architecture
**Visual entries (display order):**
- 01 · plan **Frontier planner** Settles architecture, constraints, review criteria and unfamiliar failure modes.
- 02 · route **Model-neutral gateway** Selects a target for each turn from task, policy, cost, latency and recent tool signals.
- 03 · execute + escalate **Specialist models → tools → frontier** Local or hosted workers handle bounded calls; repeated errors, loops or uncertainty route back up.
Current local worker
### Nemotron 3.5 Lightning
NVIDIA’s 30B-total / 3B-active hybrid MoE targets high-volume tool calls, validation and delegated execution. The official NVFP4 checkpoint is about 20.1 GiB before runtime and cache, making 32 GB the sensible personal-system floor.
Keep the seam stable
### Route to roles, not vendor IDs.
Name targets such as `planner` , `private-worker` and `frontier-review` . Map those aliases to a local Gemma/Muse/Qwen-class endpoint, a European hosted service or a frontier API without rewriting application logic.
Measure useful work
### Tokens are an input—not the score.
Benchmark task success, total workflow cost, completion time, retries and frontier-escalation frequency. Include human review and failed runs so a cheap but unreliable route cannot look efficient.
Current maturity boundary
### Switchyard is an experiment, not a production default.
NeMo Switchyard provides provider translation plus classifier, stage, escalation and custom routing. Its repository labels the software pre-alpha and not for production use; isolate it behind a gateway contract, pin a revision and keep a tested fallback.
**Promising vendor and partner evidence—not a universal saving.** NVIDIA reports up to 4× the output speed of similar-sized models and a 30% faster 10,000-task agent run at comparable accuracy. On a separate 145-task internal suite, LangChain reported that an escalation route between Lightning and Claude Opus 4.8 cut total cost 74% while sending 7% of calls to the frontier model, but accuracy fell from 86% to 80%. Reproduce the comparison on your own task mix, prices, providers and failure costs.
- [NVIDIA Lightning launch and benchmark context →](https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/)
- [NVIDIA Switchyard architecture →](https://developer.nvidia.com/blog/route-ai-agent-workloads-across-models-with-nvidia-nemo-switchyard/)
- [LangChain’s 145-task benchmark →](https://www.langchain.com/blog/switchyard-agent-routing-benchmark)
- [Switchyard source, licence and maturity warning →](https://github.com/NVIDIA-NeMo/Switchyard)
- [Plan a hosted or data-centre Lightning route →](https://isaiuseful.com/cloud-models.html.md#nemotron-35-lightning)
Pocket control
## Let the phone drive without pretending it hosts the model.
Remote access and local inference are separate decisions. The phone can carry the conversation and approvals while an authenticated gateway, model and tools stay on an always-on machine you control.
> Visual: Phone-to-home-model control path
**Visual entries (display order):**
- 01 · phone **Browser · app · messaging** Collects the request, shows progress and asks for approval. It may be only a control surface.
- 02 · gateway **Identity · sessions · tool policy** Authenticates the client, owns the agent loop and decides which backend and capabilities are available.
- 03 · model host **Ollama · LM Studio · vLLM** Loads weights and serves inference on the home machine, workstation or private server.
Easiest browser route
### Open WebUI + Ollama
Run both on the host, make only the authenticated WebUI reachable over a private route and open it in the phone browser. Keep Ollama itself on localhost or an internal network.
- [Open WebUI documentation →](https://docs.openwebui.com/)
Native companion route
### OpenClaw + home gateway
Keep the gateway and session state on the always-on host, then pair an iPhone or Android companion. Start with remote chat; add device-node capabilities only when the workflow needs them.
- [OpenClaw remote access →](https://docs.openclaw.ai/gateway/remote)
Coding-agent cockpit
### Orca mobile + desktop
Pair Orca's iPhone or Android companion with its desktop runtime to see worktrees, agent status and session output, then reply or dictate from the phone. The desktop remains the source of truth.
- [Orca mobile companion →](https://www.onorca.dev/docs/mobile)
Use an app you already have
### Hermes + messaging
Reach the agent through a supported messaging channel while its model runs elsewhere. Treat that channel provider as part of the data path, even when inference stays at home.
- [Hermes messaging gateway →](https://hermes-agent.nousresearch.com/docs/user-guide/messaging/)
**Private-by-default rule:** expose the authenticated interface or gateway through a Tailnet or SSH tunnel. Do not forward a raw model, retrieval or agent-control port through the router. Test from cellular data, confirm the execution-location label and verify that a denied tool still fails. A model that actually runs on the phone is a different architecture, with device limits on model fit, sustained speed, memory, heat and battery.
- [OpenClaw iPhone setup →](https://docs.openclaw.ai/platforms/ios)
- [OpenClaw Android setup →](https://docs.openclaw.ai/platforms/android)
- [Orca private remote server →](https://www.onorca.dev/docs/remote-servers)
- [LM Studio API authentication →](https://lmstudio.ai/docs/developer/core/authentication)
- [See an always-on host topology →](https://isaiuseful.com/remote-spark.html.md#topology)
Agent-safe memory
## How much memory is left after the model loads?
These defaults assume one interactive agent, a Q4-class model, the OS and runtime, plus useful KV-cache headroom. “Chat fit” can be larger; “agent fit” must survive tool use and growing context.
16 GB agent fit: 7–12B Q4 with 8–16K context. A 14 GB model may load, but leaves almost no runway for a harness.
Shortlist for this machine
### Start with these 16 GB candidates.
These are working configurations, not weight-only fits. Context is a conservative configured starting point for one interactive agent; increase it only after watching memory and prompt-processing time.
Multilingual worker
### Gemma 4 12B
**QAT Q4 · start at 8–16K**
Best first candidate for multilingual drafting, retrieval-grounded work and bounded code changes.
- [Open family record →](#gemma)
Tool-aware worker
### Qwen3.5 9B
**Q4 · start at 8–16K**
A compact coding and tool-use baseline with broad runtime packaging.
- [Open family record →](#qwen)
Moonshot research model
### Moonlight 16B‑A3B
**Community 4-bit · 8K maximum**
The small Moonshot option. It is useful for experiments, but it is not a miniature Kimi K3.
- [Open family record →](#kimi)
Older edge baseline
### GPT‑OSS 20B
**14 GB MXFP4 · short context**
Retained for compatibility testing. It can load in 16 GB, but leaves almost no margin and is not a current first pick.
- [Open family record →](#gpt-oss)
**Looking for the viral Kimi demo?** Kimi K3 is a 2.8T server-scale model, not a 16 GB download. The Moonshot card above is an older small research checkpoint. [Plan the provider tier →](https://isaiuseful.com/cloud-models.html.md)
16 GB · worker
### 7–12B Q4 · 8–16K
**Agent-safe start:** Gemma 4 12B QAT Q4 (about 6.7–7.4 GB) or Qwen3.5 9B Q4 (6.6 GB). Moonlight 16B‑A3B is a small Moonshot research option in community 4-bit builds, but its model limit is only 8K.
**Chat fit is not agent fit:** GPT‑OSS 20B is designed to run within 16 GB and its Ollama package is 14 GB, but that leaves almost no margin for macOS, long context or tool results.
24 GB · strong worker
### 12–20B Q4 · 16–32K
**Current start:** Gemma 4 12B at Q6/Q8 or Muse Glimmer’s official 17 GB quant. Gemma 4 26B‑A4B Q4 at 14.4 GB is also viable when the runtime supports it well; treat 24 GB Muse as a tighter agent configuration.
**Tight:** Qwen3.8‑27B FP8 (16.38 GB; supported serving backends), Qwen3.6 27B Q4 (17 GB; broader desktop packaging), GLM‑4.7‑Flash (19 GB) and 32B Q4 packages load on paper but sacrifice the context and tool runway that makes a harness useful.
32 GB · everyday agent
### 20–31B low-bit · 16–32K
**Ranked start:** Gemma 4 26B‑A4B Q4 first, Muse Glimmer 30B second and Qwen3.8‑27B FP8 third. Qwen remains the coding-heavy choice on a supported backend and leaves about 15 GB before runtime allocation.
**Fallback route:** GLM‑4.7‑Flash Q4 and older broadly packaged models remain compatibility options where current Gemma, Muse or Qwen builds are unavailable.
64 GB · choose your bias
### Context route or capability route
**Ranked start:** Gemma 4 31B for the strongest current family fit, Muse Glimmer for agent tools and review, then Qwen3.8‑27B FP8 for coding-heavy work and parallel sequences. Use the older DeepSeek R1 Distill 70B only when a reasoning acceptance set proves the trade.
**Do not spend memory for its own sake:** Qwen says the FP8 package’s performance is nearly identical to the original model. The 55.56 GB full checkpoint is a fidelity test route, not the default 64 GB configuration.
128 GB · planner or team
### One 120B or several specialists
**Ranked start:** run Gemma 4 at high precision, Muse Glimmer for one or more local agents, or Qwen3.8‑27B FP8 with parallel headroom. Keep GPT‑OSS 120B as an older compatibility baseline only when the exact acceptance set rewards its 65 GB package.
Qwen3.5 122B‑A10B Q4 (81 GB) remains a capability-first stretch, not the automatic Qwen default. Choose it only when your tasks beat Qwen3.8‑27B by enough to justify the lost context and concurrency runway.
01
### Weights share; context multiplies
Parallel agents can share one loaded model, but each active sequence adds KV-cache pressure and its own growing history.
02
### Set context deliberately
A model may advertise 128K–1M while the runtime loads 8K. Raising it costs memory and prompt-processing time.
03
### Prefer Q4_K_M or official QAT
Start here, then compare Q6/Q8 or a larger Q3 only on a repeatable tool, code and language test set.
04
### Measure completed work
Task success, retries and human corrections matter more than tokens per second or a leaderboard point.
Context budget
## 128K advertised is not 128K for your files.
The configured window must hold the harness, system and project instructions, MCP schemas, conversation, retrieved files, tool results, hidden reasoning where applicable and the next answer.
01
### Advertised
**What the model card allows**
The training limit: often 128K, 256K or more. It says nothing about your runtime setting, speed or memory fit.
02
### Configured
**What the runtime actually loads**
A runtime can be set to 8K even when the model supports far more. The inference server and agent client may also impose separate limits.
03
### Usable
**What remains for the task**
Configured window minus instructions, tool definitions, history, output reserve and safety buffer. This is the number to plan around.
Coding fit
**Fits with 6.2K tokens of runway**
The initial slice fits. The remaining range is roughly 1 additional 16 KiB tool-result batch, or 445–1,039 more nonblank code lines before compaction.
> Visual: 32K configured: 10K harness, 8.3K supplied material, 8K output reserve, 6.2K free.
Harness
**10K**
*Working set*
**8.3K**
Output reserve
**8K**
Runway
**6.2K**
**4**
files
**480**
relevant lines
**12 KiB**
tool output
**350**
plan words
**A coding task is not “the repo”**
Load a deliberate working set
Count the files and relevant lines the agent must see now. Search and retrieval can fetch more later; dumping the full tree spends context before work begins.
**Tool rounds eat the runway**
Diffs, tests and errors accumulate
A long command result may cost more than the prompt that caused it. Prefer focused tests, clipped logs and fresh reads over carrying every failed attempt.
**Plan → save → compact → build**
Use capability where it propagates
Let a stronger model resolve ambiguity and write the plan plus acceptance criteria. Start the worker with that artifact and only the files needed for the next task.
**This is a token-occupancy estimator, not a tokenizer or RAM calculator.** Coding mode uses transparent planning midpoints: 10 tokens per relevant nonblank line, 4 bytes per token for logs/diffs, and 0.75 prose words per token. Real repositories, languages, tokenizers, harness prompts and hidden model work vary widely. Use the range as a loading plan, then inspect the harness’s live context view.
- [LM Studio context fields →](https://lmstudio.ai/docs/developer/rest/list)
- [Copilot context management →](https://docs.github.com/en/copilot/concepts/agents/copilot-cli/context-management)
Family deep dive
## Inspect the families behind your memory shortlist.
Families are ordered by current recommendation, not download size: **Gemma 4 first, Muse Glimmer second and Qwen3.8 third.** An **older** label marks the relevant local line as pre‑2026 and keeps it for compatibility or research—not as a first pick. Fit chips still show where a package can be practical, not where its maximum context will fit.
room for useful context
tight or runtime-specific
workstation, multi-GPU or cloud
01
Gemma 4
Multilingual European default with unusually strong Estonian proof.
**16 GB**
12B QAT Q4 · about 6.7–7.4 GB weights
**24 GB**
26B‑A4B Q4 · official estimate 14.4 GB
**32 GB**
31B Q4 · official estimate 17.5 GB
**64 GB**
31B SFP8 · official estimate 34.9 GB
**128 GB**
31B BF16 · official estimate 69.9 GB
01 · Current local default
### A European model choice with a demanding Estonian proof point.
Google reports pre-training across 140+ languages and out-of-the-box support for 35+. Independent TartuNLP evaluation of **Gemma 3** found strong Estonian instruction following, grammar and word-meaning results at 12B and 27B. That makes Estonian an excellent stress test for the family—not a promise for every language or automatic evidence for Gemma 4.
**All Gemma 4 sizes**
E2B
E4B
12B
26B · 4B active
31B
Base and instruction-tuned checkpoints; Google also publishes quantization-aware-trained variants.
**Scale inside one family**
Google model-card scores · higher is better
- [MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**60.0**
E2B
- [MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**77.2**
12B
- [MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**85.2**
31B
- [LiveCodeBench v6](https://isaiuseful.com/benchmarks.html.md#livecodebench)
**44.0**
E2B
- [LiveCodeBench v6](https://isaiuseful.com/benchmarks.html.md#livecodebench)
**72.0**
12B
- [LiveCodeBench v6](https://isaiuseful.com/benchmarks.html.md#livecodebench)
**80.0**
31B
- [GitHub →](https://github.com/google-deepmind/gemma)
- [Hugging Face collection →](https://huggingface.co/collections/google/gemma-4)
- [Ollama tags →](https://ollama.com/library/gemma4)
- [LM Studio family →](https://lmstudio.ai/models/gemma-4)
- [Open 12B in LM Studio →](https://lmstudio.ai/deeplink?name=gemma-4-12b&owner=google)
- [Gemma 4 paper →](https://arxiv.org/abs/2607.02770)
- [Official model card and scores →](https://ai.google.dev/gemma/docs/core/model_card_4)
- [TartuNLP Estonian evaluation →](https://huggingface.co/tartuNLP/llama-estllm-prototype-0825)
- [Gemma Scope, ShieldGemma and companion models →](https://ai.google.dev/gemma/docs)
02
Muse Glimmer
Meta’s current 30B local-agent launch for tools, coding and image input.
**24 GB**
Official K‑Quant‑17GB · 16.76 GB weights · start below maximum context
**32 GB**
K‑Quant‑17GB plus optional 1.40 GB vision encoder and 1.63 GB DFlash drafter
**64 GB**
Official K‑Quant‑Dynamic · 19.65 GB weights · target hardware 64 GB
**128 GB**
DGX Spark owner runs · 5–10.5 tok/s base · about 23–38 tok/s with DFlash
02 · Current agent launch
### A real 30B release for one device—with unusually clear package boundaries.
**Muse Glimmer 30B** is a dense model distilled from Muse Spark and released under Apache 2.0 for agentic work, tool use, coding, multilingual tasks and text-plus-image input. Meta publishes a 131,072+ model context and training across more than 100 languages. The official GGUF repository contains a 16.76 GB K‑Quant‑17GB build and a 19.65 GB dynamic build; image input needs the separate 1.40 GB perception file, while the optional 1.63 GB DFlash drafter accelerates generation without changing accepted output. Two early DGX Spark owner runs measured 5–10.5 tok/s without DFlash and about 23–38 tok/s with the official drafter set to 15 speculative tokens; useful evidence, but not a standardized cross-engine benchmark. Those file sizes make 24–32 GB systems credible targets, but they do not make the full context cheap or guarantee that every launch-day client understands the new architecture.
**Official released artifacts**
29.6B total · dense
17 GB local quant
19.65 GB dynamic quant
1.40 GB vision encoder
1.63 GB DFlash drafter
131,072+ model context
Meta targets the two quants at 24/32 GB and 64 GB respectively. Its announcement says Ollama and LM Studio integrations are coming in the days after launch, so pin a current llama.cpp or supported Transformers build and test the complete scaffold.
**Muse Glimmer vendor scorecard**
High reasoning; harnesses differ; compare the full table
- [MCP‑Atlas public](https://isaiuseful.com/benchmarks.html.md#mcp-atlas)
**75.5**
Tool and server use
- [SWE‑bench Pro](https://isaiuseful.com/benchmarks.html.md#swe-bench)
**51.2**
Agentic coding
- [Terminal‑Bench 2.1](https://isaiuseful.com/benchmarks.html.md#terminal-bench)
**51.7**
With Terminus‑2
- [OSWorld Verified](https://isaiuseful.com/benchmarks.html.md#osworld)
**65.9**
Qwen3.6‑27B: 75.6
- [Meta launch, local-speed evidence and packaging timeline →](https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model)
- [Full-precision model card and benchmark table →](https://huggingface.co/meta-models/Muse-Glimmer-30B)
- [Official GGUF files and exact sizes →](https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF/tree/main)
- [Meta evaluation methodology →](https://research.meta.ai/static/muse-glimmer-methodology)
- [Meta developer documentation →](https://developer.meta.com/ai/models/muse-glimmer/)
- [Usage policy and agent safety guidance →](https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/main/USAGE_POLICY.md)
- [DGX Spark llama.cpp, DFlash and extended-context owner test →](https://www.reddit.com/r/LocalLLaMA/comments/1vl9adk/i_ran_muse_glimmer_1m_context_all_tests_passed/)
- [DGX Spark vLLM DFlash owner test →](https://www.reddit.com/r/LocalLLM/comments/1vm4j0i/muse_glimmer_30b_on_dgx_spark_using_dflash_is/)
- [Spark capacity and performance boundary →](https://isaiuseful.com/remote-spark.html.md#spark-reviews)
03
Qwen3.8 / Qwen
Current all-rounder for multilingual, coding and agentic work.
**16 GB**
Qwen3.5 9B Q4 · 6.6 GB
**24 GB**
Qwen3.8 27B FP8 · 16.38 GB · short-context fit
**32 GB**
Qwen3.8 27B FP8 · useful context runway
**64 GB**
Qwen3.8 27B FP8 · long context or parallel agents
**128 GB**
Qwen3.8 27B FP8 · high-concurrency agent team
03 · Current all-rounder
### Qwen3.8‑27B is the practical Qwen pick for a 32 GB machine.
It handles text, images and video, can switch thinking on or off, and ships in a 16.38 GB FP8 build. Start with 32 GB for one useful local agent; 64 GB gives you more context or parallel work. Qwen reports that the FP8 build performs close to the full 55.56 GB checkpoint, but test it on your own tasks.
**Qwen3.8‑27B release**
27B · dense
16.38 GB · FP8
55.56 GB · full
262K · native context
1M · extensible
Text + image + video
Controllable thinking
Official serving recipes: Transformers, vLLM, SGLang and TokenSpeed.
**Released Qwen3.8 scorecard**
Qwen model card · vendor-reported · higher is better
- [LiveCodeBench v6](https://isaiuseful.com/benchmarks.html.md#livecodebench)
**90.3**
Competitive coding
- [OSWorld‑Verified](https://isaiuseful.com/benchmarks.html.md#osworld)
**84.3**
Computer use
- [Terminal‑Bench 2.1](https://isaiuseful.com/benchmarks.html.md#terminal-bench)
**73.0**
Terminus scaffold
CoWorkBench
**70.7**
Long-horizon office work
- [SWE‑bench Pro](https://isaiuseful.com/benchmarks.html.md#swe-bench)
**61.7**
Agentic coding
IFBench
**79.5**
Instruction following
- [Qwen3.8‑27B model card, scores and full checkpoint →](https://huggingface.co/Qwen/Qwen3.8-27B)
- [Official Qwen3.8‑27B FP8 package →](https://huggingface.co/Qwen/Qwen3.8-27B-FP8)
- [SGLang Qwen3.8 cookbook →](https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8-27B)
- [vLLM Qwen3.8 recipe →](https://recipes.vllm.ai/Qwen/Qwen3.8-27B)
- [TokenSpeed Qwen3.8 recipe →](https://lightseek.org/tokenspeed/recipes/models#qwen3-8)
- [Qwen Agent →](https://github.com/QwenLM/Qwen-Agent)
- [Qwen Code →](https://github.com/QwenLM/qwen-code)
04
Nemotron
Lightning for routed local execution; larger models for planning and orchestration.
**16 GB**
Nano 4B Q4 · 2.8 GB
**32 GB**
Lightning 30B‑A3B NVFP4 · 20.1 GiB before runtime and cache
**64 GB**
Nano 30B‑A3B Q8 · 34 GB; Lightning BF16 is too tight for agent use
**128 GB**
Super 120B‑A12B low-bit · backend-specific; verify support
04 · Current NVIDIA specialist
### Lightning is the local execution worker; Super and Ultra move up the planning ladder.
**Nemotron 3.5 Lightning** is a 30B-total / 3B-active hybrid MoE released for high-volume agent execution. Its official NVFP4 repository is about 20.1 GiB, and NVIDIA documents DGX Spark, Jetson and GeForce RTX 5090 routes plus data-centre serving. Treat 32 GB as the practical personal-system floor: runtime state, its optional 1.26 GiB DSpark drafter and useful context still need memory. The model’s OpenMDW 1.1 terms are not the earlier NVIDIA Open Model License.
**Current family**
Nano · 4B
Nano · 30B / 3B active
Lightning 3.5 · 30B / 3B active
Nano Omni · 30B / 3B active
Super · 120B / 12B active
Ultra · 550B / 55B active
Lightning supports up to a one-million-token model window; available memory and the serving recipe determine the usable context.
**Lightning → Super**
NVIDIA model-card evaluations; harness settings matter
- [MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**81.9**
Lightning BF16
- [MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**83.6**
Super 120B FP8
- [GPQA](https://isaiuseful.com/benchmarks.html.md#gpqa)
**75.4**
Lightning BF16
- [GPQA](https://isaiuseful.com/benchmarks.html.md#gpqa)
**79.4**
Super 120B FP8
- [SWE‑bench Verified](https://isaiuseful.com/benchmarks.html.md#swe-bench)
**51.6**
Lightning BF16
- [Terminal‑Bench 2.1](https://isaiuseful.com/benchmarks.html.md#terminal-bench)
**24.6**
Lightning BF16
- [Lightning launch and local routes →](https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/)
- [Lightning NVFP4 card and recipe →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4)
- [Lightning BF16 card and scores →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16)
- [Use Lightning in a hybrid route →](#routing)
- [Nemotron recipes and data →](https://github.com/NVIDIA-NeMo/Nemotron)
- [Nemotron 3 Nano card →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16)
- [Super card and hardware floor →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8)
- [Ultra deployment tier →](https://isaiuseful.com/cloud-models.html.md#nemotron-3-ultra)
- [NeMo Guardrails →](https://github.com/NVIDIA-NeMo/NeMo-Guardrails)
05
GLM
GLM‑4.7‑Flash for personal machines; GLM‑5.2 for self-hosting; GLM‑5.3 for current hosted long-horizon work.
**16 GB**
Older GLM‑4 9B Q4 · about 5–6 GB
**24 GB**
GLM‑4.7‑Flash 30B‑A3B Q4 · 19 GB · tight
**32 GB**
GLM‑4.7‑Flash Q4 · 19 GB · 16–32K starting context
**64 GB**
GLM‑4.7‑Flash Q8 or BF16 · package-dependent
**748 GB**
GLM‑5.2 · 753B / 40B active · official 465 GB NVIDIA NVFP4 checkpoint
05 · Current mixed deployment
### Choose GLM‑5.3 for hosted work and GLM‑5.2 for self-hosting.
GLM‑5.3 is Z.ai’s newer model for long-running coding and agent tasks. It is available through Coding Plan, while the general API and downloadable weights are still pending. If you need to run GLM yourself today, use GLM‑5.2; its official 465 GB NVFP4 build is aimed at DGX Station-class hardware. On a normal local machine, GLM‑4.7‑Flash is the practical choice.
**Relevant family members**
GLM‑4 · 9B
GLM‑4.7‑Flash · 30B / 3B active
GLM‑5.1 · 744B / 40B active
GLM‑5.2 · 753B / 40B active · weights live
GLM‑5.3 · same base · hosted release
GLM‑5.3 weights, licence and serving recipes remain pending; do not reuse GLM‑5.2 hardware results as GLM‑5.3 measurements.
**GLM‑5.3 launch scorecard**
Z.ai-reported · protocols differ · higher is better
- [Terminal‑Bench 2.1](https://isaiuseful.com/benchmarks.html.md#terminal-bench)
**88.2**
Claude Code 2.1.207; six-hour timeout
- [Terminal‑Bench 3.0](https://isaiuseful.com/benchmarks.html.md#terminal-bench)
**28.3**
avg@3; max effort; 400K context
- [Agents’ Last Exam CLI](https://isaiuseful.com/benchmarks.html.md#agents-last-exam)
**28.5**
Official protocol; max effort; 1M context
- [HLE + tools](https://isaiuseful.com/benchmarks.html.md#hle)
**62.5**
300K context; context management; LLM judge
**GLM‑5.3 migration recipe**
Do not change only the model ID. Z.ai says thinking must remain enabled: replace `thinking.type: "disabled"` with `"enabled"` , then choose `reasoning_effort` `low` , `high` or `max` . Start existing non-thinking workloads at `low` ; use `max` for coding acceptance tests. A disabled-thinking request will fail.
- [GLM‑5.3 announcement, benchmarks and methods →](https://z.ai/blog/glm-5.3)
- [GLM‑5.3 API status and migration guide →](https://docs.z.ai/guides/llm/glm-5.3)
- [GLM Coding Plan and coding-agent setup →](https://docs.z.ai/devpack/overview)
- [GLM‑5 family repository →](https://github.com/zai-org/GLM-5)
- [GLM‑5.2 model card, weights and licence →](https://huggingface.co/zai-org/GLM-5.2)
- [Official NVIDIA GLM‑5.2 NVFP4 checkpoint →](https://huggingface.co/nvidia/GLM-5.2-NVFP4)
- [GLM‑5.2 DGX Station measurements →](https://isaiuseful.com/dgx-station.html.md#sweet-spot)
- [GLM‑4.7‑Flash model card →](https://huggingface.co/zai-org/GLM-4.7-Flash)
06
DeepSeek
Older distilled reasoning baselines locally; current frontier releases belong on servers.
**16 GB**
R1‑0528 Qwen3 8B Q4 · 5.2 GB
**24 GB**
R1 Distill 32B Q4 · 20 GB · tight
**32 GB**
R1 Distill 32B Q4 · 20 GB
**64 GB**
R1 Distill Llama 70B Q4 · 43 GB
**Server**
V3.2 and full R1 · 671B total · multi-GPU/cloud
06 · Older reasoning baseline
### Keep the distills as reasoning baselines, not first picks.
The local R1 distills predate the current generation. Use R1‑0528‑Qwen3‑8B for compatibility tests, 32B for the useful middle and 70B on a 64 GB-class machine only when your acceptance set rewards them. DeepSeek‑V3.2 and full R1 are 671B MoE deployments and belong on multi-GPU servers or hosted endpoints.
**R1 releases**
1.5B
7B
8B
14B
32B
70B
671B full
Distilled variants are based on Qwen or Llama checkpoints; V3.2 is a separate 671B MoE family.
**Original R1 distill scaling**
DeepSeek-reported pass@1
- [AIME 2024](https://isaiuseful.com/benchmarks.html.md#aime)
**28.9**
1.5B
- [AIME 2024](https://isaiuseful.com/benchmarks.html.md#aime)
**69.7**
14B
- [AIME 2024](https://isaiuseful.com/benchmarks.html.md#aime)
**72.6**
32B
- [LiveCodeBench](https://isaiuseful.com/benchmarks.html.md#livecodebench)
**16.9**
1.5B
- [LiveCodeBench](https://isaiuseful.com/benchmarks.html.md#livecodebench)
**53.1**
14B
- [LiveCodeBench](https://isaiuseful.com/benchmarks.html.md#livecodebench)
**57.2**
32B
- [GitHub →](https://github.com/deepseek-ai/DeepSeek-R1)
- [Hugging Face 8B →](https://huggingface.co/deepseek-ai/DeepSeek-R1-0528-Qwen3-8B)
- [Ollama size ladder →](https://ollama.com/library/deepseek-r1/tags)
- [LM Studio model →](https://lmstudio.ai/models/deepseek/deepseek-r1-0528-qwen3-8b)
- [Open 8B in LM Studio →](https://lmstudio.ai/deeplink?name=deepseek-r1-0528-qwen3-8b&owner=deepseek)
- [DeepSeek‑R1 paper →](https://arxiv.org/abs/2501.12948)
- [V3.2 model card and scores →](https://huggingface.co/deepseek-ai/DeepSeek-V3.2)
- [DeepEP communication library →](https://github.com/deepseek-ai/DeepEP)
07
Mistral
Older compact local fallback; Small 4 is a separate server-class current option.
**16 GB**
Ministral 3 14B Q4 · roughly 9–10 GB
**24 GB**
Ministral 3 14B Q8 or Devstral 24B Q4
**32 GB**
Devstral Small 24B Q6/Q8 · package-dependent
**64 GB**
High-precision 24B; Small 4 119B needs aggressive low-bit quant
**128 GB**
Small 4 119B low-bit · runtime-specific; server-first
07 · Older compact fallback
### Ministral 3 is a compatibility fallback, not the local default.
The 3B, 8B and 14B Ministral models are compact, multimodal and permissively licensed. Mistral Small 4 combines instruct, reasoning and coding modes at 119B total/6.5B active; treat local low-bit builds as experimental until your runtime lists the exact architecture.
**Current main families**
Ministral 3 · 3B
8B
14B
Small 4 · 119B / 6.5B active
Large 3 · 675B / 41B active
Base, instruct and reasoning variants exist; Devstral is the coding-specialist companion line.
**Current GPQA evaluation metadata**
Separate Hugging Face model-card runs; not a controlled head-to-head
- [GPQA Diamond](https://isaiuseful.com/benchmarks.html.md#gpqa)
**71.2**
Small 4 119B
- [GPQA Diamond](https://isaiuseful.com/benchmarks.html.md#gpqa)
**67.2**
Large 3 675B
- [GitHub inference →](https://github.com/mistralai/mistral-inference)
- [Hugging Face Ministral 14B →](https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512)
- [Ollama tags →](https://ollama.com/library/ministral-3/tags)
- [LM Studio model →](https://lmstudio.ai/models/mistralai/ministral-3-14b)
- [Open 14B in LM Studio →](https://lmstudio.ai/deeplink?name=ministral-3-14b&owner=mistralai)
- [Ministral 3 paper →](https://arxiv.org/abs/2601.08584)
- [Small 4 card and score →](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603)
- [Large 3 card, score and deployment →](https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512)
- [Mistral fine-tuning toolkit →](https://github.com/mistralai/mistral-finetune)
08
Kimi / Moonshot
Older small research checkpoints locally; Kimi K3 is a current server-scale release.
**16 GB**
Moonlight 16B‑A3B · community 4-bit · 8K maximum context
**32 GB**
Moonlight 16B‑A3B at higher precision; still an 8K research model
**64 GB**
Kimi Linear 48B‑A3B community 4-bit · runtime-specific · test 32–128K first
**128 GB**
Kimi Linear 48B‑A3B BF16-class fit · leave room for context and runtime
**K3**
2.8T · 1M model limit · weights live · multi-GPU
08 · Older local research line
### Small Moonshot models are older research choices; Kimi K3 is server-scale.
**Kimi K3** now has released weights, but its 2.8T parameters, native vision and one-million-token model limit still make it a multi-GPU or hosted deployment—not a personal-machine recommendation. Moonlight 16B‑A3B and experimental Kimi Linear 48B‑A3B are the smaller downloadable Moonshot options; both predate the current generation and should be treated as research checkpoints, not miniature K3 substitutes.
**Do not collapse these into one size ladder**
Moonlight · 16B / 3B active · 8K
Kimi Linear · 48B / 3B active · 1M model limit
K2.6 · 1T / 32B active
K3 · 2.8T · 1M model limit
Kimi Linear uses custom code and a specialist long-context architecture. Its one-million-token support does not make one million tokens practical on a 64 or 128 GB personal machine.
**Small checkpoints are research choices**
Vendor-reported; not a K3 capability proxy
- [Moonlight MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**42.4**
16B‑A3B; 8K context
- [Moonlight HumanEval](https://isaiuseful.com/benchmarks.html.md#humaneval)
**48.1**
16B‑A3B
- [Kimi Linear MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**51.0**
48B‑A3B; evaluated at 4K
- [Kimi Linear RULER](https://isaiuseful.com/benchmarks.html.md#ruler)
**84.3**
48B‑A3B; evaluated at 128K
- [Official Kimi K3 launch and weight-release date →](https://www.kimi.com/blog/kimi-k3)
- [Play the embedded Kimi K3 introduction →](https://isaiuseful.com/cloud-models.html.md#kimi-k3)
- [Moonshot AI model overview →](https://www.moonshot.ai/)
- [Kimi K2.6 weights →](https://huggingface.co/moonshotai/Kimi-K2.6)
- [Moonlight 16B‑A3B model card →](https://huggingface.co/moonshotai/Moonlight-16B-A3B-Instruct)
- [Kimi Linear 48B‑A3B model card →](https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct)
- [Independent Estonian benchmark including Kimi K2 →](https://huggingface.co/tartuNLP/llama-estllm-prototype-0825)
09
Llama 3 + 4
Older ecosystem baseline with broad quant and compatibility support.
**16 GB**
Llama 3.1 8B Q8 or 3.2 3B high precision
**24 GB**
Llama 3.2 Vision 11B Q8 · package-dependent
**64 GB**
Llama 3.3 70B Q4 · 43 GB
**128 GB**
Llama 4 Scout 109B total · community low-bit quant
09 · Older ecosystem baseline
### Use Llama only when compatibility matters more than current capability.
Llama 3.1 8B remains a safe runtime test and Llama 3.3 70B is the clean 64 GB step. Llama 4 Scout is 109B total/17B active and may fit at 128 GB in community quants, but Meta’s official full-precision path is multi-GPU. Check the custom license and exact runtime.
**Shipped family sizes**
3 · 8B
70B
3.1 · 8B
70B
405B
3.2 · 1B
3B
11B vision
90B vision
3.3 · 70B
4 Scout · 109B / 17B active
4 Maverick · 400B / 17B active
**Llama 4 instruct scores**
Meta model card
- [MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**74.3**
Scout
- [MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**80.5**
Maverick
- [GPQA](https://isaiuseful.com/benchmarks.html.md#gpqa)
**57.2**
Scout
- [GPQA](https://isaiuseful.com/benchmarks.html.md#gpqa)
**69.8**
Maverick
- [LiveCodeBench](https://isaiuseful.com/benchmarks.html.md#livecodebench)
**32.8**
Scout
- [LiveCodeBench](https://isaiuseful.com/benchmarks.html.md#livecodebench)
**43.4**
Maverick
- [GitHub and model cards →](https://github.com/meta-llama/llama-models)
- [Hugging Face Llama 3.3 →](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct)
- [Ollama 70B tags →](https://ollama.com/library/llama3.3/tags)
- [LM Studio 8B →](https://lmstudio.ai/models/meta-llama/llama-3.1-8b)
- [Open 8B in LM Studio →](https://lmstudio.ai/deeplink?name=llama-3.1-8b&owner=meta-llama)
- [Llama 4 Scout card and scores →](https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct)
- [Llama Stack →](https://github.com/meta-llama/llama-stack)
- [Llama Guard 4, Prompt Guard 2 and LlamaFirewall →](https://ai.meta.com/blog/ai-defenders-program-llama-protection-tools/)
10
GPT‑OSS
Older MoE compatibility baseline; useful footprints, no longer a first recommendation.
**16 GB**
20B MXFP4 · 14 GB package · short-context edge fit
**24 GB**
20B MXFP4 · 14 GB · comfortable
**32 GB**
20B MXFP4 · space for longer context
**64 GB**
120B package is 65 GB; stay on 20B here
**128 GB**
120B MXFP4 · 65 GB package
10 · Older compatibility baseline
### Keep 20B and 120B for compatibility tests, not as defaults.
Both prior-year models use native MXFP4 MoE weights and the Harmony response format. Prefer the current top-three families unless GPT‑OSS wins the exact runtime, tool or footprint test. OpenAI says 20B can run within 16 GB and 120B within 80 GB; a 16 GB personal machine still has almost no margin after loading the 14 GB 20B package.
**All GPT‑OSS sizes**
20B · 3.6B active
120B · 5.1B active
**20B → 120B**
Current Hugging Face evaluation metadata
- [MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**73.6**
20B
- [MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro)
**80.8**
120B
- [GPQA Diamond](https://isaiuseful.com/benchmarks.html.md#gpqa)
**58.6**
20B
- [GPQA Diamond](https://isaiuseful.com/benchmarks.html.md#gpqa)
**80.8**
120B
- [GitHub and reference runtimes →](https://github.com/openai/gpt-oss)
- [Hugging Face 20B and scores →](https://huggingface.co/openai/gpt-oss-20b)
- [Hugging Face 120B and scores →](https://huggingface.co/openai/gpt-oss-120b)
- [Ollama tags →](https://ollama.com/library/gpt-oss/tags)
- [LM Studio model →](https://lmstudio.ai/models/openai/gpt-oss-20b)
- [Open 20B in LM Studio →](https://lmstudio.ai/deeplink?name=gpt-oss-20b&owner=openai)
- [Official release and evaluation →](https://openai.com/index/introducing-gpt-oss/)
- [Model-card paper →](https://arxiv.org/abs/2508.10925)
- [GPT‑OSS Safeguard companion models →](https://openai.com/index/introducing-gpt-oss-safeguard/)
Also high-profile
## Four more families—and one separate specialist.
The four language-model families fill useful gaps in reasoning, enterprise licensing, open research and Estonian work. H3 sits outside the agent ranking because it generates audio and video.
16 GB · 14B Q4 9.1 GB
### Microsoft Phi‑4
A compact MIT-licensed reasoning model. Microsoft reports [MMLU](https://isaiuseful.com/benchmarks.html.md#mmlu-pro) 84.8, [GPQA](https://isaiuseful.com/benchmarks.html.md#gpqa) 56.1 and [HumanEval](https://isaiuseful.com/benchmarks.html.md#humaneval) 82.6 for the 14B base instruction model; reasoning, mini and multimodal Phi‑4 variants also ship.
**Sizes**
14B
mini
multimodal
reasoning
- [Phi Cookbook →](https://github.com/microsoft/PhiCookBook)
- [Hugging Face + paper + scores →](https://huggingface.co/microsoft/phi-4)
- [Ollama tags →](https://ollama.com/library/phi4/tags)
- [LM Studio model →](https://lmstudio.ai/models/microsoft/phi-4)
8–32 GB · Apache 2.0
### IBM Granite 4
Enterprise-oriented models with governance disclosures, hybrid Mamba/Transformer variants and compact footprints. Shipped language sizes include Micro, H‑Micro, H‑Tiny, H‑Small and dense 8B; evaluate the exact card because scores vary by variant.
**Sizes**
Micro
H‑Micro
H‑Tiny
H‑Small
8B
- [GitHub, sizes and disclosures →](https://github.com/ibm-granite/granite-4.0-language-models)
- [Hugging Face collection →](https://huggingface.co/collections/ibm-granite/granite-40-models)
- [Ollama tags →](https://ollama.com/library/granite4/tags)
- [Granite Guardian →](https://github.com/ibm-granite/granite-guardian)
16 / 32 GB · fully open research
### AI2 OLMo 3
A rare option with training code, data, checkpoints and detailed recipes. The family ships 7B and 32B Base, Instruct and Think variants with 65,536-token context; use 7B Q4/Q8 at 16 GB and 32B Q4 at 24–32 GB.
**Sizes**
7B
32B
Base
Instruct
Think
- [GitHub training code →](https://github.com/allenai/OLMo)
- [Hugging Face card and scores →](https://huggingface.co/allenai/Olmo-3-1025-7B)
- [OLMo 3 paper →](https://allenai.org/papers/olmo3)
- [Ollama tags →](https://ollama.com/library/olmo3/tags)
Estonian specialist · prototype
### TartuNLP EstLLM 8B
A locally testable Estonian-focused research checkpoint and evaluation suite. The authors clearly label it an early prototype with 4K context and no multi-turn chat; use it as a benchmark and fine-tuning reference, not an automatic production default.
**Size**
8B
BF16 checkpoint
community quants
- [Understand the Estonian benchmark suite →](https://isaiuseful.com/benchmarks.html.md#ifeval)
- [Model card and Estonian tables →](https://huggingface.co/tartuNLP/llama-estllm-prototype-0825)
- [EstLLM paper →](https://arxiv.org/abs/2603.02041)
- [Open the EKI benchmark →](https://mõõdupuu.eki.ee/benchmark/bib_bench)
128 GB Spark specialist · video + stereo audio
### MiniMax H3 is local—but it is not an assistant LLM or a simple BF16 load.
Its local H3‑Base takes text, images, video and audio and generates 4–15-second, 24 fps video with native 32 kHz stereo audio. One released BF16 FL2VA task package is about 134.2 GiB on disk—already beyond Spark’s 128 GB unified capacity before activations. NVIDIA’s single-Spark path fits by using a pruned FP8 DiT, FP8 Qwen3‑VL conditioner and the released VAEs; its vendor benchmark cuts one 5-second, 480p, 50-step job from 710.6 to 181.3 seconds. That is strong fit and optimization evidence, not real-time generation.
The fully local release stops at 768p H3‑Base. MiniMax’s recommended 2K result still calls hosted Context‑IR and Regenerate‑2K modules, and the initial open release omits its native sparse-attention implementation. The community license also excludes the EU, UK, United States and South Korea; users there need a separate MiniMax license rather than treating the public weights as permission to run them.
**Released route**
33B dense DiT
4–15 seconds
768p local base
24 fps
32 kHz stereo
181.3 s Spark benchmark
- [Official model card and local workflows →](https://huggingface.co/MiniMaxAI/MiniMax-H3)
- [NVIDIA GB10 runtime, precision and config →](https://github.com/NVlabs/Sana/tree/sol-engine/models/minimax_h3/GB10)
- [NVIDIA on-device benchmark →](https://nvlabs.github.io/Sana/Sol-Engine/H3-OnDevice/)
- [Read the territorial and commercial license terms →](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE)
- [Request a license for an excluded territory →](https://platform.minimax.io/h3-license)
- [Read the Spark-specific decision note →](https://isaiuseful.com/remote-spark.html.md#minimax-h3-spark)
Capability evidence
## Benchmarks shortlist. Work samples decide.
Compare sizes inside the same family and table first. Then test the failure that costs you money: wrong tool calls, incomplete edits, weak language, invented citations or human cleanup.
Good comparison
### Same family, same card
Gemma 4 E2B → 12B → 31B on the same [MMLU‑Pro](https://isaiuseful.com/benchmarks.html.md#mmlu-pro) table shows a meaningful capability curve.
Qualified comparison
### Same named benchmark
Check version, prompt, pass@1 versus consensus, tool access and whether the result is vendor-reported.
Required comparison
### Your 30–100 tasks
Score multilingual quality, tool selection, argument accuracy, code edits, recovery and the amount of human correction required.
European multilingual proof
### Estonian is the demo, not the border.
Estonian is a useful low-resource, morphologically rich stress test. TartuNLP’s independent table reports [IFEval‑et](https://isaiuseful.com/benchmarks.html.md#ifeval) of 0.756 for Gemma 3 12B and 0.766 for 27B; Llama 3.3 70B scores 0.771 and Kimi K2 0.789. Gemma 3 27B also posts 0.817 on [Grammar‑et](https://isaiuseful.com/benchmarks.html.md#ifeval) and 0.953 on [Word‑Meanings‑et](https://isaiuseful.com/benchmarks.html.md#ifeval) . That is impressive evidence for the candidate list, while every production language pair still needs its own prompts.
**0.756**
Gemma 3 12B ·
- [IFEval‑et](https://isaiuseful.com/benchmarks.html.md#ifeval)
**0.766**
Gemma 3 27B ·
- [IFEval‑et](https://isaiuseful.com/benchmarks.html.md#ifeval)
**0.817**
Gemma 3 27B ·
- [Grammar‑et](https://isaiuseful.com/benchmarks.html.md#ifeval)
**0.953**
Gemma 3 27B ·
- [Word‑Meanings‑et](https://isaiuseful.com/benchmarks.html.md#ifeval)
Inference layer
## First decide how the model will run.
The inference runtime loads weights, allocates context and exposes an API. It is not the agent harness. Confirm acceleration, chat template, structured output and tool-call format before connecting an agent.
Inspect + prototype
### LM Studio
Best first stop for discovery, MLX/GGUF downloads, chat-template inspection and an OpenAI-compatible local API. Check both configured and model-maximum context; they are separate values.
- [LM Studio →](https://lmstudio.ai/)
Simple local service
### Ollama
Fast route to a reproducible pull command and local API for coding harnesses. Inspect the exact tag and context setting: “latest” can hide quantization and file-size differences.
- [Ollama →](https://ollama.com/)
Apple-native
### MLX + llama.cpp
Use MLX when an Apple-silicon-native package exists; use llama.cpp for explicit GGUF control, broader quant choices and portable serving.
- [MLX LM →](https://github.com/ml-explore/mlx-lm)
- [llama.cpp →](https://github.com/ggml-org/llama.cpp)
Workstation/server
### vLLM or SGLang
Prefer these on supported NVIDIA systems when concurrency, continuous batching, tensor parallelism and an always-on endpoint matter.
- [vLLM →](https://docs.vllm.ai/)
- [SGLang →](https://docs.sglang.ai/)
Agent + specification layers
## Then choose who plans, acts and preserves intent.
A harness owns the loop around the model: tools, memory, permissions and retries. A specification workflow stores an agreed plan outside the conversation. Both can call a local inference API; neither makes a weak model reason like a frontier model.
> Visual: Separation between inference, harness and specification workflow
**Visual entries (display order):**
- 01 · inference **LM Studio · Ollama · llama.cpp · vLLM** Loads the model and serves tokens.
- 02 · agent harness **Hermes Agent · OpenClaw** Runs tools, sessions, memory and approval gates.
- 03 · durable intent **OpenSpec · Specflow** Keeps plans, tasks and acceptance criteria outside chat history.
Agent harness
### Hermes Agent
Use Hermes when you want a general agent loop with tools, skills, memory and multiple model-provider endpoints. Point it at the inference server; then evaluate the model through Hermes rather than assuming chat quality transfers.
- [Hermes Agent repository →](https://github.com/nousresearch/hermes-agent)
Agent harness
### OpenClaw
Use OpenClaw for an always-on personal agent, channels and remote clients. Keep tool authority narrower than model capability, authenticate the gateway and treat local inference as one backend—not the harness itself.
- [OpenClaw →](https://openclaw.ai/)
Specification workflow
### OpenSpec
Stores a proposal, tasks and spec deltas in the repository before implementation. This is a clean hand-off: use a stronger model to settle the plan, then give a smaller worker the approved artifacts and relevant files.
- [OpenSpec repository →](https://github.com/Fission-AI/OpenSpec)
Planning methodology
### Specflow
SpecStory’s open methodology moves from intent to roadmap, workplans, execution and refinement. It helps preserve decisions between agent sessions; it does not serve a model or enforce an inference format.
- [Specflow repository →](https://github.com/specstoryai/specflow)
Scale route
## Move from personal memory to a deskside frontier tier.
RTX Spark is offered for slim laptops and small desktops; DGX Spark is a compact dedicated system; DGX Station adds a 748 GB coherent tier. What matters is available memory, memory bandwidth, backend support and whether the model runs on the client or a remote host.
Personal computer
**16–64 GB Mac, laptop or desktop**
Interactive use, private documents and one-user local APIs. LM Studio, Ollama, MLX or llama.cpp.
Spark-class computer · native NVFP4
**128 GB · ≈150–165B planning fit**
The speculative band reserves 20–25% for runtime and cache; NVIDIA's up-to-200B figure is a tighter capacity ceiling. Validate the exact GB10 kernel.
Linked workstations · native NVFP4
**256 GB · ≈295–330B planning fit**
NVIDIA’s ConnectX path advertises up to 405B and demonstrates Qwen‑235B NVFP4. The lower band keeps practical headroom; software and model architecture still decide whether it works.
Station, cluster or cloud · native support varies
**≈860–960B Station planning fit**
DGX Station can plausibly hold a mixed-NVFP4 model in that band across coherent memory. Blackwell datacenter nodes and racks scale further; H100/H200 do not provide native NVFP4 W4A4.
- [Compare licensable models and NVFP4 servers →](https://isaiuseful.com/cloud-models.html.md#nvfp4)
Separate tier · DGX Station
### 748 GB makes the official NVFP4 GLM‑5.2 a single-node fit—not a guaranteed fast one.
Station combines 252 GB of 7.1 TB/s HBM3e with 496 GB of 396 GB/s LPDDR5X. NVIDIA's 465 GB GLM‑5.2 NVFP4 checkpoint fits across the coherent pool, but spills beyond HBM and still needs runtime and cache headroom. The linked Station measurements belong to GLM‑5.2; GLM‑5.3 weights and serving recipes were not public when checked. Station is most compelling when 70–300B models live in HBM, or when local access to a larger low-bit model is more valuable than cloud-scale throughput.
**Sweet-spot rule**
Use Spark for 20–35B dense models and sparse 30–120B MoEs. Consider Station for daily 70–120B high-precision work, 200–400B low-bit work, or controlled 400B–1T experiments. Rent first when that top tier is occasional.
- [Open the Station buyer's guide →](https://isaiuseful.com/dgx-station.html.md)
- [Compare Spark and Station →](https://isaiuseful.com/dgx-station.html.md#comparison)
- [Check the independent evidence gap →](https://isaiuseful.com/dgx-station.html.md#reviews)
**Parameter ceilings are not memory guarantees.** NVFP4 planning bands reserve 20–25% of advertised memory and assume roughly 5.0–5.2 effective bits per stored parameter. Architecture, mixed-precision layers, KV cache, multimodal towers, memory tier and backend support can move the result. A model package can load while leaving too little memory for useful context, tool traffic or concurrency.
- [NVIDIA NVFP4 format and Blackwell support →](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)
- [Official RTX Spark laptops and desktops →](https://www.nvidia.com/en-us/products/rtx-spark/)
- [Official DGX Spark specifications →](https://www.nvidia.com/en-us/products/workstations/dgx-spark/)
- [Official DGX Station specifications →](https://www.nvidia.com/en-eu/products/workstations/dgx-station/)
- [Price NVFP4-capable servers →](https://isaiuseful.com/cloud-models.html.md#nvfp4)
- [Operate a remote Spark safely →](https://isaiuseful.com/remote-spark.html.md)
Method
## What “compatible” means here.
The core catalogue was checked against official model cards, repositories and current Ollama/LM Studio packages in July 2026; dated launch records and specialist routes were rechecked through 15 August 2026.
### Agent-safe
The default package leaves roughly 20–25% of system memory plus a practical starting context target. It is a planning rule, not a throughput guarantee.
### Tight
The weights can load, but context, parallel requests or other applications may trigger swap or out-of-memory errors. Tight fits are labelled in amber and should start with short context.
### Server
The official model card expects multi-GPU or vendor-specific infrastructure, or no verified LM Studio/Ollama local package was found. Cloud-only Ollama listings are not counted as local support.
### Quantization
Q4 sizes are package-level approximations unless the vendor publishes a memory table. QAT, GGUF Q4_K_M, MLX 4-bit, NVFP4 and MXFP4 are not interchangeable. Native NVFP4 W4A4 requires Blackwell-generation Tensor Cores plus a supported kernel; Hopper fallback is not the same compute path.
**Nemotron 3.5 Lightning** weights, licence, model-card scores, repository sizes, DGX Spark recipe and routing evidence were checked on 12 August 2026; NVIDIA’s speed figures and the separate LangChain routing study remain workload-specific evidence. Muse Glimmer’s official model card and GGUF files were rechecked on 15 August 2026; Meta’s RTX 5090 speed data and two early DGX Spark owner reports remain workload- and configuration-specific, with no standardized third-party Muse benchmark yet. Other fast-moving names and the NVFP4 hardware boundary were rechecked on 26 July 2026. **GLM‑5.2** is the available open 753B server-class release; NVIDIA's official 465 GB NVFP4 checkpoint and the creator-run Station measurements are the current local evidence. **GLM‑5.3** availability, API migration requirements and vendor benchmark methods were checked on 15 August 2026; its weights, licence and local serving recipes were not yet public. **Kimi K3** now has downloadable weights as well as product and API access, but its 2.8T scale still keeps it in the hosted or multi-GPU tier; the smaller Moonshot checkpoints are older research lines, not K3 substitutes. Benchmark scores are vendor-reported unless an independent source is named.
The short answer
Gemma 4 first for the strongest current local family. Muse Glimmer second for portable agents. Qwen3.8 third for coding and breadth. Nemotron and GLM cover specialist routes. Keep GPT‑OSS, Llama and the older local lines as baselines—not defaults. Keep room for the work.
- [Choose an agent role](#roles)
- [Design the hybrid route](#routing)
- [Budget the context](#context)
- [Scale to a model cloud](https://isaiuseful.com/cloud-models.html.md)
- [Choose a workflow](https://isaiuseful.com/guides.html.md#chooser)
---
## https://isaiuseful.com/cloud-models (`/cloud-models.html.md`)
# How Do You Build a Private Cloud for Open-Weight AI Models?
Canonical source: [https://isaiuseful.com/cloud-models](https://isaiuseful.com/cloud-models)
Provider build guide · checked 15 August 2026
Start with downloadable weights and a serving stack you can operate inside your own security boundary. Then buy for **memory, interconnect, power, cooling and measured token throughput** —not just peak FLOPS.
- [Choose models](#models)
- [Price servers](#hardware)
- [Estimate users + ROI](#economics)
**14 + 1**
open-weight lines + hosted preview
GLM‑5.3 ACCESS LIVE; WEIGHTS PENDING
**2 custom**
frontier licences need review
QWEN3.8 + K3 LARGE-SCALE MAAS TRIGGERS
**14–600 kW**
NVIDIA system planning range
8-GPU SERVER TO RUBIN ULTRA ROADMAP RACK
Two-minute guide
## Which private AI deployment route should you choose?
This page is collapsed into decision-sized sections. If the workload itself is not proven, start with [one measured adoption workflow](https://isaiuseful.com/adoption.html.md#walk) ; if the gap may need retrieval or changed weights, use the [model-intervention chooser](https://isaiuseful.com/training-models.html.md#chooser) before sizing infrastructure.
- [One private service **Start with a current model that fits one node.** *Nemotron 3.5 Lightning is the cleaner execution launch; benchmark your own workload.*](#nemotron-35-lightning)
- [Owned infrastructure **Buy memory, interconnect and a facility—not peak FLOPS.** *One air-cooled node is a different business from an NVL rack.*](#hardware)
- [Fastest operating route **Rent first when demand, model fit or utilization is uncertain.** *Use a managed API or dedicated capacity before carrying idle hardware.*](#cloud-routes)
- [Procurement gate **Replace every default with a quote and measured throughput.** *Proceed only when the privacy, capacity or support premium pays for ownership.*](#economics)
Recommendation catalogue
## Deploy now, validate next, or retain only as a baseline.
Cards are ordered by present recommendation, not chronology or parameter count. **Current launch** means a 2026 release with usable access and a credible deployment route. **Current option** marks a specialized, superseded or review-gated 2026 line. **Older baseline** marks a pre‑2026 release retained for compatibility or comparison—not as the default. GLM‑5.3 remains a hosted preview until its checkpoint and licence ship; Qwen3.8‑Max and Kimi K3 require their custom commercial terms to be reviewed.
current launch · preferred current release
current option · secondary or specialized
review gate or hosted preview · incomplete route
current self-host · newer hosted model exists
older baseline · retained, not default
✓ Current launch
MiniMax Community License
### MiniMax M3
**A current launch choice for frontier coding and agents.** A native-multimodal 428B-total / 23B-active MoE with a 1M-token model limit and downloadable weights. The repositories total about 795.5 GiB for BF16, 413.3 GiB for MXFP8 and 232.9 GiB for NVIDIA's NVFP4 build. The calculator uses a measured 4× B200 FP4/MTP profile; keep the official 8× B200 recipe as the compatibility fallback and validate the exact four-GPU build before quoting.
**First node** — Calculator: 4× B200 measured cell; official fallback: 8× B200
**Serve with** — vLLM or SGLang; pin MSA, parser and quant paths
**Licence gate** — Attribution + notice; authorization above $20M product/service revenue
- [Official repository →](https://github.com/MiniMax-AI/MiniMax-M3/)
- [Weights, licence and deployment routes →](https://huggingface.co/MiniMaxAI/MiniMax-M3)
**Cloud first; DGX Station second**
The published large-node route is mature enough to trial: vLLM verified M3 on H200, GB200, B300 and AMD MI300/MI350 systems, and InferenceX now publishes a current 4× B200 FP4/MTP throughput sweep. The official NVFP4 recipe is newer and does not yet prove one-Station comfort. Rent the reference shape, lock an acceptance set, then repeat it on the quoted workstation.
- [NVIDIA NVFP4 build and current recipe →](https://huggingface.co/nvidia/MiniMax-M3-NVFP4)
- [InferenceX B200 throughput sweep →](https://inferencex.semianalysis.com/)
- [Assess the single-Station fit →](https://isaiuseful.com/dgx-station.html.md#minimax-m3)
✓ Current launch
Custom Qwen3.8‑Max licence
### Qwen3.8‑Max / 2.4T‑A95B
Qwen has released the self-hostable 2.4T-total / 95B-active MoE in full and block-scaled FP8 packages: about 4.89 TB and 2.50 TB of model files respectively. The open checkpoint is text-only, requires thinking mode, has a native 262,144-token context and can extend to 1,010,000. Do not treat it as a byte-for-byte copy of hosted Qwen3.8‑Max, which adds vision input, non-thinking mode, built-in tools and a default 1M context.
**First FP8 service cell** — 16 GPUs; SGLang verifies 4× GB300 NVL4 trays
**Serve with** — SGLang, vLLM or TokenSpeed release recipes
**Weight licence** — Custom terms—not Apache 2.0
- [Full checkpoint, model card and licence →](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B)
- [Official block-scaled FP8 checkpoint →](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8)
- [vLLM Qwen3.8 deployment recipe →](https://recipes.vllm.ai/Qwen/Qwen3.8-2.4T-A95B)
- [SGLang verified GB300 FP8 recipe →](https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8#hw=gb300&variant=default&quant=fp8&strategy=balanced&nodes=multi-4)
- [TokenSpeed 16‑GPU recipes →](https://lightseek.org/tokenspeed/recipes/models#qwen3-8)
- [Compare the hosted Max product →](https://www.qwencloud.com/models/qwen3.8-max)
- [Video: Meet Qwen3.8-Max: A New Bar for Coding and Cowork.](https://www.youtube.com/watch?v=CKlK-KDFKjM)
- [Video: 6 days autonomous coding](https://www.youtube.com/watch?v=VG2OWqBk0Gk)
- [Video: Dynamic workflows to quant strategies](https://www.youtube.com/watch?v=oJgZivsyKO4)
- [Video: Visual agentic intelligence](https://www.youtube.com/watch?v=ByRx4xSm-tM)
- [Video: Cowork: Any role, build beyond](https://www.youtube.com/watch?v=XSWOREA1d9I)
✓ Current launch
OpenMDW 1.1
### NVIDIA Nemotron 3.5 Lightning
A 30B-total / 3B-active hybrid Mamba-2, attention and MoE model for high-volume agent execution. The official NVFP4 weights total about 20.1 GiB before runtime and cache; NVIDIA documents one DGX Spark or one H100 as deployment paths, alongside Jetson, GeForce RTX 5090 and hosted/data-centre routes. Its one-million-token model limit is not a practical context promise for every device.
**First node** — 1× DGX Spark for the documented local recipe; 1× H100 for data-centre serving
**Serve with** — vLLM; broader ecosystem support still needs exact-version tests
**Licence fee** — €0 / $0; review OpenMDW 1.1 obligations
- [NVIDIA launch, local routes and vendor benchmarks →](https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/)
- [Official NVFP4 card and deployment recipe →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4)
- [BF16 reference weights and evaluation table →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16)
- [Design the local-to-frontier route →](https://isaiuseful.com/local-models.html.md#routing)
✓ Current launch
NVIDIA Open Model License
### NVIDIA Nemotron 3 Super
A fully open 120B total / 12B active hybrid MoE for multi-agent work, with weights, datasets, recipes and a 1M-token model limit. Its native NVFP4 build loads on one B200; a matched 8K/1K vLLM run provides a stronger starting profile than a generic model-size estimate.
**First node** — 1× B200 NVFP4; 2× H100/B200 for FP8
**Serve with** — vLLM, SGLang, TensorRT‑LLM or NIM
**Licence fee** — €0 / $0
- [NVIDIA release and deployment cookbooks →](https://developer.nvidia.com/blog/introducing-nemotron-3-super-an-open-hybrid-mamba-transformer-moe-for-agentic-reasoning/)
- [Review the measured B200 profile →](https://lambda.ai/inference-models/nvidia/nvidia-nemotron-3-super-120b-a12b)
- [Current OpenRouter benchmark price →](https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b/pricing)
✓ Current launch
OpenMDW 1.1
### NVIDIA Nemotron 3 Ultra
NVIDIA’s 550B total / 55B active frontier orchestration model. The official mixed-precision NVFP4 checkpoint is about 352.3 GB and supports up to 1M context. Blackwell runs native W4A4; Hopper lacks native FP4 Tensor Cores and automatically uses a W4A16 fallback.
**First node** — 4–8 H200/B200-class GPUs
**Serve with** — vLLM or NVIDIA-optimized stack
**Licence fee** — €0 / $0
- [Official NVFP4 model card →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4)
- [NVIDIA quantization and footprint notes →](https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/)
- [Review the four-GPU decode profile →](https://lambda.ai/inference-models/nvidia/nemotron-3-ultra)
- [Current OpenRouter benchmark price →](https://openrouter.ai/nvidia/nemotron-3-ultra-550b-a55b/pricing)
✓ Current launch
MIT
### DeepSeek V4 Flash 0731
The official Flash release supersedes the Preview checkpoint. Its 166.9 GB mixed FP4/FP8 repository includes the DSpark speculative module, supports 1M context and up to 384K output, and exposes low, high and max reasoning effort. DeepSeek reports large agentic-benchmark gains—including results above V4 Pro Preview—but its code-agent scores use a model-specific harness and two listed tests are internal, so validate the gain on your own stack.
**First node** — Official reference: 4× GB300; profile other layouts
**Serve with** — vLLM or SGLang with DSpark enabled
**Licence fee** — €0 / $0
- [Official model card and MIT licence →](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731)
- [Official API features and pricing →](https://api-docs.deepseek.com/quick_start/pricing)
- [vLLM deployment recipe →](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Flash?hardware=b300&features=tool_calling,reasoning)
- [SGLang deployment cookbook →](https://docs.sglang.io/cookbook/autoregressive/DeepSeek/DeepSeek-V4)
✓ Current launch
Kimi K3 License
### Kimi K3
Moonshot’s released 2.8T-parameter multimodal MoE has 104B active parameters, selects 16 of 896 experts, supports a 1M context and uses MXFP4 weights with MXFP8 activations. Its 96-shard weight index totals 1.56 TB before runtime and cache headroom.
**First node** — 8× B300 or 8× MI350X/MI355X
**Serve with** — SGLang, vLLM or TokenSpeed; validate recipes
**Licence trigger** — Separate deal for large Model-as-a-Service operators
- [Official weights, model card and licence →](https://huggingface.co/moonshotai/Kimi-K3)
- [Official technical report and repository →](https://github.com/MoonshotAI/Kimi-K3)
- [SGLang hardware recipes and validation notes →](https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K3)
- [Current Kimi K3 market-price reference →](https://openrouter.ai/moonshotai)
- [Video: Meet Kimi K3](https://www.youtube.com/watch?v=bn0atstgavo)
△ Current option
Apache 2.0
### Qwen3.5 397B‑A17B
A still-current alternative now ranked behind Qwen3.8. Multimodal MoE with 397B total / 17B active parameters. The calculator now uses NVIDIA's roughly 251 GB NVFP4 repository and a measured 4× B300 SGLang/MTP profile at 8K input / 1K output, replacing the much larger 807 GB BF16 planning route.
**First service cell** — 4× B300 NVFP4 for the measured profile
**Serve with** — SGLang with the pinned NVFP4/MTP path
**Licence fee** — €0 / $0
- [Official model card →](https://huggingface.co/Qwen/Qwen3.5-397B-A17B)
- [NVIDIA NVFP4 build →](https://huggingface.co/nvidia/Qwen3.5-397B-A17B-NVFP4)
- [Inspect the measured InferenceX profile →](https://inferencex.semianalysis.com/)
! Hosted preview
Licence pending
### GLM‑5.3
A post-trained successor that uses the same base model as GLM‑5.2. Z.ai reports large gains on long-horizon coding and agent benchmarks, but the launch status is split: Coding Plan access is live, the general API is marked coming soon and weights are targeted for release within two weeks of 14 August. Keep GLM‑5.2 for self-hosted planning until the 5.3 checkpoint, licence and engine recipes can be inspected.
**Available now** — GLM Coding Plan and ZCode
**API change** — Thinking required; low, high or max effort
**Self-hosting** — Wait for weights, licence and measured recipes
- [Official launch, score table and evaluation methods →](https://z.ai/blog/glm-5.3)
- [Official API status and migration requirements →](https://docs.z.ai/guides/llm/glm-5.3)
- [Coding Plan and coding-agent setup →](https://docs.z.ai/devpack/overview)
- [Compare 5.3 capability with 5.2 deployment →](https://isaiuseful.com/local-models.html.md#glm)
↻ Current self-host
MIT
### GLM‑5.2
A 753B model with up to 1M context. The BF16 repository is about 1.51 TB and needs at least 2 TB aggregate HBM with useful headroom; the calculator instead uses the released FP8 build and a matched 8× B200 profile.
**First node** — 8× B200 for FP8; ≥2 TB HBM for BF16
**Serve with** — vLLM or SGLang
**Licence fee** — €0 / $0
- [Official model card, files and licence →](https://huggingface.co/zai-org/GLM-5.2)
- [Review the measured B200 FP8 profile →](https://lambda.ai/inference-models/zai-org/glm-5.2)
△ Current fallback
Apache 2.0
### Mistral Small 4
A current compact fallback rather than the catalogue default; validate its vendor comparisons on your own tasks. 119B total / 6.5B active, multimodal, 256K context and 24-language support. Its official 70.8 GB NVFP4 build is a clean Blackwell fit; SGLang documents TP1 on B200/B300, while FP8 needs TP2 on H100/H200. Native W4A4 acceleration requires Blackwell.
**First node** — 1× B200 with NVFP4; benchmark before scaling
**Serve with** — SGLang or vLLM; validate the exact NVFP4 kernel
**Licence fee** — €0 / $0
- [Official model card →](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603)
- [Official NVFP4 build →](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603-NVFP4)
- [SGLang deployment topology →](https://docs.sglang.io/cookbook/autoregressive/Mistral/Mistral-Small-4)
! Current · review
Modified MIT
### Kimi K2.7 Code
A downloadable trillion-parameter coding and agent model with 32B active parameters, native INT4 and 256K context. The calculator uses a matched 8× B200 vLLM run; the modified terms, nightly runtime and model-specific kernels still need review.
**First node** — 8× B200 at native INT4
**Serve with** — vLLM, SGLang or KTransformers
**Licence fee** — €0 / $0
- [Official model card and modified licence →](https://huggingface.co/moonshotai/Kimi-K2.7-Code)
- [Review the measured B200 profile →](https://lambda.ai/inference-models/moonshotai/kimi-k2.7-code)
↓ Older · validate
MIT
### DeepSeek V3.2
A prior-year reasoning baseline retained for comparison and existing deployments. A 685B FP8 reasoning and agent model with a 690 GB repository. The card documents vLLM and SGLang; production tool parsing still needs your own robustness tests.
**First node** — 8× H200-class; pilot other stacks
**Serve with** — vLLM or SGLang
**Licence fee** — €0 / $0
- [Official model card and MIT licence →](https://huggingface.co/deepseek-ai/DeepSeek-V3.2)
- [Review the measured H200 sweep →](https://docs.gpustack.ai/2.0/performance-lab/deepseek-v3.2/h200/)
↓ Older baseline
Apache 2.0
### Mistral Large 3
A prior-year Mistral baseline, not a first recommendation against current releases. 675B total / 41B active, multimodal and 256K context. Mistral documents FP8 on one 8× H200 or B200 node and an NVFP4 checkpoint on one 8× H100 or A100 node. H100 and A100 do not have native FP4 Tensor Cores, so that older-GPU route depends on model-specific fallback or dequantization—not native NVFP4 W4A4 acceleration.
**First node** — 8× B200 for native NVFP4; 8× H200 for FP8
**Serve with** — vLLM; pin the hardware-specific path
**Licence fee** — €0 / $0
- [Official model card and deployment notes →](https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512-BF16)
↓ Older baseline
Apache 2.0
### gpt-oss-120b
Retained as an older compatibility and footprint baseline; start with current candidates unless this exact package wins your acceptance tests. About 117B total / 5.1B active, tool-capable and packaged in MXFP4 to fit one 80 GB H100. Minimum fit is not the same as an economical service profile: the calculator uses a measured 2× B200, 8K/1K TensorRT‑LLM run so packed capacity is grounded in observed throughput.
**First service cell** — 2× B200 for the measured capacity profile; H100 fits the build
**Serve with** — TensorRT‑LLM, vLLM or supported reference stack
**Licence fee** — €0 / $0
- [Official model card and licence →](https://huggingface.co/openai/gpt-oss-120b)
- [Inspect the measured InferenceX profile →](https://inferencex.semianalysis.com/)
**Commercially usable does not mean restriction-free.** Before listing any model, retain the exact licence and notice files, scan model code, document provenance limits, red-team the served build and define an abuse policy. Qwen3.8’s custom licence requires the notice to travel with copies; commercial products or services above 100 million monthly active users or $20 million monthly revenue must display the model name, while Model-as-a-Service or AI Work Assistant businesses above $50 million aggregate revenue in any consecutive 12 months need a separate Qwen licence. MiniMax M3 is also open-weight under custom terms, not Apache 2.0 or MIT: commercial deployments require visible attribution and a one-time notice, and products or services above the licence's $20M annual-revenue threshold require prior written authorization. Architecture, quantization, context and concurrency determine the real memory requirement.
NVIDIA platform ladder
## Start with platform scale—then see what it takes to become a provider.
The ten current and roadmap tiers keep node and rack choices legible. The provider-scale view then connects those purchases to research, open-model labs, frontier fleets and the EU and US public routes that can help bridge the gaps.
The ten-name map
### Five node tiers. Four NVL72 racks. One NVL144 rack.
Read the node class first, then the rack class. Open the card for the NVIDIA platform map, videos, NVFP4 guidance, prices, Rubin analysis and OEM alternatives.
H100 through GB300 are shipping systems; **Vera Rubin NVL72** is in its production ramp with preliminary published specifications. **V300** uses the later Rubin Ultra roadmap's 576 GB-per-GPU assumption; **VB300 NVL72** is a 72‑GPU comparison domain derived from half a Kyber rack, not an announced NVIDIA SKU. Kyber NVL144 remains the actual Rubin Ultra roadmap rack label.
8-GPU node · Hopper
**H100 → H200**
The mature entry tier. H200 keeps the same node shape and raises accelerator memory from 640 GB to 1.13 TB.
8-GPU node · newer
**B200 → B300 → V300**
More memory and newer low-precision engines; V300 is a derived roadmap planning node rather than an orderable SKU.
72-GPU NVLink rack
**GB200 → GB300 → Vera Rubin → VB300**
Vera Rubin is the current production-ramp generation. VB300 is a later, derived Rubin Ultra comparison—not the name of Vera Rubin NVL72.
144-GPU roadmap rack
**Kyber NVL144**
The Rubin Ultra/V300 roadmap tier changes rack density, power delivery and price class again.
From number format to facility
### See the platform NVIDIA is describing—then translate the spectacle into a bill of materials.
These vendor videos connect NVFP4, Vera, Rubin, DSX facilities, a scientific workload and the full GTC keynote. Use them to understand the intended system shape; use the tables below for procurement questions, facility constraints and explicit planning caveats.
- [Video: What Is NVFP4? Faster LLM Inference Without Losing Quality](https://www.youtube.com/watch?v=UTfg-_EGurw)
Low precision · NVIDIA Developer
### NVFP4 can make large models smaller—but the GPU matters
NVFP4 stores most values in four bits and uses fine-grained scaling to preserve more accuracy than a crude four-bit conversion. The catch is simple: native W4A4 acceleration starts with Blackwell Tensor Cores. Hopper can use selected checkpoints through a W4A16 fallback, but that is not the same speed path.
- [Video: NVIDIA Vera Rubin Platform Ramping into Full Production | Built for the Era of Agents](https://www.youtube.com/watch?v=jMZgjAVR7bo)
Platform overview · NVIDIA
### Vera Rubin is a full system, not just a GPU
Use the overview to see how NVIDIA frames CPUs, GPUs, networking and software as one agent platform. The delivered configuration still needs an exact vendor quote and acceptance test.
- [Video: NVIDIA Vera—The CPU for Agents](https://www.youtube.com/watch?v=vLfrBembjsk)
Host architecture · NVIDIA
### The CPU still shapes the service
Vera is NVIDIA's host-side story for agent systems. Translate that promise into memory bandwidth, data movement, storage and software requirements for the workload you will actually serve.
- [Video: NVIDIA DSX Powers Gigawatt‑Scale AI Factories at Maximum Efficiency](https://www.youtube.com/watch?v=cf40vNN5_Js)
Facility blueprint · NVIDIA
### At rack scale, the building joins the stack
DSX makes the facility-level ambition visible. Power delivery, cooling, networking, operations and recovery are part of the product long before a gigawatt becomes relevant.
- [Video: Advancing Scientific Discovery in the Agentic AI Era](https://www.youtube.com/watch?v=Il4dhCv0Li0)
Workload context · NVIDIA
### Start with the scientific job, not the rack
The discovery story is a useful demand-side counterweight to hardware spectacle. Define the models, data, latency and evaluation first; only then choose the infrastructure tier.
- [Video: NVIDIA GTC Keynote 2026](https://www.youtube.com/watch?v=jw_o0xr8MWU)
Full keynote · NVIDIA
### See the complete platform story in one sitting
The full GTC 2026 keynote connects Vera Rubin, agents, networking and AI factories in NVIDIA’s own long-form narrative. Use it for context, then return to the evidence tables for procurement decisions.
NVFP4 in plain English
### A promising format with one hard boundary: native acceleration needs newer NVIDIA GPUs.
NVFP4 is not a universal “make any model four times faster” switch. It reduces stored weights and memory traffic, but the checkpoint, serving engine, kernels, context and GPU generation must all agree.
Promising · hardware-gated
### Smaller weights can mean a larger model, more replicas or more cache.
NVFP4 groups 16 four-bit values under a higher-precision scale. NVIDIA reports about 4.5 bits per quantized value including block-scale overhead, roughly 3.5× less storage than FP16 and 1.8× less than FP8. Real checkpoints stay mixed precision: Nemotron 3 Ultra is 352.3 GB rather than the 309.4 GB that pure 4.5-bit math would suggest.
#### Read the labels this way
- **Native NVFP4:** Blackwell or Blackwell Ultra can execute W4A4 through FP4 Tensor Cores, subject to a supported kernel.
- **No native NVFP4:** Hopper H100/H200 lacks FP4 Tensor Cores. A model-specific W4A16 fallback may preserve weight-memory savings, but activations use 16-bit math.
- **Roadmap assumption:** the capacity arithmetic is useful, but support is not a shipping-product promise.
Same B300 core · different integration boundary
### HGX vs DGX: choose who integrates, licenses and supports the node.
**HGX B300** is NVIDIA’s eight-GPU baseboard and interconnect; an OEM turns it into a server with the CPU, RAM, storage, chassis, cooling, firmware, BMC and hardware-support route. **DGX B300** is NVIDIA’s integrated system baseline with DGX OS preinstalled. Compare the delivered configuration, entitlement certificate, facility fit and support path—not just the GPU label.
Why OEM HGX
#### Fit the node into the fleet you already operate.
OEMs can supply liquid cooling and rack density, preferred CPU, memory, storage and networking choices, and standard fleet controls such as iLO, iDRAC or XClarity where applicable. They can also offer regional procurement, warranty, spares and existing vendor-contract terms—sometimes at a lower or differently structured price. NVIDIA certification validates systems, but hardware support normally comes through the OEM or channel.
Why DGX
#### Buy a more tightly integrated NVIDIA baseline.
DGX B300 arrives with Ubuntu, NVSM, DCGM, driver/CUDA, Docker, Container Toolkit and DOCA-OFED/MST in its DGX OS stack; partner or NVIDIA field installation and entitlement registration are part of the supported path. That can simplify accountability, updates and escalation, while an OEM build needs more configuration and entitlement diligence.
| Boundary | DGX B300 | OEM HGX B200/B300 vs OEM GB200/GB300 NVL72 |
| --- | --- | --- |
| Base software | DGX OS stack is preinstalled with the system. | OEM image and fleet tooling vary; NVIDIA AI Enterprise can be supported, but enterprise support is an optional purchase. |
| NVIDIA AI Enterprise | Blackwell DGX licenses are purchased separately. Hopper DGX systems include NVIDIA AI Enterprise in the DGX software bundle; this does not carry into Blackwell. | Purchase and metric are separate unless the quote explicitly includes them. The five-year exceptions are only eligible H100 PCIe/NVL and H200 NVL GPUs in NVIDIA-Certified systems (A800 40 GB Active: three years), not generic HGX SXM or this Blackwell hardware; activation ties the term to the selected GPU. |
| Mission Control | Separately purchased or entitled unless the exact order says otherwise. It is recommended for DGX B200/B300 and required for GB200/GB300 NVL72 management. | The current Mission Control 2.3.x support matrix excludes OEM HGX B200/B300. Its OEM/partner support applies to listed GB200/GB300 NVL72 deployments; verify the exact release and configuration. |
| Mission Control footprint | For supported B300 deployments, the current requirements/support matrix lists ten separate control-plane nodes; include them in facility and TCO planning. | The same ten-node control-plane requirement is listed for supported GB-series deployments—not for the excluded OEM HGX B200/B300 path. |
| Support accountability | NVIDIA support requires the applicable active entitlement; exact inclusions depend on the order. | Hardware support is through OEM/channel; software, integration and escalation can be split across vendors. For an included eligible-GPU term, start is the OEM board ship date plus 90 days. |
- [NVIDIA Certified Systems validation →](https://docs.nvidia.com/certification-programs/latest/nvidia-certified-systems.html)
- [NVIDIA HGX AI Factory reference architecture →](https://docs.nvidia.com/enterprise-reference-architectures/hgx-ai-factory/latest/overview.html)
- [DGX B300 base stack and installation →](https://docs.nvidia.com/dgx/dgxb300-user-guide/introduction-to-dgxb300.html)
- [NVIDIA AI Enterprise licensing guide →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/licensing.html)
- [NVIDIA Mission Control requirements and support matrix →](https://docs.nvidia.com/nvidia-mission-control/)
Commercial platform path · checked 10 August 2026
### Hardware is only one layer of the NVIDIA commercial stack.
Open the stack to see which layers are software products, which are included system software and which require a separate commercial entitlement.
Order-line warning
### Blackwell DGX hardware alone does not grant a lifetime NVIDIA AI Enterprise or Mission Control license.
The exact order and entitlement control the software, support level and term. Selected H100 PCIe/NVL and H200 NVL GPUs include a five-year NVIDIA AI Enterprise subscription, but it is time-limited—not lifetime—and does not establish an included entitlement for HGX SXM or current Blackwell DGX/HGX. Hopper-generation DGX systems include NVIDIA AI Enterprise in the DGX software bundle; Blackwell DGX B200/B300 and GB200/GB300 systems require separate purchase unless the exact order or EC says otherwise. NVIDIA AI Enterprise is per GPU: subscriptions include support during their term; perpetual use is indefinite with five years of Business Standard support, renewable thereafter. Current list pricing is $4,500 per GPU for one year, $18,000 per GPU for a discounted five-year subscription, or $22,500 per GPU perpetual with five years of Business Standard support: for eight GPUs, that is list-price arithmetic of $36,000/year, $144,000/5 years, or $180,000 perpetual plus five-year support.
There is no public free-forever entitlement for the complete supported production suite. The general production trial is 90 days, includes Omniverse, excludes Run:ai and has support governed by its offer and Entitlement Certificate (EC); it is an evaluation route, not a production purchase.
- [NVIDIA AI Enterprise licensing exceptions and terms →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/licensing.html)
- [NVIDIA AI Enterprise pricing and trial routes →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/pricing.html)
> Visual: Layered architecture · links open NVIDIA documentation. One possible commercial stack—not a bundled bill of materials. Products can be selected independently, and support or entitlement depends on the exact order and supported configuration.
Layered architecture · links open NVIDIA documentation
**One possible commercial stack—not a bundled bill of materials.**
Products can be selected independently, and support or entitlement depends on the exact order and supported configuration.
Outcomes and workloads
- Applications
- Agents
- Model APIs
- Analytics
- Training
Application development
#### Build, customize and prepare workloads
Developer-facing tools; availability may depend on the chosen software entitlement.
- [**NIM** Model-serving APIs Packages model-serving APIs for deploying supported models.](https://docs.nvidia.com/nim/)
- [**NeMo Framework** Model development tools Builds, customizes and evaluates models and related workflows.](https://docs.nvidia.com/nemo-framework/user-guide/latest/overview.html)
- [**RAPIDS** Accelerated data science Accelerates GPU-backed data science and analytics workflows.](https://docs.rapids.ai/)
- [**Omniverse** Development and simulation platform Free for development, production and redistribution; community support is free, while enterprise support needs the applicable subscription.](https://docs.omniverse.nvidia.com/dev-guide/latest/common/NVIDIA_Omniverse_License_Agreement.html)
Inference serving
#### Optimize, expose and coordinate model responses
Serving components are separate choices; a workload does not need all of them.
- [**TensorRT-LLM** Optimized LLM inference Optimizes large-language-model inference on NVIDIA GPUs.](https://nvidia.github.io/TensorRT-LLM/)
- [**Triton Inference Server** Model-server endpoints Exposes model servers through production-oriented inference endpoints.](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/)
- [**Dynamo** Distributed inference coordination Coordinates distributed inference across GPUs and serving components.](https://docs.nvidia.com/dynamo/)
Infrastructure management
#### Schedule workloads and operate supported infrastructure
Entitlement-sensitive layers; confirm the configuration, term and support route.
- [**Run:ai** GPU scheduling and governance Schedules and governs shared GPU workloads.](https://docs.nvidia.com/run-ai/index.html)
- [**Mission Control** Supported AI-factory operations Operates supported AI-factory deployments; entitlement and support are configuration-specific.](https://docs.nvidia.com/nvidia-mission-control/)
- [**Base Command Manager (BCM)** Cluster provisioning Provisions and manages supported compute clusters.](https://docs.nvidia.com/base-command-manager/index.html)
- [**UFM** InfiniBand fabric management Manages InfiniBand fabrics used by supported systems.](https://networking-docs.nvidia.com/ufmenterpriseum/6221)
- [**NetQ** Network observability Observes and troubleshoots network state and faults.](https://docs.nvidia.com/networking-ethernet-software/cumulus-netq/)
System software
#### Operating system, GPU platform and container access
DGX OS is a system foundation; it is not itself a production-software entitlement.
- [**DGX OS** Supported DGX OS stack The supported operating-system stack for DGX systems.](https://docs.nvidia.com/dgx/dgx-os-7-user-guide/)
- [**DCGM** GPU health and telemetry Supplies GPU telemetry, health checks and diagnostics.](https://docs.nvidia.com/datacenter/dcgm/latest/index.html)
- [**CUDA** GPU programming platform Provides the GPU programming platform and libraries.](https://docs.nvidia.com/cuda/)
- [**NVIDIA Container Toolkit** GPU-enabled containers Makes NVIDIA GPUs available to container workloads.](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/)
System and hardware foundation
- DGX or certified systems
- NVIDIA GPUs
- Networking
- Storage
- Active support where purchased
Deploy anywhere
- Cloud
- Data center
- Edge
- Local workstation
- [Licensing, support and Blackwell DGX distinction →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/licensing.html)
- [Current NVIDIA AI Enterprise list pricing →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/pricing.html)
- [NVIDIA DGX platform purchase boundary →](https://www.nvidia.com/en-us/data-center/dgx-platform/)
Operator-assembled alternative
### An open-source control plane can lower license fees—but moves operations to you.
Open the diagram for a realistic surrounding stack, including the vendor dependencies that remain at the hardware boundary.
> Visual: Operator-assembled architecture · examples, not a required bill of materials. Open components, explicit operator ownership. Each layer is replaceable. The engineering team owns the integration, upgrades, recovery path and production support boundary.
Operator-assembled architecture · examples, not a required bill of materials
**Open components, explicit operator ownership.**
Each layer is replaceable. The engineering team owns the integration, upgrades, recovery path and production support boundary.
Outcomes and workloads
- Applications
- Agents
- Model APIs
- Analytics
- Training
Applications and agents
#### Compose the user-facing workload and its stateful agent paths
These tools supply interfaces and orchestration; the operator still owns identity, permissions and durable business state.
- [**LibreChat** Self-hosted model interface Provides a multi-user interface for local and hosted models, agents and MCP tools.](https://www.librechat.ai/docs)
- [**LangGraph** Stateful agent orchestration Builds explicit stateful agent graphs, checkpointing and human-in-the-loop paths.](https://langchain-ai.github.io/langgraph/)
Model serving and gateways
#### Expose models through an API contract you own
Choose a serving engine and gateway deliberately; the workload does not need every option.
- [**vLLM** GPU model serving Serves supported models through a high-throughput, OpenAI-compatible API.](https://docs.vllm.ai/en/stable/)
- [**SGLang** Distributed model serving Runs language and multimodal models from one GPU to distributed clusters.](https://docs.sglang.io/)
- [**LiteLLM Proxy** OpenAI-compatible gateway Routes, authenticates and meters requests across local and hosted endpoints.](https://docs.litellm.ai/)
Training, evaluation and lifecycle
#### Train, evaluate and promote reproducible model artifacts
Model code, adapters, datasets, evaluation and registry evidence are separate operating responsibilities.
- [**PyTorch** Training framework Provides the core tensor and training framework for model and adapter development.](https://pytorch.org/docs/stable/index.html)
- [**Transformers + PEFT + TRL** Models, adapters and training loops Supplies model implementations, parameter-efficient adapters and supervised or preference-training routes.](https://huggingface.co/docs/transformers/index)
- [**MLflow** Experiment and artifact tracking Tracks run parameters, metrics and artifacts for comparison and promotion evidence.](https://mlflow.org/docs/latest/)
Workload control
#### Pick one primary scheduling and queueing model
Combining schedulers without a clear owner creates competing admission and recovery rules.
- [**Kubernetes** Container orchestration Schedules GPU-enabled container workloads through device plugins and resource requests.](https://kubernetes.io/docs/tasks/manage-gpus/scheduling-gpus/)
- [**Kueue** Kubernetes job queueing Adds quota-aware admission and queueing for batch workloads on Kubernetes.](https://kueue.sigs.k8s.io/docs/)
- [**Volcano** Batch scheduling Adds batch and high-performance workload scheduling to Kubernetes.](https://volcano.sh/docs/home/introduction/)
- [**Slurm** HPC workload manager Allocates accelerators and other generic resources for queued compute jobs.](https://slurm.schedmd.com/gres.html)
- [**NVIDIA GPU Operator + KAI Scheduler** GPU enablement and gang scheduling Installs and manages Kubernetes GPU software while KAI Scheduler adds AI workload and gang scheduling.](https://docs.nvidia.com/gpu-operator/latest/)
Provisioning and configuration
#### Rebuild nodes from documented state
Bare-metal lifecycle and desired configuration are separate responsibilities.
- [**MAAS** Bare-metal provisioning Discovers, commissions and provisions physical machines.](https://maas.io/docs)
- [**Foreman** Host lifecycle management Provisions and manages physical and virtual host lifecycles.](https://docs.theforeman.org/)
- [**Ansible** Configuration automation Applies repeatable configuration and operational automation across nodes.](https://docs.ansible.com/)
Delivery and supply chain
#### Build, promote and retain deployable artifacts
Workflow automation and a private registry still need access control, signing policy and retention rules.
- [**Argo Workflows** Kubernetes-native workflows Runs DAG and step-based workflows as Kubernetes resources.](https://argoproj.github.io/argo-workflows/)
- [**Harbor** Artifact registry Provides a private registry with project-level distribution and retention controls.](https://goharbor.io/docs/)
Observability
#### Own telemetry, dashboards, logs and alerts
Keep sensitive prompts and identifiers out of labels and unbounded log streams.
- [**OpenTelemetry** Telemetry instrumentation Standardizes collection and export of traces, metrics and logs.](https://opentelemetry.io/docs/)
- [**Prometheus** Metrics and alerting Scrapes time-series metrics and evaluates alerting rules.](https://prometheus.io/docs/introduction/overview/)
- [**Grafana** Dashboards and exploration Visualizes and explores operational data from configured sources.](https://grafana.com/docs/grafana/latest/)
- [**Loki** Log aggregation Indexes labels around log streams for storage and investigation.](https://grafana.com/docs/loki/latest/)
Data, network and security
#### Assemble infrastructure controls as separate services
Storage, policy and secrets still need backups, upgrades and tested recovery.
- [**Ceph** Distributed storage Provides object, block and file storage across a managed cluster.](https://docs.ceph.com/en/latest/)
- [**MinIO / AIStor** S3-compatible object storage The current vendor documentation is for AIStor; verify product terms and the archived OSS route.](https://docs.min.io/aistor/)
- [**Cilium** Networking and policy Provides eBPF-based networking, policy and observability for cloud-native workloads.](https://docs.cilium.io/en/stable/)
- [**Vault** Secrets management Controls access to secrets and dynamic credentials under HashiCorp's current terms.](https://developer.hashicorp.com/vault/docs)
Hardware and vendor boundary
#### Open control software does not make the accelerator stack open
Pin and validate the OEM, driver, CUDA, collectives and firmware compatibility chain.
- [**Redfish** Hardware management API Standardizes server-management APIs; OEM implementations and extensions still vary.](https://www.dmtf.org/standards/redfish)
- [**NVIDIA Driver** GPU host driver Connects the operating system and supported NVIDIA accelerator hardware.](https://docs.nvidia.com/datacenter/tesla/driver-installation-guide/)
- [**CUDA** GPU programming platform Provides the NVIDIA GPU programming platform and libraries.](https://docs.nvidia.com/cuda/)
- [**NCCL** Collective communication Coordinates collective communication across supported GPUs and networks.](https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/)
- [**Platform firmware** OEM-specific lifecycle Firmware updates remain platform-specific and inside the vendor support boundary.](https://docs.nvidia.com/dgx/dgxb300-fw-update-guide/)
**Ownership trade:** this route can avoid NVIDIA AI Enterprise and Mission Control fees, but it transfers integration, upgrade validation, incident response, security patching, recovery automation and support ownership to the operator. It is not a claim of feature parity with Mission Control.
- [Find every component in the tools catalogue →](https://isaiuseful.com/tools.html.md#tools-ai-infrastructure)
- [NVIDIA HGX AI Factory boundary →](https://docs.nvidia.com/enterprise-reference-architectures/hgx-ai-factory/latest/overview.html)
- [NVIDIA Certified Systems validation →](https://docs.nvidia.com/certification-programs/latest/nvidia-certified-systems.html)
Vendor explainers · not independent evidence
### PNY Pro explains the commercial software story; verify entitlements against the order.
These vendor videos are useful orientation for NVIDIA AI Enterprise, not proof that a particular DGX or OEM quote includes a license, support term or Mission Control.
- [Video: Introducing NVIDIA AI Enterprise](https://www.youtube.com/watch?v=duemyiNxl4M)
Vendor explainer · PNY Pro
### Introducing NVIDIA AI Enterprise
Use PNY Pro's overview to understand the vendor's platform framing, then confirm the license metric, term and support on the order.
- [Video: NVIDIA AI Enterprise | End-to-End Platform for Production AI](https://www.youtube.com/watch?v=gJ5OvKFStIs)
Vendor explainer · PNY Pro
### Production AI is an entitlement question
The explainer describes the commercial platform. It does not establish a paid production entitlement or its duration for any hardware purchase.
Strong recommendation
### Host the gear in a datacenter.
A 120–600 kW liquid-cooled rack is a facility project before it is a model project. Shortlist colocation providers that already support high-density direct-liquid cooling, redundant power, coolant distribution, carrier-neutral networking, remote hands, physical security and the client’s compliance regime. Get written facility acceptance for the exact NVIDIA/OEM rack before signing the hardware order.
- [NVIDIA provider facility requirements →](https://docs.nvidia.com/dsx/ncp/inference-provider-requirements/home)
- [NVIDIA NVL72 reference architecture →](https://docs.nvidia.com/enterprise-reference-architectures/nvl72-ai-factory/latest/overview.html)
| Platform family | Scale + accelerator memory | Facility envelope | Best planning use | Ballpark, ex VAT |
| --- | --- | --- | --- | --- |
| H100 Available 8‑GPU Hopper node
[Official H100/H200 system specs →](https://docs.nvidia.com/dgx/dgxh100-user-guide/introduction-to-dgxh100.html) | 8× H100 · 640 GB HBM3 ✕ No native NVFP4
Speculative fit: ≈0.74–0.82T parameters through a supported W4A16-style fallback. | **10.2 kW max · air** 8U rack server; ordinary high-density datacenter deployment. | Mature, lower-capex CUDA node for inference, adapters and evaluation. | **≈€263–403k / $300–460k** Broad 2026 complete-system quote band. |
| H200 Available 8‑GPU Hopper node
[Official H100/H200 system specs →](https://docs.nvidia.com/dgx/dgxh100-user-guide/introduction-to-dgxh100.html) | 8× H200 · 1.13 TB HBM3e ✕ No native NVFP4
Speculative fit: ≈1.30–1.45T parameters through a supported W4A16-style fallback. | **10.2 kW max · air** Same DGX node envelope as H100. | Memory-first Hopper choice for larger checkpoints without Blackwell migration. | **≈€350–438k / $400–500k** Broad 2026 complete-system quote band. |
| B200 Available 8‑GPU Blackwell node
[Official B200 system specs →](https://docs.nvidia.com/dgx/dgxb200-user-guide/introduction-to-dgxb200.html) | 8× B200 · 1.44 TB HBM3e ✓ Native NVFP4
Speculative single-checkpoint fit: ≈1.65–1.85T parameters. | **14.3 kW max · air** 1,550 CFM and 48,794 BTU/hr at the system ceiling. | General Blackwell serving and post-training node. | **€438–569k / $500–650k** Public-reseller and integrator planning range. |
| B300 Available 8‑GPU Blackwell Ultra node
[Official B300 system specs →](https://docs.nvidia.com/dgx/dgxb300-user-guide/introduction-to-dgxb300.html) | 8× B300 · 2.3 TB HBM3e ✓ Native NVFP4
Speculative single-checkpoint fit: ≈2.65–2.95T parameters. | **14.5 kW max · air** 49,476 BTU/hr published system ceiling. | Largest current single-node memory tier before NVL72. | **≈€569–744k / $650–850k** Editable 2026 planning band; require an exact OEM quote. |
| V300 Derived 8‑GPU Rubin Ultra roadmap node
[Inspect the V300/Kyber roadmap estimate →](https://wccftech.com/nvidia-rubin-ultra-rack-estimated-to-cost-21-million-hbm4e-swelling-to-1-5m-per-unit/) | 8× V300 · 4.6 TB HBM4e ◇ Roadmap assumption
≈5.3–5.9T parameters if this derived node retains NVFP4-class support. | **≈33 kW · full DLC planning** One eighteenth of the Kyber power target; not a vendor system specification. | Roadmap node comparison before deciding whether NVL72 scale is justified. | **≈€0.9–1.3M / $1.0–1.5M** Derived planning band, not a list price or announced node. |
| GB200 NVL72 Available 72‑GPU Blackwell rack
[Official GB200 NVL72 specs →](https://www.nvidia.com/en-us/data-center/gb200-nvl72/) | 72× Blackwell · 13.4 TB HBM3e ✓ Native NVFP4
≈15.5–17.2T parameter-equivalent capacity; real services normally use replicas. | **≈120 kW · direct liquid** Residual air remains for networking and storage. | First rack-scale tier for large distributed inference and training. | **€2.45–2.98M / $2.8–3.4M** Reported 2026 purchase-quote range. |
| GB300 NVL72 Available 72‑GPU Blackwell Ultra rack
[Official NVL72 component design →](https://docs.nvidia.com/enterprise-reference-architectures/nvl72-ai-factory/latest/components.html) | 72× B300 · 20 TB HBM3e ✓ Native NVFP4
≈23–25.6T parameter-equivalent capacity; usually spent on replicas, cache and test-time compute. | **Up to 142 kW · full DLC** Facility CDU and rack leak-detection integration required. | High-memory NVL72 reasoning, inference and post-training. | **€5.25–5.69M / $6.0–6.5M** Reported purchase quotes; lower figures are often component/BOM estimates. |
| Vera Rubin NVL72 2026 production ramp · preliminary specs
[Official Vera Rubin NVL72 specs →](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/) | 72× Rubin + 36× Vera · 20.7 TB HBM4 ✓ Native NVFP4 + 3-bit LUT path
1,580 TB/s aggregate HBM bandwidth and 260 TB/s NVLink 6 switch bandwidth. | **≈240 kW comparison input · full DLC** 72 × 3.337 kW all-in per-GPU power used by the SemiAnalysis model; not a published rack nameplate. | High-interactivity, long-context and multitrillion-parameter inference where decode bandwidth and scale-up latency dominate. | **Quote required** SemiAnalysis owner-TCO assumption: $3.57/GPU-hour, not a rental rate or purchase quote. |
| VB300 NVL72 Derived 72‑GPU V300 comparison domain
[Inspect the V300/Kyber roadmap estimate →](https://wccftech.com/nvidia-rubin-ultra-rack-estimated-to-cost-21-million-hbm4e-swelling-to-1-5m-per-unit/) | 72× V300 · 41.5 TB HBM4e ◇ Roadmap assumption
≈48–53T parameter-equivalent if the derived rack retains NVFP4-class support. | **≈300 kW · 800VDC + full DLC** Half-Kyber planning envelope; not an announced NVIDIA rack. | Like-for-like NVL72 comparison before the jump to Kyber NVL144. | **≈€8.75–10.50M / $10–12M** Half-Kyber planning band with contingency; not a quote. |
| Kyber NVL144 Rubin Ultra/V300 analyst roadmap
[NVIDIA 800VDC rack-power direction →](https://developer.nvidia.com/blog/?p=100571) | 144× V300 · 82.9 TB HBM4e ◇ Roadmap assumption
≈96–106T parameter-equivalent if the roadmap system retains NVFP4-class support. | **≈600 kW · 800VDC + full DLC** Megawatt-class facility design around the compute rack. | Frontier scale where power architecture becomes a first-order design. | **≈€18.38M / $21M** Bank of America analyst ASP estimate reported July 2026. |
**These are ballparks, not list prices or model-fit promises.** NVFP4 capacity bands reserve 20–25% of advertised accelerator memory and assume roughly 5.0–5.2 effective bits per stored parameter, reflecting mixed-precision checkpoints rather than pure four-bit arithmetic. They exclude unusually large embeddings, multimodal towers and cache. H100 through GB300 prices use broad complete-system or reported rack quote bands; V300, VB300 and Kyber use explicitly labelled roadmap arithmetic. Actual draw follows utilization and power caps; network, storage, CDU/pumps, PUE overhead, spares, deployment and support are additional. Every price excludes VAT.
- [NVIDIA NVFP4 format and Blackwell support →](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)
- [Nemotron mixed-precision checkpoint evidence →](https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/)
- [2026 H100/H200 system quote bands →](https://intuitionlabs.ai/articles/data-center-gpu-pricing-2026)
- [Reported GB200 and GB300 purchase quotes →](https://www.tomshardware.com/tech-industry/artificial-intelligence/price-of-nvidias-vera-rubin-nvl72-racks-skyrockets-to-as-much-as-usd8-8-million-apiece-but-server-makers-margins-will-be-tight-nvidia-is-moving-closer-to-shipping-entire-full-scale-systems)
- [Morgan Stanley GB300 component estimate →](https://finance.yahoo.com/sectors/technology/articles/morgan-stanley-estimate-says-single-114101104.html)
- [V300 and Kyber roadmap estimate →](https://wccftech.com/nvidia-rubin-ultra-rack-estimated-to-cost-21-million-hbm4e-swelling-to-1-5m-per-unit/)
Early silicon · evidence checked 28 July 2026
### Vera Rubin is clearly ahead—but “10×” describes one old baseline at one speed.
CoreWeave measured a pre-production Vera Rubin NVL72 rack; SemiAnalysis then normalized it against July 2026 GB200 and GB300 InferenceX results. The useful question is not whether Rubin wins, but which Blackwell software vintage, user interactivity and cost boundary you compare.
100 tok/s/user
**2.00× vs GB200**
Rubin delivers 1,115,625 output tok/s/MW versus 558,569 on the July 2026 GB200 recipe; the corresponding GB300 ratio is 2.15×.
150 tok/s/user
**10.18× vs 2025 GB200**
This is the headline region. Against current July 2026 recipes, the lead is 2.74× over GB200 and 2.58× over GB300—not tenfold.
300 tok/s/user
**5.39× vs GB300**
Rubin delivers 96,446 output tok/s/MW. GB300 reaches 17,892 at its last viable frontier point; GB200 cannot reach this target.
Owner TCO · 300 tok/s/user
**5.03× cheaper**
The model produces $3.076 per million Rubin output tokens versus $15.456 for GB300. This is modeled ownership cost—not cloud rental pricing.
! Early external result
#### What was actually tested
DeepSeek R1 0528 671B at FP4, single-turn 8K input / 1K output. CoreWeave says both sides used NVFP4, speculative decoding with MTP, disaggregated prefill/decode with Dynamo, wide expert parallelism and TensorRT‑LLM. “Interactivity” is output tokens per second per user.
#### How the comparison was normalized
Throughput counts output tokens only, while power and cost include all prefill and decode GPUs. SemiAnalysis uses all-in power inputs of 2.168 kW/GPU for GB200, 2.553 for GB300 and 3.337 for Rubin, with the same PUE across the direct-liquid-cooled systems.
#### What remains unproven
The Rubin result came from a Dell engineering-sample rack without scale-out fabric and was supplied by CoreWeave; SemiAnalysis says it did not independently verify it. The workload is one older model and one short, single-turn recipe—not a multi-turn coding-agent trace.
Memory system
**20.7 TB HBM4 · 1,580 TB/s**
Each Rubin GPU carries up to 288 GB and 22 TB/s. Capacity helps model and KV-cache residency; bandwidth attacks the token-by-token decode bottleneck.
Scale-up fabric
**260 TB/s NVLink 6**
The 72-GPU rack provides 3.6 TB/s per GPU of all-to-all bandwidth. Counted writes reduce synchronization traffic for device-initiated communication.
Tensor Core path
**2× work per clock**
Rubin doubles the K dimension handled by its matrix instruction. Fewer K-loop iterations reduce overhead in throughput-, memory- and latency-bound kernels.
MoE data movement
**Inline TMA overrides**
A shared tensor descriptor can override addresses and strides in the instruction, avoiding an in-memory descriptor rewrite when switching experts with the same layout.
Weight compression
**3.125 bits/weight raw**
The new 3-bit LUT-B path stores an index plus an eight-entry E4M3 codebook per 512-weight block and resolves it inside the matrix operation. Accuracy depends on fitting and calibration.
Software transition
**SM100 kernels can start on SM107**
Rubin can reuse important Blackwell-family kernels for bring-up, unlike the Hopper-to-Blackwell break. Peak performance still needs Rubin-specific tuning; CUDA 13.4 support is a developer preview.
Planning arithmetic
#### Kimi K3 shows why 3-bit LUT support matters
For a 2.8T-parameter model, a 4.25-bit raw payload is about 1.49 TB; 3.125 bits is about 1.09 TB, roughly 394 GB less. At 288 GB HBM4 per Rubin GPU, raw weights alone move from about six packages to four. This excludes KV cache, activations, runtime buffers, sensitive higher-precision layers and parallelism replication—and it is not a released Kimi K3 LUT checkpoint.
**Do not bank the saving yet.**
One codebook serves 512 weights, whereas NVFP4 adapts a scale over groups of 16. The programmable codebook may preserve more useful values, but calibration, model quality and production kernels decide whether the paper capacity saving becomes a service gain.
> Visual: Reconstructed in HTML from SemiAnalysis data. Output throughput per all-in utility MW. Bars are scaled within each interactivity target; longer is better. Labels retain the exact output-token rate.
Reconstructed in HTML from SemiAnalysis data
#### Output throughput per all-in utility MW
Bars are scaled within each interactivity target; longer is better. Labels retain the exact output-token rate.
Vera Rubin · Jul 2026
GB300 · Jul 2026
GB200 · Jul 2026
CoreWeave GB200 · 2025
##### 100 tok/s/user
Vera Rubin
1,115,625
**1,115,625**
GB300
517,848
**517,848**
GB200
558,569
**558,569**
GB200 · 2025
300,000
**300,000**
##### 150 tok/s/user
Vera Rubin
786,208
**786,208**
GB300
305,023
**305,023**
GB200
287,289
**287,289**
GB200 · 2025
77,224
**77,224**
##### 200 tok/s/user
Vera Rubin
300,000
**300,000**
GB300
66,799
**66,799**
GB200
77,218
**77,218**
GB200 · 2025
50,000
**50,000**
##### 250 tok/s/user
Vera Rubin
140,264
**140,264**
GB300
37,049
**37,049**
GB200
36,394
**36,394**
GB200 · 2025
39,123
**39,123**
##### 300 tok/s/user
Vera Rubin
96,446
**96,446**
GB300
17,892
**17,892**
GB200
*Not feasible*
GB200 · 2025
*Not feasible*
> Visual: Reconstructed in HTML from SemiAnalysis TCO data. Modeled owner cost per million output tokens. Bars are scaled within each target; shorter is better. Values include prefill and decode GPUs under the source's ownership assumptions.
Reconstructed in HTML from SemiAnalysis TCO data
#### Modeled owner cost per million output tokens
Bars are scaled within each target; shorter is better. Values include prefill and decode GPUs under the source's ownership assumptions.
Vera Rubin · Jul 2026
GB300 · Jul 2026
GB200 · Jul 2026
CoreWeave GB200 · 2025
##### 100 tok/s/user
Vera Rubin
$0.266
**$0.266**
GB300
$0.497
**$0.497**
GB200
$0.423
**$0.423**
GB200 · 2025
$0.786
**$0.786**
##### 150 tok/s/user
Vera Rubin
$0.378
**$0.378**
GB300
$0.842
**$0.842**
GB200
$0.799
**$0.799**
GB200 · 2025
$3.046
**$3.046**
##### 200 tok/s/user
Vera Rubin
$0.991
**$0.991**
GB300
$3.837
**$3.837**
GB200
$3.016
**$3.016**
GB200 · 2025
$4.715
**$4.715**
##### 250 tok/s/user
Vera Rubin
$2.099
**$2.099**
GB300
$6.859
**$6.859**
GB200
$6.471
**$6.471**
GB200 · 2025
$6.013
**$6.013**
##### 300 tok/s/user
Vera Rubin
$3.076
**$3.076**
GB300
$15.456
**$15.456**
GB200
*Not feasible*
GB200 · 2025
*Not feasible*
Inspect the complete reconstructed data tables
**Output tokens per second per all-in utility MW**
| Recipe | 50 tok/s | 100 tok/s | 150 tok/s | 200 tok/s | 250 tok/s | 300 tok/s | 350 tok/s |
| --- | --- | --- | --- | --- | --- | --- | --- |
| GB200 NVL72 · Jul 2026 | 705,543 | 558,569 | 287,289 | 77,218 | 36,394 | Not feasible | Not feasible |
| GB300 NVL72 · Jul 2026 | 707,037 | 517,848 | 305,023 | 66,799 | 37,049 | 17,892 | Not feasible |
| CoreWeave GB200 · 2025 | 500,000 | 300,000 | 77,224 | 50,000 | 39,123 | Not feasible | Not feasible |
| CoreWeave Vera Rubin · Jul 2026 | 1,330,000 | 1,115,625 | 786,208 | 300,000 | 140,264 | 96,446 | 70,703 |
**Modeled owner cost per million output tokens**
| Recipe | 50 tok/s | 100 tok/s | 150 tok/s | 200 tok/s | 250 tok/s | 300 tok/s | 350 tok/s |
| --- | --- | --- | --- | --- | --- | --- | --- |
| GB200 NVL72 · Jul 2026 | $0.333 | $0.423 | $0.799 | $3.016 | $6.471 | Not feasible | Not feasible |
| GB300 NVL72 · Jul 2026 | $0.364 | $0.497 | $0.842 | $3.837 | $6.859 | $15.456 | Not feasible |
| CoreWeave GB200 · 2025 | $0.472 | $0.786 | $3.046 | $4.715 | $6.013 | Not feasible | Not feasible |
| CoreWeave Vera Rubin · Jul 2026 | $0.223 | $0.266 | $0.378 | $0.991 | $2.099 | $3.076 | $4.179 |
**Decision rule:** use the July 2026 GB200 and GB300 recipes for a current buy/no-buy comparison; keep the 2025 GB200 line only as a software-maturity lesson. Re-run the exact model, context distribution, concurrency, service-level target and power boundary before procurement. The charts above are original HTML/CSS reconstructions of the published values, with unavailable frontier points shown as “not feasible.”
- [SemiAnalysis inference TCO and architecture analysis →](https://newsletter.semianalysis.com/p/vera-rubin-nvl72-vs-gb200-nvl72-inference)
- [CoreWeave original engineering-sample result →](https://www.coreweave.com/blog/nvidia-vera-rubin-nvl72-on-coreweave-10x-more-tokens-per-megawatt-than-blackwell)
- [NVIDIA Vera Rubin NVL72 preliminary specs →](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/)
- [NVIDIA Rubin architecture detail →](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/)
- [CUDA 13.4 developer-preview notes →](https://docs.nvidia.com/cuda/developer-preview/13.4/cuda-toolkit-release-notes/index.html)
- [InferenceX methods and results →](https://github.com/SemiAnalysisAI/InferenceX)
OEM alternatives
### AMD or Intel host CPUs can still feed native-CUDA NVIDIA GPU servers.
The CPU vendor affects host-memory bandwidth, data preparation, storage and operational standardization. CUDA compatibility follows the NVIDIA accelerator, so all four configurations below remain native CUDA systems.
| Orderable OEM server | Host + accelerator | Power + cooling | Good first use | Ballpark, ex VAT |
| --- | --- | --- | --- | --- |
| Dell PowerEdge XE9685L [Official Dell configuration →](https://www.dell.com/support/manuals/en-us/poweredge-xe9685l/pexe9685l_ism_pub/system-overview?guid=guid-997e27f8-549c-42e4-8d17-41c79b40b0c5&lang=en-us) | 2× AMD EPYC 9005 + 8× B200 · 1.44 TB HBM ✓ Native CUDA + NVFP4 | **6× 3 kW PSU capacity** 4U direct liquid cooling; reserve residual air. | B200 service node with an AMD-standardized CPU fleet. | **€438–569k / $500–650k** Accelerator-led planning range. |
| HPE ProLiant Compute XD685 [Official HPE configurations →](https://www.hpe.com/us/en/collaterals/collateral.a00073553enw.html) | AMD EPYC + 8× H200, B200 or B300 H200: ✕ no native NVFP4
B200/B300: ✓ native | **Up to 12 high-watt PSUs** Direct-liquid-cooled B200/B300 configurations; quote actual max draw. | A supported AMD-host path from H200 through Blackwell Ultra. | **€438–744k / $500–850k** Wide range spans GPU generation and support. |
| Dell PowerEdge XE9680L [Official Dell configuration →](https://www.dell.com/support/manuals/en-us/poweredge-xe9680l/pexe9680l_ism_pub/system-overview?guid=guid-eef3f3fc-57c9-4939-958a-7128c53241b4) | 2× 5th-gen Intel Xeon + 8× H200 or B200 H200: ✕ no native NVFP4
B200: ✓ native | **6× 3 kW PSU capacity** 4U direct liquid cooling; reserve residual air. | B200 service node for Intel-standardized estates. | **€438–569k / $500–650k** Accelerator-led planning range. |
| Lenovo ThinkSystem SR680a V3 [Official Lenovo B200 guide →](https://lenovopress.lenovo.com/lp2247-thinksystem-sr680a-v3-with-b200) | 2× 5th-gen Intel Xeon + 8× B200 · 1.44 TB HBM ✓ Native CUDA + NVFP4 | **3.2 kW PSU modules** 8U air-cooled design; validate rack airflow and redundancy. | Air-cooled B200 where liquid plumbing is unavailable. | **€438–569k / $500–650k** Accelerator-led planning range. |
Capacity, training and market scale
### What each hardware tier can actually do—and how it becomes a model provider.
A server is a useful product unit. A model lab is a fleet, software organization, power contract and customer pipeline. Open the card to connect hardware ownership with the organizations and services it can support.
> Visual: AI model organization scale ladder
**Visual reading order:**
1. 01 · 1–8 accelerators Research project **One workstation or node** Evaluate, fine-tune and serve existing weights; pretrain compact models; prove a bounded workflow. A failed experiment costs hours or days, not a datacenter programme. **People on the work** ≈2–30 direct **Value / market cap** N/A or <$50m org
2. 02 · 1k–20k+ H100e Independent model lab **Mistral · DeepSeek · Kimi class** Pretrain large models, run ablations and post-training, then operate an API. Mistral disclosed a 3,000-H200 run; DeepSeek’s reported H100/H800 pool is around 20,000. Moonshot’s fleet is not disclosed. **People on the work** ≈200–2,000 direct **Value / market cap** ≈$12–75bn lab value
3. 03 · 600k–2m H100e Frontier lab **xAI · Anthropic · OpenAI** Run several training and post-training programmes, generate synthetic data, absorb failed runs and serve global products. Capacity is usually spread across clouds and campuses. **People on the work** ≈2,000–10,000 direct **Value / market cap** ≈$0.85–1.0tn lab value
4. 04 · 2m–5m+ parent pool Hyperscaler-backed frontier **Amazon · Microsoft · Meta · Google · Alibaba** The parent pool feeds cloud tenants, first-party models, recommenders and other AI. Alibaba’s total is not separately disclosed; the whole pool is never one model run. **People on the work** ≈10k–100k+ cross-stack **Value / market cap** ≈$0.29–3.9tn parent
| Owned shape | What it can actually do | Service reality | Organization it supports |
| --- | --- | --- | --- |
| 1–8 accelerators Workstation to one node | Evaluation, LoRA/Q‑LoRA, small-model pretraining and one production replica of a model that fits. | **Pilot to ≈1,200 heavy users** Model and node dependent; benchmark the exact build. | Research group, internal platform team or specialist service with cloud overflow. |
| 72–144 linked accelerators One NVL72/NVL144-class domain | Large distributed inference, multiple replicas, long-context cache, RL and substantial post-training. | **≈2,500–80,000+ heavy users** A planning range, not a vendor benchmark. | Regional model cloud or a major enterprise AI platform; the facility is now part of the product. |
| 1k–20k+ H100e Many racks and training domains | Pretrain competitive large open models, maintain several model lines and operate a public API. | **Provider-scale, still capacity constrained** Training and inference compete for the same fleet. | Mistral/DeepSeek/Kimi-class lab with dedicated infra, capital and model operations. |
| 600k–2m H100e accessible Multi-campus and multi-cloud | Parallel frontier programmes, large synthetic-data and RL pipelines, global inference and rapid retraining. | **Global consumer + enterprise products** No single training job consumes the whole estate. | Frontier lab backed by hyperscaler contracts, very large financing and energy commitments. |
| 2m–5m+ H100e owned Parent pool across regions and products | Rent compute to other labs, train first-party models, run internal AI and absorb multi-year silicon supply. | **Cloud platform + model catalogue** Customer workloads and first-party teams compete for allocations. | Amazon/AWS, Microsoft/Azure, Google and Alibaba Cloud; Meta has the owner scale without a public cloud. |
**People and value are scale context, not qualification rules.** “People on the work” estimates direct model, product, platform, silicon and datacenter effort—not every employee in the parent company. Independent labs are valued by private funding rounds; only listed parents have a market cap. July 2026 reference points include Mistral at 900+ employees, DeepSeek at an implied ≈$52bn, OpenAI at ≈$852bn and Anthropic at ≈$965bn.
- [Mistral headcount →](https://mistral.ai/careers/)
- [Mistral valuation →](https://mistral.ai/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai/)
- [DeepSeek implied valuation →](https://www.investing.com/news/stock-market-news/chinese-filing-implies-deepseek-valuation-of-around-52-billion-4796314)
- [OpenAI valuation →](https://openai.com/index/accelerating-the-next-phase-ai/)
- [Anthropic valuation →](https://www.anthropic.com/news/series-h)
Rent, customize or build
### AWS, Azure, Google Cloud and Alibaba Cloud are both model shelves and infrastructure routes.
A hyperscaler can sell a managed API, provision dedicated throughput, rent a training cluster and train its own model family. Open the card to compare the four routes.
AWS · AMAZON
#### Bedrock + SageMaker AI
Use Bedrock for managed access to hundreds of foundation models; move to SageMaker AI and HyperPod for customization, distributed training and controlled deployment.
**First-party models** — Amazon Nova 2, Nova Sonic and Nova multimodal embeddings
**Other providers** — OpenAI and open/model-partner catalogues, Anthropic (CSP accounts not supported)
**Silicon route** — NVIDIA GPU plus Amazon Trainium and Inferentia
**Owner pool**
≈2.45m H100e
**Parent market cap**
≈$2.54tn · 24 Jul 2026
- [Amazon Bedrock →](https://aws.amazon.com/bedrock/)
- [Amazon Nova models →](https://aws.amazon.com/nova/models/)
- [Nova on SageMaker →](https://docs.aws.amazon.com/nova/latest/nova2-userguide/nova-model.html)
AZURE · MICROSOFT
#### Microsoft Foundry + Azure OpenAI
Use Azure-hosted model APIs with standard or provisioned deployment; move to managed compute and Azure ML when the workload needs custom weights, capacity or training control.
**First-party models** — MAI, Phi, Model Router and healthcare AI models
**Other providers** — OpenAI, Cohere, DeepSeek, Meta, Mistral, xAI, Anthropic (CSP accounts not supported) and other models.
**Infrastructure route** — Azure GPU estates, Foundry managed compute and Azure ML
**Owner pool**
≈3.42m H100e
**Parent market cap**
≈$2.84tn · 24 Jul 2026
- [Foundry model catalogue →](https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
- [Azure ML compute routes →](https://learn.microsoft.com/en-us/azure/machine-learning/concept-compute-target)
GOOGLE CLOUD · ALPHABET
#### Gemini Enterprise Agent Platform
Use the platform for managed Gemini and partner APIs, provisioned throughput, tuning and agent workflows; use Managed Training, GKE, Cloud TPU or NVIDIA GPU capacity when the workload needs infrastructure control.
**First-party models** — Gemini, open-weight Gemma, Imagen, Veo and Google embeddings
**Other providers** — Meta, Mistral, Qwen, Anthropic (CSP accounts not supported) and other Gemini Enterprise Agent Platform Model Garden options
**Silicon route** — Cloud TPU including Trillium/Ironwood, plus NVIDIA GPU estates
**Owner pool**
≈5.05m H100e
**Parent market cap**
≈$3.89tn · 24 Jul 2026
- [Gemini Enterprise Agent Platform overview →](https://docs.cloud.google.com/gemini-enterprise-agent-platform/overview)
- [Gemini Enterprise Agent Platform models →](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models)
- [Google Cloud TPU route →](https://docs.cloud.google.com/tpu/docs/intro-to-tpu)
ALIBABA CLOUD · QWEN
#### Model Studio + PAI
Use Model Studio for OpenAI-compatible Qwen and partner APIs; use Platform for AI for dedicated deployment, fine-tuning and distributed training through DLC, EAS and Lingjun resources.
**First-party models** — Qwen text, reasoning, coding and multimodal families; Wan media models
**Other providers** — DeepSeek, Kimi and GLM availability varies by region
**Open-weight route** — Deploy Qwen weights with vLLM/SGLang or buy a managed Qwen API
**Owner pool**
Not separately public
**Parent market cap**
≈$0.29tn · 24 Jul 2026
- [Alibaba Model Studio →](https://www.alibabacloud.com/help/en/model-studio/what-is-model-studio)
- [PAI training and deployment →](https://www.alibabacloud.com/help/en/pai/use-cases/llm1/)
Scaling effect
#### More compute improves the search, the model and the service—but not in a straight line.
When model size, training data and compute stay balanced, language-model loss tends to improve along a power law: every extra unit buys a smaller marginal gain. The practical leap is broader than one benchmark score. Larger fleets buy more parallel experiments, longer post-training, more inference-time reasoning, replicas, recovery capacity and faster iteration. Better algorithms can compress the gap; weak data, networking or utilization can waste it.
- [OpenAI scaling-laws research →](https://openai.com/index/scaling-laws-for-neural-language-models/)
- [Chinchilla compute-optimal training →](https://arxiv.org/abs/2203.15556)
Best public comparison · latest like-for-like baseline is end-2025
### Lab-access compute, parent-owner pools and a fleet-cost proxy.
Open the card for a logarithmic or linear comparison of accessible compute and parent-owned pools. “H100e” compares peak AI operations with an NVIDIA H100 and does not equal real model throughput.
Start with the logarithmic order-of-magnitude view, then switch to linear to make equal chart width mean equal H100e—and therefore an equal step in the simple fleet-cost proxy. This is not an inventory audit. Solid bars estimate operational compute a lab could access; striped bars mark a disclosed run, floor or bound; outlined bars are parent-owned pools.
Axis spacing
Log view: each equal step is roughly ten times more accessible compute.
| Provider | Category | Fleet-cost proxy | Accessible compute / disclosure |
| --- | --- | --- | --- |
| Mistral AI | OPEN WEIGHTS | Fleet proxy $0.08–0.15bn | ≥3k observed H200 run |
| Moonshot / Kimi | OPEN WEIGHTS | Fleet cost not public | Fleet not disclosed |
| DeepSeek | OPEN WEIGHTS | Fleet proxy $0.5–1.0bn | ≈20k reported H100 + H800 |
| Qwen / Alibaba | HYPERSCALER MODEL LAB | Fleet cost not isolated | Qwen allocation not disclosed |
| Meta AI / MSL | OPEN + FRONTIER LAB | Fleet cost not isolated | <1.7m lab bound; parent pool below |
| **Frontier API labs** | | | |
| xAI | FRONTIER | Fleet proxy $15–35bn | ≈600–700k lab access |
| Anthropic | FRONTIER | Fleet proxy ≥$25–50bn | ≥1m lab access |
| OpenAI | FRONTIER | Fleet proxy $43–85bn | ≈1.7m lab access |
| Google DeepMind | FRONTIER LAB | Fleet cost not isolated | <2m lab bound; parent pool below |
| **Hyperscaler parent owner pools · not one lab allocation** | | | |
| Alibaba Cloud | CLOUD OWNER | Fleet cost not public | Not split out · China total ≈1.16m ex-smuggling |
| Meta | PARENT OWNER | Fleet proxy $58–115bn | ≈2.30m Q4 2025 owner pool |
| Amazon / AWS | CLOUD OWNER | Fleet proxy $61–122bn | ≈2.45m Q4 2025 owner pool |
| Microsoft / Azure | CLOUD OWNER | Fleet proxy $85–171bn | ≈3.42m Q4 2025 owner pool |
| Google | CLOUD + MODEL OWNER | Fleet proxy $126–252bn | ≈5.05m Q4 2025 owner pool |
**Fleet-cost proxy**
Plotted quantities × $25k–$50k per H100e: a broad accelerator, server and fabric replacement band derived from current complete-system pricing. It is not an accounting estimate of what the company paid.
It excludes buildings, grid connection, cooling, storage, energy, spares, financing and operations. TPU and Trainium economics differ; leases shift capex to a cloud owner; lab bounds are not costed when the parent allocation is unknown.
**Read DeepSeek correctly.** The ≈20,000 figure is accelerator-class chips, not 20,000 eight-GPU servers. Its V3 paper documents one 2,048-H800 training cluster and 2.788M H800 GPU-hours; wider fleet figures and additional H20/A100 hardware are not converted into the plotted number.
**Read Qwen and Alibaba separately.** Qwen is Alibaba’s model family; Alibaba Cloud is the infrastructure and API parent. Neither Qwen’s allocation nor Alibaba’s owner total is public. Epoch estimates identified Chinese owners together at ≈1.16m H100e in Q4 2025 before its optional smuggling estimate, so assigning that number to Alibaba would be wrong.
**Ownership is not availability.** Epoch AI’s Q4 2025 owner estimates are Google 5.05m, Microsoft 3.42m, Amazon 2.45m and Meta 2.30m H100e, but those pools also serve cloud customers and internal products. Its July 2026 campus directory provides a newer site check without a like-for-like company total.
- [Epoch AI owner-pool data and method →](https://epoch.ai/data/ai-chip-owners)
- [Epoch AI lab-access estimates →](https://epoch.ai/gradient-updates/frontier-labs-dont-use-most-ai-compute)
- [Operational campus directory →](https://epoch.ai/data/ai-data-centers)
- [Mistral 3,000-H200 run →](https://mistral.ai/news/mistral-3/)
- [DeepSeek fleet estimate →](https://www.csis.org/analysis/deepseek-deep-dive)
- [DeepSeek V3 report →](https://arxiv.org/abs/2412.19437)
- [System-price basis for fleet proxy →](https://intuitionlabs.ai/articles/data-center-gpu-pricing-2026)
- [Kimi capacity constraint →](https://apnews.com/article/4c66a2e0f557ce79d3cc2d769c9a6226)
Public bridges to scale
### Use public access, capital and demand before carrying the full infrastructure risk.
The European Union and United States offer different ladders from research access to provider-scale infrastructure. Open the card to compare the regional routes and their limits.
European Union
##### Shared compute, scale-up finance and sovereign demand
01 · ACCESS
##### Borrow the first large run
EuroHPC’s open industrial call offers AI Factory allocations above 50,000 GPU-hours with a stated ten-working-day approval target. Nineteen AI Factories and thirteen antennas are being assembled to give startups and SMEs compute plus technical support.
- [Apply for large-scale AI Factory access →](https://www.eurohpc-ju.europa.eu/large-scale-access-ai-factories_en)
- [Check the factory and antenna network →](https://www.eurohpc-ju.europa.eu/eurohpc-ju-launches-call-proposals-strengthen-european-ai-ecosystem-2026-04-28_en)
02 · FINANCE
##### Bridge startup to infrastructure company
TechEU says it will deploy €70bn of EIB Group equity, loans and guarantees in 2025–2027 to mobilize €250bn with partners across the innovation lifecycle, including AI and digital infrastructure. It is a financing channel, not an automatic subsidy.
- [Review the TechEU finance route →](https://www.eib.org/en/press/all/2025-314-europe-s-innovative-companies-get-boost-as-eib-group-launches-techeu-platform-to-simplify-financing)
03 · BUILD
##### Finance the leap to gigafactory scale
InvestAI plans to mobilize €20bn for several AI Gigafactories, with the EIB exploring advisory support and loans. The target shape is about 100,000 advanced chips per site. The formal call is due in summer 2026 and first construction is scheduled for 2027: this is future capacity, not online supply.
- [Track the AI Gigafactory call →](https://commission.europa.eu/topics/competitiveness/competitiveness-coordination-tool-projects/ai-gigafactories_en)
- [Read the EIB financing mandate →](https://www.eib.org/en/press/all/2025-491-eib-group-and-european-commission-join-forces-to-finance-ai-gigafactories)
04 · DEMAND
##### Turn sovereignty into a buying criterion
The Commission’s proposed Cloud and AI Development Act would streamline sites, energy and finance, create an EU sovereignty framework and establish common public-sector procurement. That can make trusted European capacity easier to specify and buy; it does not guarantee a contract.
- [Read the proposed Cloud and AI Development Act →](https://digital-strategy.ec.europa.eu/en/policies/cloud-and-ai-development-act)
United States
##### Research access, competitive funding, sites and procurement
01 · ACCESS
##### Use the national research resource
The NSF-led National Artificial Intelligence Research Resource connects US researchers, educators, startups and small businesses to public and private compute, models, data and expertise. By March 2026 it had supported more than 600 research teams and 6,000 students.
- [Explore current NAIRR access →](https://www.nsf.gov/focus-areas/ai/nairr)
- [Read the two-year progress update →](https://www.nsf.gov/cise/updates/nairr-2-years-advancing-american-artificial-intelligence)
02 · FUND
##### Fund efficiency and regional capacity
NSF-backed STRIDE can invest up to $21m over two years in deployment-ready AI infrastructure efficiency, with awards up to $3.5m. NSF Regional Innovation Engines can fund regional technology coalitions up to $160m over ten years. Both are competitive programmes, not project debt.
- [Review the AI Efficiency Challenge →](https://www.nsf.gov/tip/updates/nsf-supported-stride-ventures-launches-ai-efficiency)
- [Explore NSF Regional Innovation Engines →](https://www.nsf.gov/funding/initiatives/regional-innovation-engines)
03 · BUILD
##### Shorten the path to 100 MW+ campuses
A federal order defines qualifying AI data center projects as requiring more than 100 MW of new load and directs expedited federal permitting. DOE has also selected four federal sites for private AI data center and energy partnerships. This lowers siting friction; it does not supply a fleet or guaranteed grid capacity.
- [Read the data center permitting order →](https://www.whitehouse.gov/presidential-actions/2025/07/accelerating-federal-permitting-of-data-center-infrastructure/)
- [Review the selected DOE sites →](https://www.energy.gov/articles/doe-announces-site-selection-ai-data-center-and-energy-infrastructure-development-federal)
04 · DEMAND
##### Use federal procurement as a first market
OMB M-25-22 directs more efficient federal AI acquisition, while GSA’s OneGov consolidates technology buying and reported 20 vendor agreements by April 2026. The route can accelerate agency adoption, but eligibility is limited and no programme guarantees a provider contract.
- [Open current OMB memoranda →](https://www.whitehouse.gov/omb/information-resources/guidance/memoranda/)
- [Review the current OneGov route →](https://www.gsa.gov/buy-through-us/purchasing-programs/multiple-award-schedule/onegov)
EU operator example
**Mistral moved from a disclosed 3,000-H200 model run to production GB200 service in February 2026 and says it is targeting 200 MW of sovereign EU capacity by 2027.** The route mixes private capital, NVIDIA supply, datacenter partners, cloud revenue and sovereign-enterprise customers. It is a useful pattern—not proof that every provider should own the same fleet.
- [Inspect Mistral Compute’s live timeline →](https://mistral.ai/products/compute/)
- [See the 200 MW Campus AI agreement →](https://presse.bpifrance.fr/mistral-securise-jusqua-200-mw-de-capacite-de-calcul-avec-campus-ai-en-france/?lang=fra)
Read the comparison carefully
**The US programmes are analogues, not one-for-one equivalents.** They split research access, innovation funding, federal land and permitting, and government demand across separate agencies. Large-fleet capital still depends mainly on commercial customers and private finance, and eligibility in one lane does not imply support in the others.
- [Read the US AI Action Plan →](https://www.whitehouse.gov/releases/2025/07/white-house-unveils-americas-ai-action-plan/)
- [Review NAIRR’s operating scope →](https://www.nsf.gov/focus-areas/ai/nairr)
**Inference fit is not training fit.** Full training holds weights, gradients, optimizer states and activations, often using several times the inference memory before checkpoint and data-pipeline overhead. LoRA and Q‑LoRA are much lighter. Use the hardware table to shortlist a tier, the scale comparison to understand the organization around it, and the simulator below for the workload, quote and payback target you actually expect.
Try before you buy
## Rent the exact workload before committing to private hardware.
Before a five- or six-figure private-cloud commitment, rent the candidate stack long enough to discover whether the model, service shape and operating cost actually fit. A rental is a test drive, not a GPU-availability, price or hardware-validation guarantee.
**Run the acceptance pack first.** Test the exact model, engine, quantization, context, concurrency and acceptance harness you would buy for; retain the container, settings, logs, latency, throughput, memory and task-quality results.
**Referral disclosure:** this isaiuseful.com link is a Runpod referral link. As checked 10 August 2026, eligible new first-time users must sign up through it with Google SSO and load their first $10: European users receive $5 credit, while non-European users receive a weighted $5–$500 credit (Runpod says most are $10 or less). Terms can change. If eligible, isaiuseful.com receives its referral bonus and earns Runpod credits on actual usage for the first six months: 3% of Pod spend and 5% of Serverless spend. Using the link supports the site.
- [Rent a Runpod test environment through this referral link →](https://runpod.io?ref=l40ix174)
- [Read Runpod’s current referral terms →](https://docs.runpod.io/accounts-billing/referrals)
Users + ROI simulator
## Turn a vendor quote into a price per user—or per token.
Start from the model’s evidence-matched serving preset or lock a platform you already own. The planner packs repeatable replica cells into that platform, scales the packed loadout toward a user target and keeps every generated assumption editable.
Fits selected system
90 GB service build in 1.44 TB HBM, with 15% reserved for runtime headroom.
Pricing view
Monthly operating cash
**—**
Selected private billing minus modeled monthly operating cost.
Simple capex payback
**—**
Capex divided by positive monthly operating cash; excludes financing and demand ramp.
Gap to capital-recovery target
**—**
Monthly billing minus operating cost and the selected capital-recovery allowance.
Selected private quote / user / month
**—**
Same-model OpenRouter cost × the selected private quote multiplier.
Same-model OpenRouter / user / month
**—**
Selected local model’s market rate × workload.
Operating break-even / served user / month
**—**
Covers the costs explicitly modeled here, without capital recovery.
Target-payback price / served user / month
**—**
Covers modeled operating cost plus capex recovery inside the selected target.
Per-user gap to target-payback price
**—**
Selected private quote minus the target-payback price, per served user per month.
Estimated supported users
**—**
Constrained by the tighter of token throughput, concurrency and per-request memory.
Auto-sized loadouts for target
**—**
Calculated from the capacity of one packed loadout.
Selected API alternative / user / month
**—**
API rate × uncached input, cached input and output.
Selected private quote multiplier
**—**
The commercial ask currently applied to the same-model OpenRouter anchor.
Quote multiplier required for target
**—**
In-range results are rounded upward to the next 0.25× slider step so the recommendation never understates the target.
Modeled monthly operating cost
**—**
Monthly private billing at selected price
**—**
Selected API cost for served target
**—**
Change any assumption to test the scenario.
- [Open selected model’s current OpenRouter pricing →](https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b/pricing)
- [Open selected API pricing →](https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b/pricing)
Planning note
**This is a workload simulation—not a benchmark, capacity promise or complete financial model.**
**The financial outputs use two different thresholds on purpose.** Operating cash is private billing minus the recurring costs entered here. Simple payback divides capex by positive operating cash. The capital-recovery target additionally asks the service to recover capex within the selected number of years. A plan can therefore make positive operating cash and have a finite payback while still missing a faster target. Financing costs, tax, working capital, depreciation, residual value and any staffing, software, insurance or network cost not entered above are excluded.
Observed coding workloads vary sharply: one 12,000-developer study reported a 51M-token monthly median and about 380M at P90, while an agent-task study found up to 30× variation between runs. The higher presets represent continuously operating agent-worker equivalents, not ordinary human seats. Start there, then replace them with your own exports. [Review the developer workload study →](https://jellyfish.co/blog/is-tokenmaxxing-cost-effective-new-data-from-jellyfish-explains/) [Review the agent-task variance study →](https://arxiv.org/abs/2604.22750)
Each model starts from a complete replica shape, not a model-size speed multiplier alone. gpt‑oss‑120b, MiniMax M3, Qwen3.5 and DeepSeek V4 Pro now use 8K/1K single-node InferenceX profiles; Nemotron 3 Super and Ultra, GLM‑5.2 and Kimi K2.7 use matched Lambda runs; DeepSeek V3.2 uses a GPUStack sweep; and Kimi K3 keeps its throughput–interactivity curve. The MiniMax and Qwen memory gates include both data-parallel checkpoint copies in their measured four-GPU cells. Mistral Small 4, Mistral Large 3 and DeepSeek V4 Flash remain explicitly labelled reference-topology proxies. A complete proxy cell is no longer divided by all eight GPUs merely because several small cells pack into one node. [Inspect the InferenceX profiles →](https://inferencex.semianalysis.com/) [Nemotron 3 Super measured profile →](https://lambda.ai/inference-models/nvidia/nvidia-nemotron-3-super-120b-a12b) [Nemotron 3 Ultra measured profile →](https://lambda.ai/inference-models/nvidia/nemotron-3-ultra) [DeepSeek V3.2 measured sweep →](https://docs.gpustack.ai/2.0/performance-lab/deepseek-v3.2/h200/) [GLM‑5.2 measured profile →](https://lambda.ai/inference-models/zai-org/glm-5.2) [Kimi K2.7 measured profile →](https://lambda.ai/inference-models/moonshotai/kimi-k2.7-code) [Kimi K3 throughput curve →](https://vllm.ai/blog/2026-07-27-k3)
Nemotron 3 Ultra needs a special translation: its public run uses 8,192 input / 65,536 output tokens. The published input rate is therefore the observed token share of a decode-heavy saturation test, not an independently saturated prefill ceiling. The preset retains that observed value for context but uses measured decode throughput, interactivity and memory as its hard capacity gates; replace the profile with a mixed-shape sweep before procurement.
The service targets separate user-visible output speed from total server throughput. Independent API measurements checked on 1 August 2026 observed about 32 output tok/s for Kimi K3, 66 for GPT-5.6 Sol at maximum effort and 74 for Claude Fable 5. The 30 and 70 targets are parity anchors, 100 is a fast interactive target and 15 is for background agents; none is a latency SLA. gpt‑oss‑120b, MiniMax M3, Qwen3.5 and DeepSeek V4 Pro load measured points from their InferenceX concurrency sweeps where available, while Kimi K3 uses its approximate published curve. High-volume workloads also reserve the parallel active requests implied by monthly output volume and measured per-request speed; the per-request KV assumption is multiplied by that request count. Every single-point benchmark and topology proxy changes aggregate throughput through a conservative cross-model envelope calibrated to those 8K/1K sweeps: 0.70×, 1.00×, 1.50× and 1.90× the normalized frontier point. Those values are planning estimates, not measurements for the selected model; whole-platform rounding or a tighter memory gate can still leave two adjacent targets at the same purchase count. [Review the measured Kimi K3 API speed →](https://artificialanalysis.ai/models/kimi-k3) [Review the measured GPT-5.6 Sol API speed →](https://artificialanalysis.ai/models/gpt-5-6-sol) [Review the measured Claude Fable 5 API speed →](https://artificialanalysis.ai/models/claude-fable-5)
Prefix reuse changes the self-hosted prefill path, not output decoding. The simulator counts uncached input at full prefill work and cached input at the editable cache-hit allowance. Moonshot reports above 90% cache hits for Kimi K3 coding traffic, so its coding presets start at 90%; that provider result is not a promise for a different scheduler or workload. Cache lookup, state transfer, retention and eviction still consume memory and network capacity. [Review Moonshot’s K3 cache claim and direct pricing →](https://www.kimi.com/blog/kimi-k3) [Review vLLM’s hybrid prefix-cache design →](https://vllm.ai/blog/2026-07-27-k3)
Every preset assigns at least 20% of tokens to output; coding profiles use a 50/50 split so billed reasoning is not hidden. Replace workload, prefix reuse, cache-read work, aggregate throughput, concurrency, interactivity and per-request memory assumptions with telemetry from the same ISL/OSL test shape.
The same-model OpenRouter rate remains the grounding benchmark even when another API alternative is selected. Rates are stored as planning inputs checked on 1 August 2026 rather than fetched live; recheck the linked model page before quoting. Direct API comparisons exclude tool calls, cache writes, cache storage, long-context or regional uplifts and negotiated discounts. **All prices exclude VAT.** [Review the platform evidence →](#hardware)
Compatibility
## CUDA compatibility is a binary boundary, not a performance grade.
Only NVIDIA hardware runs CUDA binaries. AMD and Intel can still be excellent inference platforms, but the application, kernels and quantization must support their native stack.
| Platform | Runs CUDA binaries? | Native stack | Portability route | Provider decision |
| --- | --- | --- | --- | --- |
| NVIDIA Ampere · A100 | ✓ Native CUDA | CUDA · NCCL · NVLink
✕ No native NVFP4 | Use BF16/FP16 or supported integer paths. Some model cards document loading an NVFP4 checkpoint through a fallback. | Treat NVFP4 as model-specific weight compatibility here—not native FP4 throughput. |
| NVIDIA Hopper · H100/H200 | ✓ Native CUDA | CUDA · NCCL · NVLink · TensorRT‑LLM
✕ No native NVFP4 | FP8 is the native low-precision path. Selected NVFP4 checkpoints can fall back to W4A16. | Buy for mature CUDA and FP8—not for native NVFP4 throughput. |
| NVIDIA Blackwell · GB10/B200/B300/GB200/GB300 and RTX PRO | ✓ Native CUDA | CUDA · NCCL · NVLink · TensorRT‑LLM
✓ Native NVFP4 | Native W4A4 path when the engine, kernel and model support it. | Current launch target for NVFP4; validate GB10 and workstation-specific kernels separately from datacenter Blackwell. |
| AMD MI300X/MI325X/MI355X and Radeon | ↻ No—port and rebuild | ROCm · HIP · RCCL · Infinity Fabric | HIP source can target AMD or NVIDIA, but one compiled binary cannot run on both. | Strong memory economics when the exact engine, model and kernels pass a pilot. |
| Intel Gaudi 3 | ✕ No—separate stack | SynapseAI · Habana libraries · RoCE | Use Gaudi-supported frameworks and Optimum Habana; CUDA binaries do not transfer. | Treat as its own product lane with an explicit supported-model catalogue. |
| Intel GPU / XPU | ↻ No—source migration | oneAPI · SYCL · OpenVINO | SYCLomatic can assist CUDA-source migration; manual work and validation remain. | Useful for selected Intel-XPU workloads, not a Gaudi substitute. |
| tinygrad runtime | ◇ Backend choice | Its own small compiler with CUDA, AMD, NV, Metal and other backends | A tinygrad program can target different backends; tinygrad does not make AMD or Intel CUDA-compatible. | Use for owned compiler research and ports, not assumed vLLM feature parity. |
**Feature support changes faster than the silicon.** As of 26 July 2026, vLLM documents NVIDIA CUDA, AMD ROCm and Intel XPU routes, but feature parity and quantization support vary. NVFP4 is an NVIDIA format; native W4A4 acceleration is a Blackwell-generation capability, while Hopper uses a model- and engine-specific fallback where supported. Pin the container, driver, firmware, engine commit and model revision used in acceptance testing.
- [NVIDIA NVFP4 hardware explanation →](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)
- [Blackwell W4A4 versus Hopper W4A16 →](https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/)
- [NVIDIA CUDA libraries →](https://docs.nvidia.com/cuda-libraries/index.html)
- [AMD HIP binary-compatibility FAQ →](https://rocm.docs.amd.com/projects/HIP/en/docs-6.2.4/faq.html)
- [Current vLLM GPU support →](https://docs.vllm.ai/en/latest/getting_started/installation/gpu/)
- [Intel Gaudi documentation →](https://docs.habana.ai/en/latest/)
- [Intel CUDA-to-SYCL migration →](https://www.intel.com/content/www/us/en/docs/oneapi/programming-guide/2025-1/migrating-from-cuda-to-sycl-for-the-dpc-compiler.html)
tinygrad route
## Own more of the stack, accept more integration work.
tinygrad’s useful idea is not “cheap CUDA.” It is a small, MIT-licensed compiler/runtime and reference hardware path that makes the stack inspectable, portable and open to operator modification.
Reference lab · AMD
### tinybox red v2
**€10,502 / $12,000 ex VAT**
4× Radeon 9070 XT, 64 GB aggregate GPU memory, 128 GB system memory and 2 TB NVMe. Best for compiler work, AMD ports and smaller sharded models—not a 400B+ production endpoint. Confirm store stock before planning delivery.
- [Official tinybox red v2 →](https://tinycorp.myshopify.com/products/tinybox-red-v2)
- [Read tinygrad’s hardware pitch →](https://tinygrad.org/pitch.pdf)
Reference lab · NVIDIA
### tinybox green v2 Blackwell
**€65,640 / $75,000 ex VAT**
4× RTX PRO 6000 Blackwell and 384 GB aggregate GPU memory. It runs CUDA, but PCIe-attached workstation GPUs are not an NVLink HBM supernode; test throughput, concurrency and current store stock.
- [Official tinybox green v2 →](https://tinycorp.myshopify.com/products/tinybox-green-v2-with-4x-rtx-pro-6000-blackwell)
- [Read tinygrad’s hardware pitch →](https://tinygrad.org/pitch.pdf)
Software ownership
### Port only where it earns control
**€0 / $0 software licence**
Use tinygrad to understand and own selected kernels or models. Keep vLLM or SGLang for production until the port meets the same correctness, batching, observability and recovery tests.
**Before promising hardware, ask what must stay compatible.** Examples include an OpenAI-compatible endpoint; vLLM, SGLang, NIM or Triton; PyTorch or JAX; tool calling and structured JSON; MCP or Hermes-style agent harnesses; embeddings, reranking, vision, PDF processing or fine-tuning; plus SSO, audit logs, retention, data residency, context, concurrency and latency requirements.
A client may not know it has a CUDA dependency. An NVIDIA container, TensorRT‑LLM, FlashAttention, bitsandbytes, NCCL or a PyTorch CUDA wheel is one—even if the business owner never says “CUDA.” Ask for the current container, lockfile and a representative acceptance test.
- [tinygrad repository and backends →](https://github.com/tinygrad/tinygrad)
- [Project status and tinybox specifications →](https://tinygrad.org/)
01
### Hardware pluralism
Keep the option to buy NVIDIA, AMD or commodity cards where the workload and kernels allow it.
02
### Inspectability
A smaller compiler makes generated kernels and backend behavior easier for an operator to understand and modify.
03
### Acceptance tests first
Port the client’s real model and critical software path before making a platform-wide compatibility promise.
04
### Honest maturity
tinygrad describes itself as alpha and says it is not faster than PyTorch for most use cases yet. Treat it as an engineering lane, not magic production parity.
Capacity ladder
## Match the checkpoint first, then reserve serving headroom.
These are procurement starting points, not throughput guarantees. Keep 15–25% aggregate memory for runtime overhead and KV cache, then load-test the actual context length and concurrency.
✕ No native NVFP4 · 640 GB–1.13 TB
### H100 / H200
Use native FP8, or a checkpoint-specific W4A16 fallback when memory capacity matters. Speculative mixed-NVFP4 weight capacity is ≈0.74–1.45T parameters across an eight-GPU node, but it does not gain native W4A4 compute.
✓ Native NVFP4 · 1.44–2.3 TB
### B200 / B300
About 1.65–2.95T parameters by conservative mixed-checkpoint memory math. Nemotron Ultra, Qwen, DeepSeek and Mistral Large are more realistic first services; leave the rest for cache, replicas and concurrency.
✓ Native NVFP4 · 13.4–20 TB
### GB200 / GB300 NVL72
Roughly 15.5–25.6T parameter-equivalent weight capacity. In practice, spend that headroom on Kimi K3-scale sharding, multiple replicas, long-context cache and test-time compute.
◇ Roadmap assumption · 4.6–82.9 TB
### V300 / VB300 / Kyber
The 5.3–106T NVFP4-equivalent range is compounded planning arithmetic across derived or roadmap systems—not a shipped support matrix, benchmark or useful single-model target.
Price method
## One exchange rate, visible caveats.
USD prices were converted at the ECB reference rate for 20 July 2026: **€1 = $1.1426** . EUR hardware figures are rounded to the nearest euro; token prices are rounded to the nearest euro cent.
### Excluded from every price
VAT, sales tax, shipping, import duties, rack, networking, storage, facility upgrades, energy, support and integration unless the row says otherwise.
### Ballpark means planning input
Each row says whether its number comes from reported purchase quotes, a component/BOM estimate or a roadmap ASP estimate. Replace it with your vendor quote before making a return case.
### Requote before buying
GPU allocations, support terms and discounts move quickly. Keep the accelerator, memory, interconnect, cooling and software acceptance criteria fixed while comparing quotes.
- [ECB reference exchange rates →](https://www.ecb.europa.eu/stats/policy_and_exchange_rates/euro_reference_exchange_rates/html/index.en.html)
- [Morgan Stanley BOM breakdown reported by Tom’s Hardware →](https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidias-memory-costs-soar-485-percent-latest-ai-systems-now-cost-usd7-8-million-to-build-memory-now-comprises-25-percent-of-the-total-cost-rubin-gpus-a-mere-usd50-000-apiece)
- [Return to server prices →](#hardware)
Provider default
Launch one permissive model on one validated NVIDIA reference system with one datacenter partner. Add a second accelerator stack only when its memory economics survive the cost of maintaining a second software product.
- [Model the first service](#economics)
- [Check the software boundary](#cuda)
- [Return to local models](https://isaiuseful.com/local-models.html.md)
---
## https://isaiuseful.com/training-models (`/training-models.html.md`)
# Should You Use RAG, Fine-Tuning or Pretraining?
Canonical source: [https://isaiuseful.com/training-models](https://isaiuseful.com/training-models)
- [Guides](https://isaiuseful.com/guides.html.md)
· Model adaptation · checked 27 July 2026
Start with the measured failure, then choose the lightest intervention that can fix it. Most organisations need retrieval or a focused tune—not a foundation model trained from scratch.
- [Choose an intervention](#chooser)
- [Follow the evaluation loop](#loop)
- [Plan compute](#planner)
**4**
distinct intervention routes
RETRIEVAL TO PRETRAINING
**1**
measured gap before training
BENCHMARK THE BASELINE
**0**
evaluation gates skipped
TEST EVERY CHECKPOINT
Interactive decision tree
## When should you use RAG instead of fine-tuning?
Select every gap that applies. The recommendation can combine retrieval and training because production assistants usually need more than one layer.
Start with RAG
### Give the model governed access to current knowledge.
Index the allowed documents, retrieve evidence into context and require citations. Measure that baseline before changing weights.
RAG
Fine-tune
Continue
Scratch
Retrieval does not fix deep language fluency or reliably teach a new output protocol.
The four rungs
## “Our own data” can mean four different systems.
The more of the foundation you change, the more data rights, compute, evaluation and rollback discipline you inherit.
01 · Inference layer
### [RAG + prompting](https://isaiuseful.com/rag.html.md)
Retrieve documents at run time. Best for changing facts, private records and answers that need citations.
**Data: governed documents**
Does not place the knowledge reliably inside the weights.
02 · Post-training
### SFT / LoRA / preference tuning
Train on demonstrations or preference pairs to change format, tone, policies and task behavior.
**Data: labelled examples**
Can overfit style or damage other capabilities without a regression suite.
03 · Foundation adaptation
### Continued pretraining
Resume the next-token objective on a licensed language or domain corpus, then post-train again.
**Data: large raw corpus**
Useful for vocabulary and cultural grounding; materially harder than SFT.
04 · Full foundation
### Training from scratch
Design the tokenizer and recipe, initialize weights and pretrain across a massive balanced corpus.
**Data: foundation-scale mix**
Only justified when sovereignty and underrepresentation outweigh the programme risk.
Method boundaries draw on [Meta’s Llama 3 report](https://ai.meta.com/research/publications/the-llama-3-herd-of-models/) , [Google’s Gemma tuning guidance](https://ai.google.dev/gemma/docs/tune) , the [InstructGPT paper](https://arxiv.org/abs/2203.02155) and public adaptation reports for [ALLaM](https://arxiv.org/abs/2407.15390) , [NorwAI](https://arxiv.org/abs/2601.03034) and [EuroLLM-22B](https://arxiv.org/abs/2602.05879) .
Model architecture
## Another scaling axis: reuse depth.
A looped or recurrent-depth Transformer stores fewer unique layers and applies some of them repeatedly. The proposition is fewer unique weights, with more sequential computation available for a problem.
Conventional Transformer
**Layer 1 *→* Layer 2 *→* Layer 3 *→* Layer 4**
Each layer normally owns a separate set of parameters.
Looped Transformer
**Prelude *→* [shared block × several passes] *→* Coda**
The hidden state is refined by repeatedly applying some of the same parameters.
Unique parameters **+** training data **+** recurrent computation
The recurrence is latent computation inside one model computation. It is not the model printing a longer chain of thought.
Why it is interesting
### Effective depth without storing every layer.
Weight tying can reduce weight memory versus an equally deep untied model. Iterative refinement may suit reasoning and algorithmic tasks; implementations with variable passes can expose adaptive test-time compute. That is worth testing on memory-constrained local or edge systems.
What it costs
### Memory saved is not compute saved.
Extra passes usually add latency and accelerator work. Fewer unique weights may hold less factual capacity, gains can plateau, and training stability, halting, KV-cache design and serving support remain active engineering problems. A model that fits can still run slowly.
What it does not replace
### Architecture is only one system layer.
Recurrence does not replace retrieval, tools, memory, agents or post-training. A national-language model still needs strong language, culture, instruction and evaluation data; the architecture choice is separate from the language-data strategy.
Three different loops
### Ask what repeats—and where.
1 · Neural recurrence
### Looped Transformer
The model reapplies shared neural-network blocks during one forward computation.
`hidden state → shared block→ refined state → shared block → output`
2 · Generated reasoning
### Reasoning-token loop
A conventional autoregressive model emits extra reasoning or scratchpad tokens before the answer.
`token → token → token → answer`
3 · Product orchestration
### Agent loop
An external harness calls a model, tools and memory repeatedly. Its persistence says nothing conclusive about recurrent blocks inside the model.
`plan → act → observe → update→ verify → repeat`
Research checkpoints you can run
### Study the architecture before betting a product on it.
Most approachable
### Ouro
`ByteDance/Ouro-1.4B` and `ByteDance/Ouro-1.4B-Thinking` are open Looped Language Models pretrained for iterative latent computation; official 2.6B base and Thinking variants are also public. Treat them as research models and compare against a mature conventional model at similar runtime cost.
- [Inspect the official checkpoints →](https://huggingface.co/collections/ByteDance/ouro)
Variable depth
### Huginn
`tomg-group-umd/huginn-0125` is an approximately 3.5B-parameter recurrent-depth proof of concept. Its official implementation exposes recurrence depth, making it useful for architecture experiments—not a polished default local assistant.
- [Inspect Huginn and its usage notes →](https://huggingface.co/tomg-group-umd/huginn-0125)
Paper only · checked 27 July
### Loopie
The July 2026 paper reports layer-level recurrence in MoE models with 6B total / approximately 0.6B active parameters and 20B total / approximately 2B active. No official weights or serving code were linked or discoverable at this check, so this is a significant research result—not yet a deployment recommendation.
- [Read the Loopie paper →](https://arxiv.org/abs/2607.16051)
**Does Fable or Mythos use recurrent depth?**
**It is not publicly known.** Anthropic describes Claude Fable 5 and Claude Mythos 5 as the same underlying model with different safeguards and access arrangements, and reports unusually strong long-horizon autonomy. Anthropic has not publicly identified a Looped Transformer, recurrent-depth block, OpenMythos-style recurrence or another hidden-state looping design. The observed persistence can also come from long-horizon training, adaptive reasoning effort, an agent harness, context compaction, persistent files and notes, sub-agents, repeated verification and training to recover after failures.
[OpenMythos](https://github.com/kyegomez/OpenMythos) describes itself as an independent, community-built theoretical reconstruction based on public research and speculation. It is not leaked Anthropic code and is not evidence of Anthropic’s architecture.
**Behaviour can suggest an architectural hypothesis, but persistence observed through an agent product is not enough to reverse-engineer the neural architecture underneath it.**
Training decision
### Fine-tuning usually cannot create native recurrent depth.
| Intervention | What it can do | Architecture boundary |
| --- | --- | --- |
| [RAG](https://isaiuseful.com/rag.html.md) | Supply current or private evidence at inference time. | Does not change model depth. |
| LoRA / SFT | Specialise the behaviour of a looped checkpoint. | Normally does not convert a conventional Transformer into native recurrent depth. |
| Continued pretraining | Adapt an existing looped checkpoint to a language or domain. | Preserves the checkpoint’s basic architecture. |
| Training from scratch | Design and pretrain a genuinely new recurrent architecture. | The cleanest route—and the highest programme burden. Retrofitting recurrence into pretrained models exists, but remains experimental. |
**Decision rule**
**Start with an existing looped checkpoint when studying the architecture.** Do not redesign a national or enterprise model around recurrence until it beats a conventional baseline on the same data, compute, latency and task suite.
Bounded experiment
### Compare completed work, not parameter labels.
1. **Pair the models.** Use Ouro 1.4B Thinking and a mature conventional 1–4B model at comparable precision and on the same hardware.
2. **Test the work.** Measure arithmetic and algorithms, multi-step instructions, retrieval-grounded QA, Estonian quality and tool-call formatting.
3. **Measure operations.** Record completed-task accuracy, latency, peak memory and total compute or energy where measurable.
4. **Vary depth carefully.** Change recurrent passes only where the official implementation supports it; record where quality improves, plateaus or falls.
**Evidence boundary.** Research indicates recurrent depth can add useful latent computation without proportional growth in stored parameters. It does not establish that every task benefits, that a small looped model universally equals a much larger conventional one, or that recurrent depth is the successor to ordinary Transformers. The 2018 Universal Transformer is an important predecessor; current systems extend the idea with modern pretraining, variable depth and serving research.
- [Read the Ouro paper →](https://arxiv.org/abs/2510.25741)
- [Read the Huginn paper →](https://arxiv.org/abs/2502.05171)
- [Read Universal Transformers →](https://arxiv.org/abs/1807.03819)
- [Review experimental retrofitting →](https://arxiv.org/abs/2605.23872)
- [Read Anthropic’s Fable/Mythos disclosure →](https://www.anthropic.com/news/claude-fable-5-mythos-5)
Sub-1B specialization
## Small models become useful when the job becomes specific.
A sub-1B model is rarely convincing as a miniature general-purpose chatbot. It can be a credible, cheap language-processing component inside ordinary software.
**Useful mental model**
**Large models solve unfamiliar problems. Micro-models perform familiar jobs extremely cheaply.** The smaller the model, the narrower and better-tested its contract should be.
Understand
### Turn messy language into known fields.
Classify intent or documents, rewrite a query, and extract entities, requirements or metadata filters.
Intent
Entities
Filters
Route
### Connect language to deterministic software.
Select a search path, API or internal function; rerank a small candidate set; and emit a validated JSON plan.
Tools
Reranking
JSON
Present + guard
### Finish a bounded, grounded artifact.
Write a short cited summary or template, flag spam or sensitive content, and support local or offline actions where the device permits.
Templates
Moderation
Local
**Architectural correction**
**Fine-tune the behavior; retrieve the facts.** Teach customer language, intent labels, filter schemas, tool traces, output contracts and tone. Keep prices, policies, compatibility rules and catalogue content in pages, databases, APIs or search indexes, then retrieve verified passages at run time.
> Visual: Reference micro-model search and document workflow
**Visual reading order:**
1. **01** **Ask** Receive the customer’s natural-language question.
2. **02** **Interpret** Classify intent, extract filters and write a search plan.
3. **03** **Retrieve** Use lexical search for exact matches and vectors for semantic candidates.
4. **04** **Rerank** Select the few passages that best support the requested decision.
5. **05** **Compose** Give the model only allowlisted evidence and stable source IDs.
6. **06** **Validate** Check schema, claims, calculations and permissions before rendering.
**Browser delivery boundary**
In-browser inference removes the server inference queue and can keep text on the device, but it shifts model download, memory, battery and compatibility costs to the visitor. Make a substantial download explicit and keep a WASM, server or non-AI fallback.
| Parameter band | 4-bit weights-only floor | Practical public-website boundary |
| --- | --- | --- |
| Up to 150M | Up to ≈75 MB | Easy to justify for classifiers, embeddings, entity extraction and specialized transformations when the measured feature earns the download. |
| 270M–360M | ≈135–180 MB | Reasonable for an explicit AI-powered feature after opt-in, progress feedback and testing on representative phones and laptops. |
| Around 500M | ≈250 MB | Viable for a valuable local feature, but the visitor should knowingly start the download and have a graceful fallback. |
| Around 1B | ≈500 MB | Technically possible, but usually too heavy for an invisible enhancement on a normal public site. |
The size column is arithmetic, not a package quote: four-bit weights require roughly 0.5 bytes per parameter. Tokenizers, metadata and runtime files increase the download; activations and the KV cache increase working memory. Real speed, memory pressure and output quality depend on the exact model, quantization, context, browser and device.
**What the sources establish.** Google describes FunctionGemma as a specialized 270M function-calling base intended for further workflow-specific tuning and local agents; its own guide says smaller models need tuning to learn intent reliably. The original RAG paper separates parametric model memory from retrieved non-parametric memory. Transformers.js documents browser pipelines for classification, token classification, feature extraction, generation and summarization; ONNX Runtime Web documents WASM and WebGPU execution; WebLLM documents browser-side decoder-model inference and caching. These capabilities make the architecture feasible—not automatically accurate on your workflow.
- [Inspect FunctionGemma’s 270M design →](https://ai.google.dev/gemma/docs/functiongemma)
- [Read Google’s tuning boundary →](https://ai.google.dev/gemma/docs/functiongemma/finetuning-with-functiongemma)
- [Read the original RAG paper →](https://papers.neurips.cc/paper_files/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html)
- [Build a retrieval system →](https://isaiuseful.com/rag.html.md)
- [Review Transformers.js task support →](https://huggingface.co/docs/transformers.js/main/pipelines)
- [Review ONNX Runtime Web →](https://onnxruntime.ai/docs/tutorials/web/)
- [Review WebLLM’s browser runtime →](https://github.com/mlc-ai/web-llm)
Workflow loop
## A useful model programme is an evaluation loop.
The deploy gate is not the end. Production failures become new tests; they do not flow directly into training data.
01
**Define the workflow**
Owner, boundary, critical failures
02
**[Benchmark the base](https://isaiuseful.com/benchmarks.html.md#database)**
Public + local gold set
03
**Curate lawful data**
Rights, provenance, deduplication
04
**Adapt one rung**
Version every recipe and artifact
05
**Evaluate regressions**
Quality, safety, cost, latency
06
**Red-team tools**
Permissions, injection, exfiltration
07
**Pilot behind a gate**
Human review and rollback
08
**Monitor drift**
Log misses; curate the next set
Inference engineering
## Training is only half the model programme.
Products, evaluations, synthetic-data generation and reinforcement-learning rollouts all run inference. They do not share the same latency, throughput, cost or reproducibility target.
Online serving
### Protect user latency under load.
Measure time to first token, inter-token latency, p95/p99 tails, throughput per accelerator, errors and cost per accepted result with realistic prompt lengths, output lengths and concurrency.
TTFT + ITL
Tail latency
Availability
Evals + synthetic data
### Maximize useful, reproducible output.
Offline generation can trade single-request latency for batching and aggregate throughput. Pin the model, dataset, prompt template, sampler and engine, then retain outputs so a score or corpus can be reproduced.
Batch throughput
Cost
Reproducibility
RL + post-training
### Treat rollout as a distributed data path.
Rollout workers must serve the intended policy and reward models, refresh weights safely and return versioned trajectories. Slow or unstable inference can idle the rest of the training loop.
Rollouts
Weight updates
Fault recovery
| Layer | Learn first | Working proof |
| --- | --- | --- |
| 1 · Model foundation | Modern Python, PyTorch tensor execution and Transformer anatomy: attention, feed-forward or expert blocks, tensor shapes, dtypes and memory use. | Load a pinned model and reproduce reference outputs and numerical tolerances. |
| 2 · Serving engines | SGLang and vLLM request paths; prefill versus decode; KV-cache allocation; continuous batching; prefix caching; structured output and speculative decoding. | Serve the same supported model through both engines, preserve the API contract and compare TTFT, inter-token latency and throughput. |
| 3 · Performance | Profiling, roofline reasoning, GPU memory and bandwidth, kernel launches and synchronization. Add CUDA or Triton for kernel work and C++ for native extensions, bindings and framework internals. | Locate one measured bottleneck, change one variable and keep correctness within an explicit tolerance. |
| 4 · Distributed scale | Processes, queues, backpressure and failure recovery; collectives and tensor, pipeline, data or expert parallelism; NCCL plus the role of NVLink and InfiniBand. | Explain when the run is compute-, memory- or communication-bound, then demonstrate multi-GPU scaling and restart behavior. |
| 5 · Production evidence | Load generation, request tracing, numerical-stability tests, regression suites, admission control, observability, release pinning and rollback. | Publish a workload-specific service target and a benchmark report with model, engine, hardware, quantization, traffic shape and failure tests. |
**Do you need C++?**
Not to begin. Python, PyTorch, Transformer inference math and disciplined benchmarking are the entry layer. C++ becomes important when you change native framework code, bindings, memory movement or kernel launch paths; CUDA or Triton matters when you change the GPU kernels themselves. An application developer consuming an inference API may never need that depth, while an inference-framework engineer eventually will.
**Start from systems mechanics, then learn an engine.** SGLang and vLLM are production-oriented serving frameworks, not substitutes for understanding prefill, decode, memory and communication. Exact model, hardware, quantization and feature support changes quickly, so pin the tested releases and compare them on your own traffic.
- [Study the systems view of model scaling →](https://jax-ml.github.io/scaling-book/)
- [Read the SGLang architecture and guides →](https://docs.sglang.io/)
- [Use SGLang’s serving benchmark guide →](https://docs.sglang.io/docs/developer_guide/bench_serving)
- [Read the vLLM documentation →](https://docs.vllm.ai/en/stable/)
- [See where PyTorch C++/CUDA extensions fit →](https://docs.pytorch.org/tutorials/advanced/cpp_extension.html)
- [Compare the serving-tool records →](https://isaiuseful.com/tools.html.md#tools-run-models-locally)
Language + culture
## A national model is more than fluent output.
It should work across local language, institutions, culture, safety norms and actual public or enterprise tasks—and preserve evidence of where its data came from.
Foundation corpus
### Language and world model
Licensed web, books, news, archives, Wikipedia, parliamentary and legal text, science, maths and code.
Parallel corpus
### Cross-language coverage
Translation memories, bilingual text and careful translation of high-value material for lower-resource coverage.
Instruction set
### Useful behavior
Local QA, summarisation, public-service workflows, document work, coding and structured-output demonstrations.
Preference + safety
### Boundaries
Ranked answers, refusal edge cases, abuse prompts, jailbreaks and culturally grounded safety judgments.
Tool traces
### Actions
Function-call schemas, tool results, recovery paths and multi-step workflow traces with permission boundaries.
Evaluation sets
### Anti-self-deception
Held-out local tests for idioms, geography, institutions, law, culture, safety, cost and production tasks.
**Open-data example · checked 14 August 2026.** [OpenWALDO](https://openwaldo.org/) is developing a public corpus and training toolchain that carries source, license, count and hash records from selected data into model artifacts. It is worth inspecting if “open weights” are not enough for your definition of open AI; its license identifiers are assertions, not legal proof, and the resulting model still needs independent quality and safety evaluation.
**European starting point**
[EuroLLM-22B](https://arxiv.org/abs/2602.05879) covers all 24 official EU languages plus 11 additional languages, including Estonian. That makes it a relevant base or benchmark candidate—not automatic proof that it passes your local tasks.
Interactive compute + memory planner
## See the order of magnitude before the purchase order.
The compute estimate uses the common dense-transformer heuristic of roughly 6 × parameters × training tokens. It is planning math—not a quote or a promise.
Total accelerator-hours
**—**
Idealized elapsed time
**—**
Compute rental
**—**
Accelerator electricity
**—**
Approximate memory per model replica
**—**
Planning estimate
Base weights
**—**
Train state
**—**
Activations/runtime
**—**
This excludes data engineering, storage, networking, checkpoints, failed runs, evaluation, staff and serving.
The memory visual assumes a 4-bit base plus 15% load overhead for QLoRA, a BF16 base for LoRA, and about 16 bytes per parameter before activation reserve for full Adam-style training. Sharding changes per-device fit; long sequences and large batches can make activation memory much higher.
Public scale references
## The final run is not the programme.
Published disclosures show orders of magnitude, not transferable price quotes. Different architectures, data mixes and clusters make direct cost comparisons approximate.
> Visual: Log-scale comparison of selected published training workloads
**Visual entries (display order):**
- Focused instruct tuning **NorwAI · 7.5B** **38.4 H100-hours**
- Continued pretraining **NorwAI disclosed runs** **4.4K–31.1K H100-hours**
- Final foundation run **BLOOM · 176B** **1.083M A100-hours**
- Wider research programme **BigScience / BLOOM** **3.46M GPU-hours**
Sources: [NorwAI technical report](https://arxiv.org/abs/2601.03034) ; [BLOOM carbon and training disclosure](https://jmlr.org/papers/volume24/23-0069/23-0069.pdf) ; and the wider programme accounting in [Estimating the Carbon Footprint of BLOOM](https://arxiv.org/abs/2211.02001) . Accelerator generations and accounting boundaries differ.
Evaluation contract
## A tuned model can improve and still be worse.
Ship only when the target gain survives general capability, safety, cost and production-like checks.
Target
### Did the intended task improve?
Held-out local prompts, exact output contract and representative languages.
Regression
### What did the model forget?
General reasoning, multilingual performance, calibration and base-model strengths.
Safety
### Did refusals or tool behavior move?
Injection, sensitive data, dangerous requests, permissions and false tool calls.
Operations
### Can you afford the new behavior?
Tokens, latency, memory, energy, concurrency, retries and human review time.
Release
### Can you reproduce and roll back?
Data manifest, code, base hash, adapter, hyperparameters, eval artifacts and owner.
- [Choose a public benchmark and build a local set →](https://isaiuseful.com/benchmarks.html.md)
- [Use the NIST AI RMF measure and governance outcomes →](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/)
What failed
## Six ways a model programme goes wrong.
These are not arguments against training. They are reasons to preserve evidence, human data, rollback paths and a cheaper baseline.
Wrong intervention
### Fine-tuning was used as a database.
Facts still became stale, citations disappeared and each update required another training run.
**Recovery: retrieval first; tune behavior only.**
Synthetic recursion
### The model learned from its descendants.
Nature experiments found that indiscriminate recursive training on model-generated data loses distribution tails and degrades later models.
- [Read the primary study →](https://www.nature.com/articles/s41586-024-07566-y)
Budget fiction
### Only the successful run was costed.
BLOOM’s final run used about 1.08M A100-hours; the wider project accounted for 3.46M GPU-hours.
- [Inspect the accounting boundary →](https://arxiv.org/abs/2211.02001)
Leaderboard overfit
### The public score rose; the local task did not.
Prompt templates, contamination and harness choices can move scores without improving the production distribution.
- [Review benchmark traps →](https://isaiuseful.com/benchmarks.html.md#failure-modes)
Rights afterthought
### The corpus could not be documented or reused.
Unclear copyright, personal-data basis or source provenance can stop release after compute has already been spent.
**Recovery: make the data manifest a release artifact.**
Serving blind spot
### The model trained successfully and failed economically.
A larger model or longer reasoning trace raised latency, memory and review cost beyond the workflow’s value.
**Recovery: benchmark total cost per accepted result.**
Open models + distillation
## Jensen Huang’s case—and the boundary around it.
Huang argues that learning from other systems is fundamental. The legal and operational question is not whether distillation exists, but what data, contract, privacy and intellectual-property permissions govern a specific use.
- [Video: Axios interview with NVIDIA CEO Jensen Huang](https://www.youtube.com/watch?v=fr1IQspixmM)
Axios · 22 July 2026
### “Learning from AI … is fundamental to intelligence.”
Huang said AI systems will increasingly learn from other AI-generated knowledge and argued that policy should target contract, privacy or other misconduct rather than prohibit the technique broadly.
- [Read Axios’s report →](https://www.axios.com/2026/07/22/nvidia-jensen-huang-china-open-source-ai)
Open-weights letter · 24 July 2026
“The world needs both frontier closed models and frontier open models.”
NVIDIA joined 26 other named organizations in a letter arguing that open weights expand access, competition, control, safety research and sovereignty. The letter also acknowledges that released weights are hard to trace or reverse and calls for targeted legal and commercial treatment of unlawful extraction.
This is an advocacy position signed by NVIDIA and others—not neutral evidence that every open release is safe or lawful.
- [Open Jensen Huang’s post →](https://x.com/JensenHuang/status/2080643682408321103)
- [Read the signed three-page letter →](https://isaiuseful.com/downloads/Open-Weights-and-American-AI-Leadership.pdf)
Technique
**Teacher outputs improve a student**
Distillation can compress capability, create training examples, evaluate or validate another model.
Permission
**A specific use is authorized**
Terms, access controls, privacy, copyright, trade-secret and competition rules still apply to the actual collection and use.
Edge cases + obligations
## The hard costs sit outside the training script.
This is operational guidance, not legal advice. Scope obligations with qualified counsel and the competent authority for the actual provider, model, system and market.
EU GPAI · provider
### Document the model and training content.
EU guidance says GPAI providers must maintain technical documentation, support downstream providers, implement a Union-copyright policy and publish a sufficiently detailed training-content summary.
- [Read Commission guidance →](https://digital-strategy.ec.europa.eu/en/faqs/guidelines-obligations-general-purpose-ai-providers)
EU Article 50 · provider + deployer
### Make AI interaction and generated content legible.
From 2 August 2026, covered providers and deployers face transparency duties including disclosure of AI interaction, machine-readable marking and notices for specified deepfakes or public-interest content.
- [Open the July 2026 guidelines →](https://digital-strategy.ec.europa.eu/en/news/commission-publishes-guidelines-transparency-obligations-providers-and-deployers-certain-ai-systems)
Personal data
### A model is not anonymous by assertion.
The EDPB says anonymity and legitimate-interest analysis are case-specific, including whether people can be identified or personal data extracted by querying the model.
- [Read Opinion 28/2024 →](https://www.edpb.europa.eu/documents/opinion-of-the-board-art-64/opinion-282024-on-certain-data-protection-aspects-related-to_en)
Copyright + provenance
### Track rights before tokenization.
The EU AI Act keeps training-content summary and copyright-policy duties relevant even for many open-weight routes. Store source, licence, collection date, restrictions and transformations.
- [Read the AI Act →](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex:32024R1689)
Energy
### Measure experiments and serving, not just the final run.
NIST calls for documented energy, water and emissions impacts. The IEA reports that AI-focused data-centre electricity use grew 50% in 2025 while per-task efficiency improved rapidly.
- [Read IEA 2026 analysis →](https://www.iea.org/reports/key-questions-on-energy-and-ai/executive-summary)
Integration
### The model is not the product.
Budget corpus pipelines, evaluation, retrieval, serving, tool permissions, observability, security review, incident response, user support and model retirement.
- [Map the lifecycle with NIST →](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/)
2 Aug 2025
**GPAI duties started applying**
2 Aug 2026
**Commission enforcement + Article 50 transparency**
2 Aug 2027
**Training summaries due for covered pre-Aug-2025 GPAI models**
The Commission’s training-content template FAQ says failure to publish a required summary can trigger enforcement from 2 August 2026, with potential fines up to 3% of prior-year worldwide turnover or €15 million, whichever is higher. [Check the official scope, transition and penalty guidance](https://digital-strategy.ec.europa.eu/en/faqs/template-general-purpose-ai-model-providers-summarise-their-training-content) .
Who builds what
## Four teams, four definitions of progress.
A model programme fails when every team thinks the deliverable is “the model.”
Managers
### Fund a measured workflow.
- [Baseline RAG](https://isaiuseful.com/rag.html.md#evaluate) before approving training.
- Choose one sovereign or language-sensitive use case.
- Own rights, risk appetite and the deploy gate.
Developers
### Own behavior and reproducibility.
- Datasets → Transformers → PEFT / TRL → evals.
- Version prompts, data, adapters and output contracts.
- Hand the inference track a representative workload, not a demo prompt.
Platform teams
### Operate an AI factory, not a GPU rack.
- Storage, network, scheduler and checkpoint recovery.
- Isolate training from production inference.
- Measure utilization, energy and cost per accepted output.
Partners
### Buy speed without surrendering evidence.
- Require exportable weights or adapters where promised.
- Keep your data manifest and evaluation set.
- Separate service claims from reproducible artifacts.
### European commercial route
Mistral markets customisation, self-hosting and work from fine-tuning through pretraining. Use it to buy delivery speed while negotiating data, evaluation, portability and operating boundaries.
- [Review Mistral custom model training →](https://mistral.ai/solutions/custom-model-training/)
### Ecosystem route
EuroLLM, Hugging Face, public corpora, universities, national libraries, media archives and design partners build more internal capability—but require stronger programme ownership.
- [Hugging Face documentation →](https://huggingface.co/docs)
- [EuroLLM model organization →](https://huggingface.co/utter-project)
AI Now Summit 2026 · video library
## See how custom AI moves from model to institution.
These 32 Mistral-hosted talks are first-party conference perspectives, organized by the decision they can inform. They are useful implementation context—not independent evidence that every deployment claim generalizes. Pair the physical-AI talk with the [task-first robotics route](https://isaiuseful.com/robotics.html.md#route) and its field acceptance gate.
Showing all 32 talks.
- [Video: Opening keynote](https://www.youtube.com/watch?v=IC0VNOzPZU8)
Strategy · 42:44
**Opening keynote**
- [Video: A CIO’s vision on the AI industrial revolution](https://www.youtube.com/watch?v=Tb7RLf9qBlc)
Strategy · 16:14
**A CIO’s vision on the AI industrial revolution**
- [Video: How EDF and Mistral are reinventing France’s electricity](https://www.youtube.com/watch?v=oypLpWb74PY)
Industry · 14:20
**How EDF and Mistral are reinventing France’s electricity**
- [Video: How HTX is scaling AI for public safety](https://www.youtube.com/watch?v=33hFF-IZ6I0)
Public systems · 17:00
**How HTX is scaling AI for public safety**
- [Video: How agentic AI is reinventing large organizations](https://www.youtube.com/watch?v=MJf_qB3waqc)
Agents · 19:41
**How agentic AI is reinventing large organizations**
- [Video: Transforming a global group with sovereign technology](https://www.youtube.com/watch?v=KlI08FYHaXc)
Industry · 9:50
**Transforming a global group with sovereign technology**
- [Video: Perspectives on AI for defense](https://www.youtube.com/watch?v=zR_jp_D7eWg)
Public systems · 20:19
**Perspectives on AI for defense**
- [Video: Scaling AI to support the energy transition](https://www.youtube.com/watch?v=yZ5g7Jic714)
Industry · 25:44
**Scaling AI to support the energy transition**
- [Video: How nations can develop resilient, citizen-centric AI systems](https://www.youtube.com/watch?v=d5jLSGn2oj8)
Public systems · 49:00
**How nations can develop resilient, citizen-centric AI systems**
- [Video: The machines behind the machines](https://www.youtube.com/watch?v=usu2729-3yA)
Models · 20:29
**The machines behind the machines**
- [Video: Scaling enterprise value with sovereign AI](https://www.youtube.com/watch?v=HICHSfDkBzs)
Industry · 22:26
**Scaling enterprise value with sovereign AI**
- [Video: A culture-driven approach to innovation](https://www.youtube.com/watch?v=gIC1ftvrGwo)
Industry · 17:44
**A culture-driven approach to innovation**
- [Video: Scaling secure and transparent workflows](https://www.youtube.com/watch?v=9IsgWBMyp50)
Agents · 21:31
**Scaling secure and transparent workflows**
- [Video: How Airbus is powering Europe’s AI industrial revolution](https://www.youtube.com/watch?v=PdeLBAwb9IQ)
Industry · 14:28
**How Airbus is powering Europe’s AI industrial revolution**
- [Video: How La Banque Postale is building a human-centric future for banking](https://www.youtube.com/watch?v=OnpTJd9ZtrM)
Industry · 24:56
**How La Banque Postale is building a human-centric future for banking**
- [Video: Talk by Benjamin Haddad, Minister Delegate for Europe](https://www.youtube.com/watch?v=wz8XFHJvmE4)
Public systems · 9:42
**Talk by Benjamin Haddad, Minister Delegate for Europe**
- [Video: Shaping France’s AI future: CDC’s roadmap to digital autonomy](https://www.youtube.com/watch?v=ImQyRxgBCho)
Public systems · 16:52
**Shaping France’s AI future: CDC’s roadmap to digital autonomy**
- [Video: Mistral models powering Alexa+](https://www.youtube.com/watch?v=nYl9_TWxN7U)
Industry · 25:12
**Mistral models powering Alexa+**
- [Video: Closing keynote: The end of AI as we know it](https://www.youtube.com/watch?v=dEWPTAXC3JA)
Strategy · 5:26
**Closing keynote: The end of AI as we know it**
- [Video: Co-developing scaffolds and models hand-in-hand](https://www.youtube.com/watch?v=2q_MTMAuu8w)
Models · 28:08
**Co-developing scaffolds and models hand-in-hand**
- [Video: AI for earth observation: Building the future with EVE](https://www.youtube.com/watch?v=DRCkXZtndcI)
Models · 28:09
**AI for earth observation: Building the future with EVE**
- [Video: AI infrastructure is the new critical infrastructure](https://www.youtube.com/watch?v=QxgcNjBO580)
Strategy · 29:52
**AI infrastructure is the new critical infrastructure**
- [Video: Luxembourg’s sovereign AI playbook for Europe](https://www.youtube.com/watch?v=0lqZpQZLGKs)
Public systems · 28:27
**Luxembourg’s sovereign AI playbook for Europe**
- [Video: Building interconnected AI for complex operations](https://www.youtube.com/watch?v=rZGW2G_CX6M)
Agents · 27:56
**Building interconnected AI for complex operations**
- [Video: Channelling the power of LLMs to unlock ancient archives](https://www.youtube.com/watch?v=H2Vnx50Oves)
Models · 28:08
**Channelling the power of LLMs to unlock ancient archives**
- [Video: Mistral robotics and physical AI](https://www.youtube.com/watch?v=qwTKCxHxakc)
Models · 14:33
**Mistral robotics and physical AI**
- [Video: Rethinking the architecture of agentic systems](https://www.youtube.com/watch?v=Q1yn6khGCBk)
Agents · 21:20
**Rethinking the architecture of agentic systems**
- [Video: Best practices for building autonomous AI workflows](https://www.youtube.com/watch?v=lluXpzkLpZo)
Agents · 20:43
**Best practices for building autonomous AI workflows**
- [Video: Advancing innovation at the European Patent Office](https://www.youtube.com/watch?v=hZQa78vKxZ4)
Public systems · 27:27
**Advancing innovation at the European Patent Office**
- [Video: The AI sovereignty paradox: Scalable ecosystems for trusted adoption](https://www.youtube.com/watch?v=xP-3iMTPav0)
Strategy · 32:05
**The AI sovereignty paradox: Scalable ecosystems for trusted adoption**
- [Video: Domain AI models fine-tuned with proprietary knowledge](https://www.youtube.com/watch?v=7evOiuXFkQo)
Models · 29:24
**Domain AI models fine-tuned with proprietary knowledge**
- [Video: Building custom code models for Ericsson proprietary silicon](https://www.youtube.com/watch?v=ArWG4pmTXPQ)
Models · 36:12
**Building custom code models for Ericsson proprietary silicon**
The practical answer
Adapt a strong open base. Preserve the data trail. Let evaluation earn the next rung.
- [Choose the intervention](#chooser)
- [Budget the programme](#planner)
- [Choose the evaluation stack](https://isaiuseful.com/benchmarks.html.md)
---
## https://isaiuseful.com/robotics (`/robotics.html.md`)
# How Do You Build a Useful Robot?
Canonical source: [https://isaiuseful.com/robotics](https://isaiuseful.com/robotics)
Robotics builder’s guide + field watch · checked 1 August 2026
The useful robotics startup is not “a humanoid for everything.” It is one expensive, repetitive or unsafe task—inside one measurable environment—with a recovery path and economics that survive contact with the real world.
- [Choose a first task](#route)
- [Browse useful robots](#mainstream-robots)
- [See the humanoid frontier](#frontier)
- [**≈90K ft²** Apptronik Robot Park VENDOR-REPORTED TRAINING FACILITY](https://apptronik.com/news-collection/welcome-to-robot-park-where-apptroniks-apollo-goes-to-work)
- [**100K+ totes** Digit at GXO VENDOR-REPORTED LIVE-DEPLOYMENT OUTPUT](https://www.agilityrobotics.com/content/digit-moves-over-100k-totes)
- [**13,361** UWORLD U1 orders UBTECH-REPORTED AT LAUNCH · NOT DELIVERIES](https://www.prnewswire.com/news-releases/ubtech-launches-uworld-u1-the-worlds-first-full-size-mass-produced-ultra-bionic-humanoid-robot-302815285.html)
- [**10 h shifts** Figure 02 at BMW CUSTOMER-REPORTED PILOT SCOPE](https://www.bmwgroup.com/en/news/general/2026/humanoid-robot-in-leipzig.html)
- [**1M+ traces** AgiBot World OPEN ROBOT-LEARNING DATASET](https://github.com/OpenDriveLab/Agibot-World)
- [**12 releases** Figure · 2023 → 2026 BODY → VLA → BOTQ → FIELD WORK](https://www.figure.ai/news/ramping-figure-03-production)
Humanoids · August 2026
## What makes a humanoid robot useful at work?
Humanoid teams are coupling dexterous hardware, vision-language-action models, whole-body control, large robot datasets and production engineering. The real dividing line is no longer a polished motion—it is sustained work with a denominator.
- [Video: Brett Adcock on Figure’s eight-hour autonomous shift](https://www.youtube.com/watch?v=og0VF9-XCSs)
Figure package-sorting endurance test ·
14 May 2026
### From an eight-hour demo to named customer work.
Watch the endurance run here; use the named BMW and Catalyst deployments below to judge whether the capability is becoming sustained customer work.
What is independently attributable
The interview follows Figure’s autonomous package-sorting endurance run. Figure did not identify that May test as work for a postal carrier or at a customer site. BMW separately says Figure 02 worked ten-hour weekday shifts during a ten-month production pilot, moving more than 90,000 parts in about 1,250 operating hours.
On 26 May, Catalyst Brands announced a commercial Figure deployment at its Reno distribution center for sorting and packing around its Joey Pouch system—the closest verified customer match to the package-sorting demo.
**Evidence boundary:** Neither Figure nor Catalyst says the livestreamed endurance test was commissioned by Catalyst. The agreement also does not disclose robot count, intervention rate, uptime or all-in cost, so the announcement and demo remain separate evidence.
- [Read BMW’s pilot account →](https://www.bmwgroup.com/en/news/general/2026/humanoid-robot-in-leipzig.html)
- [Read Catalyst’s customer announcement →](https://corporate.jcpenney.com/2026/05/26/catalyst-brands-taps-figure-ai-for-humanoid-automation/)
- [Inspect Figure’s package-sorting note →](https://www.figure.ai/news/helix-logistics)
01
**Whole-body autonomy**
Coordinate feet, torso, arms, hands and balance from pixels instead of stitching together isolated skills.
02
**Dexterous sensing**
Add touch, force and palm-level vision for contact-rich work, occlusion and deformable objects.
03
**Data engines**
Scale teleoperation, human-motion transfer, simulation, failure capture and reusable robot datasets.
04
**Long-horizon work**
Preserve task state, recover from variation and complete minutes or shifts—not edited single attempts.
05
**Industrialization**
Drive cost, reliability, charging, serviceability, safety integration and manufacturing volume together.
**Do not stop at the video.**
1. Demo What capability is being claimed?
2. Technical note What data, architecture and test were used?
3. Open artifact Can code, weights, data or a benchmark be inspected?
4. Customer account Does the operator confirm the task and duration?
5. Field denominator Success, interventions, uptime, safety and cost per accepted cycle.
Figure · release timeline
## Three years of iteration, in order.
The sequence matters more than any single clip: body, factory hardware, upper-body VLA, data scaling, deformable objects, home hardware, full-body autonomy and a return to production. Every result below is vendor evidence unless a customer source is explicitly linked.
**2023**
- [Video: Introducing Figure](https://www.youtube.com/watch?v=b37rQZ4maPo)
01 Mar 2023
Mission
### Introducing Figure
The starting thesis: put a commercially useful general-purpose humanoid into human environments.
- [Figure release archive →](https://www.figure.ai/news)
**2024**
- [Video: Introducing Figure 02](https://www.youtube.com/watch?v=0SRVJaOg9Co)
06 Aug 2024
Hardware
### Introducing Figure 02
A second-generation body aimed at the factory: actuation, hands, sensing, compute and cable routing move toward deployment.
- [Read the later BMW results →](https://www.figure.ai/news/production-at-bmw)
**2025**
- [Video: Introducing Helix](https://www.youtube.com/watch?v=Z3yQHYNXPws)
20 Feb 2025
Generalist control
### Introducing Helix
One VLA connects language and vision to 200 Hz continuous upper-body control; Figure reports training on roughly 500 hours of teleoperated data.
- [Read the technical note →](https://www.figure.ai/news/helix)
- [Video: Scaling Helix Logistics](https://www.youtube.com/watch?v=lkc2y0yb89U)
07 Jun 2025
Data scaling
### Scaling Helix · logistics
More demonstrations plus visual memory, state history and force feedback target package variety, speed and barcode orientation.
- [Inspect results + ablations →](https://www.figure.ai/news/scaling-helix-logistics)
- [Video: Scaling Helix Laundry](https://www.youtube.com/watch?v=HOoRnv3lA0k)
12 Aug 2025
Deformables
### Scaling Helix · laundry
The same architecture, with a new dataset, moves from rigid parcels to cloth that folds, slips and changes geometry.
- [Read the method note →](https://www.figure.ai/news/helix-learns-to-fold-laundry)
- [Video: Scaling Helix Dishes](https://www.youtube.com/watch?v=8gfuUzDn4Q8)
03 Sep 2025
Household task
### Scaling Helix · dishes
Bimanual grasping, object placement and recovery move the learning system into fragile, cluttered household work.
- [Read the evaluation note →](https://www.figure.ai/news/helix-loads-the-dishwasher)
- [Video: Introducing Figure 03](https://www.youtube.com/watch?v=Eu5mYMavctM)
09 Oct 2025
Co-designed body
### Introducing Figure 03
Palm cameras, tactile fingertips, softer coverings, wireless charging and manufacturing design bring the body closer to Helix and the home.
- [Read the hardware account →](https://www.figure.ai/news/introducing-figure-03)
**2026**
- [Video: Introducing Helix 02](https://www.youtube.com/watch?v=lQsvTrRTBRs)
27 Jan 2026
Full body
### Introducing Helix 02
Figure reports a unified pixels-to-actions system controlling locomotion, manipulation and balance, including a four-minute 61-action kitchen sequence.
- [Read architecture + tests →](https://www.figure.ai/news/helix-02)
- [Video: Helix 02 Living Room Tidy](https://www.youtube.com/watch?v=CAdTjePDBfc)
09 Mar 2026
Long-horizon home
### Helix 02 · living room tidy
Room-scale navigation and object return test whether the full-body policy can preserve task state across a changing scene.
- [Read Figure’s test account →](https://www.figure.ai/news/helix-02-living-room-tidy)
- [Video: BotQ Ramping Figure 03 Production](https://www.youtube.com/watch?v=YZH1csMhnDo)
29 Apr 2026
Production system
### BotQ ramps Figure 03
Figure reports more than 350 third-generation robots delivered from BotQ and a demonstrated line cadence increase from one robot per day to one per hour. That is factory throughput—not proof of autonomous customer utilization.
- [Inspect yield, testing + fleet claims →](https://www.figure.ai/news/ramping-figure-03-production)
- [Video: Helix 02 Bedroom Tidy](https://www.youtube.com/watch?v=8xEuFQz4E4A)
08 May 2026
Generalization
### Helix 02 · bedroom tidy
A second room and new object distribution probe whether the same learned system transfers beyond one staged environment.
- [Read Figure’s test account →](https://www.figure.ai/news/helix-02-bedroom-tidy)
- [Video: Figure 03 at BMW](https://www.youtube.com/watch?v=tv90GFM9RAo)
30 Jun 2026
Return to factory
### F.03 at BMW
Helix 02 and Figure 03 meet a sequencing workflow that combines thin-part picking, stepping, placement and cart manipulation.
- [Read Figure’s account →](https://www.figure.ai/news/f-03-at-bmw)
- [Check BMW’s media record →](https://www.press.bmwgroup.com/global/video/detail/PF0010224/Figure-03-at-BMW-Group-Plant-Spartanburg)
Humanoid builders · different bets
## Home, factory, cost, data—and talent.
These companies are not racing on one scoreboard. The humanoid field spans supervised home service, vertically integrated manufacturing, cheaper hardware, commercial logistics, robot-data factories, consumer companions and the builder networks that supply the next generation of talent.
- [Video: Introducing Calvin-40](https://www.youtube.com/watch?v=YddS-aI097Q)
Wandercraft · Calvin‑40
### Turn medical-grade balance and torque into industrial work.
Founded in 2012, Wandercraft built hands-free, self-balancing medical exoskeletons before adding an upper body for Calvin. The company says that platform is engineered to move users weighing up to 100 kg; its safety-critical history is relevant provenance, not proof that Calvin is the industry’s most stable humanoid. Renault reports testing Calvin at Douai on a roughly 30 kg tire-handling task and targets about ten robots by the end of 2026 and 350 across French and Spanish factories by the end of 2027. Those are customer roadmap figures, not completed deployment, reliability or task economics.
- [Inspect the exoskeleton lineage →](https://en.wandercraft.eu/)
- [Inspect the current Calvin product route →](https://en.wandercraft.eu/calvin-40)
- [Read Renault’s factory account →](https://www.renaultgroup.com/en/magazine/technology/calvin-a-new-generation-robot-is-born/)
- [Video: NEURA Gym: The First Physical AI Training Center for Robots](https://www.youtube.com/watch?v=hUNujYlRmZU)
NEURA Robotics · NEURA Gym + 4NE1
### Make real-world training infrastructure part of the product.
NEURA announced a Series C with a total round size of up to $1.4 billion in June 2026; that is a maximum round size, not a $1.4 billion lifetime total or proof that every dollar has already translated into deployed robots. The supplied film presents NEURA Gym as a place for robots to collect multimodal real-world data and for companies to train applications. NEURA’s current route describes partner-specific training, validation and integration facilities opening across Europe and the United States. The infrastructure and funding are substantial, but neither supplies 4NE1 uptime, intervention rate or customer throughput.
- [Inspect the current NEURA Gym route →](https://neura-robotics.com/neuragym/)
- [Read the exact Series C disclosure →](https://neura-robotics.com/record-series-c/)
- [Video: UMA unveils its vision for the next generation of robots, Remi Cadene at Machina Summit](https://www.youtube.com/watch?v=0vc-dAQDIQc)
UMA · Northstar + real-time learning
### Combine open robot-learning talent with a new body.
UMA’s founding team includes LeRobot builders from Hugging Face and a founding member of Google DeepMind’s robotics team. The supplied July 2026 talk presents its Northstar design and a learning approach intended to acquire skills from human demonstration rather than task-by-task programming. UMA’s current site says logistics and manufacturing pilot programs will launch during 2026. A prototype and planned pilots are not product shipments: no public unit count, customer operating record, uptime or task economics is yet available.
- [Verify the team + pilot status →](https://uma.bot/)
- [Video: GENE.01 Design Concept at CES 2026](https://www.youtube.com/watch?v=zYRz1uIIpwQ)
Generative Bionics · GENE.01
### Make touch part of the body—then prove the industrial task.
The supplied CES film shows the January 2026 design concept; Generative Bionics now presents GENE.01 as a sensorized platform with tactile, force and vision sensing and publishes an open URDF model. Fincantieri independently confirms a four-year program to adapt the platform for shipyard welding, with first on-site tests scheduled by the end of 2026. That is credible current development and an inspectable artifact—not certification, deployment uptime or welding performance.
- [Inspect the current platform + robot model →](https://gbionics.ai/gene01/)
- [Read Fincantieri’s validation plan →](https://www.fincantieri.com/en/newsroom/press-releases/2026/fincantieri-and-generative-bionics-launch-an-industrial-partnership-to-develop-a-humanoid-welding-robot-for-shipyards)
- [Video: PAL Robotics Kangaroo Pro](https://www.youtube.com/watch?v=G7riD4yb1tg)
PAL Robotics · KANGAROO Pro
### Expose a configurable humanoid research platform.
Barcelona-based PAL Robotics still markets KANGAROO Pro and lists ROS 2, MuJoCo, Gazebo and Isaac Lab support across Lite, Standard and Plus configurations. The platform is positioned for research and advanced manipulation rather than verified production work; the published specifications and film do not supply uptime, intervention or customer-throughput denominators.
- [Inspect the current product →](https://pal-robotics.com/robot/kangaroo/)
- [Read the current datasheet →](https://pal-robotics.com/wp-content/uploads/2025/05/2025-Datasheet_ENG_KANGAROO-PRO.pdf)
- [Video: Meet Genesis AI Eno](https://www.youtube.com/watch?v=zab62_u-a5U)
- [Video: Announcing Genesis World 1.0](https://www.youtube.com/watch?v=ggfsGD1PCQA)
Genesis AI · Eno + Genesis World 1.0
### Co-design the body, hand, model and evaluation engine.
Eno uses a wheeled, height-adjustable body that folds compactly and carries proprietary human-scale hands with 20 active, back-drivable degrees of freedom. Genesis plans targeted customer deployments by the end of 2026, so this is a product roadmap rather than field proof. Its open Genesis World 1.0 stack makes a complementary bet: use GPU simulation for closed-loop evaluation, while training on real data. Genesis reports reducing a typical 200-plus-hour hardware evaluation to under 0.5 simulated hours; the company’s own sim-to-real calibration and task suite still set that comparison.
- [Read Eno’s launch + deployment plan →](https://www.genesis.ai/press/meet-eno)
- [Inspect Genesis World’s method →](https://www.genesis.ai/blog/the-role-of-simulation-in-scalable-robotics-genesis-world-10-and-the-path-forward)
- [Open the simulator repository →](https://github.com/Genesis-Embodied-AI/genesis-world)
- [Video: ▶ Europe’s robotics clubs](https://www.youtube.com/watch?v=OHCf75KS8i4)
Europe · builder network
### The frontier is also who gets a robot into their hands.
Prototype’s March 2026 account describes the European Student Robotics Association as 13 clubs across eight countries with more than 2,500 students. That is an ecosystem claim, not a product benchmark, but it matters: cheaper platforms, shared labs and cross-border teams determine how quickly Europe converts research talent into repeatable hardware practice.
- [Read the sourced club list →](https://updates.prototypecap.com/p/the-kids-are-building)
- [Open ESRA →](https://www.studentrobotics.eu/)
- [Video: Physical AI in Action at NVIDIA GTC 2026 | Humanoid Robotics Demo](https://www.youtube.com/watch?v=0oZAw6rryIE)
Humanoid · HMND 01 + KinetIQ
### Scale the company and fleet brain before commercial rollout.
Founded in 2024, Humanoid disclosed a $152 million Series A in July 2026, bringing its company-reported total raised to $270 million. The supplied GTC film shows two wheeled, gripper-equipped robots in a simulated store responding to voice requests, allocating work and coordinating handoffs through KinetIQ. That is a live multi-robot demonstration, not evidence of bipedal field autonomy. Humanoid says Beta robot rollout begins in Q4 2026, so the funding and industrial agreements are capacity signals rather than delivered-unit, uptime or cost-per-task evidence.
- [Read the GTC demonstration account →](https://thehumanoid.ai/humanoid-brings-voice-controlled-robotic-fulfillment-powered-by-kinetiqto-life-at-nvidia-gtc/)
- [Check funding + rollout timing →](https://thehumanoid.ai/humanoid-raises-152-million-at-1-35-billion-post-money-valuation-becoming-europes-first-pure-play-humanoid-robotics-unicorn/)
- [Video: Reflect v1.0 - The Path Towards Long-Horizon Autonomous Humanoid Work](https://www.youtube.com/watch?v=6dme3JYj3Hg)
Flexion · Reflect v1.0
### Compose navigation, manipulation and recovery into one mission.
The supplied film puts one humanoid through a multi-floor parcel mission involving stairs, an elevator, doors, tool use and recovery from errors after a single instruction. Flexion reports 90% end-to-end completion with reinforcement learning versus 38% with supervised fine-tuning alone on one 16-step internal evaluation. Its technical note also names the limits: a bounded task distribution, difficult grasps, visual assumptions and incomplete recovery coverage. This is a useful integrated vendor test of the intelligence stack—not independent replication or a customer production denominator.
- [Inspect architecture, evaluation + limits →](https://flexion.ai/news/flexion-reflect-v1.0)
- [Video: WSJ tries NEO](https://www.youtube.com/watch?v=f3c4mQty_so)
- [Video: 1X launch film](https://www.youtube.com/watch?v=LTYMWadOW7c)
1X · NEO
### Ship the home service before full autonomy.
1X’s own product page says early owners receive basic autonomy while unfamiliar complex tasks can be handled through scheduled remote expert supervision. That is a real product architecture—not a footnote. Redwood is described as a 160M-parameter onboard policy running at roughly 5 Hz; privacy, teleoperation availability and the rate at which remote help declines are central performance metrics.
- [Inspect NEO + Expert Mode →](https://www.1x.tech/neo)
- [Read the Redwood technical note →](https://www.1x.tech/discover/redwood-ai)
- [Video: Tesla Bot update](https://www.youtube.com/watch?v=XiQkeWOFwmk)
- [Video: Optimus Gen 2](https://www.youtube.com/watch?v=cpraXaw7dyc)
Tesla · Optimus
### Reuse the real-world AI and manufacturing machine.
The strategic bet is vertical: vision-based autonomy, in-house compute, actuators, factories and manufacturing scale under one roof. But the supplied public videos are from 2023. Tesla’s 2025 annual filing still describes Optimus as a humanoid “in development,” while its 2025 year-end update lists California Optimus manufacturing capacity as under construction. Treat production timing and capability forecasts as targets until deployed task data appears.
- [Read Tesla’s 2025 annual filing →](https://ir.tesla.com/_flysystem/s3/sec/000162828026003952/tsla-20251231-gen.pdf)
- [Check the capacity disclosure →](https://ir.tesla.com/_flysystem/s3/sec/000162828026003837/tsla-20260128-gen.pdf)
- [Video: Form and function of enterprise humanoid design: Boston Dynamics Atlas](https://www.youtube.com/watch?v=wL0-Pu_8F0U)
Boston Dynamics · Atlas
### Design the humanoid around industrial work.
Spun out of MIT’s Leg Lab by Marc Raibert in 1992, Boston Dynamics is based in Waltham, Massachusetts. Hyundai Motor Group owns 80% after a 2021 transaction that valued the company at $1.1 billion; SoftBank retains 20%. Production Atlas is 1.9 m tall, weighs 90 kg and lists 56 degrees of freedom, a four-hour battery and 30 kg sustained payload. Boston Dynamics says its 2026 production is committed to Hyundai’s RMAC and Google DeepMind. These are company specifications and scheduled deployments—not customer-reported uptime, intervention or cost-per-task.
- [Inspect Atlas specifications →](https://bostondynamics.com/products/atlas/)
- [Read the design brief →](https://bostondynamics.com/webinars/form-function-enterprise-humanoid-design/)
- [Review the company history →](https://bostondynamics.com/about/)
- [Check the deployment status →](https://bostondynamics.com/blog/boston-dynamics-unveils-new-atlas-robot-to-revolutionize-industry/)
- [Video: Gemini Robotics 2 brings whole body intelligence to robots](https://www.youtube.com/watch?v=4lSQnrMC6nY)
Apptronik · Apollo 2 + Robot Park
### Build a data factory around the fleet.
Apptronik’s expanded Austin Robot Park is a nearly 90,000 ft² training facility for fleets of Apollo 2 robots in bipedal and wheeled configurations. The company says related workflows run at Google DeepMind and customer sites including Mercedes‑Benz and GXO, mixing teleoperation, autonomous execution and simulation to train Gemini Robotics. Mercedes‑Benz independently confirms using Apollo in production-data collection and intralogistics tests. Neither source publishes fleet size, autonomy share, intervention rate or customer throughput.
- [Read Apptronik’s Robot Park report →](https://apptronik.com/news-collection/welcome-to-robot-park-where-apptroniks-apollo-goes-to-work)
- [Inspect the Apollo 2 platform →](https://apptronik.com/apollo/apollo-2)
- [Review Gemini Robotics 2’s results →](https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/)
- [Check Mercedes‑Benz’s account →](https://group.mercedes-benz.com/company/production/procuction-network/mbdfc-humanoid-robots.html)
- [Video: What Will You Build With Us?](https://www.youtube.com/watch?v=LM3hTXUShvA)
Agility Robotics · Digit
### Turn one logistics task into a commercial service.
Digit’s evidence is unusually operational for a humanoid: Agility reports more than 100,000 totes moved in GXO’s live facility, and Toyota Motor Manufacturing Canada signed a Robots‑as‑a‑Service agreement in February 2026 after a pilot. The Toyota agreement confirms a commercial step but omits robot count, price and deployment-wide reliability. The tote milestone is also vendor-reported and does not disclose the time period, intervention rate or number of robots behind the total.
- [Inspect the 100,000-tote claim →](https://www.agilityrobotics.com/content/digit-moves-over-100k-totes)
- [Read the Toyota agreement →](https://www.agilityrobotics.com/content/agility-robotics-announces-commercial-agreement-with-toyota-motor-manufacturing-canada)
- [Video: Introducing Generalist GEN-1](https://www.youtube.com/watch?v=SY2xyrmV44Y)
Generalist · GEN‑1
### Pretrain on human interaction, then adapt to the robot.
Generalist says GEN‑1 was pretrained on more than 500,000 hours of human physical-interaction data with no robot data, then adapted to each shown task with about one hour of robot data. Its own six-task evaluation reports 99% average success versus 64% for the previous GEN‑0 and roughly three-times-faster completion on selected tasks. The company also names unsolved tasks and offers early access; the figures remain vendor-run, task-specific evaluations rather than customer production metrics.
- [Inspect runs, baselines + limitations →](https://generalistai.com/blog/gen-1)
- [Video: Unitree R1](https://www.youtube.com/watch?v=mTMYfVD4zCw)
- [Video: Unitree GD01](https://www.youtube.com/watch?v=oWOyUMJWptc)
Unitree · R1 + GD01
### Compress the hardware cost—and widen the form factor.
R1 starts at a vendor-listed $4,900 before tax and shipping, weighs roughly 27–29 kg and lists about one hour of battery life. At the other extreme, GD01 is a 500 kg-class rideable biped/quadruped machine. Together they show Unitree’s hardware breadth, not general autonomy. Unitree itself warns that humanoids remain early-stage and that some demonstrated functions are still under development.
- [Check R1 configurations + caveats →](https://www.unitree.com/R1/)
- [Review Unitree’s release record →](https://www.unitree.com/about/)
- [Video: First batch of Henan-made EngineAI T800 robots rolls off the line](https://www.youtube.com/watch?v=1bTC_qfF7sA)
EngineAI · T800
### Pair athletic hardware with a manufacturing ramp.
Established in October 2023, headquartered in Shenzhen and led by founder and CEO Zhao Tongyang, EngineAI reports RMB 1 billion in cumulative Pre‑A++ and A1 funding, followed by A1+ and A2 rounds whose amounts it did not disclose. It lists four T800 editions from $40,500. The official video shows a first Henan-made batch leaving the Zhengzhou line; separately, the company says its 12,000 m² Shenzhen base can complete one humanoid every 15 minutes and supports a 10,000-unit scalable-delivery capability. Those are vendor capacity claims—not shipment totals, utilization data or customer-verified autonomous work.
- [Check T800 editions + caveats →](https://en.engineai.com.cn/product-t800.html)
- [Review the company profile →](https://en.engineai.com.cn/about-us)
- [Inspect the funding disclosure →](https://en.engineai.com.cn/about-news-media/38.html)
- [Inspect the capacity claim →](https://www.prnewswire.com/news-releases/engineai-launches-shenzhen-intelligent-manufacturing-base-as-first-batch-of-t800-humanoid-robots-roll-off-the-production-line-to-begin-mass-delivery-302785446.html)
- [Video: ▶ AGIBOT in precision manufacturing](https://www.youtube.com/watch?v=h6rCRa8qUFw)
AGIBOT · deployment + open data
### Scale the data engine and the factory rollout together.
AGIBOT’s strongest signal is the pairing of artifacts and operations. AgiBot World exposes more than one million robot trajectories across 217 tasks and five deployment scenarios. Separately, the company reports about 100 cumulative hours of a G2 factory livestream and a 15,000th robot production milestone in June 2026. Production count is not autonomous-task success, so keep shipment, utilization and intervention metrics separate.
- [Read AGIBOT’s production report →](https://www.agibot.com/article/231/detail/82.html)
- [Inspect data + GO-1 code →](https://github.com/OpenDriveLab/Agibot-World)
- [Read the AgiBot World paper →](https://arxiv.org/abs/2503.06669)
- [Video: Meet LimX Luna](https://www.youtube.com/watch?v=-lgo5xqgVko)
LimX Dynamics · Luna
### Package motion intelligence for public spaces.
Luna is a 160 cm, 56 kg interactive humanoid aimed at malls, museums, theme parks and live stages—not a factory generalist. LimX lists 27 active body degrees of freedom, up to 3 kg arm payload and about four hours of battery life, measured in its lab. The standard configuration uses fist-shaped ends; a six-DoF five-finger hand is optional. Video imitation, a no-code task editor and 200-plus-unit synchronized control are vendor capabilities, not independently measured autonomy or uptime.
- [Read the Luna launch →](https://www.limxdynamics.com/en/news/BK000062)
- [Check configurations + lab specs →](https://www.limxdynamics.com/en/products/luna/spec)
- [Video: Why XPENG IRON is so human-like](https://www.youtube.com/watch?v=ssxdSVQN14I)
XPENG · next-generation IRON
### Reuse an EV company’s chips, AI and production discipline.
XPENG’s 2025 model is specified at 82 whole-body degrees of freedom, including 22 per hand, with three in-house Turing AI chips. The company says IRON is being tested in its own production environment and targets mass production by the end of 2026. That target remains prospective: public material does not give delivered units, sustained autonomous cycles, intervention rate or all-in task cost.
- [Inspect XPENG’s technical claims →](https://www.xiaopeng.com/news/company_news/5511.html)
- [Read XPENG’s 2025 ESG disclosure →](https://ir.xiaopeng.com/static-files/10e70f1b-cbed-4fb1-9f2f-f6be521d72a2)
- [Video: Astribot Lumo-2](https://www.youtube.com/watch?v=Q0WMUDF3JKI)
Astribot · Lumo-2
### Reason over future dynamics without rendering the future.
Founded in Shenzhen in December 2022, Astribot pairs its S1 manipulation platform with Lumo‑2, a latent world-action model introduced in a July 2026 paper. The authors report gains over named VLA and world-action baselines on long-horizon and dexterous real-world tasks by aligning actions with learned latent dynamics, vision and language. This is author-reported research evidence; the release does not establish customer throughput, deployment uptime or broadly available model weights.
- [Read the Lumo‑2 paper →](https://arxiv.org/abs/2607.11270)
- [Review Astribot’s company account →](https://www.astribot.com/en/about)
- [Video: UBTECH Launches UWORLD U1](https://www.youtube.com/watch?v=jXRbNaFqByo)
UBTECH · UWORLD U1
### Move the humanoid pitch from industry to companionship.
UBTECH’s June 2026 launch introduces three U1 models spanning a semi-torso Lite and two full-body editions, with vendor pricing from RMB 119,800. The company reports 88 degrees of freedom and 13,361 cumulative orders at launch. Those are demand and specification claims—not delivered units, independent capability tests or evidence that the “mass-produced” positioning has translated into sustained work. Consumer interaction also creates a different evidence burden around privacy, reliability and safety.
- [Inspect UBTECH’s launch claims →](https://www.prnewswire.com/news-releases/ubtech-launches-uworld-u1-the-worlds-first-full-size-mass-produced-ultra-bionic-humanoid-robot-302815285.html)
**Evidence policy:** company pages and product films establish what each vendor claims; customer accounts from BMW, Mercedes‑Benz, GXO, Renault, Fincantieri and Toyota’s quoted manufacturing leadership corroborate specific pilots, development programs or agreements. Filings constrain Tesla and XPENG’s public status; Humanoid and NEURA’s funding disclosures, Flexion’s internal evaluation, UMA’s pilot timing, Apptronik’s fleet size, Agility’s tote denominator, UBTECH’s orders, EngineAI’s capacity, Figure’s production ramp and the new model benchmarks remain vendor claims or incompletely specified roadmaps. Open repositories and papers make AgiBot, Astribot, GENE.01 and Genesis World artifacts more inspectable. None substitutes for a site-specific safety case or your own uptime and cost measurements.
Mainstream robotics · narrow jobs
## Most useful robots do not look human.
Surgical systems, warehouse fleets, service carts, robot arms, construction printers and autonomous mowers trade generality for a bounded job. “Mainstream” describes this mixed field—not equal maturity, autonomy or return on investment.
- [Video: GEBHARDT InstaPick warehouse robotics](https://www.youtube.com/watch?v=NuI9ACH0Gjk)
GEBHARDT · InstaPick · commercial
### Move standard bins without building a humanoid.
GEBHARDT continues to market InstaPick as a modular robot-based container storage system and integrates it with picking and conveyor technology. The meaningful comparison is against alternative storage and material-flow designs: capacity, fire strategy, replenishment, failure isolation, service access and peak throughput.
- [Inspect current warehouse robotics →](https://gebhardt-group.com/index.php/en/products/warehouse-technology-systems/warehouse-robotics.html)
- [Video: Timo Boll versus a KUKA robot](https://www.youtube.com/watch?v=lv6op2HHIuM)
KUKA · industrial automation · established
### A spectacular film is not the industrial use case.
The table-tennis film demonstrates motion and marketing, not a production benchmark. KUKA is an active global automation group with industrial robots, mobile robots, software and integration services; evaluate the exact cell, payload, reach, safety design, cycle time, tooling and integrator support instead of generalizing from the stunt.
- [Verify the current company →](https://www.kuka.com/en-us/company)
- [Browse the current portfolio →](https://www.kuka.com/en-us)
- [Video: Brightpick robotic warehouse](https://www.youtube.com/watch?v=U2AGLeJBFNg)
Brightpick · fulfilment robots · commercial
### Bring picking to the aisle instead of people to the goods.
Brightpick remains active and launched Gridpicker in 2026 after Autopicker and Giraffe. Its films and customer announcements show a coherent commercial system; the buyer still needs SKU-level success, human touches, exception labor, storage density, peak throughput and service response measured on the proposed design.
- [Inspect current systems →](https://brightpick.ai/)
- [Check current customer + product news →](https://brightpick.ai/news-releases/)
- [Video: Ocado Smart Platform warehouse automation](https://www.youtube.com/watch?v=Eu3jgy2-tL8)
Ocado Group · fulfilment platform · deployed
### Make the fleet, grid, arms and software one operating system.
Ocado still sells the Smart Platform and now also describes a mobile-robot system with Chuck and Porter. The film is useful precisely because no single robot explains the outcome: storage geometry, orchestration, picking, packing, software and site operations determine throughput together.
- [Inspect the current robotics stack →](https://www.ocadogroup.com/about-us/our-technology)
- [Review the commercial platform →](https://www.ocadogroup.com/our-solutions/online-grocery)
- [Video: Robotic hernia repair surgery](https://www.youtube.com/watch?v=KAvQsRL-jeo)
Intuitive · da Vinci · commercial
### A mature robot can be highly specialized and human-controlled.
Sutter Health’s film shows robotic-assisted hernia repair; the da Vinci system is made by Intuitive, which continues to manufacture, place and support surgical systems in 2026. This is not autonomous surgery: a trained surgeon controls the instruments, and clinical procurement requires procedure-specific evidence, training, cybersecurity and service planning.
- [Inspect the current systems →](https://www.intuitive.com/)
- [Check the 2026 filing →](https://isrg.intuitive.com/static-files/8d536618-2150-41fb-b463-a6d6691d59bd)
- [Video: Introducing the ICON Titan Print System](https://www.youtube.com/watch?v=Ebir6dl_cxc)
ICON · Titan · commercial rollout
### Turn 3D printing into a supported construction system.
ICON commercially launched Titan in March 2026 with reservations, training planned for Q3 2026 and first deliveries anticipated in early 2027. The vendor lists the printer, pump, mixer, software, materials and support as one program; cost and speed claims still need validation against local design, permitting, foundations, reinforcement, finishes and crew requirements.
- [Verify rollout + delivery timing →](https://www.iconbuild.com/newsroom/icon-announces-first-commercial-rollout-of-its-3d-printing-construction-technology-for-builders)
- [Inspect the current program →](https://www.iconbuild.com/technology/program)
- [Video: AmbiSort AI-powered robotic parcel sorting](https://www.youtube.com/watch?v=iYkg3woXB1A)
Ambi Robotics · AmbiSort · commercial
### Automate one parcel touch inside an existing operation.
AmbiSort combines a robot arm, vision, grasp planning, barcode scanning and configurable destinations for small-parcel sortation. Ambi Robotics still markets the A‑Series and announced a 2026 integration with Pickle Robot; throughput gains shown by the vendor remain configuration- and parcel-mix-dependent.
- [Inspect the current A‑Series →](https://www.ambirobotics.com/ambisort-a-series/)
- [Check current deployments + partnerships →](https://www.ambirobotics.com/media-room/press/)
- [Video: Introducing Almond Axol](https://www.youtube.com/watch?v=WjcBwSufr7Q)
Almond · Axol · early commercial
### Sell a dual-arm data and manipulation platform.
Y Combinator lists Almond as active, and Almond currently offers Axol with an open SDK, published configurations, pricing and short stated lead time. It is a young four-person company, so treat availability, support, repair, payload and integration claims as procurement questions—not the same maturity class as an established industrial vendor.
- [Inspect Axol + current ordering →](https://www.almond.bot/)
- [Verify the active YC profile →](https://www.ycombinator.com/companies/almond-2)
- [Video: Scythe autonomous commercial mower](https://www.youtube.com/watch?v=dysTBGW6g_A)
Scythe · M.52 · acquired, supported
### Acquisition did not end the product.
ASI acquired Scythe Robotics in March 2026, and the company says M.52 remains an operating equipment brand with continuing customer support, field service and development. This is landscaping rather than crop agriculture; assess terrain, boundary handling, transport, charging, blade service, public exposure and operator supervision.
- [Verify the acquisition + support path →](https://scytherobotics.com/blog-posts/scythe-acquired-by-asi)
- [Video: KEENON XMAN-R1 robot barista](https://www.youtube.com/watch?v=oTZFSlYEBSk)
- [Video: KEENON DINERBOT T9](https://www.youtube.com/watch?v=9FQzaSydT-c)
KEENON · service robots · marketed
### Separate beverage preparation from table delivery.
KEENON currently markets XMAN‑R1 alongside DINERBOT, BUTLERBOT, KLEENBOT and industrial-delivery products. The two films show different automation layers: a robot barista and an indoor delivery platform. Buyer evidence should cover site mapping, queue behavior, cleaning, refill labor, food safety and human recovery—not only the serving motion.
- [Review KEENON’s current range →](https://www.keenon.com/en/product/XMAN-F1/)
- [Video: ROKAE robotic welding](https://www.youtube.com/watch?v=DlHE3zcauBk)
ROKAE · welding robots · active
### Package vision and force control around a skilled trade.
ROKAE still sells industrial and collaborative welding systems and published updated 2026 manuals, certifications and product material. The vendor’s trade-show demonstrations establish current activity, not weld qualification on your parts; test seam extraction, fixturing, consumables, spatter, rework and operator training.
- [Inspect the current welding line →](https://www.rokae.com/en/product/show/726.html?ly=serch)
- [Check the 2026 company update →](https://www.rokae.com/en/news/show/2458/ROKAE-Robotics-at-ITES-Shenzhen-2026.html)
- [Video: Meet Beni by Mondo Robotics](https://www.youtube.com/watch?v=rBETwu7ssVw)
Mondo Robotics · Beni · crowdfunding
### A live company is not yet a mainstream product.
Mondo Robotics is active and reports more than 30 beta users, but Beni launched through Kickstarter in July 2026. The camera rover belongs on the watchlist with a clear crowdfunding label: independent durability, privacy, repair, software-support and delivery evidence are still thin.
- [Check the company timeline →](https://mondorobotics.com/pages/about-us)
- [Review current updates →](https://mondorobotics.com/blogs/news)
- [Video: temi personal robot](https://www.youtube.com/watch?v=jNriaEZkNJI)
temi · mobile telepresence · active
### Mobility, video and reminders can be enough.
temi remains an active self-navigating service and telepresence platform, with current deployments in care, education and commercial settings. It has no arms and should not be sold as a general household worker; evaluate navigation coverage, privacy, remote administration, fall or obstruction handling, and the staff workflow around it.
- [Review the current company →](https://www.robotemi.com/about/)
- [Inspect the product routes →](https://www.robotemi.com/)
**Current-only policy:** every included company has a current official product, support, filing, acquirer or accelerator record checked on 28 July 2026. A compelling film without a verified current company and product route is omitted.
Agriculture robotics · field + greenhouse
## Useful farm robots are crop- and task-specific.
Weeding, harvesting, planting, spraying and pest control fail in different ways. Every card below represents a current company and a current product or active development program.
- [Video: Agrobot robotic strawberry harvester](https://www.youtube.com/watch?v=M3SGScaShhw)
Agrobot · strawberry harvesting · development
### Give each berry its own perception-and-cut problem.
Agrobot remains active in Spain and California and describes a newer solar-powered, modular-arm harvester. Its public material does not establish broad commercial availability or customer throughput, so the machine stays in the development lane until field yield, damage, miss rate, speed and service data are available.
- [Review the current company account →](https://www.agrobot.com/copia-de-about-us)
- [Video: AgXeed AgBot field use case](https://www.youtube.com/watch?v=ATqTBUedyjM)
AgXeed · AgBot · commercial
### Automate the tractor route and keep standard implements.
AgXeed actively markets four AgBot models, TraXwise planning software and a dealer network, with 2026 demonstrations and product news. Vendor savings and utilization claims require farm-specific validation across implements, soil, weather, refueling, road transport, connectivity, supervision and local machinery rules.
- [Inspect the current AgBot range →](https://www.agxeed.com/)
- [Check 2026 activity →](https://www.agxeed.com/news-and-events/)
- [Video: VitiBot Bakus vineyard robot](https://www.youtube.com/watch?v=v5oig0P9X2Y)
VitiBot · Bakus · SDF group
### Design the vehicle around vineyard rows.
VitiBot remains active inside SDF Group and markets narrow- and wide-vineyard Bakus models with dealer and demonstration routes. Its electric autonomous straddle form is specific to vineyard geometry; validate slopes, row width, tools, battery duty, safety zoning and dealer support locally.
- [Inspect the current Bakus range →](https://vitibot.fr/)
- [Video: PATS-X bat-like pest-control drones](https://www.youtube.com/watch?v=4g5Sb-UVn6I)
PATS · pest-control drones · final testing
### Target airborne pests inside the greenhouse.
PATS remains active, but PATS‑X is still described as undergoing final testing in Dutch and Belgian crops and is offered via a waitlist. The film is a development demonstration; effectiveness, non-target impact, crop coverage, maintenance and operating supervision need field evidence.
- [Check current test + waitlist status →](https://www.pats-drones.com/pats-x)
- [Review the company →](https://www.pats-drones.com/pats)
- [Video: Ridder and MetoMotion GRoW tomato harvesting robot](https://www.youtube.com/watch?v=-FB7mjc9eQ8)
Ridder + MetoMotion · GRoW · pre-order
### Pick, collect and box vine tomatoes in one pass.
Ridder and MetoMotion remain active, and Ridder currently accepts reservations for GRoW. The two-arm robot has named greenhouse trials and a customer scale-up account, but published labor and cost reductions remain vendor estimates; require current cycle, damage, intervention, variety and service data.
- [Check current reservation status →](https://info.ridder.com/reserve-your-grow-tomato-harvesting-robot)
- [Inspect the current system →](https://metomotion.com/robotic-worker/)
- [Video: TTA-ISO CuttingPlanter 2.0](https://www.youtube.com/watch?v=QcqKduExS0U)
TTA‑ISO · nursery automation · commercial
### Automate cuttings without pretending the whole greenhouse is autonomous.
TTA‑ISO is active and introduced CuttingPlanter 2.0 in July 2026, adding vision software, a faster arm and an electric tilting gripper. It sells a broader nursery-automation portfolio; compare plant material, tray formats, changeover, rejects, sanitation and upstream/downstream labor.
- [Verify the July 2026 launch →](https://tta-iso.com/updates/introducing-our-newest-cuttingplanter-20)
- [Browse the current equipment range →](https://www.tta-iso.com/)
- [Video: FarmBot open-source farming robot](https://www.youtube.com/watch?v=yHy8rJpbxH4)
FarmBot · garden CNC · commercial + open source
### Make small-plot automation inspectable and hackable.
FarmBot continues to sell hardware and update its app and operating system in 2026. Its CAD, software and developer routes are open, which improves inspectability and education; it does not turn a garden-scale gantry into a field-scale commercial farming system.
- [Inspect software + open-source routes →](https://farm.bot/pages/software)
- [Check current product updates →](https://farm.bot/blogs/news)
- [Video: AGRIST cucumber harvesting robot](https://www.youtube.com/watch?v=bylYcrRl9ik)
AGRIST · greenhouse harvesting · active
### Combine harvesting with routine crop care.
AGRIST is active and in May 2026 announced a funded program to extend its cucumber robot from harvesting into leaf removal and fruit thinning. That is forward development, not proof the three-task system is already deployed; ask for crop variety, greenhouse geometry, success rate, cycle time and human reset data.
- [Check the May 2026 program →](https://agrist.com/archives/14976)
**Current-only policy:** every entry must have a current company plus a current product or active development program. Superseded machines, acquired IP, research-only prototypes, category mismatches and unverifiable businesses stay off the public catalogue.
Food + drink robotics
## The hard part is the whole kitchen shift.
Cooking motion is only one layer. Ingredients, cold chain, sanitation, allergen controls, replenishment, waste, recipes, service, approvals and human recovery decide whether the system is useful.
- [Video: goodBytz robotic kitchen](https://www.youtube.com/watch?v=GiG6Kmz_FfE)
goodBytz · autonomous kitchen · commercial
### Use deployments, not awards, as the stronger signal.
goodBytz is active and announced a July 2026 U.S. defence-site delivery after deployments in Europe and South Korea. Its system cooks, portions and serves; site buyers still need accepted-meal throughput, ingredient labor, cleaning, allergens, downtime, local approvals and service data.
- [Check current products + deployments →](https://www.goodbytz.com/)
- [Inspect the current press record →](https://www.goodbytz.com/de/press)
- [Video: Eatch robotic kitchen for large-scale cooking](https://www.youtube.com/watch?v=qtkONuUDwHI)
Eatch · production kitchen · active
### Scale individually cooked meals from a central kitchen.
Eatch remains active and markets its Robotic Kitchen Technology for central production and white-label meals. Its claim is flexible meal production rather than a public-facing robot restaurant; evaluate recipe range, batch planning, input prep, packaging, sanitation, maintenance and delivered-food quality.
- [Inspect the current platform →](https://eatch.me/)
- [See the operating food route →](https://maaltijden.eatch.me/pages/onze-robot)
- [Video: Moley Robotics showroom](https://www.youtube.com/watch?v=WzeGT3m5oMU)
Moley Robotics · luxury kitchen · marketed
### A showroom product needs a home-service model.
Moley remains active, showed the system at Salone del Mobile in 2026 and markets a built-in robotic kitchen. This is a high-end fitted installation, not a countertop appliance; procurement hinges on kitchen integration, recipe authoring, cleaning, utensils, child safety, remote support and long-term parts availability.
- [Check current company activity →](https://www.moley.com/news/)
- [Inspect the current kitchen →](https://www.moley.com/)
- [Video: Richtech Robotics ADAM at the Vegas Golden Knights](https://www.youtube.com/watch?v=cPXQFNWisAQ)
Richtech Robotics · ADAM · public company
### Make the robot both service equipment and a visible attraction.
Richtech is an active Nasdaq filer and continues to promote ADAM in 2026. The arena film proves a branded event installation, not beverage-unit economics; verify drink quality, queue time, consumables, cleaning, staffing, uptime and the difference between entertainment value and operational savings.
- [Check current company releases →](https://ir.richtechrobotics.com/news-events/news-releases)
- [Read the 2026 annual filing →](https://ir.richtechrobotics.com/static-files/778957e1-6ab8-40e9-a349-a16482278166)
- [Video: BOTINKIT OMNI cooking robot](https://www.youtube.com/watch?v=Kcw0kw2v8ps)
BOTINKIT · OMNI · commercial
### Standardize wok cooking while people run the kitchen.
BOTINKIT is active, has a Japanese subsidiary and says OMNI is deployed across Asian and North American markets. The automated cooking station is one component of a kitchen; test recipes, ingredient prep, chef controls, extraction, cleaning, certification, spare parts and regional service.
- [Verify company + market presence →](https://www.botinkit.co.jp/about)
- [Inspect the current platform →](https://www.botinkit.ai/)
**Current-only policy:** every entry must have a current operating company and a current supported product. Liquidated operations, acquired technology without a current product route and stale operating evidence stay off the public catalogue; duplicate films are consolidated.
Market scale · counts with caveats
## A small builder field can serve a very large machine economy.
There is no audited global census of “robotics companies,” and databases classify vendors, products and integrators differently. The order of magnitude is still revealing: specialist supplier landscapes number in the hundreds, while deployed machines already number in the millions.
**≈700**
warehouse-automation ecosystem companies
LogisticsIQ’s landscape, cited by ITIF in 2023, is broad: it includes robots, software, components, integrators, infrastructure and related services—not 700 interchangeable robot startups.
- [Inspect the category definition →](https://itif.org/publications/2023/06/26/policymakers-should-support-robotic-automation-to-solve-productivity-crunch-in-logistics/)
**≈200**
companies developing humanoids
A 2026 industry account cites Gartner’s “nearly 200” estimate. Gartner’s own outlook is sharper: by 2028, it expects fewer than 100 companies to move proofs of concept beyond experiments and fewer than 20 to reach production in manufacturing or supply chain.
- [Trace the ≈200 estimate →](https://www.machinedesign.com/markets/robotics/article/55363632/physical-ai-hype-vs-reality-kung-fu-robots-are-coolbut-should-you-hire-one)
- [Read Gartner’s production forecast →](https://www.gartner.com/en/newsroom/press-releases/2026-01-21-gartner-predicts-fewer-than-20-companies-will-scale-humanoid-robots-for-manufacturing-and-supply-chain-to-production-stage-by-2028)
**15,384**
commercial martech solutions in 2025
This is a product catalog, not a company count, so it is not an apples-to-apples market share calculation. It is a useful competition benchmark: one mature software function supports tens of thousands of discoverable products.
- [Open the 2025 landscape method →](https://chiefmartec.com/2025/05/2025-marketing-technology-landscape-supergraphic-100x-growth-since-2011-but-now-with-ai/)
Deployment is the bigger number
**1,000,000+**
robots deployed by Amazon across more than 300 facilities
Operator-reported · June 2025
**542,000**
industrial robots installed worldwide in 2024 alone
International Federation of Robotics · 2025 release
The startup opening is not “build every robot.” A small team can own one stubborn task, instrument its failures and make a measurable improvement that repeats across sites.
- [Verify Amazon’s fleet milestone →](https://www.aboutamazon.com/news/operations/amazon-million-robots-ai-foundation-model)
- [Read the IFR installation data →](https://ifr.org/ifr-press-releases/global-robot-demand-in-factories-doubles-over-10-years)
**How to read this:** the ≈700 and ≈200 figures are directional landscape estimates, while the martech figure counts solutions. Amazon’s total includes multiple robot types, and IFR counts industrial installations rather than humanoids. None of the figures alone is a market-size forecast or a claim about startup success.
The last centimetre · hands + actuation
## The robot often stops at the wrist.
Many robot bodies are sold or configured without a dexterous hand. The base, arm and end effector are often separate procurement decisions, so “buying the robot” may still leave the hardest contact problem—and a second SDK—to solve.
Separate is normal
### A hand is a subsystem, not an accessory.
LimX lists a fist-shaped end as standard on Luna and a five-finger hand as optional. Unitree publishes separate removal and installation guides for Dex3‑1 and Inspire hands on different G1 configurations. PSYONIC sells its Ability Hand with an open API for robot integration and shows it on Apptronik’s Apollo. These are concrete examples of a broader integration pattern: the body may provide a wrist interface while the hand, sensors, control electronics and policy data come from another supplier.
**Compatibility boundary:** a removable hand is not automatically a compatible hand. Confirm flange geometry, handedness, mass and inertia, power, communication, control mode and rate, collision model, safety limits, calibration, firmware and the training embodiment before ordering.
- [See Luna’s optional hand →](https://www.limxdynamics.com/en/products/luna/spec)
- [See G1 hand swap guides →](https://www.unitree.com/mobile/app/g1/)
- [Inspect PSYONIC’s robot API →](https://www.psyonic.io/robots)
> Visual: Hand integration layers
1. 01 **Wrist** Flange · cable route · mass · inertia
2. 02 **Hand** Fingers · transmission · force · speed
3. 03 **Touch** Taxels · force/torque · calibration · wear
4. 04 **Control** SDK · protocol · rate · safety limits
5. 05 **Policy** Retargeting · data · sim model · evaluation
- [Video: Wuji Hand 2](https://www.youtube.com/watch?v=7QYedp3ozjw)
Wuji Technology · Hand 2 Beta 1
### Twenty independently driven joints, exposed as a developer platform.
Wuji documents 20 active, back-drivable direct-drive rotary joints, a 1,000 Hz × 20-axis control rate, Ethernet, 12 V input and a 745 ± 10 g hand with soft body. The Python SDK, ROS 2 route and URDF, MJCF and USD assets make the integration surface unusually inspectable. The current documentation is also explicit that this is Beta 1 and that its external connector form will change; published durability language is not a substitute for your load, impact and lifecycle test.
**Active DoF** — 20
**Control** — 1 kHz × 20 axes
**Mass** — 745 ± 10 g
**Status** — Beta 1
- [Inspect product claims →](https://wuji.tech/en/hand2)
- [Read the current integration docs →](https://docs.wuji.tech/docs/en/wuji-hand/latest/overview/)
- [Video: ORCA Dexterity announces three new open source robotic hands](https://www.youtube.com/watch?v=WNtlUViSrPg)
ORCA Dexterity · open-source hand family
### Choose touch, simplicity or the full research platform.
ORCA’s announcement introduces three open-source routes: orcahand touch, orcahand lite and the classic orcahand. The classic hand is a 17-DoF tendon-driven platform with integrated tactile sensing, published design files, control code, bill of materials and assembly documentation. Open access makes the system easier to inspect, reproduce and repair; it does not make the three variants equivalent or remove the need to validate load, sensing, maintenance and licensing for your deployment.
**Family** — Touch · Lite · Classic
**Classic DoF** — 17
**Architecture** — Tendon-driven
**Artifacts** — CAD · code · BOM
- [Watch the ORCA announcement →](https://www.youtube.com/watch?v=WNtlUViSrPg)
- [Inspect the open platform →](https://www.orcahand.com/paper)
- [Video: Sharpa Wave dexterous robotic hand](https://www.youtube.com/watch?v=GcTUlOHvdOs)
Sharpa · Wave
### Make touch part of the control stack.
Wave is a one-to-one human-scale hand with 22 active degrees of freedom and a proprietary tactile array. Sharpa reports 0.02 N tactile sensitivity, more than 20 N fingertip force, greater than 4 Hz full-gesture speed and 2.5 million press-test cycles. It also publishes ROS 2, SDK, firmware, URDF, tactile-simulation and mechanical-CAD resources. Those figures are vendor tests without the complete protocols needed for cross-vendor comparison; validate payload, repeatability, tactile drift, wear, repair time and performance on your objects.
**Active DoF** — 22
**Touch** — 0.02 N claimed
**Fingertip** — >20 N claimed
**Assets** — ROS 2 · Isaac · CAD
- [Inspect specs + test claims →](https://www.sharpa.com/pages/wave)
- [Open software + CAD resources →](https://www.sharpa.com/pages/downloads)
- [Video: HW1 Launch & Feature Showcase | Gesture Platforms](https://www.youtube.com/watch?v=LPbB96fT_uw)
Gesture Platforms · HW1
### A lightweight hand aimed at desktop research and repair.
HW1 is an ESP32-S3 robotic-hand platform with 10 actively controlled degrees of freedom across 19 joints. Gesture states a mass below 500 g, repeatability within 1 mm, 100 Hz control, USB-C and Bluetooth Low Energy, plus motor-angle, current and temperature telemetry. The launch video also emphasizes replaceable fingers, accessible electronics and an offline desktop app. These are vendor specifications for a crowdfunding-stage, pre-delivery product—not independent lifetime or task evidence. Validate grasp payload, play, repeatability, heat, impact survival, spare-part supply, SDK maturity and delivery status before treating it as lab infrastructure.
**Active DoF** — 10 · 19 joints
**Controller** — ESP32-S3
**Mass** — <500 g claimed
**Status** — Pre-delivery
- [Inspect the current product page →](https://gestureplatforms.com/)
- [Read the independent technical summary →](https://www.cnx-software.com/2026/05/28/gesture-hw1-10-dof-esp32-s3-robotic-hand-with-high-dexterity-manipulation/)
- [Video: Introducing the mimic hand M1](https://www.youtube.com/watch?v=ikjPRgE8WLM)
mimic robotics · hand M1
### Match the data-capture hand to the robot hand.
mimic unveiled M1 and its U1 wearable in July 2026 as a full-stack dexterous-manipulation platform. The tendon-driven M1 has 15 active degrees of freedom across 21 joints and moves its primary actuators into the forearm. It is a current vendor launch—not yet independent durability or production evidence—so validate payload, backdrivability, collision recovery, retargeting, integration and availability on the intended task.
**Active DoF** — 15 · 21 joints
**Architecture** — Tendon-driven
**Data capture** — U1 wearable
**Status** — July 2026 launch
- [Inspect the launch + specifications →](https://www.mimicrobotics.com/blog/solving-dexterity-a-full-stack-approach)
- [Review the current company →](https://www.mimicrobotics.com/about)
- [Video: No Gears. No Belts. The Future of Robotics Starts Here](https://www.youtube.com/watch?v=WwxNc5Bh74E)
Genesis Advanced Technology · LiveDrive direct-drive actuator
### LiveDrive replaces the gearbox and belt in a Delta robot.
The video argues that mechanical motion has not advanced as quickly as robotics software, AI, sensing and vision. Genesis presents LiveDrive as a high-torque direct-drive actuator that removes the gearbox and belt, then demonstrates it powering a Delta robot for high-speed pick-and-place work. The company says the oil-free design reduces drivetrain complexity, maintenance and contamination risk while allowing denser layouts and precise control. Those are vendor claims: validate torque, heat, energy use, cycle time, repeatability, uptime and service life on the real workload.
Actuation lens · direct-drive Delta robot
- [Watch the full LiveDrive demonstration →](https://www.youtube.com/watch?v=WwxNc5Bh74E)
- [Jump to the direct-drive actuator →](https://www.youtube.com/watch?v=WwxNc5Bh74E&t=74s)
- [Jump to the Delta robot →](https://www.youtube.com/watch?v=WwxNc5Bh74E&t=127s)
- [Review the current company →](https://www.genesisadvancedtechnology.com/)
Mechanical fit
### Mount the real mass.
Check flange, cable bend, wrist workspace, self-collision, inertia, center of mass and payload after the hand and tool are attached.
Contact envelope
### Test the object distribution.
Measure fingertip and grasp force, speed, compliance, impact survival, wear, contamination and the smallest stable grasp on your objects.
Control surface
### Time the whole loop.
Verify voltage, current peaks, bus, command mode, feedback fields, update rate, latency, watchdogs and safe behavior after packet loss.
Learning surface
### Match simulation to hardware.
Pin SDK and firmware versions, inspect URDF/USD/MJCF assets, calibrate touch, retarget demonstrations and preserve failed grasps in evaluation.
**Evidence policy:** degrees of freedom count motion axes, not useful dexterity. Force, speed and tactile sensitivity use different fixtures and definitions across vendors. Ask for the test protocol, lifecycle curve, spare-finger or module process, repair turnaround and customer task data before comparing headline numbers.
The first deployment
## Start with a task you can count.
A narrow wedge gives the team a fixed environment, a repeatable dataset and a customer metric. Generality can be earned later by adding nearby tasks.
01 · Observe
### Shadow the work.
Record task frequency, handling variation, travel, exceptions, safety controls and the human skills that make recovery look easy.
**Output: task + exception map**
02 · Price
### Find the costly constraint.
Quantify labour hours, injury exposure, downtime, scrap, throughput, night coverage and the cost of doing nothing.
**Output: value per successful cycle**
03 · Bound
### Shape the environment.
Standardize bins, lighting, floor markers, approach angles, fixtures or handoff points before demanding more intelligence.
**Output: operating design domain**
04 · Prototype
### Teleoperate before autonomy.
Use remote operation or a scripted controller to prove reach, payload, cycle time and customer workflow before training a policy.
**Output: working data collector**
05 · Learn
### Simulate, collect and evaluate.
Version scenes, demonstrations, checkpoints and tests. Include recoveries and off-nominal states—not only polished success episodes; reuse the [model evaluation loop](https://isaiuseful.com/training-models.html.md#loop) for every learned component.
**Output: reproducible policy evidence**
06 · Pilot
### Keep a human recovery path.
Run inside a fenced scope with stop controls, incident logging and an operator who can recover the task without improvisation.
**Output: field reliability + economics**
Good first wedge
### One object family, one site, one shift.
Examples: machine tending for a named part, tote movement on a mapped route, visual inspection at a fixed station, or cleaning one repeatable floor type.
Bad first wedge
### “Replace any worker anywhere.”
No stable task distribution, no realistic acceptance set, no bounded safety case and no credible denominator for unit economics.
System architecture
## The robot is a chain of assumptions.
Choose the simplest component at every layer that preserves the field requirement. A spectacular policy cannot repair the wrong gripper, missing stop circuit or brittle site integration.
Job + environment
### Operating design domain
Task, objects, people, floor, lighting, weather, network, shift and allowed exceptions.
Acceptance starts here.
Body + tooling
### Embodiment
Wheels, legs, arm, gripper, payload, reach, speed, durability and maintainability.
Use a humanoid only when its form earns access.
Sensing + compute
### Perception at the edge
RGB, depth, LiDAR, force, proprioception, synchronization, bandwidth and thermal budget.
Redundancy is a safety and reliability decision.
Planning + control
### Deterministic and learned layers
State estimation, navigation, motion planning, low-level control, learned policy and safety supervisor.
Do not let a VLA bypass hard limits.
Data + simulation
### Learning loop
Teleoperation, scene generation, demonstrations, failure capture, policy training and regression evaluation.
Version the environment with the model.
Fleet + service
### Production system
Deployment, observability, remote assist, spares, updates, access control, incident response and customer support.
The service margin lives here.
- [Robot Operating System →](https://www.ros.org/)
- [Gazebo simulation →](https://gazebosim.org/)
- [Drake planning and control →](https://drake.mit.edu/)
- [LeRobot real-world loop →](https://huggingface.co/docs/lerobot/main/en/getting_started_real_world_robot)
- [NVIDIA Isaac Sim documentation →](https://docs.nvidia.com/isaacsim/latest/index.html)
Model + edge computer
## Robostral and Thor solve different layers.
One is a task-specific navigation model; the other is a compute platform. A product still needs sensors, localization, control, stop logic, integration and a measured recovery loop. Use the [local-model memory guide](https://isaiuseful.com/local-models.html.md#memory) to shortlist adjacent edge candidates before measuring them on-device.
- [Video: Introducing Robostral Navigate](https://www.youtube.com/watch?v=7dpLB9NoY1A)
Embodied navigation · Mistral
### Language instruction → a route through the world
Mistral presents an 8B navigation model trained in simulation that uses one RGB camera. Its reported R2R‑CE success is benchmark evidence for the named task—not proof that any robot can safely navigate any site.
- [Read Mistral’s technical account →](https://mistral.ai/news/robostral-navigate/)
- [Video: Getting Started with the NVIDIA Jetson AGX Thor Developer Kit for Physical AI](https://www.youtube.com/watch?v=iYT2haVIgSM)
Edge compute · NVIDIA Developer
### Put a larger multimodal stack at the edge
Jetson AGX Thor combines 128 GB of memory, a Blackwell GPU and extensive robot I/O in a 130 W-class developer platform. Peak FP4 compute is not application latency; test the exact model, sensors, thermal envelope and control loop.
- [Open the current developer guide →](https://docs.nvidia.com/jetson/agx-thor-devkit/user-guide/latest/)
- [Verify NVIDIA’s published specifications →](https://nvidianews.nvidia.com/news/nvidia-blackwell-powered-jetson-thor-now-available-accelerating-the-age-of-general-robotics)
Navigation policy
### Robostral Navigate
**Input** — RGB history + language instruction
**Output** — Image point/orientation or local displacement
**Reported scale** — 8B · 2.4M simulated trajectories · 350k scenes
**Evidence** — Vendor report on R2R‑CE and office demonstrations
Reproduce the named benchmark and then test your camera placement, route geometry, people, glare, obstacles and recovery states. A model score is not a site safety case.
- [Inspect the release evidence →](https://mistral.ai/news/robostral-navigate/)
Robot computer
### Jetson AGX Thor
**Memory** — 128 GB unified LPDDR5X
**Vendor peak** — Up to 2,070 FP4 TFLOPS
**Developer power mode** — Up to 130 W; supplied adapter constraints apply
**I/O** — USB, camera, Ethernet and QSFP28 routes
Start from the current JetPack release, pin containers and measure end-to-end sensor-to-action latency under sustained heat. The developer kit is not the final production carrier or certification plan.
- [Open the current Thor guide →](https://docs.nvidia.com/jetson/agx-thor-devkit/user-guide/latest/)
Learned layer
**Interpret, perceive, propose**
Language, scene grounding, task planning and candidate actions.
Supervised control
**Validate, constrain, execute**
Workspace limits, collision checks, speed limits, watchdogs and stop circuits.
Field evidence
**Log, review, improve**
Interventions, near misses, task failures, recovery time and accepted cycles.
Practitioner experience · sentdex
## Follow one humanoid from unboxing to learned locomotion.
This series is valuable because it exposes the integration work between the product film and a custom behavior: networking, LiDAR, SLAM, arm and hand control, external compute, simulation and reinforcement learning.
- [Video: Unboxing the Unitree G1 EDU Humanoid](https://www.youtube.com/watch?v=pPTo62O__CU&t=2859s)
01 · Hardware baseline
**Unboxing the Unitree G1 EDU Humanoid**
*Inventory what arrives, what is EDU-specific and what still needs integration.*
- [Video: LiDAR, SLAM, navigation and control](https://www.youtube.com/watch?v=sJYlJlIEBpg&t=2251s)
02 · Mobility stack
**LiDAR, SLAM, navigation and control**
*Connect perception, mapping and the first controllable movement loop.*
- [Video: Moving the arms and hands](https://www.youtube.com/watch?v=Uc1nhT8beTU)
03 · Manipulation
**Moving the arms and hands**
*See where command interfaces meet joint limits and end-effector reality.*
- [Video: A bigger brain for the G1](https://www.youtube.com/watch?v=cmnJhOWp2z4)
04 · External compute
**A bigger brain for the G1**
*Trace the networking and architecture questions behind off-board intelligence.*
- [Video: Vibe coding the Inspire robot hands](https://www.youtube.com/watch?v=MeHWIXLV3Zo)
05 · Hand interface
**Vibe coding the Inspire robot hands**
*Turn a vendor interface into a small, testable manipulation experiment.*
- [Video: Make a robotic hand crawl](https://www.youtube.com/watch?v=57cPmzwCqd4)
06 · Creative control
**Make a robotic hand crawl**
*A compact study in actuation, iteration and unexpected embodiments.*
- [Video: Reinforcement learning with G1](https://www.youtube.com/watch?v=wiIUF9pIDYw&t=93s)
07 · Policy training
**Reinforcement learning with G1**
*Move from built-in behavior to a reproducible simulation and training route.*
- [Video: Train a G1 to walk](https://www.youtube.com/watch?v=FGnAeUXRZ4E)
08 · Sim to real
**Train a G1 to walk**
*Connect a learned locomotion policy to the physical deployment boundary.*
Vendor baseline
### Know the exact G1.
Unitree lists configurations ranging from 23 to 43 joint motors. EDU options, hands, compute and sensors change the software surface and the price.
- [Check current G1 configurations →](https://www.unitree.com/g1/)
Real robot interface
### Pin the SDK and firmware.
The official SDK2 route uses DDS interfaces for G1 and other Unitree robots. Confirm the exact messages and services available on your firmware before building control around them.
- [Inspect Unitree SDK2 →](https://github.com/unitreerobotics/unitree_sdk2)
Learning loop
### Rehearse before deployment.
Unitree publishes MuJoCo, Isaac Lab and reinforcement-learning routes plus a LeRobot bridge for G1 data, training and real-world evaluation.
- [Unitree MuJoCo →](https://github.com/unitreerobotics/unitree_mujoco)
- [Unitree RL Lab →](https://github.com/unitreerobotics/unitree_rl_lab)
- [Unitree LeRobot →](https://github.com/unitreerobotics/unitree_lerobot)
Deployment gate
## Measure the whole service, not the demo.
Choose thresholds before the pilot. Report a distribution and its failure denominator—not only the cleanest successful video. The site’s [evidence boundary](https://isaiuseful.com/evidence.html.md#claim-boundary) explains why a strong result on one task must stay scoped to that task.
Task
### Successful cycles
**Successes ÷ all attempted cycles**
Segment by object, route, lighting, operator and exception type.
Intervention
### Human recovery
**Interventions per operating hour**
Include remote assist, resets, falls, stuck states and manual completion.
Time
### Useful throughput
**P50 / P95 cycle + recovery time**
Compare with the actual human or machine process—not a lab ideal.
Reliability
### Availability
**Uptime, MTBF and repair time**
Track batteries, sensors, joints, networking, software and consumables separately.
Safety
### Leading indicators
**Stops, near misses and limit violations**
Treat “no injury” as insufficient when the pilot is small.
Economics
### Cost per accepted cycle
**Robot + service + people + site changes**
Include financing, support, spares, teleoperation and customer integration.
Safety boundary
### Risk assessment belongs to the robot system.
The end effector, program, power, sensors, communication interfaces, cell, people and maintenance workflow all matter. OSHA’s industrial-robot guidance distinguishes manufacturer, integrator and user responsibilities and points to task-based risk assessment and safeguarding.
- [Read OSHA’s robot-system guidance →](https://www.osha.gov/otm/section-4-safety-hazards/chapter-4)
European route
### Plan conformity before the hardware freezes.
The EU Machinery Regulation replaces the Machinery Directive and applies from January 2027. An AI system used as a safety component of machinery can also enter the AI Act’s high-risk route when the legal criteria are met. Scope the actual product with qualified safety and legal specialists.
- [EU Machinery Regulation →](https://eur-lex.europa.eu/legal-content/en/ALL/?uri=CELEX:32023R1230)
- [AI Act Article 6 classification →](https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-6)
Company layer · Europe
## Build the cap table for a hardware journey.
Robotics crosses borders early: suppliers, certification, pilots, investors and hires rarely sit in one Member State. Company structure is infrastructure—but it does not replace local tax, labour, product-safety or operating obligations.
EU—INC
### One European standard is a proposal, not today’s incorporation route.
EU‑INC is a grassroots campaign for a pan-European corporate standard. Its July 2026 position paper argues for free choice of registered office, one central registry, broad company access, standardized stock options, and local labour law and taxes tied to real activity.
This is an advocacy position. Verify the final law, national implementation and your actual incorporation, employment, tax, IP and fundraising facts with qualified advisers.
Before the first institutional round
- IP assignment from every founder, contractor and university partner
- Hardware, dataset, model and open-source licence inventory
- Supplier terms, export controls and product-liability allocation
- Employee option plan that works where the team actually lives
- Pilot contracts that separate experiments from production commitments
- [Read the EU‑INC position paper](https://www.eu-inc.org/position-paper)
Prototype Capital starter pack
## Use the link list as a route, not a reading pile.
Every substantive external resource linked in Prototype Capital’s “Why you should start a robotics startup” post is organized below. The post is an investor’s editorial case for the opportunity; the primary courses, tools and documentation are the stronger sources for implementation details.
- [Video: Why you should start a robotics startup](https://www.youtube.com/watch?v=FqfTQFuSalY)
Opportunity thesis · Prototype Capital
### Why 2026 may be a strong moment to start
The video argues that cheaper components, improving vision-language-action models and macro demand are expanding the design space. Treat the market counts and timing as an investor thesis; validate the named customer task and procurement reality yourself.
- [Read the source post →](https://updates.prototypecap.com/p/why-you-should-start-a-robotics-startup)
- [Prototype Capital’s YouTube channel →](https://www.youtube.com/@prototypecap?sub_confirmation=1)
Jump within the embedded video
16 chapters
- [Video chapter: 00:00 Intro](https://www.youtube.com/watch?v=FqfTQFuSalY)
- [Video chapter: 00:34 Robotics are exploding](https://www.youtube.com/watch?v=FqfTQFuSalY&t=34s)
- [Video chapter: 01:58 Why start in 2026?](https://www.youtube.com/watch?v=FqfTQFuSalY&t=118s)
- [Video chapter: 03:56 Component costs](https://www.youtube.com/watch?v=FqfTQFuSalY&t=236s)
- [Video chapter: 04:42 The “spaghetti moment”](https://www.youtube.com/watch?v=FqfTQFuSalY&t=282s)
- [Video chapter: 06:11 The unsettled stack](https://www.youtube.com/watch?v=FqfTQFuSalY&t=371s)
- [Video chapter: 06:53 The data problem](https://www.youtube.com/watch?v=FqfTQFuSalY&t=413s)
- [Video chapter: 07:46 Macro tailwinds](https://www.youtube.com/watch?v=FqfTQFuSalY&t=466s)
- [Video chapter: 08:40 Robotics as the new SaaS](https://www.youtube.com/watch?v=FqfTQFuSalY&t=520s)
- [Video chapter: 09:14 Structured environments](https://www.youtube.com/watch?v=FqfTQFuSalY&t=554s)
- [Video chapter: 10:05 Global niches](https://www.youtube.com/watch?v=FqfTQFuSalY&t=605s)
- [Video chapter: 10:32 General vs vertical](https://www.youtube.com/watch?v=FqfTQFuSalY&t=632s)
- [Video chapter: 11:54 Small-team speed](https://www.youtube.com/watch?v=FqfTQFuSalY&t=714s)
- [Video chapter: 12:32 What about humanoids?](https://www.youtube.com/watch?v=FqfTQFuSalY&t=752s)
- [Video chapter: 14:52 Open opportunities](https://www.youtube.com/watch?v=FqfTQFuSalY&t=892s)
- [Video chapter: 15:37 Summary](https://www.youtube.com/watch?v=FqfTQFuSalY&t=937s)
Start here
### Touch the learning loop.
- [LeRobot robot-learning tutorial →](https://huggingface.co/spaces/lerobot/robot-learning-tutorial)
- [Mythbusting physical AI →](https://planet-a.medium.com/robots-in-the-real-world-mythbusting-physical-ai-f9eb1688ee35)
- [The physical-AI deployment gap →](https://www.a16z.news/p/the-physical-ai-deployment-gap)
- [The new labour economy thesis →](https://newsletter.semianalysis.com/p/america-is-missing-the-new-labor-economy-robotics-part-1)
Learn deeply
### Build the foundations.
- [Springer Handbook of Robotics →](https://link.springer.com/referencework/10.1007/978-3-319-32552-1)
- [MIT Robotic Manipulation →](https://manipulation.csail.mit.edu/)
- [MIT Underactuated Robotics →](https://underactuated.csail.mit.edu/)
- [NVIDIA Robotics learning path →](https://www.nvidia.com/en-us/learn/learning-path/robotics/)
Build something
### Pick one implementation spine.
- [Robot Operating System →](https://www.ros.org/)
- [Gazebo →](https://gazebosim.org/)
- [Drake →](https://drake.mit.edu/)
- [Awesome Robotics index →](https://github.com/kiloreux/awesome-robotics)
Keep current
### Follow deals and practice.
- [Robotics Roundup →](https://roboticsroundup.substack.com/)
- [Robotics for Software Engineers →](https://newsletter.pragmaticengineer.com/p/robotics-for-software-engineers)
- [Lukas M. Ziegler →](https://www.linkedin.com/in/zieglerr/)
- [Lukas Ziegler’s newsletter →](https://ziegler.substack.com/)
Track the ecosystem
### Add funding, standards and workforce views.
- [Paulina Szyzdek →](https://www.linkedin.com/in/paulina-szyzdek)
- [Paulina Szyzdek’s newsletter →](https://paulinaszyzdek.substack.com/)
- [Aaron Prather →](https://www.linkedin.com/in/amprather)
The practical answer
Choose one physical job. Instrument every failure. Make reliability earn scale.
- [Define the wedge](#route)
- [Assemble the stack](#stack)
- [Set the deployment gate](#scoreboard)
- [Browse AI Now engineering talks](https://isaiuseful.com/training-models.html.md#ai-now-summit)
---
## https://isaiuseful.com/remote-spark (`/remote-spark.html.md`)
# How to Run a Private Remote AI Assistant on DGX Spark
Canonical source: [https://isaiuseful.com/remote-spark](https://isaiuseful.com/remote-spark)
Remote operator · Linux compute
Run OpenClaw, NemoClaw, Hermes or Open WebUI with Ollama on DGX Spark, then operate it from Windows 11, macOS, Linux, a browser or a messaging channel—without exposing the gateway, model or retrieval ports to the public network.
- [Choose a control surface](#operators)
- [Route a sensitive file](#hybrid-privacy)
- [Build the Spark side](#spark-setup)
Official NVIDIA starting points
## Pick a DGX Spark recipe.
These four official NVIDIA playbooks adapt our [workflow recipes](https://isaiuseful.com/guides.html.md#chooser) , [local-model stack](https://isaiuseful.com/local-models.html.md#runtimes) and [remote-operation runbook](#operators) to DGX Spark. Pick one here, then use our guides for model choice, private access and operating safeguards.
- [Run OpenClaw with a local LLM 30 min Install the local-first agent and connect it to a private OpenAI-compatible model endpoint.](https://build.nvidia.com/spark/openclaw)
- [Run NemoClaw with a local LLM 30 min Build an OpenClaw assistant in an OpenShell sandbox with local vLLM inference.](https://build.nvidia.com/spark/nemoclaw)
- [Run Hermes Agent with a local LLM 30 min Connect the terminal-first Nous Research agent to a model served locally with vLLM.](https://build.nvidia.com/spark/hermes-agent)
- [Open WebUI with Ollama 15 min Use the browser-based chat interface with Ollama and a model running on Spark.](https://build.nvidia.com/spark/open-webui)
- [Official DGX Spark collection **Explore every NVIDIA Spark playbook** Open build.nvidia.com/spark](https://build.nvidia.com/spark)
**Pocket-ready**
phone steers; state stays on Spark
OPENCLAW OR HERMES
**Insider**
for the newest Windows isolation
MXC SESSION PREVIEW
**Localhost**
services cross an SSH tunnel
NO OPEN ADMIN PORTS
Compatibility first
## How can you control DGX Spark remotely?
The operator device can be a phone, laptop or desktop without moving the gateway, memory, models or jobs off the always-on Linux host.
Native companion
### Windows 11
Use OpenClaw Windows Hub for Command Center, notifications and an optional Windows node. The current bleeding-edge containment route adds Windows Insider and MXC; ordinary remote chat does not move the gateway onto Windows.
Native companion
### macOS
Use the OpenClaw menu bar app in Remote mode. It can own the SSH tunnel, health checks and Web Chat while the gateway remains loopback-bound on Spark.
CLI or desktop
### Linux
Use the Linux companion, CLI, browser Control UI or an SSH tunnel. It is also the closest match to Spark when you need to reproduce commands locally.
**iPhone and Android:** use a paired companion, the browser UI or a configured messaging channel. OpenClaw documents iOS and Android as clients of the Gateway—not replacements for it—so Spark keeps the session and compute while the phone carries chat, status, approvals and only the device capabilities you explicitly enable.
- [iPhone companion →](https://docs.openclaw.ai/platforms/ios)
- [Android companion →](https://docs.openclaw.ai/platforms/android)
- [Compare every control surface →](#operators)
### Hardware + price warning · checked 24 July 2026 Do not buy DGX Spark on the assumption that Windows support will arrive later.
DGX Spark and the pre-release RTX Spark family publish strikingly similar headline numbers—up to one petaflop and up to 128 GB unified memory—but NVIDIA currently lists them as different product categories: DGX Spark is a Linux companion system, while RTX Spark is a Windows primary system. Public documentation does not establish that their drivers, firmware or boards are interchangeable, and NVIDIA has announced no Windows driver, Windows image or upgrade path for DGX Spark.
Buy DGX Spark only if DGX OS works for the machine's useful life. Future Windows support is possible in theory, but today it is speculation. There is also no public evidence for claims about NVIDIA's commercial motive; the support gap itself is the decision-relevant fact.
Indicative MSRP · NVIDIA Marketplace Germany
**€4,800**
NVIDIA's own DGX Spark listing displayed this price and was out of stock when checked. Use it as a reference for the NVIDIA-branded system—not a ceiling for partner products. OEMs choose their own configurations and selling prices.
- [Check the current NVIDIA Marketplace price →](https://marketplace.nvidia.com/de-de/enterprise/personal-ai-supercomputers/dgx-spark/)
- [NVIDIA platform comparison →](https://developer.nvidia.com/local-ai)
- [DGX Spark software requirements →](https://docs.nvidia.com/dgx/dgx-spark-porting-guide/porting/software-requirements.html)
- [Pre-release RTX Spark example →](https://www.microsoft.com/en-us/surface/devices/surface-rtx-spark-dev-box)
### NVIDIA AI Enterprise check · GB10 / DGX Spark Do not treat preinstallation, NIM access or a Spark badge as a production entitlement.
A 90-day NVIDIA AI Enterprise—DGX Spark evaluation exists only when it is purchased, requested or explicitly issued on the NVIDIA Entitlement Certificate (EC); not every OEM offer includes it. Record the EC/order line, GPU metric, start and end date, support route and renewal before deployment. It is an evaluation, not a public free-forever entitlement to the complete supported production suite.
Free components remain useful for development: Omniverse and NVIDIA AI Workbench are free, and the standard NVIDIA Developer Program NIM route is for development, research and test up to 16 GPUs with community support. Production self-hosting generally needs NVIDIA AI Enterprise.
- [DGX Spark 90-day evaluation and EC activation →](https://docs.nvidia.com/dgx/dgx-spark/nvaie-quickstart.html)
- [NVIDIA AI Enterprise licensing guide →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/licensing.html)
- [NIM developer-program route and production boundary →](https://forums.developer.nvidia.com/t/nvidia-nim-faq/300317)
- [Omniverse license and support route →](https://docs.omniverse.nvidia.com/dev-guide/latest/common/NVIDIA_Omniverse_License_Agreement.html)
- [AI Workbench introduction →](https://docs.nvidia.com/ai-workbench/user-guide/latest/overview/introduction.html)
See the category boundary
### RTX Spark is being introduced as a Windows PC family—not as a DGX Spark software update.
These NVIDIA and Microsoft videos show the platform framing, one announced product and one example workload. Treat them as vendor demonstrations: they clarify positioning, but they do not prove interchangeability, delivery timing or performance for your workflow.
- [Video: NVIDIA RTX Spark Reinvents Windows PCs for the Age of Personal AI](https://www.youtube.com/watch?v=H4nJo-oqAro)
Platform overview · NVIDIA
### A Windows PC category for local AI
Use this overview to understand NVIDIA's intended product category. Keep the announced positioning separate from tested application support on a shipping system.
- [Video: Introducing Surface RTX Spark Dev Box](https://www.youtube.com/watch?v=VlAI1_JkXL4)
Product example · Microsoft
### Surface makes the Windows distinction concrete
The Surface RTX Spark Dev Box is a separate announced Windows product. Its existence is not evidence that a DGX Spark can be converted into one.
- [Video: Architectural Design With Agents on NVIDIA RTX Spark](https://www.youtube.com/watch?v=a6fUvL9gYAQ)
Workload demo · NVIDIA
### Judge the workflow, then test the stack
The architectural-design demo shows how NVIDIA imagines agents using the platform. It is a useful workflow reference, not a benchmark or compatibility matrix.
- [Video: Announcing NVIDIA RTX Spark | GTC Taipei 2026 Keynote by CEO Jensen Huang](https://www.youtube.com/watch?v=11Y3B33oCLE)
Launch context · NVIDIA
### Hear the promise—then compare DGX Station.
The keynote gives the broad RTX Spark context. If 128 GB is the constraint rather than the solution, the DGX Station guide maps the 748 GB tier, realistic model fits and the current benchmark gap.
- [Compare DGX Spark with Station →](https://isaiuseful.com/dgx-station.html.md)
Independent reviews + benchmarks
## Spark is a capacity machine with a measured bandwidth ceiling.
Hands-on third-party tests agree on the shape even when engines, quants and prompts differ: 128 GB unlocks models that ordinary small systems cannot load, while 273 GB/s LPDDR5X limits dense single-stream decode.
Hands-on review
### LMSYS measured the trade.
Its early-access Spark ran GPT‑OSS 20B MXFP4 at 49.7 decode tok/s, but Llama 3.1 70B FP8 at 2.7 decode tok/s. Llama 3.1 8B scaled to 368 aggregate decode tok/s at batch 32. NVIDIA supplied early access, and the authors warn that software results can age.
- [Read the LMSYS methodology and tables →](https://www.lmsys.org/blog/2025-10-13-nvidia-dgx-spark/)
Independent comparison
### Dense 70B was capacity-first, not fast.
A separate published comparison reported 4.67 tok/s for Llama 3.3 70B, 38.03 tok/s for Qwen3 Coder and 60.33 tok/s for GPT‑OSS 20B on Spark. Treat these as workload-specific results, not universal product scores.
- [Inspect the comparison and conditions →](https://www.pcgamer.com/hardware/graphics-cards/nvidias-little-gold-box-of-pure-ai-power-the-dgx-spark-is-finally-out-and-the-comparison-with-amds-much-cheaper-strix-halo-chip-is-looking-a-little-fugly/)
Cluster evidence
### Two nodes need the right parallelism.
StorageReview tested Dell, GIGABYTE and HP pairs over the 200 Gb fabric and found OEM performance within a narrow band. For batched inference at practical concurrency, its pipeline-parallel layout mattered more than small chassis differences.
- [Read the two-node review →](https://www.storagereview.com/review/nvidia-dgx-spark-cluster-review-distributed-inference-on-dell-gigabyte-and-hp)
Independent video reviews
### Two longer practitioner views to put beside the numbers.
These videos add independent system context. Keep their software versions, workloads and methodology attached to any performance observation, and use the normalized table below for direct comparisons.
- [Video: Deep Dive into Nvidia's DGX Spark GB10](https://www.youtube.com/watch?v=Lqd2EuJwOuw)
Independent deep dive · Level1Techs
### Put the GB10 platform under a practitioner’s lens.
Level1Techs examines DGX Spark as a system rather than a specification sheet. Treat its observations as independent context and keep software versions and workload conditions attached to any result.
- [Video: NVIDIA DGX Spark: From “Inference Box” to Dev Rig (What It Actually Is) | Ep 2](https://www.youtube.com/watch?v=0CI19dXmOws)
Independent perspective · Domesticating AI
### Frame Spark as a development rig, not only an inference box.
This longer independent discussion focuses on what the machine is and how its development role differs from a simple inference appliance. Pair the framing with the measured limits below before buying.
Measured snapshot · do not mix rows
### Engine, precision, batch and model architecture explain the spread.
“Tokens per second” without those fields is not a benchmark. Prefill and decode are separate, and aggregate batch throughput is not the speed each interactive user sees.
| Source | Model and format | Workload | Measured Spark result | What it supports |
| --- | --- | --- | --- | --- |
| LMSYS · Oct 2025 | GPT‑OSS 20B · MXFP4 · Ollama | Single-stream decode | 49.7 tok/s | A small sparse model can be comfortably interactive. |
| LMSYS · Oct 2025 | Llama 3.1 70B · FP8 · SGLang | Batch 1 decode | 2.7 tok/s | Loading a large dense model is not the same as serving it quickly. |
| LMSYS · Oct 2025 | Llama 3.1 8B · FP8 · SGLang | Batch 32 aggregate decode | 368 tok/s | Batching can use compute that one bandwidth-bound stream leaves idle. |
| Third-party comparison · Oct 2025 | Llama 3.3 70B | Single prompt | 4.67 tok/s | A second setup reproduces the slow dense-70B shape, not the exact LMSYS number. |
| NVIDIA SANA · Aug 2026 | MiniMax H3 · pruned FP8 · Sol Engine | 5 s video · 480p · 124 frames · 50 steps | 181.3 s optimized · 3.92× | A specialized FP8 runtime makes H3 fit and materially faster; generation remains far from real time. |
**Buying implication:** Spark's best LLM fit is usually a 20–35B daily model or a larger sparse MoE with few active parameters—not the largest dense checkpoint that can be forced into memory. Re-run your exact engine after every major software update. [See where DGX Station changes the boundary →](https://isaiuseful.com/dgx-station.html.md#comparison)
### Spark owner tests · checked 15 August 2026 Muse Glimmer fits one Spark—and DFlash lifts decode into the 20s and 30s.
Meta’s Apache‑2.0 **Muse Glimmer 30B** is a dense local-agent model with tool use, coding and optional image input. Its official GGUF release provides 16.76 GB and 19.65 GB language-model builds, plus a separate 1.40 GB perception encoder and optional 1.63 GB DFlash drafter. That leaves ample room inside a 128 GB Spark-class system for runtime and a useful context budget, although the advertised 131,072+ model limit is not a promise that the maximum context will be fast or memory-efficient.
Early owner runs provide a useful speed range, not a controlled benchmark. On one DGX Spark, a llama.cpp run reported about 10.5 tok/s conventional decode and 36–38 tok/s with the official DFlash drafter at 15 speculative tokens; a separate vLLM run rose from 5–8 tok/s to about 23 tok/s at the same setting. Workload, draft acceptance, quant, context depth and engine build can move those numbers.
The llama.cpp test also reported about 700 tok/s prefill at short context, about 390 tok/s deep into an 832K-token prompt, and 3/3 needle retrieval at 97K, 188K, 415K and 832K after an 8× YaRN and metadata override. Treat that as an extended-context experiment rather than a new guaranteed model limit. Its two-Spark RPC layer split slowed decode to 25–28 tok/s, so a model this size is better run as one instance per Spark unless a measured workload proves otherwise. Pin the exact engine revision and rerun your own tool-call, recovery, context and latency harness.
- [Inspect Meta’s official GGUF files →](https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF/tree/main)
- [Read Meta’s launch and RTX 5090 speed conditions →](https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model)
- [Inspect the DGX Spark llama.cpp and extended-context owner test →](https://www.reddit.com/r/LocalLLaMA/comments/1vl9adk/i_ran_muse_glimmer_1m_context_all_tests_passed/)
- [Inspect the DGX Spark vLLM DFlash owner test →](https://www.reddit.com/r/LocalLLM/comments/1vm4j0i/muse_glimmer_30b_on_dgx_spark_using_dflash_is/)
- [Compare memory fits and vendor benchmarks →](https://isaiuseful.com/local-models.html.md#muse-glimmer)
### Vendor video benchmark · checked 15 August 2026 MiniMax H3 proves the Spark fit—but 3.92× still means three minutes for five seconds.
NVIDIA’s SANA team ran the 33B dense **MiniMax H3** audio-video generator on one DGX Spark at 832×480, 24 fps, 124 frames and 50 denoising steps. Sol Engine reduced end-to-end wall time from 710.6 to 181.3 seconds, a reported 3.92× speedup. The full recipe combines kernel work with approximate Sol‑Attn sparse attention and cross-step caching, so the accompanying near-lossless quality assessment is part of the vendor experiment—not proof that every prompt is unchanged. NVIDIA’s separate GeForce RTX 5090 result uses 1344×768 and must not be treated as an RTX Spark benchmark or compared by raw time.
The ordinary BF16 FL2VA package is about 134.2 GiB before activations, so NVIDIA’s GB10 runtime uses a pruned FP8 DiT, FP8 Qwen3‑VL conditioner and the released VAEs, with one resident model process. MiniMax’s downloadable local system produces 768p H3‑Base output; the recommended Context‑IR preprocessing and Regenerate‑2K stage remain hosted. There is also a procurement-level license gate: the community license excludes the EU, UK, United States and South Korea, with a separate application route for those territories.
- [Inspect NVIDIA’s H3-on-device results →](https://nvlabs.github.io/Sana/Sol-Engine/H3-OnDevice/)
- [Inspect the GB10 runtime and reproducible config →](https://github.com/NVlabs/Sana/tree/sol-engine/models/minimax_h3/GB10)
- [Read the MiniMax H3 model card →](https://huggingface.co/MiniMaxAI/MiniMax-H3)
- [Read the MiniMax H3 license →](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE)
- [Request an excluded-territory license →](https://platform.minimax.io/h3-license)
- [Compare the local-model route →](https://isaiuseful.com/local-models.html.md#minimax-h3)
### Native NVFP4 hardware · speculative capacity DGX Spark has Blackwell FP4 hardware. The serving kernel is still a gate.
NVFP4 is promising here because it can shrink weights and reduce memory traffic while preserving more quality than a crude four-bit conversion. For planning, reserving 20–25% of Spark's 128 GB and assuming a mixed checkpoint at roughly 5.0–5.2 bits per parameter gives an estimated **150–165B parameter** single-node fit. That is below NVIDIA's “up to 200B” capacity ceiling because this estimate leaves useful room for the runtime and cache.
Two linked Sparks provide a speculative **295–330B parameter** planning band with the same reserve. NVIDIA has demonstrated Qwen‑235B in NVFP4 on two systems and markets an up-to-405B model ceiling, but neither statement guarantees your model, context, kernel or interactive speed.
- [NVIDIA's dual-Spark NVFP4 example →](https://developer.nvidia.com/blog/?p=111120)
- [Understand NVFP4 and Blackwell support →](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)
- [Compare Spark, Station and datacenter capacity →](https://isaiuseful.com/cloud-models.html.md#nvfp4)
DGX Spark-class OEM systems
## Eight GB10 listings. One workload test.
The active products share the compact GB10, 128 GB coherent-memory class, but not necessarily NVIDIA's DGX Spark product name, factory image, storage, regional price or support terms. HP advised directly that ZGX Nano is discontinued, so it remains here as a clearly labelled historical reference.

### Veriton GN100
GB10 AI mini workstation.
- [Open Acer system page →](https://www.acer.com/gb-en/desktops-and-all-in-ones/veriton-workstations/veriton-gn100-ai-mini-workstation?utm_source=nvidia)

### ASUS Ascent GX10
GB10 desktop AI supercomputer.
- [Open ASUS system page →](https://www.asus.com/networking-iot-servers/desktop-ai-supercomputer/ultra-small-ai-supercomputers/asus-ascent-gx10/?utm_source=nvidia)

### Dell Pro Max with GB10
GB10 micro workstation.
- [Open Dell system page →](https://www.dell.com/en-ie/shop/desktop-computers/dell-pro-max-with-gb10/spd/dell-pro-max-fcm1253-micro/xcto_fcm1253_emea#features_section)

### GIGABYTE AI TOP ATOM
GB10 AI TOP system.
- [Open GIGABYTE system page →](https://www.gigabyte.com/AI-TOP-PC/GIGABYTE-AI-TOP-ATOM)

### HP ZGX Nano AI Station
GB10 nano AI station.
- [Open the HP reference page →](https://www.hp.com/us-en/workstations/zgx-nano-ai-station.html?utm_source=nvidia)

### Lenovo ThinkStation PGX
GB10 AI development workstation.
- [Open Lenovo system page →](https://www.lenovo.com/ie/en/p/workstations/thinkstationp/lenovo-thinkstation-pgx-sff/len102s0023)

### MSI EdgeXpert MS-C931
GB10 AI supercomputer.
- [Open MSI system page →](https://ipc.msi.com/product_detail/Industrial-Computer-Box-PC/AI-Supercomputer/EdgeXpert-MS-C931?utm_source=nvidia)

Offer checked 10 August 2026
### PNY DGX Spark
Authorized DGX Spark channel route. The current NVIDIA/PNY offer advertises an optional free 90-day NVIDIA AI Enterprise—DGX Spark evaluation after registration and entitlement; it is time-limited, offers community-driven support only, and is not a production-support entitlement.
- [Open PNY system page →](https://www.pny.com/dgx-spark)
- [NVIDIA AI Enterprise—DGX Spark 90-day evaluation →](https://docs.nvidia.com/dgx/dgx-spark/nvaie-quickstart.html)
- [NVIDIA Marketplace DGX Spark offer →](https://marketplace.nvidia.com/en-us/enterprise/personal-ai-supercomputers/dgx-spark/)
- [NVIDIA trial support caveat →](https://www.nvidia.com/content/dam/en-zz/Solutions/dgx-spark/workstation-print-gtc26-nvaie-spark-solution-overview-5004550-r7.pdf)
**Shortlist + quote check:** first compare the exact model, engine, quant, context and concurrency. Then verify the configured SSD, DGX OS or factory image, ConnectX accessories and two-node kit, thermals, power supply, regional availability, support lifecycle, warranty and VAT-inclusive price. Start with the [independent Spark benchmark shape](#spark-reviews) , then repeat the same harness on the OEM unit you can actually buy.
Try before you buy
## Rent the remote workload before buying a Spark or larger system.
Before committing four, five or six figures to a Spark or larger deployment, rent the candidate workload and prove the whole operator path. A rental is a test drive, not a guarantee of GPU availability, equivalent hardware, benchmark transfer or price.
**Exercise the real route.** Run the exact model, engine, quantization, context, concurrency and acceptance harness, including the remote-control and tunnel pattern you need; retain the container, settings, logs, latency, throughput, memory and task-quality results.
**Referral disclosure:** this isaiuseful.com link is a Runpod referral link. As checked 10 August 2026, eligible new first-time users must sign up through it with Google SSO and load their first $10: European users receive $5 credit, while non-European users receive a weighted $5–$500 credit (Runpod says most are $10 or less). Terms can change. If eligible, isaiuseful.com receives its referral bonus and earns Runpod credits on actual usage for the first six months: 3% of Pod spend and 5% of Serverless spend. Using the link supports the site.
- [Rent a Runpod test environment through this referral link →](https://runpod.io?ref=l40ix174)
- [Read Runpod’s current referral terms →](https://docs.runpod.io/accounts-billing/referrals)
DGX Spark product views
## Hear NVIDIA's platform promise. Keep it separate from measured results.
The keynote explains the platform story, the PNY film shows one product implementation, and the assistant and [GraphRAG](https://isaiuseful.com/rag.html.md#graph) sessions demonstrate intended workflows. All four are vendor presentations—use the independent review block above for measured model performance.
- [Video: NVIDIA GTC Spring 2025 Keynote: Introducing NVIDIA DGX Spark](https://www.youtube.com/watch?v=6p4U1kSiegg)
Launch context · NVIDIA
### Introducing DGX Spark at GTC
The keynote segment sets out NVIDIA's original product positioning and target users. Compare those claims with current documentation and the independent throughput results on this page.
- [Video: NVIDIA DGX Spark | A Grace Blackwell AI Supercomputer on your desk](https://www.youtube.com/watch?v=kZRMshaNrSA)
PNY product view · NVIDIA
### PNY DGX Spark in NVIDIA's product film
The short film shows the PNY implementation and intended compact appliance experience. It does not establish model quality, sustained tokens per second, thermals or equivalence with other OEM systems.
- [Video: Build Your Own AI Assistant with Hugging Face on NVIDIA DGX Spark](https://www.youtube.com/watch?v=dMpLCGvE2A0)
Build tutorial · NVIDIA
### Turn the appliance into an assistant workflow.
This concise NVIDIA walkthrough shows a Hugging Face assistant build on DGX Spark. Use it for workflow ideas and product setup context—not as independent performance evidence.
- [Video: DGX Spark Live: Process Text for GraphRAG With Up to 120B LLM](https://www.youtube.com/watch?v=uQtzjAvJMlE)
GraphRAG demo · NVIDIA Developer
### Use the memory pool for a larger retrieval pipeline.
The NVIDIA Developer session demonstrates text processing for [GraphRAG](https://isaiuseful.com/rag.html.md#architecture) with an LLM of up to 120B parameters. It shows an intended capacity-led workflow, not a standardized throughput benchmark.
Choose the operator
## Use the device already in your hands.
A phone is enough for chat, status and approvals. Desktop- or mobile-node permissions are optional additions for workflows that genuinely need to touch the operator device.
Windows 11
### Windows Hub + MXC preview
Best native OpenClaw diagnostics and Windows-node experience. Use the Insider path below when you want the current session-isolation work; use SSH, Web Chat or messaging when the Windows machine is only an operator.
- [Follow the Insider setup →](#windows-setup)
- [Windows Hub →](https://github.com/openclaw/openclaw-windows-node)
macOS
### Menu bar app in Remote mode
Point the signed macOS companion at `user@SPARK_IP` . Its default remote mode manages a strict-host-key SSH tunnel, health checks and Web Chat without starting a second gateway on the Mac.
- [Official remote-mode guide →](https://docs.openclaw.ai/platforms/mac/remote)
Linux
### Companion, CLI or browser
Use the Linux desktop companion when you want a tray and Canvas, or forward port 18789 and open the Control UI on localhost. The CLI is the simplest path on a minimal workstation.
- [Official Linux guide →](https://docs.openclaw.ai/platforms/linux)
**Phone or desktop companion**
Best local notifications, status and optional device-node capabilities.
**Web or CLI**
Universal path through an SSH tunnel; no desktop-node permissions required.
**Messaging**
Telegram, Discord, WhatsApp or another configured channel for everyday remote conversations.
**Hermes surfaces**
Use its TUI, dashboard or messaging gateway from any client while Hermes stays on Spark.
Remote access · private by default
### Use the narrowest path that fits the job.
Keep Spark services loopback-bound and authenticate every client. NVIDIA Sync is the managed route for Windows, macOS and Ubuntu; direct SSH is the universal fallback; remote desktop is for work that truly needs the full Linux interface.
**01 · Sync** — Managed SSH, application launches and tunnels.
**02 · SSH** — Direct terminal access and explicit port forwards.
**03 · Desktop** — Full remote interface on a private network only.
- [NVIDIA Sync user guide →](https://docs.nvidia.com/sync/latest/index.html)
- [DGX Spark access options →](https://docs.nvidia.com/dgx/dgx-spark/system-overview.html)
- [Set up remote access in step 2 →](#spark-remote-access)
Windows 11 · bleeding edge
## Insider is mandatory for the newest containment path.
OpenClaw Windows Hub can run on ordinary Windows builds. This guide targets the newer MXC session-isolation and agent-policy work Microsoft is still shipping through Windows Insider.
01
### Give the preview a recoverable Windows installation.
Use a secondary device or separate system image when possible. Enable BitLocker or device encryption, Secure Boot, Windows Hello and Defender; create a tested recovery drive and backup before enrolling. Run the companion as a standard user, not from a shared or daily administrator account.
02
### Join Windows Insider Experimental on a retail-aligned core.
Experimental is where Microsoft says actively developed features appear first. Stay on the 25H2 or 26H1 line offered for your hardware rather than the separate Future Platforms option. After updating, confirm the actual build with `winver` ; MXC currently documents build 26300.8553 as the minimum for its `isolation_session` backend.
- [Current channel definitions →](https://blogs.windows.com/windows-insider/2026/04/10/improving-your-windows-insider-experience/)
- [MXC build matrix →](https://github.com/microsoft/mxc)
03
### Verify the containment feature—not merely the OS label.
Install the current feature and platform updates, then check the OpenClaw release notes for MXC support and confirm that the requested backend is available. An Insider badge alone proves nothing. Microsoft currently describes OpenClaw's Windows node and gateway as an MXC integration, but the MXC repository also warns that its early-preview profiles are still overly permissive in known cases.
**Preview rule**
If MXC or the expected session backend is unavailable, stop. Do not silently fall back to unrestricted execution and call the result hardened.
04
### Use the canonical signed installer and verify it.
Download the x64 or ARM64 asset and checksum file from the project's latest release. Compare the SHA-256 value before opening it, retain SmartScreen and Defender checks, and reject a binary whose publisher or digest does not match.
```
Get-FileHash .\OpenClawTray-Setup-x64.exe -Algorithm SHA256
```
- [Latest release →](https://github.com/openclaw/openclaw-windows-node/releases/latest)
- [Official setup guide →](https://github.com/openclaw/openclaw-windows-node/blob/main/docs/SETUP.md)
05
### Connect to the existing Spark gateway.
Choose the remote or existing-gateway route in Windows Hub. Keep the gateway on Spark bound to loopback and let the Hub manage an SSH tunnel, or use a private Tailnet with authentication. The local WSL gateway is a fallback for people without an always-on host—not the default topology here.
06
### Pair, then allow only harmless Windows commands.
Approve the Windows node from the Spark gateway. Begin with notifications and device information/status. Leave command execution, screen capture, camera, location, speech and browser control disabled until a named workflow needs them and both the gateway policy and MXC policy deny everything else.
```
openclaw devices list
openclaw devices approve
```
**Starter allowlist**
`system.notify · device.info · device.status`
OpenClaw requires exact command names. If `system.run` is later enabled, keep the separate Windows node execution policy default-deny as well.
07
### Test the deny path before startup automation.
Run ten harmless actions, inspect the activity and MXC diagnostics, disconnect the gateway, reject an unexpected pairing request, attempt access to a deliberately denied file and domain, and confirm that every disabled node capability fails. Re-run this after Windows, Hub, MXC, OpenShell or gateway updates.
Insider policy
### Required for this bleeding-edge path; not proof of safety.
Stable Windows can run OpenClaw and MXC's lighter process backend, but the current session-isolation backend and several Windows AI connector/workspace policies are Insider-era features. That makes Insider non-optional for the setup described here. It also makes the setup a lab: Microsoft explicitly says current MXC profiles have known over-permissive cases and should not yet be treated as complete security boundaries. Keep Windows permissions, OpenClaw allowlists, network isolation, backups and human approval in place.
- [OpenClaw + MXC announcement →](https://blogs.windows.com/windowsdeveloper/2026/06/02/windows-platform-security-for-ai-agents/)
- [MXC preview warning →](https://github.com/microsoft/mxc)
- [Insider policy surface →](https://learn.microsoft.com/en-us/windows/client-management/mdm/policy-csp-windowsai)
Quality-of-life layer
## Add convenience without hiding authority.
OpenClaw extensions execute inside a trust boundary. Install fewer, inspect them and pin versions where the package route supports it.
Native companions
### Hub, menu bar or Linux tray
Use the platform companion for health, chat and notifications without moving the gateway. Enable a desktop node only when the agent must act on that device; operator access and node authority are separate choices.
- [Compare platforms →](https://docs.openclaw.ai/platforms)
Remote access
### Managed SSH tunnel or Tailscale
Prefer the Hub's SSH-tunnel support or Tailscale Serve over firewalling the gateway port open. Keep the gateway bound to loopback and authenticate every client.
- [OpenClaw remote access →](https://docs.openclaw.ai/gateway/remote)
High-trust browser tool
### OpenClaw Chrome extension
Use only when the workflow must operate an already signed-in browser tab. Share an explicit OpenClaw tab group, keep the relay on loopback and assume the agent can act with that tab's account permissions.
- [Extension security model →](https://docs.openclaw.ai/tools/chrome-extension)
Skills + plugins
### ClawHub, workspace-local first
Search by the exact capability you need, inspect source and scan state, then install into one workspace before promoting it. A popular package is still executable code—not a permission boundary.
- [ClawHub quickstart →](https://docs.openclaw.ai/clawhub/quickstart)
- [Manage plugins →](https://docs.openclaw.ai/plugins/manage-plugins)
Execution isolation
### OpenShell + platform controls
Put tool execution behind NVIDIA OpenShell where its backend fits, then layer the host controls: MXC preview on Windows, Bubblewrap or LXC on Linux, and Seatbelt-backed containment on macOS. Test denied files and destinations on the actual host.
- [OpenShell →](https://github.com/NVIDIA/OpenShell)
- [MXC status →](https://github.com/microsoft/mxc)
Operational habit
### One capability ledger
Record the package, version, owner, data it can read, actions it can take, token scopes and removal test. Re-run the ten-action test after every gateway, node or plugin update.
- [Use the ten-run method →](https://isaiuseful.com/guides.html.md#ten-run-evaluation)
Preferred topology
## Your devices operate. Spark remembers and computes.
The agent does not disappear when a laptop sleeps. Gateway state, model serving and scheduled work remain together on the always-on host.
> Visual: Remote operator and DGX Spark architecture
**Visual reading order:**
1. **01 · Any operator** **Windows · macOS · Linux** Companion app, browser, CLI, mobile or messaging; desktop-node powers remain optional
2. SSH tunnel 18789 + 8000
3. **02 · DGX Spark** **OpenClaw or Hermes** Loopback-bound, authenticated, always-on sessions, memory, channels and tool orchestration
4. **03 · DGX Spark** **vLLM + optional Qdrant** OpenAI-compatible inference and private retrieval services; no direct LAN exposure
5. **04 · DGX Spark** **Training workspace** Stopped while serving needs the memory; versioned datasets, adapters, evaluations and manifests
### Default placement: gateway on Spark
OpenClaw's own model is one gateway with many clients: the gateway owns sessions, authentication profiles, channels and state. Hermes is similarly comfortable on a remote Linux host with its dashboard and messaging gateway exposed only through the private access layer. A local laptop gateway is the fallback for someone without an always-on server, not the target architecture here.
Local first · cloud by exception
## Scrub here. Escalate only the derivative.
When the laptop model cannot finish the hard part, keep the original and the identity map local. Send a stronger model only the smallest reviewed artifact that policy permits.
Vendor case study
### Bayer taught Phi the exceptions.
Microsoft says Bayer fine-tuned a small Phi model on proprietary crop-protection labels, regulatory rules and expert-authored Q&A. Labels can exceed 100 pages; Bayer reports early complex questions falling from days or weeks to under 30 seconds.
- [Read the Microsoft customer story →](https://www.microsoft.com/en/customers/story/25255-bayer-azure-phi)
Vendor case study
### Discovery Bank split one job into five.
Discovery fine-tuned five variants across Azure OpenAI 4o-mini and 4.1-mini for company language, SQL shape and workflow templates. It reports average response time dropping from five or six seconds to 1.5–2 seconds.
- [Read the Microsoft customer story →](https://www.microsoft.com/en/customers/story/26157-discovery-bank-azure-openai-in-foundry-models)
Laptop implication
### Teach the boundary before the model.
A downloaded model can flag candidate names, secrets, clauses and contextual identifiers without giving the source file to a model provider. It can also finish simple extraction or comparison locally. Escalation begins only when a harder model adds enough value to justify a reviewed data crossing.
- [Run the one-document test →](#offline-scrub)
> Visual: Local scrub and cloud escalation workflow
**Visual reading order:**
1. **01 · Local** **Classify** Route the source red, amber or green before any model sees it.
2. **02 · Local** **Detect twice** Use deterministic checks plus a local model for direct and contextual identifiers.
3. **03 · Local** **Replace + minimize** Create stable typed tokens; prefer a task brief over a redacted full copy.
4. **04 · Human gate** **Review the exact artifact** Approve recipient, purpose, region, retention and unresolved spans.
5. **05 · Route** **Finish or escalate** Use local output when it passes; otherwise send only an approved derivative.
6. **06 · Local** **Rejoin carefully** Inspect the answer and restore approved values without exposing the token map.
**The boundary that matters**
Redaction is not deletion and pseudonymization is not anonymity. The original, findings manifest and re-identification map stay local. Credentials are removed rather than tokenized, and the cloud service never receives the map. If context still identifies the person, client, transaction or project, the file remains red.
The offline setup
### One document. One local model. No fallback.
During a connected maintenance window, install LM Studio and download one instruction-tuned local model, its runtime and any local embedding model the attachment workflow requires. Test with a synthetic file, close the app, move the authorized copy outside synced folders, then switch off Wi-Fi and unplug Ethernet before reopening it.
1. 01 **Load locally.** Choose only the downloaded model; start with an 8,192-token context and temperature 0.
2. 02 **Remove side doors.** Turn off tools, MCP, web search, plugins, cloud models and automatic fallback.
3. 03 **Keep the server closed.** Leave it off; if an integration needs it, bind only to `127.0.0.1` with authentication.
4. 04 **Return locations, not prose.** Ask first for exact span, page, category, reason, confidence and a proposed token.
**Local candidate-finding prompt**
`Work only with the attached local document. Do not use tools, web search, external sources or a remote model. Find DIRECT_IDENTIFIER, CREDENTIAL_OR_SECRET, REGULATED_DATA, COMMERCIAL_CONFIDENTIAL, QUASI_IDENTIFIER and UNCERTAIN spans. Return exact text, page or section, reason, confidence and a stable token such as [PERSON-001]. Do not rewrite yet. Do not infer missing identities.`
**Do not certify from document chat**
[RAG](https://isaiuseful.com/rag.html.md#model) may retrieve only selected passages. Long documents require page-by-page or overlapping-chunk inspection, deterministic secret and identifier checks, a fresh rescan of the derivative and human approval. Any unreadable or skipped content stops cloud routing.
- [LM Studio offline-operation documentation →](https://lmstudio.ai/docs/app/offline)
- [LM Studio network binding warning →](https://lmstudio.ai/docs/developer/core/server/serve-on-network)
Red · never ordinary cloud
### Keep the job local.
Credentials, identity evidence, health or KYC records, privileged advice, protected investigations, safety-critical material, raw client or board files, and anything with unclear authority or processor terms. Use approved private infrastructure or no model.
Amber · reviewed derivative
### Minimize, rescan, approve.
Internal contracts, narratives or reports that can lose direct identifiers and distinctive combinations without losing the task. Name the exact service, purpose, region, retention, logging, training and deletion terms before sending.
Green · approved content
### Still send less.
Published material, genuinely synthetic tests, checked templates or content explicitly approved for this external purpose. Inspect comments, metadata, tracked changes, hidden sheets, notes and embedded objects too.
Grok Build · July 2026
### The model obeyed. The product still moved the repository.
An independent wire analysis of Grok Build 0.2.93 recovered a never-read tracked file and Git history from a separate uploaded bundle after the prompt said not to open files. The researcher later reported that xAI disabled the path server-side. The test established transmission and storage in that setup—not training, employee access or universal behavior. Its durable lesson is that a prompt controls the task, not necessarily packaging, traces, sync or upload: “no training,” “no retention,” “no human review,” “no upload” and “local-only” are five different claims.
- [Inspect the original wire analysis →](https://gist.github.com/cereblab/dc9a40bc26120f4540e4e09b75ffb547)
Microsoft's lock-in test
### Exclusive is private. It is not automatically portable.
Microsoft says Azure-hosted direct-model inputs, outputs, embeddings and training data are not made available to model providers or used to improve foundation models without permission, and that a customer's fine-tuned model is exclusive to that customer. Those are useful privacy commitments. They do not by themselves make the resulting weights exportable or the workflow provider-independent.
- [Read the current Foundry data terms →](https://learn.microsoft.com/en-us/azure/foundry/responsible-ai/openai/data-privacy)
### Keep the learning outside the endpoint.
Store rules, source provenance, prompts, schemas, corrections, permission policy and an untouched evaluation set in exportable formats. Treat a fine-tune as a replaceable build artifact, not the only copy of what the company learned.
**The 30-day question**
Could we replace the model in 30 days using only the artifacts we can export today, then prove the replacement against the same holdout?
The guide + skill
## Take the boundary with you.
The walkthrough includes the full LM Studio pilot, red/amber/green router, verification gates, laptop-to-enterprise thresholds and portability test. The companion Codex skill prepares the derivative and review package, then stops before transmission.
- [Download the full guide](https://isaiuseful.com/downloads/sensitive-document-hybrid-guide.md)
- [Download the router skill](https://isaiuseful.com/downloads/scrub-sensitive-documents.zip)
- [Build local document search](https://isaiuseful.com/guides.html.md#private-documents)
DGX Spark build
## Build a service you can recover.
Use NVIDIA's Spark-specific playbooks. Generic CUDA, x86 container or desktop instructions can fail on ARM64/Blackwell even when the project itself supports NVIDIA GPUs.
01
### Baseline DGX OS.
Record the installed DGX OS release, driver, CUDA stack and firmware. Apply NVIDIA-recommended updates through DGX Dashboard, then verify health before adding services. Do not replace the OS with generic Ubuntu or Windows.
- [Update guide →](https://docs.nvidia.com/dgx/dgx-spark/os-and-component-update.html)
- [Release notes →](https://docs.nvidia.com/dgx/dgx-spark/release-notes.html)
- [CUDA libraries reference →](https://docs.nvidia.com/cuda-libraries/index.html)
02
### Set up private remote access.
After updating DGX OS, create a named non-root account, install an SSH key and choose the least-powerful route that covers the operator's work. Keep it on a trusted LAN or Tailnet, retain host firewall rules and never publish SSH, RDP or VNC ports directly to the internet.
**NVIDIA Sync**
Managed SSH, application launches and forwarding; optional Tailscale off-LAN. No jump or bastion hosts.
**Direct SSH**
Terminal administration and explicit forwards. Confirm the host fingerprint and keep strict verification.
**Remote desktop**
Full DGX OS interface only. Keep the chosen tool behind the LAN or Tailnet and fully updated.
- [Set up NVIDIA Sync →](https://docs.nvidia.com/sync/latest/index.html)
- [Review supported access options →](https://docs.nvidia.com/dgx/dgx-spark/system-overview.html)
- [OpenClaw remote access →](https://docs.openclaw.ai/gateway/remote)
03
### Serve one supported model with vLLM.
Use NVIDIA's DGX Spark vLLM recipe and one model from its current compatibility list. Pin the container or environment, set a context and concurrency budget, and run a fixed latency/quality test before adding model routing.
- [Spark vLLM recipe →](https://build.nvidia.com/spark/vllm/instructions)
- [Community Spark vLLM Docker setup →](https://github.com/eugr/spark-vllm-docker)
- [vLLM docs →](https://docs.vllm.ai/en/stable/)
04
### Add retrieval only when the workflow needs it.
Start with Qdrant only if filters, hybrid search, persistence or multiple collections justify a service. Store the original document and page metadata with every chunk; back up source documents and collection configuration, then test a restore.
- [Use the private-document recipe →](https://isaiuseful.com/guides.html.md#private-documents)
05
### Place OpenClaw or Hermes on Spark.
Install one primary agent through its supported Linux path, point it at vLLM's loopback OpenAI-compatible endpoint and keep its sessions, memory and channels on persistent storage. OpenClaw uses its gateway service; Hermes can expose its dashboard and messaging gateway. Run both only when you have intentionally separated their state, ports and permissions.
- [OpenClaw on Linux →](https://docs.openclaw.ai/platforms/linux)
- [Hermes Agent →](https://github.com/NousResearch/hermes-agent)
06
### Lock down tools and data services.
Bind the gateway, vLLM and Qdrant to loopback unless a private container network requires otherwise. Run high-authority agent tools through an OpenShell sandbox where supported, then test its filesystem and outbound-network denies. Expose only the application endpoint the chosen operator route needs.
- [OpenShell →](https://github.com/NVIDIA/OpenShell)
- [Secure Qdrant →](https://qdrant.tech/documentation/operations/security/)
07
### Separate serving from training.
Do not let a training job silently evict or starve the always-on model. Use explicit service and training modes, drain requests, stop vLLM when the recipe needs the memory, checkpoint to persistent storage and restore serving from a known configuration.
08
### Prove recovery.
Back up only what cannot be recreated: gateway configuration and keys, source data, dataset manifests, adapters, evaluation sets and service definitions. Rebuild one clean service from the manifest before calling the system production-ready.
Fine-tuning on Spark
## Use the smallest recipe that answers the question.
NVIDIA publishes PyTorch and NeMo fine-tuning playbooks for Spark. Their example model sizes describe tested recipes, not universal capacity or speed guarantees.
| Question | First method | NVIDIA Spark example | Artifact to keep | Gate |
| --- | --- | --- | --- | --- |
| Does the data pipeline work? | Small LoRA run | 8B LoRA path | Adapter + run manifest | Loss is sane and sample outputs improve without obvious memorization. |
| Can a larger model learn the behavior? | LoRA or QLoRA | 70B LoRA / 70B QLoRA paths | Adapter, tokenizer config and holdout results | Beats the unchanged base model on the untouched test set. |
| Is full-weight training justified? | Full SFT only after adapter evidence | 3B full SFT path | Checkpoint + reproducible environment | Material gain over LoRA, retention tests pass and operating cost is acceptable. |
**Official starting points:** [PyTorch fine-tuning](https://build.nvidia.com/spark/pytorch-fine-tune/instructions) · [NeMo fine-tuning](https://build.nvidia.com/spark/nemo-fine-tune/instructions) · [dataset and evaluation checklist](https://isaiuseful.com/guides.html.md#finetuning) .
No Spark yet
## Keep the topology; shrink the host.
Run the gateway on the most reliable machine you already own: WSL 2 on Windows, a launchd service on macOS, or a systemd user service on Linux. Keep it loopback-bound, use a small local model or hosted endpoint and operate it from the same companion, browser and messaging surfaces. This proves the workflow; it does not reproduce Spark's unified memory or sustained service role.
- [Choose an OS path](https://docs.openclaw.ai/platforms)
- [WSL networking](https://learn.microsoft.com/en-us/windows/wsl/networking)
- [Compare hardware](https://isaiuseful.com/guides.html.md#hardware)
Quick answers
## Frequently asked questions about DGX Spark
Power, unified memory, networking, remote access and troubleshooting answers checked against current NVIDIA documentation.
### What is the expected DGX Spark power draw?
Budget for the included **240W** power supply. NVIDIA specifies a **140W TDP** for the GB10 SoC and reserves **100W** for ConnectX‑7, Wi‑Fi, storage, USB‑C and other system components. Its EU technical disclosure reports **233.2W maximum** and **38W idle** under the stated test method. The power shown by `nvidia-smi` is not whole-system power at the wall.
- [NVIDIA power requirements →](https://docs.nvidia.com/dgx/dgx-spark/hardware.html#power-requirements)
- [NVIDIA measured-power disclosure →](https://docs.nvidia.com/dgx/dgx-spark/compliance.html)
### Why does `nvidia-smi` report “Memory-Usage: Not Supported”?
This is expected. DGX Spark’s integrated GPU shares unified system memory instead of having dedicated framebuffer memory, so `nvidia-smi` does not provide the usual aggregate VRAM-usage field. Use `top` , `htop` , `free -h` or DGX Dashboard for system-memory monitoring. Open the dashboard locally at `http://localhost:11000` ; remote access needs NVIDIA Sync or an SSH tunnel.
- [NVIDIA known issue →](https://docs.nvidia.com/dgx/dgx-spark/known-issues.html#nvidia-smi-reports-memory-usage-not-supported)
- [DGX Dashboard access →](https://docs.nvidia.com/dgx/dgx-spark/dgx-dashboard.html)
### Why does an application run out of memory below the 128 GB capacity?
Unified-memory applications can report less allocatable memory than the system can reclaim, and some software does not yet account correctly for swap or cache. Start with `free -h` , the process allocations and swap use. Linux normally reclaims clean cache automatically. For a controlled diagnostic only—not routine memory management—sync writes and drop reclaimable cache with:
```
sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
```
Dropping caches can create significant I/O and CPU work; stop the affected workload first and remeasure before treating cache as the cause.
- [NVIDIA unified-memory guidance →](https://docs.nvidia.com/dgx/dgx-spark/known-issues.html#guidance-for-reporting-memory-resources-with-unified-memory-architecture)
- [Linux drop-caches warning →](https://docs.kernel.org/admin-guide/sysctl/vm.html#drop-caches)
### Why does NVIDIA Sync say “Host already exists” after I deleted a device?
A stale SSH `Host` alias may remain. Back up the file, remove only the exact stale `Host` block, then add the device again. Current Sync documentation uses `~/.ssh/config` on macOS and Ubuntu and `C:\Users\\.ssh\config` on Windows. Older Sync builds may also have a managed file at the paths below.
```
Windows: C:\Users\\AppData\Local\NVIDIA Corporation\Sync\config\ssh_config
macOS: /Users//Library/Application Support/NVIDIA/Sync/config/ssh_config
Linux: /home//.config/NVIDIA/Sync/config/ssh_config
```
- [Review NVIDIA Sync SSH aliases →](https://docs.nvidia.com/sync/latest/direct-connections.html#importing-an-existing-ssh-configuration)
### Is GPUDirect RDMA supported on DGX Spark?
No. DGX Spark’s unified-memory architecture does not support GPUDirect RDMA, `nvidia-peermem` , `dma-buf` or GDRCopy for CUDA device allocations. A portable application should query `CU_DEVICE_ATTRIBUTE_GPU_DIRECT_RDMA_SUPPORTED` and `CU_DEVICE_ATTRIBUTE_DMA_BUF_SUPPORT` , then use a supported fallback. For Linux `ibverbs` applications, NVIDIA suggests `cudaHostAlloc` memory registered with `ibv_reg_mr` .
- [NVIDIA DGX Spark porting guidance →](https://docs.nvidia.com/dgx/dgx-spark-porting-guide/porting/cuda.html#gpudirect-rdma)
### How do I set up DGX Spark after moving it to another location?
Connect Ethernet first. Try the Spark’s `spark-hostname.local` address; if mDNS is unavailable, find its wired IP in the router. Connect with NVIDIA Sync or SSH, identify the Wi‑Fi device, join the new network without putting the password in shell history, and note the Wi‑Fi address before unplugging Ethernet:
```
nmcli device status
sudo nmcli --ask device wifi connect ""
ip -4 address show
```
- [NVIDIA Sync direct connections →](https://docs.nvidia.com/sync/latest/direct-connections.html)
- [NVIDIA network setup guidance →](https://docs.nvidia.com/dgx/dgx-spark/first-boot.html)
### Why are CPU-bound NVCC processes slow?
CUDA sources built through CMake can miss OpenMP compiler support. Pass `-Xcompiler=-fopenmp` for CUDA compilation, rebuild and remeasure the CPU-bound section.
```
# CMakeLists.txt
target_compile_options(mytarget PRIVATE
$<$:-Xcompiler=-fopenmp>
)
# Command line
nvcc -Xcompiler=-fopenmp ...
```
- [NVIDIA DGX Spark compiling guide →](https://docs.nvidia.com/dgx/dgx-spark-porting-guide/porting/compilation.html)
### Can I reset a lost BIOS or UEFI Administrator password?
There is no documented user reset for a lost firmware Administrator password. Contact NVIDIA hardware support for a Founders Edition system or the device manufacturer for an OEM model. Do not assume that an OS recovery or “Restore Defaults” in UEFI clears the security credential.
- [Choose the correct support route →](https://docs.nvidia.com/dgx/dgx-spark/support.html)
- [DGX Spark UEFI security settings →](https://docs.nvidia.com/dgx/dgx-spark-uefi/security-tab.html)
### Why is the ConnectX‑7 module missing from `lspci` ?
The January 2026 DGX OS release added ConnectX‑7 hot-plug power management. With it enabled, the adapter can stay powered down until a QSFP cable is inserted, saving up to 18W. Connecting the cable activates the adapter and increases system power and temperature. To keep ConnectX‑7 active without hot-plug power saving, remove the marker; recreate it to re-enable the feature:
```
sudo rm -f /etc/nvidia/cx7-hotplug-enabled
sudo touch /etc/nvidia/cx7-hotplug-enabled
```
- [NVIDIA January 2026 release note →](https://docs.nvidia.com/dgx/dgx-spark/release-notes.html#january-2026-release)
### How many DGX Spark systems can I cluster?
Current NVIDIA Sync supports **two or three systems** with direct 200 Gb/s QSFP cabling, or **up to four systems through a switch** . Use an approved cable and the NVIDIA Sync Cluster Assistant or the matching two-node, three-node or switched NVIDIA playbook; four devices are not supported as a direct-cabled topology.
- [ConnectX‑7 cabling and playbooks →](https://docs.nvidia.com/dgx/dgx-spark/spark-clustering.html)
- [NVIDIA Sync Cluster Assistant →](https://docs.nvidia.com/sync/latest/cluster-assistant.html)
The operating rule
Keep the agent on the server. Grant every client capability separately.
- [Choose your operator](#operators)
- [Build the Spark host](#spark-setup)
- [Choose the rest of the stack](https://isaiuseful.com/tools.html.md)
---
## https://isaiuseful.com/dgx-station (`/dgx-station.html.md`)
# Is NVIDIA DGX Station Right for Your AI Workload?
Canonical source: [https://isaiuseful.com/dgx-station](https://isaiuseful.com/dgx-station)
DGX Station buyer's guide · checked 15 August 2026
DGX Station is a desk-side NVIDIA system where official NVFP4 builds of the 233 GiB-class MiniMax M3 and 465 GB GLM‑5.2 can fit in one coherent memory space. That does not make every token fast, every framework ready or a roughly €100,000 workstation the right first machine.
- [Compare with Spark](#comparison)
- [Find the model sweet spot](#sweet-spot)
- [Check independent evidence](#reviews)
Official NVIDIA starting points
## Pick a DGX Station recipe.
These four official playbooks expose Station-specific strengths: a full chat-model training run, Blackwell NVFP4 quantization, large-batch robotics fine-tuning and profiler-led GPU kernel development. Use our [training](https://isaiuseful.com/training-models.html.md) , [robotics](https://isaiuseful.com/robotics.html.md) and [acceptance](#acceptance) guides to define the evaluation before you follow the commands.
- [Train a Chat Model with NanoChat 12 hours Run the tokenizer, pretraining and supervised fine-tuning pipeline, then chat with the resulting checkpoint in a web UI or CLI.](https://build.nvidia.com/station/nanochat)
- [Quantize Models to NVFP4 60 min Use NVIDIA Model Optimizer to compress an 8B checkpoint, validate quality and serve it through an OpenAI-compatible endpoint.](https://build.nvidia.com/station/nvfp4-quantization)
- [Isaac GR00T N1.6 Fine-Tuning 45 min Fine-tune the robotics action stack on LIBERO Spatial, then run open-loop evaluation and measure inference latency.](https://build.nvidia.com/station/gr00t)
- [Profiler-Driven Kernel Optimization 2 hrs Profile Llama 3.1 8B fine-tuning, then build and benchmark fused RMSNorm and cross-entropy kernels in Triton.](https://build.nvidia.com/station/kernel-dev-ft)
- [Official DGX Station collection **Explore every NVIDIA Station playbook** Open build.nvidia.com/station](https://build.nvidia.com/station)
**252 GB**
GPU HBM3e
7.1 TB/S VENDOR SPECIFICATION
**496 GB**
CPU LPDDR5X
396 GB/S VENDOR SPECIFICATION
**20 PFLOPS**
peak FP4 Tensor compute
WITH SPARSITY · NOT APPLICATION SPEED
Read the architecture correctly
## How does DGX Station's 748 GB unified memory work?
The Blackwell Ultra GPU and Grace CPU can access a coherent pool, but the pool contains two physically different memory tiers. Weight placement, KV cache, precision and runtime support still decide useful performance.
Vendor specification
### A real 252 GB HBM fast lane.
The B300 GPU has 252 GB of HBM3e at 7.1 TB/s. Models whose weights, cache and runtime allocations stay here are the cleanest fit for low-latency work.
Capacity tier
### Another 496 GB is coherent, but slower.
Grace contributes LPDDR5X at 396 GB/s. Models larger than HBM can remain local without PCIe staging, but decode can slow when active weights repeatedly cross the lower-bandwidth tier.
Vendor ceiling
### “Up to 1T” means aggressively quantized.
One trillion BF16 parameters need about 2 TB before cache or runtime overhead. A mixed NVFP4 checkpoint at roughly 5.0–5.2 effective bits per parameter would use about 625–650 GB before runtime and cache, so the 1T claim is a tight capacity boundary—not a comfortable default.
### Do not collapse the numbers 20 PFLOPS does not predict tokens per second.
NVIDIA's peak figure is sparse FP4 Tensor arithmetic. Autoregressive decode is often constrained by memory movement, kernels, batch size and the number of active MoE parameters. Ask for time to first token, decode speed per user, aggregate throughput, power at the wall and the exact model/precision/context—not a single peak-compute number.
- [Official product specifications →](https://www.nvidia.com/en-eu/products/workstations/dgx-station/)
- [DGX Station development guide →](https://docs.nvidia.com/dgx/dgx-station-development-guide/Intro.html)
### Windows is a distinct configuration Do not treat “DGX Station for Windows” as an OS swap for every Ubuntu Station.
As checked 24 July 2026, NVIDIA labels the Windows product “Coming in Q4.” It keeps the GB300 platform, adds Windows infrastructure and WSL support, and can be configured with an additional RTX PRO GPU. The current Linux OEM listings describe Ubuntu with NVIDIA AI Developer Tools. Buy against the exact Windows or Ubuntu SKU, driver branch, expansion hardware and OEM support entitlement; do not assume a conversion path unless that OEM documents one.
- [Read the official Windows product notice →](https://www.nvidia.com/en-us/products/workstations/dgx-station-for-windows/)
DGX Spark vs DGX Station
## Prototype box versus deskside AI node.
Spark maximizes affordable memory capacity in a tiny power envelope. Station adds a datacentre-class GPU memory tier, much larger coherent capacity, enterprise management and team-serving options.
| Decision | DGX Spark | DGX Station | What changes | Practical reading |
| --- | --- | --- | --- | --- |
| Compute | GB10 Grace Blackwell; 20-core Arm CPU | GB300 Grace Blackwell Ultra; 72-core Arm CPU | Up to 1 versus 20 sparse FP4 PFLOPS | A 20× peak ratio is not a universal 20× workload ratio. |
| NVFP4 | ✓ Native Blackwell
Validate the current GB10 kernel and engine path. | ✓ Native Blackwell Ultra
Native W4A4 when the serving stack supports the model. | Both are newer-generation FP4 platforms; Station adds much more fast and coherent memory. | H100/H200 servers are different: Hopper lacks native FP4 Tensor Cores and uses a fallback where supported. |
| Coherent memory | 128 GB LPDDR5X | 748 GB total: 252 GB HBM3e + 496 GB LPDDR5X | 5.8× total capacity plus a dedicated HBM tier | Station moves 70–400B models into a far healthier memory envelope. |
| Memory bandwidth | 273 GB/s across unified LPDDR5X | 7.1 TB/s HBM3e; 396 GB/s CPU LPDDR5X; 900 GB/s C2C | A fast lane and a capacity lane | Know which tier holds the active weights and cache. |
| Vendor model ceiling | Up to 200B; up to 405B with two | Up to 1T; two systems can link | Frontier open-weight capacity becomes a one-node experiment | Ceilings describe fit, not quant quality or interactive speed. |
| Networking | ConnectX‑7 at 200 Gb/s | ConnectX‑8 up to 800 Gb/s | Faster two-node and storage fabric | Optics, cables, storage and switching are separate costs. |
| Power | 140 W GB10 TDP; 240 W supply | 1,600 W total-system specification | Office appliance becomes facilities-aware equipment | Confirm circuit, heat, acoustics and OEM configuration before ordering. |
| Operations | DGX OS; single-user companion or small endpoint | Ubuntu 24.04 with NVIDIA AI Developer Tools/CUDA-X plus BMC, Redfish and DCGM capabilities; NVIDIA AI Enterprise is a distinct entitlement unless the order states otherwise. | Central team node and fleet management become credible | A [separate Windows Station](https://www.nvidia.com/en-us/products/workstations/dgx-station-for-windows/) is listed as coming in Q4; verify the exact SKU, NVIDIA AI Enterprise order line and support branch. |
| Best role | PoC, model evaluation, private single-user agents and ARM/CUDA development | Large-model development, local frontier inference, shared lab service and migration rehearsal | From proving a workflow to reproducing a larger production model class | Neither is automatically a highly available production service. |
**Comparison basis:** NVIDIA's published specifications and NVFP4 technical notes as checked 26 July 2026. [DGX Spark](https://www.nvidia.com/en-eu/products/workstations/dgx-spark/) · [DGX Station](https://docs.nvidia.com/dgx/dgx-station-development-guide/Intro.html) · [Blackwell W4A4 versus Hopper W4A16](https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/) . “PFLOPS” figures are peak sparse FP4 vendor claims, not measured end-to-end results.
The actual sweet spot
## Buy for the model's working set, not its parameter headline.
These bands are planning estimates for one active model with useful context headroom. Package metadata, quantization scales, multimodal towers, speculative models, cache and concurrency can move the boundary.
DGX Spark · native NVFP4
### 20–35B sweet spot; ≈150–165B planning ceiling
Everyday agents fit well. The speculative ceiling reserves 20–25% of 128 GB and assumes a mixed checkpoint at 5.0–5.2 bits per parameter. NVIDIA's “up to 200B” figure is a tighter vendor ceiling, not the comfortable target.
Station HBM lane · native NVFP4
### ≈290–320B parameter planning ceiling
This is the distinctive performance tier: the estimate leaves 20–25% of 252 GB HBM3e for runtime and cache. Larger checkpoints can still fit, but they cross into the slower coherent-memory lane.
Station coherent lane · native NVFP4
### ≈860–960B parameter planning ceiling
This is speculative mixed-checkpoint memory math across 748 GB, not a benchmark. A 1T model may load only with less headroom, more aggressive precision choices or offload.
| Model class | Approximate weight math | DGX Spark | DGX Station | Verdict |
| --- | --- | --- | --- | --- |
| 20–35B dense | ≈13–23 GB at 5.0–5.2 bits/parameter before runtime | Comfortable with useful context | Easy, but usually poor capital efficiency for one stream | Spark, RTX or rented GPU is normally the sweet spot. |
| 30–120B sparse MoE | ≈19–78 GB at 5.0–5.2 bits/parameter; all experts stay stored | Best balance when the NVFP4 build and GB10 kernels are mature | Strong high-concurrency or higher-precision tier | Spark for one developer; Station for a team or heavier evaluation. |
| 70–120B dense | ≈44–78 GB at mixed NVFP4; ≈70–120 GB at ideal eight-bit | Fits, but dense decode can expose the 273 GB/s limit | Clean HBM-resident target with room for cache | Station begins to make performance sense if this is the daily workload. |
| 200–405B | ≈125–263 GB at mixed NVFP4 before runtime and cache | A 200B fit is tight; two-node planning is ≈295–330B with headroom versus NVIDIA's up-to-405B ceiling | Roughly 200–300B can remain in HBM with useful reserve; larger builds spill into coherent memory | Station's clearest single-box advantage. |
| MiniMax M3 · 428B / 23B active | Official repositories: 232.9 GiB NVFP4; 413.3 GiB MXFP8; 795.5 GiB BF16 | One is out; two leave too little reserve for the official NVFP4 package plus runtime and useful cache | NVFP4 fits across coherent memory, but nearly fills the 252 GB HBM tier before runtime and cache | Plausible local candidate; the published NVIDIA recipe is nightly vLLM, TP8 on B200—not a measured single-Station run. |
| GLM‑5.2 · 753B / 40B active MoE | Official NVIDIA NVFP4 repository: 465 GB; BF16: about 1.51 TB, before runtime | Out of scope; even two are below the official NVFP4 package size | Fits across coherent memory with about 283 GB left before runtime and cache; not HBM-resident | Measured in the linked Station hands-on at about 24 decode tok/s for one stream; retain its exact runtime, prompt and quantization when comparing. |
| 1T class | ≈625–650 GB at mixed NVFP4; about 2 TB at BF16 | Out of scope | Tight capacity demonstration with ≈98–123 GB left before runtime and cache | Do not interpret the vendor ceiling as comfortable serving, full-precision training or million-token context. |
GLM‑5.2 · official NVFP4 verdict
### It loads and runs. The memory-tier penalty is visible.
GLM‑5.2 has 753B total and 40B active parameters. NVIDIA's official NVFP4 repository is 465 GB, so the weights fit inside 748 GB but exceed the 252 GB HBM tier. The linked Station hands-on reports about 24 decode tok/s for one stream and about 243 prefill tok/s; use those as video-reported workload results, not a universal speed claim.
**The first acceptance run**
Pin the exact NVFP4 checkpoint, runtime/container, prompt length, output length and concurrency. Record cold and warm time to first token, per-user decode, aggregate throughput, peak memory in each tier, power at the wall and task accuracy against the unquantized or hosted reference.
**Keep the GLM versions separate**
GLM‑5.3 is available through the Coding Plan and uses the same base model as GLM‑5.2, but Z.ai had not released its weights or local serving recipes when checked on 15 August. Keep the 465 GB fit and every Station speed figure labelled GLM‑5.2 until a 5.3 checkpoint is measured on the same setup.
- [Official GLM‑5.2 model card →](https://huggingface.co/zai-org/GLM-5.2)
- [Official NVIDIA NVFP4 checkpoint →](https://huggingface.co/nvidia/GLM-5.2-NVFP4)
- [Official GLM‑5.3 availability and benchmarks →](https://z.ai/blog/glm-5.3)
- [Inspect the GLM‑5.2 Station measurements →](#station-benchmarks)
- [NVFP4 format and memory math →](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)
- [Compare both GLM deployment states →](https://isaiuseful.com/cloud-models.html.md#glm-53)
- [Compare the GLM family →](https://isaiuseful.com/local-models.html.md#catalog)
MiniMax M3 · official NVFP4 candidate
### Plausible to run locally. Benchmark before you call it a GPT‑5.4 mini replacement.
NVIDIA's MiniMax M3 NVFP4 repository totals about 232.9 GiB. It fits in Station's 748 GB coherent pool, but is approximately all of the advertised 252 GB HBM tier after unit conversion, leaving no comfortable HBM-only room for the runtime, multimodal tower, speculative model or KV cache. The current model card requires a nightly vLLM image and shows TP8 on B200; it does not establish single-GB300 Station support or speed.
**The replacement test**
Replay 50–100 real GPT‑5.4 mini jobs through the same agent scaffold. Compare passed tasks, reviewer edits, wall-clock time, tool-call failures, input/cache/output tokens, TTFT, decode, peak HBM and LPDDR use, and wall power. Current API prices make M3 2.5× cheaper on input and 3.75× on output below 512K, but only cost per accepted task tells you whether the switch saves money.
- [Official MiniMax M3 repository →](https://github.com/MiniMax-AI/MiniMax-M3/)
- [NVIDIA NVFP4 build and nightly recipe →](https://huggingface.co/nvidia/MiniMax-M3-NVFP4)
- [vLLM support matrix and B300 validation →](https://vllm-project.github.io/2026/06/12/minimax-m3-vllm.html)
- [Independent M3 versus GPT‑5.4 mini comparison →](https://isaiuseful.com/benchmarks.html.md#artificial-analysis)
- [Independent terminal-agent run →](https://isaiuseful.com/benchmarks.html.md#terminal-bench)
- [Compare cloud price, licence and topology →](https://isaiuseful.com/cloud-models.html.md#minimax-m3)
Independent reviews + benchmarks
## The first Station hands-on evidence needs careful scope.
As of 3 August 2026, Alex Ziskind's ASUS Station hands-on adds agent-concurrency testing. Hosted-model comparisons and non-Station local MiniMax M3 measurements also exist, but procurement still needs reproducible model-level tokens-per-second, latency, memory-tier placement and wall-power tables on the exact Station configuration.
Independent model check
### M3 is competitive—not uniformly equal.
Artificial Analysis scores M3 44 versus GPT‑5.4 mini at xhigh 40 on its composite. On the much narrower English word-sense SenseBench, M3 is at 90.62% and several GPT‑5.4 mini low runs are at 89.49–90.89%. Those are hosted-model results, not proof of an NVFP4 Station build.
- [Artificial Analysis comparison and caveats →](https://isaiuseful.com/benchmarks.html.md#artificial-analysis)
- [SenseBench scope and leaderboard →](https://isaiuseful.com/benchmarks.html.md#sensebench)
Independent agent run
### 31.5% versus 13.5%—with 13× the input tokens.
On ClawProBench's 89-task OpenCode TerminalBench 2.1 run, M3 solved 28 tasks and GPT‑5.4 mini solved 12. M3 used 302.0M input tokens versus 22.6M, so the result supports capability but warns against translating cheaper tokens directly into cheaper completed work.
- [Inspect the task, harness and token totals →](https://suyoumo.github.io/terminal-bench/)
Station hands-on scope
### A real workload test is useful—if you keep its boundary.
The video below tests agent concurrency on an ASUS ExpertCenter Pro ET900N G3. Do not generalize that run to every checkpoint, quantization, context or concurrency. Separately, a community 4-bit MLX run on a 512 GB Mac Studio M3 Ultra reported 27.2 tok/s at a 1K prompt and 16.6 tok/s at 65K, with 226.6–238.1 GB peak memory.
- [Watch the Station hands-on →](#station-hands-on)
- [Inspect the community M3 run →](https://www.reddit.com/r/LocalLLaMA/comments/1u7q046/minimax_m3_4_bit_mlx_initial_benchmark_on_mac/)
- [Video: This was a data center a year ago… Now it's on my desk](https://www.youtube.com/watch?v=qV_K0nTF6gY)
Creator hands-on · Alex Ziskind · 30 July 2026
### Agent-concurrency evidence on an ASUS DGX Station
Ziskind tests NVFP4 model throughput, continuous batching, an active agent swarm, power and thermals on the ExpertCenter Pro ET900N G3. The charts below reconstruct the video-reported values; preserve the exact runtime, prompts and logs before transferring them to a purchase decision.
Video-reported benchmark snapshot
### Single-stream latency and batch throughput answer different questions.
Every bar keeps its metric and concurrency attached. Approximate ranges are shown as ranges; their bars use the midpoint only for visual scale. The 4,096 tok/s Nemotron burst is separated from the roughly 2,600 tok/s result reported at 128 concurrent requests.
NVFP4 · video-reported
#### Model throughput at one request
Decode and prefill use separate axes. Longer is better within each block only.
##### Single-stream decode · tok/s
- [Video chapter: 04:52 Play decode results](https://www.youtube.com/watch?v=qV_K0nTF6gY&t=292s)
##### Prompt prefill · tok/s
- [Video chapter: 07:53 Play prefill results](https://www.youtube.com/watch?v=qV_K0nTF6gY&t=473s)
Nemotron 3 Super 120B · total output
#### Continuous-batching scale
Aggregate tok/s across parallel requests—not the speed seen by each user.
Recorded burst
**up to 4,096 tok/s**
- [Video chapter: 09:20 Play concurrency run](https://www.youtube.com/watch?v=qV_K0nTF6gY&t=560s)
Active local agents + system load
#### Swarm throughput and power
Throughput and watts use separate axes. Power values are reported at the superchip boundary.
##### Agent-swarm output · tok/s
- [Video chapter: 14:34 Play swarm run](https://www.youtube.com/watch?v=qV_K0nTF6gY&t=874s)
##### Reported superchip power · W
Heavy-load temperature
**≈55°C peak**
- [Video chapter: 15:08 Play load and thermals](https://www.youtube.com/watch?v=qV_K0nTF6gY&t=908s)
Qwen3 235B-A22B
### >5,000 tok/s
Reported just above this level at 128 concurrent requests. Nemotron and Qwen both showed a dip around concurrency 32.
- [Video chapter: 10:59 Play result](https://www.youtube.com/watch?v=qV_K0nTF6gY&t=659s)
DeepSeek‑V4‑Flash
### 64-request sweet spot
The video identifies concurrency 64 as the efficiency peak. Treat this as an optimum marker rather than a throughput comparison because no exact tok/s value accompanies this observation.
- [Video chapter: 11:24 Play result](https://www.youtube.com/watch?v=qV_K0nTF6gY&t=684s)
GLM‑5.2
### 35 → 56 tok/s
Aggregate decode rises from about 35 tok/s at concurrency 4 to 56 tok/s at 16; prompt processing reaches about 1,800 tok/s at 16.
- [Video chapter: 11:31 Play result](https://www.youtube.com/watch?v=qV_K0nTF6gY&t=691s)
**How to read these charts:** values are transcribed from one creator-run video on one ASUS configuration and are approximate where the video reports a range. Official model repositories verify the model names and NVFP4 checkpoints; they do not independently verify the Station measurements.
- [Open the embedded benchmark video →](#station-hands-on)
- [Official Nemotron 3 Super NVFP4 checkpoint →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4)
- [Official Qwen3 235B NVFP4 checkpoint →](https://huggingface.co/nvidia/Qwen3-235B-A22B-NVFP4)
- [Official DeepSeek‑V4‑Flash NVFP4 checkpoint →](https://huggingface.co/nvidia/DeepSeek-V4-Flash-NVFP4)
- [Official GLM‑5.2 NVFP4 checkpoint →](https://huggingface.co/nvidia/GLM-5.2-NVFP4)
GB300 architecture evidence · not Station evidence
### [MLPerf](https://isaiuseful.com/benchmarks.html.md#mlperf) shows the rack-scale platform advancing. It does not benchmark this one-GPU workstation.
Third-party submission
### Lambda reported 1.26 minutes for Llama 2 70B LoRA.
Lambda's [MLPerf Training](https://isaiuseful.com/benchmarks.html.md#mlperf) v5.1 table reports a 72× GB300 NVL72 cluster at 1.26 minutes for Llama 2 70B LoRA and 14.25 minutes for Llama 3.1 8B. Its 1.27× comparison is against the best GB200 NVL72 result from the prior MLPerf round, and Lambda attributes gains to both hardware and a newer software stack.
- [Inspect Lambda's methods and table →](https://lambda.ai/blog/lambda-mlperf-training-benchmarks-v5.1)
Vendor submission
### NVIDIA reports large gains versus Hopper—at rack scale.
NVIDIA says GB300 NVL72 delivered more than 4× Llama 3.1 405B pretraining and nearly 5× Llama 2 70B LoRA performance versus Hopper with the same GPU count. That is useful training-system evidence, but it combines architecture, methods, networking and software.
- [Read NVIDIA's MLPerf account →](https://blogs.nvidia.com/blog/mlperf-training-benchmark-blackwell-ultra/)
Transfer limit
### Do not divide the rack result by 72.
Distributed training does not scale linearly down to one desktop GPU, and a Station's 252 GB HBM3e differs from the 279 GB accelerators listed in Lambda's cluster. Only a run on the quoted Station SKU answers the purchase question.
- [Carry the distinction into acceptance →](#acceptance)
What a reproducible follow-up must contain
### Six numbers, one model manifest and the raw result file.
The creator video supplies useful measurements. A procurement-grade comparison still needs the complete manifest, repeated runs and raw result file below.
Responsiveness
### TTFT + decode
Cold and warm time to first token plus tokens per second per user at 1, 8 and 32 concurrent requests.
Context
### 8K · 64K · 256K
Prefill throughput, cache size, memory-tier placement and decode degradation at useful—not merely advertised—contexts.
Efficiency
### Wall watts + accuracy
Energy per million generated tokens and task-quality regression versus the reference checkpoint at the same harness settings.
**Evidence status:** one creator-run ASUS Station workload test is included here. Treat it as early third-party evidence, not a substitute for the reproducible table above or an invitation to fill remaining gaps with rack-scale GB300 marketing results.
DGX Station OEM systems
## Choose the implementation and support contract—not only the GB300 badge.
NVIDIA defines and brands the DGX Station reference platform, but the official route is to contact a partner rather than use a first-party checkout. These manufacturer pages are buying routes, not endorsements; storage, added RTX PRO graphics, cooling, acoustics, rack conversion, regional availability, warranty and software support can differ.
Procurement boundary · checked 10 August 2026
### Separate the preconfigured Station base from the production software entitlement.
**Preconfigured base stack:** NVIDIA documents Ubuntu 24.04 with NVIDIA AI Developer Tools/CUDA-X, plus BMC, Redfish and DCGM capabilities.
**GB300 is Blackwell:** Hopper-generation DGX systems include NVIDIA AI Enterprise in the DGX software bundle; DGX Station GB300 needs a separate NVIDIA AI Enterprise purchase unless its exact order or Entitlement Certificate (EC) says otherwise. There is no public free-forever entitlement for the complete supported production suite. A general 90-day production evaluation can be requested for compatible infrastructure; it includes Omniverse, excludes Run:ai, and support is governed by the offer and EC. Require an EC or order line naming the product, GPU metric, term, support level, start date and renewal.
**Free developer components are not that production entitlement:** Omniverse and AI Workbench are free components; NIM through the NVIDIA Developer Program is for development, research and test (up to 16 GPUs on the standard route, community support). Production self-hosting generally needs NVIDIA AI Enterprise.
**OEM-specific claims:** Supermicro's material establishes native support/compatibility and Ubuntu 24.04 with NVIDIA AI Developer Tools; ASUS says NVIDIA AI Enterprise can be deployed and supported. Neither statement by itself proves a paid production entitlement or how long it lasts.
- [NVIDIA DGX Station and partner order route →](https://www.nvidia.com/en-us/products/workstations/dgx-station/)
- [DGX Station Ubuntu and software requirements →](https://docs.nvidia.com/dgx/dgx-station-development-guide/porting/software-requirements.html)
- [NVIDIA AI Enterprise licensing guide →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/licensing.html)
- [NVIDIA AI Enterprise product and evaluation route →](https://www.nvidia.com/en-us/data-center/products/ai-enterprise/)
- [NIM developer-program route and production boundary →](https://forums.developer.nvidia.com/t/nvidia-nim-faq/300317)
- [Omniverse license and support route →](https://docs.omniverse.nvidia.com/dev-guide/latest/common/NVIDIA_Omniverse_License_Agreement.html)
- [AI Workbench introduction →](https://docs.nvidia.com/ai-workbench/user-guide/latest/overview/introduction.html)
- [Supermicro compatibility and base-stack claim →](https://www.supermicro.com/datasheet/datasheet_Supermicro_Super_AI_Station.pdf)
- [ASUS deployment and support claim →](https://www.asus.com/us/displays-desktops/workstations/performance/expertcenter-pro-et900n-g3/)

### ExpertCenter Pro ET900N G3
DGX Station / GB300 tower.
- [Open ASUS system page →](https://www.asus.com/displays-desktops/workstations/performance/expertcenter-pro-et900n-g3/)

### Dell Pro Max with GB300
GB300 AI development workstation.
- [Open Dell system page →](https://www.dell.com/en-us/lp/dell-pro-max-nvidia-ai-dev)

### VWS-158270643
GB300 workstation configuration.
- [Open Exxact system page →](https://www.exxactcorp.com/Exxact-VWS-158270643-E158270643?utm_source=web%20referral&utm_medium=backlink&utm_campaign=NVIDIA%20DGX%20Station%20NVIDIA%20Page&utm_term=NVIDIA%20Station%20Page)

### W775-V10-L01
Deskside AI supercomputer.
- [Open GIGABYTE system page →](https://www.gigabyte.com/Enterprise/Tower-Server/W775-V10-L01)

### HP ZGX Fury
DGX Station-class workstation.
- [Open HP system page →](https://reinvent.hp.com/ZGX-FURY)

### XpertStation WS300
NVIDIA DGX Station.
- [Open MSI system page →](https://www.msi.com/Landing/NVIDIA-DGX-STATION)

### Super AI Station
ARS-511GD-NB-LCC · tower or 5U.
- [Open Supermicro system page →](https://www.supermicro.com/en/accelerators/nvidia/super-ai-station)
**Shortlist + quote check:** send every OEM the same workload manifest and facilities questionnaire. Require the exact chassis and bill of materials, SSD and PCIe population, optional RTX PRO GPU, OS/driver/support branch, network accessories, power, cooling, acoustics, rack kit, regional service, lead time and acceptance result. Compare a signed configuration and repeatable workload result—not seven differently worded product claims. [Use the shared acceptance gates →](#acceptance)
Try before you buy
## Rent the workload before approving a DGX Station or other six-figure system.
Before a five- or six-figure hardware commitment, rent the candidate workload and establish whether the model’s memory tier, serving behavior and acceptance result justify ownership. A rental is a test drive, not a GPU-availability, hardware-validation, benchmark-equivalence or price guarantee.
**Make the comparison reproducible.** Run the exact model, engine, quantization, context, concurrency and acceptance harness you would deploy; retain the container, settings, logs, latency, throughput, tiered-memory use, power and task-quality results.
**Referral disclosure:** this isaiuseful.com link is a Runpod referral link. As checked 10 August 2026, eligible new first-time users must sign up through it with Google SSO and load their first $10: European users receive $5 credit, while non-European users receive a weighted $5–$500 credit (Runpod says most are $10 or less). Terms can change. If eligible, isaiuseful.com receives its referral bonus and earns Runpod credits on actual usage for the first six months: 3% of Pod spend and 5% of Serverless spend. Using the link supports the site.
- [Rent a Runpod test environment through this referral link →](https://runpod.io?ref=l40ix174)
- [Read Runpod’s current referral terms →](https://docs.runpod.io/accounts-billing/referrals)
OEM product views
## See the systems. Keep the evidence label attached.
These ASUS, MSI, Dell and HP introductions establish product shape and positioning. They are vendor demonstrations, not independent reviews or benchmark runs.
- [Video: ASUS ExpertCenter Pro ET900N G3 – Deskside AI Supercomputer built on NVIDIA DGX Station](https://www.youtube.com/watch?v=z70xOVy3V40)
OEM product film · ASUS
### ExpertCenter Pro ET900N G3
Use the film to inspect ASUS's system framing. Confirm storage, added RTX PRO GPU, support, acoustics, lead time and final memory figure on the quoted SKU.
- [Video: NVIDIA DGX Station XpertStation WS300 Trailer](https://www.youtube.com/watch?v=VWZB-W_H7VE)
OEM product film · MSI
### XpertStation WS300
The trailer shows MSI's DGX Station implementation. Treat its performance language as positioning until the exact system is independently tested.
- [Video: The AI supercomputer built for your desk](https://www.youtube.com/watch?v=T-del6FlwnA)
OEM product film · Dell
### A desk-side system for the larger model tier
Dell's short overview shows the intended form factor and enterprise role. Ask Dell for the same acceptance evidence as any other OEM.
- [Video: Introducing the HP Z8 Fury G6i AI Workstation | HP Z](https://www.youtube.com/watch?v=M7k9LLjp5qo)
OEM product film · HP
### Z8 Fury G6i AI Workstation
HP's introduction shows its DGX Station implementation and intended workstation role. Confirm the exact memory, storage, graphics, support and acceptance result on the quoted configuration.
The €100K question
## Capacity is cheap only when the workflow uses it.
One European NVIDIA Elite Partner listed DGX Station at €99,888 excluding VAT on 24 July 2026. OEM configuration, warranty, delivery, storage, networking and regional pricing can change the actual project total.
Buy Station
### The large model is part of daily work.
You repeatedly run or adapt 120–700B models, the data cannot use an ordinary cloud endpoint, several developers share the node, and queueing or egress already has a measured cost.
Buy Spark first
### The workflow is still the uncertain part.
You need private PoC, agent evaluation, ARM/CUDA development or sparse MoE inference below 128 GB. Prove demand before multiplying capital cost and facilities work.
Rent first
### The frontier model is occasional.
Use a short controlled rental or OEM evaluation unit to establish precision, quality, throughput and utilization. Buy only when the repeatable workload beats the fully loaded alternative.
Indicative price
### Do not compare €99,888 with bare GPU rental alone.
Compare three-year capital, financing, support, power, cooling, networking, storage, rack or office work, staff time, downtime and residual value with reserved and on-demand cloud—including data transfer and the cost of an approved private environment. Then attach the calculation to one measured workload volume.
- [European listing checked 24 July 2026 →](https://dgxstation.ai/)
- [Model the hosted alternative →](https://isaiuseful.com/cloud-models.html.md#economics)
Before the purchase order
## Accept the system against your workload.
DGX Station is an OEM platform, so the NVIDIA architecture does not settle storage, support, acoustics, delivery, additional GPU or operating-system details for a particular quote.
| Gate | Ask the supplier | Run during acceptance | Keep | Reject when |
| --- | --- | --- | --- | --- |
| Exact configuration | Memory, SSDs, added RTX PRO GPU, firmware, OS/support branch, warranty and lead time | Inventory, health, ECC, storage and network checks | Signed bill of materials and support entitlement | The delivered SKU or software branch differs from the tested quote. |
| Software entitlement | EC or order line for NVIDIA AI Enterprise: GPU metric, term, support level, start date, renewal and any trial conditions | Register only the stated entitlement and record its support contact | EC, activation record and named escalation path | Preinstallation, a NIM download or an OEM compatibility claim is offered instead of the stated entitlement. |
| Model result | The exact checkpoint, quant, container, engine, context and concurrency | Your cold/warm benchmark and quality holdout | Raw logs, manifest and result summary | Only peak PFLOPS, rack results or an undisclosed prompt are offered. |
| Facilities | Maximum and typical draw, connector/circuit, heat, sound pressure and service clearance | Sustained load in the intended room | Power, thermal and acoustic readings | The office, circuit or cooling cannot sustain the quoted configuration. |
| Recovery | Firmware/OS recovery, BMC access, spare parts and response times | Rebuild one model service from a clean manifest and restore its data | Offline recovery material and tested runbook | A failed SSD, update or image leaves no supported recovery route. |
| Scale path | Validated cables, optics, two-node software, storage fabric and migration support | Only the topology you expect to buy | Network diagram, compatibility list and measured delta | “Up to 800 Gb/s” substitutes for an end-to-end result. |
**ARM64 still matters.** DGX Station uses a Grace Arm CPU. NVIDIA lists PyTorch, Jupyter, vLLM, SGLang, Ollama and its own stack, but your private packages, security agents, databases and binary extensions still need an ARM64 test. Container support is not the same as every dependency being portable.
The short answer
Spark proves the workflow. Station proves the larger model class. Neither proves the business case for you.
- [Choose the model tier](#sweet-spot)
- [Open the DGX Spark guide](https://isaiuseful.com/remote-spark.html.md)
- [Compare model families](https://isaiuseful.com/local-models.html.md)
- [Place the lab node](https://isaiuseful.com/diy-palantir.html.md#paths)
---
## https://isaiuseful.com/diy-palantir (`/diy-palantir.html.md`)
# How to Build a Sovereign Operational Intelligence Stack
Canonical source: [https://isaiuseful.com/diy-palantir](https://isaiuseful.com/diy-palantir)
Build · sovereign operational intelligence ·
reviewed 28 July 2026
Palantir's practical advantage is not a magic model. It is a governed ontology over integrated data, connected to the workflows, permissions and analytical services that turn records into decisions.
- [Choose components](#stack)
- [See the architecture](#architecture)
- [Pick a deployment path](#paths)
**11**
component layers
REPLACEABLE INTERFACES
**3**
deployment paths
LOCAL · CONTROL · ELASTIC
**1**
governed decision loop
SOURCE TO AUDIT EVENT
01 · The premise
## What is an operational intelligence stack?
Treat this as enterprise data management with a modern operational interface, not as an ML research programme.
Palantir packages familiar enterprise disciplines—data integration, governed semantic models, workflow applications and model deployment—into one coherent system. Its public architecture documents expose the same recognisable layers: [Ontology](https://www.palantir.com/docs/foundry/ontology/overview/) , [data integration](https://www.palantir.com/docs/foundry/data-integration/overview/) , [Workshop applications](https://www.palantir.com/docs/foundry/workshop/overview/) and [model integration](https://www.palantir.com/docs/foundry/model-integration/overview/) . The difficult value is the coherence between them, not a novel foundation-model layer.
The builder's opportunity is not to reproduce Foundry feature for feature. It is to recover perhaps 80% of the operational value—an explicit design target, not a measured universal result—by joining mature open components around one bounded use case. The hard part remains identifiers, data quality, permissions, ownership and an interface that fits the work.
This guide is for data engineers, operations teams and government or enterprise teams that need to prove the shape before making a platform-scale commitment. The weekend exercise is intentionally narrow: one redacted or synthetic source, one graph domain, one decision screen and one measurable outcome.
**Takeaway:** define one operational decision and its owner before installing anything.
02 · Core components
## Buy coherence with interfaces, not with one vendor.
These are composable options, not a shopping list: the weekend build needs only one choice per required row.
| Layer | Palantir equivalent | OSS option A | OSS option B | Notes |
| --- | --- | --- | --- | --- |
| Graph / ontology | [Foundry Ontology](https://www.palantir.com/docs/foundry/ontology/overview/) | [JanusGraph](https://janusgraph.org/) on [Apache Cassandra](https://cassandra.apache.org/_/index.html) | [Apache TinkerPop](https://tinkerpop.apache.org/) | JanusGraph supplies the graph layer and can use Cassandra as its distributed storage backend; TinkerPop supplies the graph API and Gremlin language, not a complete database. |
| Ingestion & pipeline | [Data Connection](https://www.palantir.com/docs/foundry/data-integration/overview/) | [Apache Kafka](https://kafka.apache.org/documentation/) | [Apache NiFi](https://nifi.apache.org/docs.html) | Kafka for durable event streams; NiFi for visible routing and provenance. For one nightly CSV, a script is enough. |
| Distributed storage | [Foundry datasets](https://www.palantir.com/docs/foundry/data-integration/datasets/) | [Apache Iceberg](https://iceberg.apache.org/) on [SeaweedFS](https://github.com/seaweedfs/seaweedfs) | [Garage](https://garagehq.deuxfleurs.fr/) or [Ceph RGW](https://docs.ceph.com/en/latest/radosgw/) | SeaweedFS is the new-project default here because it is active, Apache-2.0 and explicitly targets S3 and Iceberg workloads. Pin the catalog, engine and object-store versions and test their exact S3 operations together. |
| Query & analytics | [Code Workbook](https://www.palantir.com/docs/foundry/code-workbook/overview/) | [Apache Spark](https://spark.apache.org/docs/latest/) | [Trino](https://trino.io/docs/current/) | Use Spark for distributed transforms and Trino for interactive SQL across sources. DuckDB can replace both in the PoC. |
| Search / mixed index | [Ontology search and exploration](https://www.palantir.com/docs/foundry/object-explorer/overview/) | [Apache Solr](https://solr.apache.org/) | JanusGraph mixed index | Use Solr for text, faceting and geospatial search, or as a JanusGraph mixed-index backend. Keep Cassandra as the graph store; an index is rebuildable, not the system of record. |
| Geospatial processing | Geospatial transforms | [Eclipse GeoTrellis](https://projects.eclipse.org/projects/locationtech.geotrellis) | [Apache Spark](https://spark.apache.org/) | GeoTrellis adds Scala/Spark raster processing and tiled geospatial data handling. Add it for terrain, imagery or coverage analysis—not for ordinary latitude/longitude markers. |
| 3D globe / map UI | Gotham map workspace | [CesiumJS](https://cesium.com/platform/cesiumjs/) | [MapLibre](https://maplibre.org/) | CesiumJS is the open-source 3D globe; Cesium ion is a separate hosted commercial service. Use MapLibre for a conventional 2D/2.5D operational map. |
| Workflow / decisions UI | [Workshop](https://www.palantir.com/docs/foundry/workshop/overview/) | [LocStat](https://www.locstat.co.za/) | [Apache Superset](https://superset.apache.org/docs/intro) | Use LocStat as a Palantir-clone reference for operational UI patterns; verify its current code availability and licence before treating it as a production dependency. Superset is an Apache-licensed dashboard and exploration layer, not a transactional app. |
| Orchestration | [Pipeline Builder](https://www.palantir.com/docs/foundry/building-pipelines/overview/) | [Apache Airflow](https://airflow.apache.org/docs/) | [Prefect](https://docs.prefect.io/) | Schedule, retry and observe batch work. Do not let two orchestrators own the same retry. |
| Model serving (optional) | [Model integration](https://www.palantir.com/docs/foundry/model-integration/overview/) | [Ollama](https://docs.ollama.com/) | [vLLM](https://docs.vllm.ai/) | Ollama is the simple local route; vLLM is a throughput-oriented server for a rented GPU. Keep models thin, task-specific and optional. |
| Security / access | [Foundry security](https://www.palantir.com/docs/foundry/security/overview/) | [Apache Ranger](https://ranger.apache.org/) | [Open Policy Agent](https://www.openpolicyagent.org/docs/latest/) | Ranger centralises policies and audit for supported data services; OPA evaluates policy as code inside applications and infrastructure. Identity and secrets remain separate jobs. |
### Choose object storage deliberately
Do not start a new community deployment on MinIO by default. Its [community repository](https://github.com/minio/minio) was archived on 25 April 2026, is read-only and now describes the community edition as source-only; the embedded web interface is an object browser rather than the former administration console. Existing installations can still run, but an unmaintained storage layer is a migration risk, not a neutral default.
| Object store | Best fit | Why choose it | Material constraint |
| --- | --- | --- | --- |
| [SeaweedFS](https://github.com/seaweedfs/seaweedfs) Default for this guide | A new self-hosted analytical stack, from a one-node proof to a distributed deployment. | Active Apache-2.0 project with an S3 endpoint, explicit Iceberg support, downloadable releases and a single-binary development mode. | Its master, volume, filer and S3 roles introduce a different operating model. Prove authentication, upgrades, failure recovery and the exact Iceberg client path before production. |
| [Garage](https://garagehq.deuxfleurs.fr/) | Lightweight storage replicated across unreliable sites or modest hardware. | Active AGPLv3 project designed for simple, resilient multi-site S3 storage, with binaries, containers, a CLI and an administration API. | Garage deliberately uses replication rather than erasure coding and does not implement every S3 feature, including ACLs and bucket policies. Run compatibility tests; do not assume drop-in parity. |
| [Ceph Object Gateway](https://docs.ceph.com/en/latest/radosgw/) | A larger datacentre that already operates Ceph or needs a broad object-storage control surface. | Mature S3-compatible gateway with user management, multisite, encryption, policy and erasure-coded storage options. | Ceph is a storage platform, not a weekend sidecar. Its cluster design, monitoring, upgrades and recovery need dedicated ownership. |
| Managed S3-compatible service | Teams prioritising support and low storage-operations burden over full infrastructure ownership. | The provider owns hardware repair, service upgrades and durability engineering. | Region, keys, administrators, subprocessors, egress and exit tooling determine sovereignty. Contract and restore tests still matter. |
| MinIO Community | Migration planning for an existing pinned deployment—not a new default. | Existing S3 compatibility and operational knowledge may justify a bounded transition period. | Archived upstream, source-only community distribution, reduced web UI and AGPLv3 obligations. Inventory the exact build, isolate it and set a dated migration plan. |
### Start with the graph, not “AI”
A relational schema is useful when rows and joins are stable. Operational questions usually cross changing relationships—shipment to vehicle, vehicle to depot, depot to incident, incident to responsible team—so a property graph can attach attributes to both entities and relationships without forcing every question into one rigid table shape; JanusGraph documents this model through [Apache TinkerPop and Gremlin](https://docs.janusgraph.org/getting-started/architecture/) .
### Keep the lake authoritative
Land raw and cleaned records in versioned Iceberg tables, then project only the operational entities and relationships needed into the graph. This makes the graph a serving model rather than the only copy of reality, and Iceberg's documented snapshots support reproducible reads and rollback.
### Put workflow before models
A queue of delayed shipments with an owner and an acknowledge action is useful without an LLM. Add Ollama or vLLM only for a narrow task such as classifying free-text incident notes; keep rules, permissions and final actions outside the model.
**Takeaway:** for the weekend build, use SeaweedFS + Iceberg with an explicit catalog, JanusGraph, Trino and Superset; add Kafka, Spark or a model only when the sample workflow proves the need.
03 · Architecture
## Three tiers, one traceable path.
Ingestion accepts source changes; storage and ontology create governed meaning; serving exposes only the decision surface.
The ingestion tier receives database changes, files and field events through Kafka, NiFi or a small batch loader. The storage/ontology tier writes immutable source records and curated Iceberg tables to object storage, then maps stable operational identifiers and links into JanusGraph. The serving/UI tier uses Trino or Spark for queries, Superset or a small LocStat-inspired application for decisions, and an optional model endpoint for bounded classification or extraction.
> Visual: Flowchart showing data moving from ingestion through storage and ontology to query and workflow services
### 01 Ingestion tier
**ERP / CSV / sensors**
Source records and events
**Kafka / NiFi**
Validate, route and retain
### 02 Storage / ontology tier
**Iceberg on an S3 API**
SeaweedFS by default; versioned records
**JanusGraph projection**
Operational objects and links
### 03 Serving / UI tier
**Trino / Spark**
Queries, transforms and features
**Ollama / vLLM**
Optional bounded inference
**Workflow app / Superset**
Views, approvals and audited actions
**Takeaway:** preserve the raw record, make ontology projection repeatable and log every decision back as an event.
04 · Operational globe
## Turn a cinematic demo into a governed operational picture.
A shared viewport is the beginning. Operational value starts when every mark carries origin, observation time, uncertainty and an accountable next action.
Bilawal Sidhu's [spy-satellite simulator](https://www.spatialintelligence.ai/p/i-built-a-spy-satellite-simulator) and follow-up idea [IronSight](https://www.spatialintelligence.ai/p/ironsight-turning-2d-videos-into) are useful interface provocations for two different jobs. WorldView combines public spatial feeds in one navigable scene. IronSight synchronizes ordinary camera footage, reconstructs a shared 3D scene and keeps human-labelled tests and visible failure cases beside the polished replay. Neither project proves production sensor access, identification accuracy or operational readiness.
- [Video: Ex-Google Maps PM Vibe Coded Palantir In a Weekend](https://www.youtube.com/watch?v=rXvU7bPJ8n4)
WorldView · public-feed fusion
### One globe, several imperfect feeds
Borrow the movement from wide-area context to a source-backed object record. Keep feed age, coverage and simulation state visible.
- [Video: Fable 5 Is Nuts. I Vibe Coded a Baby Anduril for My Range.](https://www.youtube.com/watch?v=FH4eS0oi4uE)
IronSight · inspectable reconstruction
### One event, several camera views
Borrow synchronization, source-frame comparison, review queues and failure displays—not the certainty implied by the HUD.
### Use public feeds as fixtures and supplements—not as authority
Public feeds make the interface concrete, but they arrive with different coverage, licences, clocks and failure modes. This first matrix separates what made the prototype visually persuasive from what an accountable production system would need instead.
Compare public feeds and production fallbacks
6 sources · caveats · owned routes
| Source | What the prototype used | Production caveat | Owned or authoritative fallback |
| --- | --- | --- | --- |
| [OpenSky Network](https://openskynetwork.github.io/opensky-api/rest.html) | Sidhu reported 7,000+ changing aircraft positions. The API exposes live state vectors, tracks and flights. | Coverage is receiver-dependent, quotas apply and operational or commercial use requires a written agreement. Treat identity and completeness as unverified. | Replay a licensed snapshot for development; for operations, use the aviation authority's approved feed or your own authorized receivers and retain the raw messages. |
| [ADS-B Exchange](https://gateway.adsbexchange.com/api/aircraft/v2/docs/index.html?url=%2Fapi%2Faircraft%2Fv2%2Fdocs%2Fopenapi.json) | Crowdsourced aircraft positions, including an API endpoint filtered to aircraft marked military. | API access is commercial and entitlement-based. A transponder flag is not authoritative identity, intent or a complete air picture. | Use it as a supplementary layer beside approved surveillance data; use synthetic tracks when real identifiers or movements are unnecessary. |
| [CelesTrak GP data](https://celestrak.org/NORAD/documentation/gp-data-formats.php) | The demo selected 180+ satellites and propagated their orbits from published element sets. | TLE or OMM records are orbital elements, not continuously observed live positions. Surface element epoch, propagation time and expected error. | Pin a dated element-set fixture for tests; use the organisation's approved space catalogue or sensor-derived track when the decision is consequential. |
| [OpenStreetMap](https://www.openstreetmap.org/copyright/attribution-guide/) | Road geometry under a particle effect that suggests vehicle flow. | OpenStreetMap supplies mapped features—not live vehicle movement. Attribute the data, and do not build production traffic on the community tile servers. | Self-host approved OSM extracts and tiles, then join the road graph to the transport authority's GIS, counters or licensed flow data. |
| [City of Austin traffic cameras](https://data.austintexas.gov/Transportation-and-Mobility/Traffic-Cameras/b4k4-adkb) | Geolocated public camera imagery projected into the 3D scene. | Austin's dataset publishes locations and, where allowed, a latest screenshot on a five-minute cadence. It explicitly disclaims survey suitability and does not retain daily video. | Use approved internal camera services, retention rules and surveyed asset locations; store fixture images for development and outage tests. |
| [Google Photorealistic 3D Tiles](https://developers.google.com/maps/documentation/tile/3d-tiles) | A high-resolution textured city mesh rendered in CesiumJS. | It requires billing, an API key and on-screen attribution; Google restricts caching, offline use, extraction and machine analysis. EEA terms and returned content can differ. | Keep the operational layers independent of the basemap. Fall back to government terrain, orthophotos and 3D city models served through the existing GIS. |
The full, searchable operational-intelligence catalogue—including owned GIS and self-hosted map routes—is in [Tools](https://isaiuseful.com/tools.html.md#tools-build-operational-intelligence-systems) .
**From feeds to governed records.** Once each source has an authority level and an outage path, the next job is to keep observations, resolved identities, derived assessments and operator actions distinct. Otherwise a polished map quietly turns uncertain reports into apparent facts.
Inspect the operational data and decision layers
5 layers · minimum records · guardrails
| Layer | Minimum record | Operator view | Guardrail |
| --- | --- | --- | --- |
| Base world | Terrain or imagery tile, provider, capture date, resolution and licence. | A 2D map or 3D globe with scale, coordinates and imagery age visible. | Do not let attractive basemaps imply that the scene is live. |
| Reported observation | Source ID, observed-at and received-at times, geometry, classification, confidence and raw-record link. | Selectable marks, trails and time controls; stale and low-confidence records look distinct. | Preserve the report separately from the entity it may describe. |
| Resolved entity | Stable internal ID, source aliases, proposed matches, reviewer and merge history. | One object dossier with every supporting and conflicting observation. | Never merge identities only because two marks overlap on screen. |
| Derived assessment | Rule, query or model version; inputs; generated time; output; uncertainty; expiry. | An overlay that can be hidden and traced back to evidence. | Label inference as inference; expiry prevents an old assessment becoming a permanent “fact”. |
| Workflow action | Case, assignee, permitted action, decision, rationale, timestamp and outcome. | Triage queue and case panel beside the map—not just more glowing layers. | Require human authorisation for consequential actions and preserve the audit event. |
### Build the time machine before the live map
Normalize each connector into an observation envelope and retain both event time and ingestion time. Record a feed heartbeat, expected update interval, last successful record and licence or redistribution constraint. Then replay a saved time window at different speeds. Deterministic replay makes late events, duplicate suppression, entity resolution and operator decisions testable without depending on a live third-party feed.
### Separate reality, simulation and presentation
Keep three explicit namespaces: **observed** records received from a source, **simulated** tracks created for training or demonstration, and **derived** interpretations produced by rules or models. Show a permanent mode banner and source legend, and prevent simulated objects from crossing into production alerts. A sensor cone, orbital path or coverage footprint is a model output unless it came from a documented source; its assumptions belong in the object panel.
### Use a thin geospatial serving path
Store raw payloads immutably, curate geometry and timestamps into Iceberg, relate identities and cases in the graph, and publish bounded vector tiles or GeoJSON through an authenticated API. CesiumJS or MapLibre should receive only the viewport, time range and fields the operator may see. Cluster and aggregate on the server; do not stream the whole lake into the browser.
**Takeaway:** public data is excellent for discovering the interface and shaping realistic fixtures. A production decision must survive the feed disappearing, changing terms or disagreeing with the authoritative system.
05 · Product case study
## Study the seams that turned a weekend build into a product.
World Monitor is useful here because its working interface, source and unusually detailed documentation expose the engineering and commercial boundaries behind the spectacle.
[World Monitor's account of its origin](https://www.worldmonitor.app/docs/about) says it began as a weekend project in January 2026. It now publishes a free dashboard, paid Pro and API plans, and an Enterprise offer. That is evidence of a commercial product and pricing structure—not evidence of revenue, profit or customer retention. The distinction matters when using a successful-looking build as a business case.
As checked 23 July 2026, the [shared dashboard URL](https://www.worldmonitor.app/dashboard?zoom=1.00&view=global&timeRange=7d&layers=conflicts%2Cbases%2Chotspots%2Cnuclear%2Csanctions%2Cweather%2Ceconomic%2Cwaterways%2Coutages%2Cmilitary%2Cnatural) restored a global seven-day view and its selected layers. The live interface exposed cached or live status, source age, coverage counts, methodology links, resizable panels and a command palette. Those small contracts make a dense map inspectable, reproducible and shareable; they are more valuable to copy than its visual drama.
| Product seam | What World Monitor documents | DIY translation | Acceptance check |
| --- | --- | --- | --- |
| State is an interface | Map view, time range and layers live in the URL; [Route Explorer](https://www.worldmonitor.app/docs/route-explorer) also serializes origin, destination, commodity and active tab. | Put viewport, time window, filters, scenario and case ID in a versioned URL or saved-view record. A link should reconstruct the same evidence window without a narrated setup. | Open the link in a clean session and obtain the same bounded working set, including the same distinction between observed and simulated data. |
| Contracts before connectors | Its newer domain APIs begin as Protocol Buffer contracts; generation produces typed clients, server interfaces and OpenAPI, while CI checks drift and breaking changes in the [endpoint workflow](https://www.worldmonitor.app/docs/adding-endpoints) . | Define Observation, Entity, Assessment, Case and Action contracts before adding adapters. Keep vendor payloads at the edge and translate them into owned schemas. | A breaking field change fails CI, and recorded source fixtures still replay through the generated client and server boundary. |
| Ingest off the request path | Independent seed jobs fetch sources on different cadences, keep the previous cache on failure and hydrate common datasets in fast and slow startup tiers. Conditional loading and adaptive polling stop work for hidden panels and disabled layers. | Schedule and deduplicate source collection separately from page requests. Land raw data first, publish a curated cache second and fetch only the layers required by the current decision. | One slow or failed provider cannot blank the interface, multiply upstream calls or delay the first useful operator view. |
| Absence is a first-class state | Per-feed circuit breakers, stale-on-error caches and source-specific freshness thresholds keep partial service available. If core inputs disappear, the risk panel says “insufficient data” instead of displaying an apparent all-clear. | Every response carries observed-at, ingested-at, last-success, expected cadence, freshness, degraded status and reason. Never encode unavailable as zero or an empty healthy list. | Pull a core feed: stale data remains visibly stale, the missing coverage is named and any dependent score is withheld or qualified. |
| Compute has an authority boundary | Local geometry lookup, clustering, selected ML fallbacks and other presentation work can run in the browser; published scores, briefs, forecasts and operational APIs remain server-authoritative. | Use browser workers for clustering, display transforms and offline convenience. Run governed identity resolution, scoring, permissions and actions in controlled services against versioned data. | The UI remains responsive or partially useful offline without creating a second, conflicting source of operational truth. |
| Agents receive tools, not the lake | REST domain endpoints and MCP tools expose bounded operations. Cache-backed tools return freshness metadata, and [JMESPath projections](https://www.worldmonitor.app/docs/mcp-jmespath) let a caller request only the fields it needs. | Expose small, read-only task contracts with field projection, row limits, provenance and stable error shapes. Keep database credentials, unrestricted queries and write authority outside the agent. | An agent answers a known operational question from cited records within a fixed payload and time budget, then hands any action to the normal approval path. |
| The paid layer is closer to a decision | The free observatory supplies broad awareness. Paid plans add [scenarios](https://www.worldmonitor.app/docs/scenario-engine) , route analysis, scheduled digests, MCP/API access and enterprise identity or deployment options. | Charge for saved monitoring, lower-latency alerts, scenario work, workflow integration, collaboration, controlled deployment and assurance—not merely for repackaging public dots on a map. | A paid feature shortens or improves a named decision loop; measure activation, retained use and operator outcome rather than map visits. |
### Borrow the product seams, not the public-dashboard topology
World Monitor's [documented architecture](https://www.worldmonitor.app/docs/architecture) —vanilla TypeScript, browser-side work, Redis-backed caches, scheduled seeders, edge functions and a separate live-data relay—is a coherent response to a public, read-heavy product. A private sovereign stack still needs an authoritative lake, an operational graph, organisation identity, row- or object-level policy, durable workflow state, audit and tested restore. Its consumer topology is evidence for specific patterns, not a reference architecture to copy whole.
> Visual: Six implementation seams to borrow from the World Monitor case study
**Visual reading order:**
1. **01** **Seed** Collect each source on its own cadence; retain the last good state.
2. **02** **Contract** Normalize into owned, typed observation and action schemas.
3. **03** **Cache** Version keys, coalesce misses and serve stale data explicitly.
4. **04** **Load by intent** Fetch the active view; pause hidden panels and unused feeds.
5. **05** **Decide** Move from map context into one scenario, route, case or approval.
6. **06** **Measure** Log the action and outcome; meter value at the workflow seam.
### Keep code, data and brand rights separate
The [published licence guide](https://www.worldmonitor.app/docs/license) says the platform is AGPL-3.0-only, while named thin client packages are MIT-licensed; it also separates commercial licensing and trademark permission. A modified public network deployment may therefore carry source-offer obligations, and upstream feed licences or API terms remain separate. For a sovereign build, either use the interfaces as learning material and implement your own bounded system, comply with the AGPL, or negotiate different terms before combining the code with a proprietary product.
**Commercial check, 23 July 2026:** the [published plan table](https://www.worldmonitor.app/docs/pricing) listed Pro at $39.99/month, API at $99.99/month, API Business at $249.99/month and custom Enterprise pricing. These are vendor-published prices, not audited revenue.
**Takeaway:** the reusable breakthrough is an honest path from unreliable signals to a shareable state, typed contract, visible freshness and paid decision workflow—not “vibe coding” or the number of map layers.
06 · Where AI fits
## Put models at uncertain edges, not at the centre.
The platform should still ingest, relate, query and route work when the model endpoint is unavailable.
AI is useful where operational data becomes ambiguous: extracting entities from incident notes, suggesting that two records refer to the same asset, classifying a message, summarising a long case or translating an operator's question into a read-only query. These are proposed interpretations of evidence, not new facts.
| AI job | Input and output | Authority | Non-AI baseline | Acceptance check |
| --- | --- | --- | --- | --- |
| Extract and classify | Unstructured notes, reports or email become typed fields, labels and source spans. | Write to a review queue or derived table; never overwrite the source record. | Rules, regular expressions and controlled forms. | Measure missed fields and false matches on a versioned, representative sample. |
| Entity-resolution suggestion | Candidate records become a proposed match with evidence and confidence. | A rule or reviewer approves graph merges; the model cannot silently join identities. | Exact identifiers and deterministic fuzzy matching. | Track false merges separately from missed matches; false merges are usually harder to undo. |
| Summarise and explain | A bounded case bundle becomes a short brief with links back to records. | Advisory only. The operator can inspect every cited record before acting. | A fixed template populated from trusted fields. | Test material omissions, unsupported statements and time saved—not writing style. |
| Natural-language query | An operator question becomes constrained SQL, Gremlin or a saved query. | Read-only service account, query allow-list, row limits and visible generated query. | Curated filters, dashboards and saved queries. | Run known questions against expected result sets and reject unsafe or unbounded queries. |
| Recommend a next step | Current state and approved policy become ranked options with reasons. | Human approval before allocation, dispatch, targeting, enforcement or any other consequential action. | Rules, thresholds and established optimisation solvers. | Compare decision quality, constraint violations and operator overrides with the baseline. |
| Accelerate delivery | Approved designs and contracts become code, tests, migrations and interface variants. | Normal code review, security scanning, tests and deployment gates still apply. | Human implementation using the same specifications. | Measure accepted change lead time and escaped defects, not generated lines of code. |
### Keep inference as a derived, traceable record
Store the source identifiers, prompt or task version, model and adapter version, timestamp, output, confidence where meaningful, reviewer decision and superseding result. This lets a later model produce a new interpretation without rewriting history. Apply the same access policy to prompts and outputs as to the source data they contain.
### Use the cheapest adequate intelligence
Start with SQL, rules, graph traversals and established statistical or optimisation models. Add a small instruction-tuned model through [Ollama](https://docs.ollama.com/) when the data must remain local, or [vLLM](https://docs.vllm.ai/) when an internally controlled GPU service needs higher throughput. The relevant deployment choices are compared in the site's [local-model guide](https://isaiuseful.com/local-models.html.md) , [DGX Station guide](https://isaiuseful.com/dgx-station.html.md) and [self-hosted model guide](https://isaiuseful.com/cloud-models.html.md) .
### Fail closed and degrade usefully
Timeouts, malformed output and model refusal should return the operator to the ordinary queue, saved query or rule-based result. Model access goes through one authenticated gateway with task-specific schemas, budgets and logs; the model does not receive database credentials or direct write access.
**Takeaway:** add one model-assisted task only after its non-model baseline, review boundary and evaluation set exist.
07 · Implementation paths
## Choose where the operational burden lives.
All three paths can use the same logical interfaces; they differ mainly in control, staffing and scale.
| Path | Cost | Ops burden | Data sovereignty | Scalability |
| --- | --- | --- | --- | --- |
| A · PoC Fully local: workstation, DGX Spark, DGX Station or small server | Lowest incremental cost if hardware exists; no managed-service bill. Do not buy Station until a 128 GB machine is a measured constraint. | Low only while single-node and disposable. Use containers, sample data and backups. | Strong physical control; still restrict local accounts, volumes and exports. | Enough for one source and a small team. Replace Kafka/Spark with files and DuckDB if sensible; add GPU capacity only for a named model, vision or simulation test. |
| B · Control Bare metal or on-premises Kubernetes | Hardware, power, backup capacity and staff time become material. | Highest: patching, certificates, storage, observability, recovery and capacity are yours. | Best placement control when residency or disconnected operation is mandatory. | Good with a capable platform team; Kubernetes does not remove stateful-system work. |
| C · Elastic Cloud-hosted managed services | Fast start, then usage and egress charges; tag the PoC and set budgets. | Lower for managed Kafka, object storage and Spark, but IAM and data governance remain yours. | Depends on provider, region, keys, subprocessors and contract; verify rather than assume. | Highest elasticity. AWS, Azure and Google Cloud each document managed streaming, object storage and Spark services. |
Local AI hardware inside Path A
### A strong lab node—not a shortcut around architecture.
DGX Spark and DGX Station can be PoC, development or test machines when local model, vision, simulation or in-memory compute is itself under test. The basic data and workflow PoC remains far smaller, and neither machine should become the ontology, object store, graph, workflow engine and recovery plan merely because it has unified memory.
PoC · capacity question
### Only buy the larger node when the experiment needs it.
Use existing gear or Spark for one replayed feed, DuckDB, a graph and a bounded 20–35B model. Use Station to prove a 200B+ local model, large vision pipeline, heavy embedding build or sensitive frontier-model workflow—not to make the hardware purchase the experiment.
- [Check Spark's measured limits →](https://isaiuseful.com/remote-spark.html.md#spark-reviews)
- [Size the Station workload →](https://isaiuseful.com/dgx-station.html.md#sweet-spot)
Development · shared lab
### Keep the model service bounded.
Host the approved model endpoint, CUDA containers, notebooks, evaluation jobs and synthetic-data runs for a small team. Station can divide its GPU into up to seven MIG instances, subject to workload memory, while the governed data services remain independently deployable.
Test · promotion rehearsal
### Make the node replaceable.
Exercise packaging, quantization, concurrent agents, failure recovery, access controls and promotion to cloud or datacentre infrastructure. Pull the model endpoint during a test and prove that the ordinary queue, saved query or rule-based result still works.
> Visual: Local AI lab node in an operational intelligence development path
**Visual reading order:**
1. **01 · replay** **Saved source window** Deterministic feeds and synthetic sensitive records.
2. **02 · data** **DuckDB / Iceberg / graph** Owned identifiers, lineage and test fixtures.
3. **03 · model** **Bounded AI endpoint** Spark, Station or a rented equivalent.
4. **04 · workflow** **Review queue** Human approval and visible source evidence.
5. **05 · evaluate** **Known cases** Accuracy, latency, overrides and cost.
6. **06 · promote** **Production target** Same manifest; separate capacity and resilience decision.
**Placement rule:** the AI machine accelerates a bounded model or data-compute seam. Kafka, the object store, graph, identity and workflow remain separate services with their own recovery. A successful lab run proves the workload and interface—not high availability or a production architecture.
Service references: [Amazon MSK](https://aws.amazon.com/msk/) , [Azure Event Hubs for Kafka](https://learn.microsoft.com/en-us/azure/event-hubs/event-hubs-for-kafka-ecosystem-overview) , [Google Managed Service for Apache Kafka](https://cloud.google.com/managed-service-for-apache-kafka/docs/overview) , [Amazon S3](https://aws.amazon.com/s3/) , [Google Dataproc](https://cloud.google.com/dataproc/docs/concepts/overview) and [Azure HDInsight Spark](https://learn.microsoft.com/en-us/azure/hdinsight/spark/apache-spark-overview) .
**Takeaway:** prove the workflow on Path A, move to Path B when sovereignty or disconnected operation is a real requirement, and choose Path C only when its managed-service trade is acceptable.
08 · Private datacentre
## Run it on VMware without pretending Kubernetes is mandatory.
For teams keeping operational data off third-party infrastructure, vSphere virtual machines are the shortest production path; Cloud Director adds tenant boundaries when an internal platform team serves several departments.
There are two credible VMware shapes. Use ordinary vSphere VMs when one team owns the stack and operational simplicity matters. Use Kubernetes through [vSphere Supervisor](https://techdocs.broadcom.com/us/en/vmware-cis/vsphere/vsphere-supervisor/8-0.html) , or tenant clusters exposed through [Cloud Director Container Service Extension](https://github.com/vmware/container-service-extension) , only when the datacentre already operates that control plane.
| VMware layer | Small production default | Scaled / tenant option | Data boundary | Decision note |
| --- | --- | --- | --- | --- |
| Compute | Separate VM groups for ingress, data services and serving; reserve memory for JanusGraph, Trino and Kafka. | Supervisor or Cloud Director tenant Kubernetes clusters with explicit resource quotas. | Keep management, storage and workload networks separate; deny direct internet egress from data services. | Do not introduce Kubernetes solely for this stack. VMs make state, failure domains and recovery easier to inspect. |
| Object storage | A tested multi-VM SeaweedFS deployment on dedicated virtual disks backed by a named vSphere storage policy; never promote the one-node development mode. | Ceph RGW when a storage team already operates Ceph; Garage when simple multi-site replication matters more than erasure coding or full S3 coverage. | Encrypt in transit and at rest; keep keys, snapshots and replicas under the organisation's control. | A VM snapshot is not an application-consistent object-store backup. Test bucket and Iceberg-catalog restoration separately. |
| Persistent volumes | Attach and document VMDKs directly for VM deployments. | [vSphere CSI driver](https://github.com/kubernetes-sigs/vsphere-csi-driver) with storage classes mapped to approved policies. | Restrict datastore, snapshot and volume permissions per tenant and service account. | Validate expansion, topology, backup and restore on the exact vSphere/CSI versions in use. |
| Network entry | Internal load balancer or reverse-proxy VMs; private DNS and organisation-issued certificates. | The datacentre's supported Kubernetes ingress and load-balancer integration. | Expose the workflow UI only to operator networks; keep Kafka, graph, object-store and model ports private. | Make firewall denies part of acceptance testing, including blocked workload egress. |
| Identity & secrets | Federate the UI with the existing identity provider; use separate machine identities and a private secrets service. | Namespace/tenant roles plus OPA or Ranger policies; never treat a Cloud Director organisation as application authorisation. | Administrators, backup operators and monitoring systems are data-access paths too. | Document who can read consoles, disks, snapshots, logs and backups before importing sensitive data. |
### A practical vSphere layout
Begin with three security zones: ingestion VMs can reach approved sources; data VMs host Kafka or NiFi, SeaweedFS/Iceberg, JanusGraph and Trino; serving VMs host Superset, the workflow API and optional Ollama. Put Airflow or Prefect in the data zone, forward only bounded telemetry to the monitoring zone, and send audit logs to an append-restricted target.
### Availability without theatre
Use vSphere anti-affinity rules to separate replicas across hosts, but test application failure rather than assuming VM restart equals service recovery. Back up configuration, graph data, the Iceberg catalog and object data on their own schedules; restore them into an isolated network and replay a known decision case before calling the design recoverable.
### Cloud Director is a control plane, not a data policy
Cloud Director can provide isolated organisations, virtual datacentres, networks and quotas; its [API documentation](https://developer.broadcom.com/xapis/vmware-cloud-director-api/latest/) is the interface to automate those boundaries. If a managed-datacentre provider operates vSphere or Cloud Director, contract terms and privileged access still determine whether “private” meets the sovereignty requirement.
**Takeaway:** default to well-separated vSphere VMs, keep every data service on private networks, and prove a full restore before considering a tenant Kubernetes layer.
09 · Weekend build
## Stop at one closed loop.
The deliverable is a working decision path, not an enterprise platform.
**Visual reading order:**
1. **Friday** **Specify** Name the decision, source, owner, SLA and acceptance measure.
2. **Saturday 09:00** **Seed** Land a redacted fixture with observed, received and last-success times.
3. **Saturday 12:00** **Contract** Type the observation, entity, assessment, case and action records.
4. **Saturday 16:00** **Relate** Create only the graph edges and exception query the decision needs.
5. **Sunday 09:00** **Decide** Build one shareable view with freshness, evidence and an acknowledge action.
6. **Sunday 15:00** **Break it** Pull a feed, replay known cases and record misses, latency and corrections.
**Takeaway:** finish when one operator can open a reproducible view, recognise missing evidence, act on one trusted exception and leave an audit event.
10 · Interface references
## Use clone projects as UI workshops, not platforms.
Foundry-style projects show how data, lineage and ontology might be navigated; Gotham-style projects show maps, alerts and a shared operational picture. Their immediate value is making interface choices tangible.
Most small “Palantir clone” repositories should be treated as rapidly assembled prototypes. That is not a dismissal: a working screen is often better than a slide deck for asking operators what they need to see, which actions belong beside an alert and what context is missing. It is not evidence that the repository should own production data, identity or workflow state.
| Reference | Archetype | What to borrow | How to use it |
| --- | --- | --- | --- |
| [koala73/worldmonitor](https://github.com/koala73/worldmonitor) | Open-source OSINT product | Shareable state, typed service contracts, freshness semantics, bounded agent tools and the transition from observation to paid decision workflows. | Start with the [engineering case study above](#world-monitor) , then read the architecture and licence before opening the code. Reimplement validated seams in the owned stack or comply with its AGPL terms; do not treat public-feed breadth as an operational data foundation. |
| [bilawalsidhu/gods-eye-view](https://github.com/bilawalsidhu/gods-eye-view) | Spatial-intelligence globe | Wide-area-to-object navigation, time-aware layers, object dossiers and the visual language separating a global overview from a selected case. | As checked 23 July 2026, the repository contains preview assets and a README stating that code is being prepared for release. It has no application code, release or licence yet, so use it only as a design reference and recheck before adoption. |
| [cherishwins/OpenFoundry](https://github.com/cherishwins/OpenFoundry) | Foundry-style data workspace | Dataset, ontology, lineage, governance and application navigation; its repository also exposes service and API boundaries. | Run it in a disposable environment or give selected screens and API contracts to a coding assistant as reference material. Ask for the same user journey against your maintained services in the language and framework your team already operates—not a line-by-line port. |
| [Przyval/openfoundry](https://github.com/Przyval/openfoundry) | Foundry SDK compatibility experiment | How an application-facing object API can be shaped around Foundry-like SDK expectations. | Turn its interfaces into contract tests and mock client flows. Do not promise compatibility until your implementation passes the calls your application actually uses. |
| [drissman/faber-foundry](https://github.com/drissman/faber-foundry) | Foundry-style architecture experiment | Vocabulary and boundaries around ontology, lineage and governance. | Compare its domain split with your own architecture. Reimplement only the bounded capability required by the pilot. |
| [simplifaisoul/osiris](https://github.com/simplifaisoul/osiris) | Gotham-style operational picture | Map composition, layers, event cards, filters, timelines and the visual hierarchy of a live operations room. | Load synthetic events and put the interface in front of operators. Record which layers affect a decision; delete the decorative ones. |
| [nabylb/aegis-intelligence](https://github.com/nabylb/aegis-intelligence) | Gotham-style feed aggregation | MapLibre views and ways to combine event, aircraft, vessel and conflict feeds. | Study feed status, stale-data handling and map interactions. Replace public feeds with synthetic records shaped like your governed sources. |
| [lluisagusti/palantir-demo](https://github.com/lluisagusti/palantir-demo) · [global-watch](https://github.com/nk10nikhil/global-watch) | Dashboard demonstrations | Camera tiles, summaries, map layouts and fast ways to communicate an operational concept. | Use screenshots and disposable prototypes in design workshops. Verify licences before copying code, assets or data connectors. |
### Give the coding assistant a job, not a repository
A useful request is: “Study this incident triage screen and implement the same operator journey in our existing application using these API contracts, design tokens, permission checks and acceptance tests.” A poor request is: “Rewrite this clone in our language.” The first preserves an outcome and constraints; the second reproduces unknown assumptions.
**Takeaway:** prototype with synthetic data, test screens with real operators, and carry only validated interaction patterns into the production codebase.
11 · Real-world uses
## Optimise a decision, not a demo.
Each starter case has observable inputs, a human owner and a result that can be checked without an AI benchmark.
**Five deployed patterns with measured outcomes.** These are not promises for a new build: they show decisions worth instrumenting, the data that had to be joined and the denominator a pilot should reproduce.
NHS · named deployment
### Fill surgical theatres
Join waiting lists, clinical priority, staff rosters, theatre sessions and booking actions so teams can fill usable capacity. NHS England reports one trust increased theatre utilisation by 13.1%, treated 8% more patients and reduced cancellations by 29% after embedding the FDP workflow.
- [Read the NHS operating evidence →](https://www.england.nhs.uk/long-read/federated-data-platform-check-and-challenge-group-minutes-and-action-notes-18-october-2024/)
NHS · named deployment
### Unblock hospital discharge
Relate beds, patients, discharge criteria, transport, pharmacy and social-care dependencies; give each blocker an owner. NHS England says North Tees and Hartlepool reduced stays of 21 days or more by 36% while admitting 7.7% more patients after introducing its data-led approach.
- [Read the NHS rollout account →](https://www.england.nhs.uk/2023/11/new-nhs-software-to-improve-care-for-millions-of-patients/)
UPS · operator report
### Sequence delivery routes
Combine stops, service commitments, road constraints and driver knowledge; optimise a route, then let the driver handle field exceptions. UPS reported ORION cut six to eight miles from each deployed route in 2014 and projected 100 million fewer miles and 10 million gallons of fuel saved at full deployment.
- [Inspect UPS’s reported denominator →](https://www.ups.com/assets/resources/media/knowledge-center/UPS-2014-Corporate-Sustainability-Report.pdf)
Airbus · named airlines
### Schedule aircraft maintenance
Fuse sensor behavior, fault history, parts, maintenance windows and fleet plans; turn an early warning into a reviewed work order. Airbus reports Bangkok Airways’ on-time performance rose from 53% to 93% and LATAM reduced mechanical issues leading to delays from 24% to 15% while using Skywise workflows.
- [Review the named airline outcomes →](https://www.aircraft.airbus.com/en/newsroom/press-releases/2019-04-skywise-community-expands)
US DOT · multi-agency study
### Dispatch paratransit trips
Continuously assign booked trips to vehicles as cancellations, delays and new requests arrive, while preserving accessibility and pickup constraints. A US DOT summary of 11 transit agencies reports 8–31% productivity gains across six agencies with before-and-after data and an average 17% improvement in on-time performance.
- [Read the deployed transit study →](https://www.itskrs.its.dot.gov/2023-b01747)
**Takeaway:** choose the case with the cleanest identifiers and shortest feedback loop, not the most impressive map.
12 · From pilot to programme
## Sell a sovereign capability, not a “DIY Palantir”.
For executives, describe the owned decision capability, data boundary and first operational outcome. Do not promise a clone of a mature product suite.
Call the project a **sovereign operational intelligence platform** or an **owned decision-support capability** . The credible proposal is that a six-person team, assisted by coding tools, can build a narrow production slice around one mission or workflow—not reproduce every Foundry or Gotham feature.
| Seat | Primary accountability | First deliverable | Why it cannot be outsourced to AI |
| --- | --- | --- | --- |
| 1 · Operational product owner / analyst | Own the decision, vocabulary, users, operating constraints and acceptance threshold. | A decision map: trigger, evidence, permitted action, owner, deadline and escalation. | The model can organise interviews; it cannot decide which trade-off the organisation is accountable for. |
| 2 · Data / ontology engineer | Identifiers, source quality, entity resolution, lineage and the operational graph. | One reproducible raw-to-curated-to-graph path with data-quality tests. | Ambiguous records require domain decisions and named ownership, not plausible mappings. |
| 3 · Integration / backend engineer | Source adapters, APIs, workflow state and auditable actions. | A bounded service contract connecting one source to one reviewed action. | Generated code still needs transaction, retry, permission and failure semantics chosen for the real system. |
| 4 · Product / frontend engineer | Operator research, interaction design, accessibility and the decision UI. | A tested exception queue or operational view using synthetic data. | Fast UI generation increases the number of screens; only observation shows which one improves work. |
| 5 · Platform / security engineer | VMware or Kubernetes deployment, identity, secrets, network policy, observability and recovery. | A private deployment with tested deny paths, backup and isolated restore. | The organisation retains the risk when generated configuration exposes data or fails during recovery. |
| 6 · Analytics / quality engineer | Decision metrics, scenario fixtures, regression tests and production feedback. | A versioned set of known cases with baseline, latency, miss and operator-correction measures. | A model judging its own output is not independent evidence that the workflow works. |
### The bottleneck moved; it did not disappear
Our operating judgement is that the scarce role is the product owner who understands both the mission and the data. Outsourcing traditionally inserts translation between operators, analysts and developers; fast AI-assisted implementation can widen that gap by producing polished software before the decision rule is understood. Keep the operational owner embedded with the team and require weekly observation of real or replayed work.
**Takeaway:** fund six accountable roles around one decision, and measure the programme by operator outcomes and controlled data—not by screens shipped or code generated.
13 · Further reading
## Read the interfaces before the pitch decks.
Primary documentation is the useful procurement surface.
### Core references
- [World Monitor documentation index →](https://www.worldmonitor.app/docs/documentation)
- [World Monitor design and architecture →](https://www.worldmonitor.app/docs/architecture)
- [World Monitor licence boundaries →](https://www.worldmonitor.app/docs/license)
- [WorldView simulator article →](https://www.spatialintelligence.ai/p/i-built-a-spy-satellite-simulator)
- [IronSight reconstruction article →](https://www.spatialintelligence.ai/p/ironsight-turning-2d-videos-into)
- [God's Eye View release repository →](https://github.com/bilawalsidhu/gods-eye-view)
- [JanusGraph →](https://janusgraph.org/)
- [LocStat →](https://www.locstat.co.za/)
- [OpenFoundry (preserved implementation) →](https://github.com/cherishwins/OpenFoundry)
- [Early OSDK-compatible OpenFoundry →](https://github.com/Przyval/openfoundry)
- [Osiris repository →](https://github.com/simplifaisoul/osiris)
### Build routes
- [Operational-intelligence tools →](https://isaiuseful.com/tools.html.md#tools-build-operational-intelligence-systems)
- [Workflow guides →](https://isaiuseful.com/guides.html.md)
- [Local model routes →](https://isaiuseful.com/local-models.html.md)
- [Cloud and rented-GPU routes →](https://isaiuseful.com/cloud-models.html.md)
### Pick a measured case
- [Use-case catalogue →](https://isaiuseful.com/use-cases.html.md)
- [How this site grades claims →](https://isaiuseful.com/evidence.html.md)
**Takeaway:** pin versions and licences, test restores and deny paths, then document the one workflow your team actually operates.
The useful 80%
Boring data foundations. One operational decision.
- [Start the weekend build](#weekend)
- [Compare other paths](https://isaiuseful.com/guides.html.md#paths)
- [Build the operating model](https://isaiuseful.com/adoption.html.md#engineer)
---