# How Do You Build a Private Cloud for Open-Weight AI Models?

Canonical source: [https://isaiuseful.com/cloud-models](https://isaiuseful.com/cloud-models)

<a id="main-content"></a>

Provider build guide · checked 15 August 2026

Start with downloadable weights and a serving stack you can operate inside your own security boundary. Then buy for **memory, interconnect, power, cooling and measured token throughput** —not just peak FLOPS.

- [Choose models](#models)

- [Price servers](#hardware)

- [Estimate users + ROI](#economics)

**14 + 1**

open-weight lines + hosted preview
GLM‑5.3 ACCESS LIVE; WEIGHTS PENDING
**2 custom**

frontier licences need review
QWEN3.8 + K3 LARGE-SCALE MAAS TRIGGERS
**14–600 kW**

NVIDIA system planning range
8-GPU SERVER TO RUBIN ULTRA ROADMAP RACK

Two-minute guide

## Which private AI deployment route should you choose?

This page is collapsed into decision-sized sections. If the workload itself is not proven, start with [one measured adoption workflow](https://isaiuseful.com/adoption.html.md#walk) ; if the gap may need retrieval or changed weights, use the [model-intervention chooser](https://isaiuseful.com/training-models.html.md#chooser) before sizing infrastructure.

- [One private service **Start with a current model that fits one node.** *Nemotron 3.5 Lightning is the cleaner execution launch; benchmark your own workload.*](#nemotron-35-lightning)

- [Owned infrastructure **Buy memory, interconnect and a facility—not peak FLOPS.** *One air-cooled node is a different business from an NVL rack.*](#hardware)

- [Fastest operating route **Rent first when demand, model fit or utilization is uncertain.** *Use a managed API or dedicated capacity before carrying idle hardware.*](#cloud-routes)

- [Procurement gate **Replace every default with a quote and measured throughput.** *Proceed only when the privacy, capacity or support premium pays for ownership.*](#economics)

<a id="models"></a>

Recommendation catalogue

## Deploy now, validate next, or retain only as a baseline.

Cards are ordered by present recommendation, not chronology or parameter count. **Current launch** means a 2026 release with usable access and a credible deployment route. **Current option** marks a specialized, superseded or review-gated 2026 line. **Older baseline** marks a pre‑2026 release retained for compatibility or comparison—not as the default. GLM‑5.3 remains a hosted preview until its checkpoint and licence ship; Qwen3.8‑Max and Kimi K3 require their custom commercial terms to be reviewed.

current launch · preferred current release
current option · secondary or specialized
review gate or hosted preview · incomplete route
current self-host · newer hosted model exists
older baseline · retained, not default

<a id="minimax-m3"></a>

✓ Current launch
MiniMax Community License

### MiniMax M3

**A current launch choice for frontier coding and agents.** A native-multimodal 428B-total / 23B-active MoE with a 1M-token model limit and downloadable weights. The repositories total about 795.5 GiB for BF16, 413.3 GiB for MXFP8 and 232.9 GiB for NVIDIA's NVFP4 build. The calculator uses a measured 4× B200 FP4/MTP profile; keep the official 8× B200 recipe as the compatibility fallback and validate the exact four-GPU build before quoting.

**First node** — Calculator: 4× B200 measured cell; official fallback: 8× B200
**Serve with** — vLLM or SGLang; pin MSA, parser and quant paths
**Licence gate** — Attribution + notice; authorization above $20M product/service revenue

- [Official repository →](https://github.com/MiniMax-AI/MiniMax-M3/)

- [Weights, licence and deployment routes →](https://huggingface.co/MiniMaxAI/MiniMax-M3)

**Cloud first; DGX Station second**

The published large-node route is mature enough to trial: vLLM verified M3 on H200, GB200, B300 and AMD MI300/MI350 systems, and InferenceX now publishes a current 4× B200 FP4/MTP throughput sweep. The official NVFP4 recipe is newer and does not yet prove one-Station comfort. Rent the reference shape, lock an acceptance set, then repeat it on the quoted workstation.

- [NVIDIA NVFP4 build and current recipe →](https://huggingface.co/nvidia/MiniMax-M3-NVFP4)

- [InferenceX B200 throughput sweep →](https://inferencex.semianalysis.com/)

- [Assess the single-Station fit →](https://isaiuseful.com/dgx-station.html.md#minimax-m3)

✓ Current launch
Custom Qwen3.8‑Max licence

### Qwen3.8‑Max / 2.4T‑A95B

Qwen has released the self-hostable 2.4T-total / 95B-active MoE in full and block-scaled FP8 packages: about 4.89 TB and 2.50 TB of model files respectively. The open checkpoint is text-only, requires thinking mode, has a native 262,144-token context and can extend to 1,010,000. Do not treat it as a byte-for-byte copy of hosted Qwen3.8‑Max, which adds vision input, non-thinking mode, built-in tools and a default 1M context.

**First FP8 service cell** — 16 GPUs; SGLang verifies 4× GB300 NVL4 trays
**Serve with** — SGLang, vLLM or TokenSpeed release recipes
**Weight licence** — Custom terms—not Apache 2.0

- [Full checkpoint, model card and licence →](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B)

- [Official block-scaled FP8 checkpoint →](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8)

- [vLLM Qwen3.8 deployment recipe →](https://recipes.vllm.ai/Qwen/Qwen3.8-2.4T-A95B)

- [SGLang verified GB300 FP8 recipe →](https://docs.sglang.io/cookbook/autoregressive/Qwen/Qwen3.8#hw=gb300&variant=default&quant=fp8&strategy=balanced&nodes=multi-4)

- [TokenSpeed 16‑GPU recipes →](https://lightseek.org/tokenspeed/recipes/models#qwen3-8)

- [Compare the hosted Max product →](https://www.qwencloud.com/models/qwen3.8-max)

- [Video: Meet Qwen3.8-Max: A New Bar for Coding and Cowork.](https://www.youtube.com/watch?v=CKlK-KDFKjM)

- [Video: 6 days autonomous coding](https://www.youtube.com/watch?v=VG2OWqBk0Gk)

- [Video: Dynamic workflows to quant strategies](https://www.youtube.com/watch?v=oJgZivsyKO4)

- [Video: Visual agentic intelligence](https://www.youtube.com/watch?v=ByRx4xSm-tM)

- [Video: Cowork: Any role, build beyond](https://www.youtube.com/watch?v=XSWOREA1d9I)

<a id="nemotron-35-lightning"></a>

✓ Current launch
OpenMDW 1.1

### NVIDIA Nemotron 3.5 Lightning

A 30B-total / 3B-active hybrid Mamba-2, attention and MoE model for high-volume agent execution. The official NVFP4 weights total about 20.1 GiB before runtime and cache; NVIDIA documents one DGX Spark or one H100 as deployment paths, alongside Jetson, GeForce RTX 5090 and hosted/data-centre routes. Its one-million-token model limit is not a practical context promise for every device.

**First node** — 1× DGX Spark for the documented local recipe; 1× H100 for data-centre serving
**Serve with** — vLLM; broader ecosystem support still needs exact-version tests
**Licence fee** — €0 / $0; review OpenMDW 1.1 obligations

- [NVIDIA launch, local routes and vendor benchmarks →](https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/)

- [Official NVFP4 card and deployment recipe →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4)

- [BF16 reference weights and evaluation table →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16)

- [Design the local-to-frontier route →](https://isaiuseful.com/local-models.html.md#routing)

✓ Current launch
NVIDIA Open Model License

### NVIDIA Nemotron 3 Super

A fully open 120B total / 12B active hybrid MoE for multi-agent work, with weights, datasets, recipes and a 1M-token model limit. Its native NVFP4 build loads on one B200; a matched 8K/1K vLLM run provides a stronger starting profile than a generic model-size estimate.

**First node** — 1× B200 NVFP4; 2× H100/B200 for FP8
**Serve with** — vLLM, SGLang, TensorRT‑LLM or NIM
**Licence fee** — €0 / $0

- [NVIDIA release and deployment cookbooks →](https://developer.nvidia.com/blog/introducing-nemotron-3-super-an-open-hybrid-mamba-transformer-moe-for-agentic-reasoning/)

- [Review the measured B200 profile →](https://lambda.ai/inference-models/nvidia/nvidia-nemotron-3-super-120b-a12b)

- [Current OpenRouter benchmark price →](https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b/pricing)

<a id="nemotron-3-ultra"></a>

✓ Current launch
OpenMDW 1.1

### NVIDIA Nemotron 3 Ultra

NVIDIA’s 550B total / 55B active frontier orchestration model. The official mixed-precision NVFP4 checkpoint is about 352.3 GB and supports up to 1M context. Blackwell runs native W4A4; Hopper lacks native FP4 Tensor Cores and automatically uses a W4A16 fallback.

**First node** — 4–8 H200/B200-class GPUs
**Serve with** — vLLM or NVIDIA-optimized stack
**Licence fee** — €0 / $0

- [Official NVFP4 model card →](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4)

- [NVIDIA quantization and footprint notes →](https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/)

- [Review the four-GPU decode profile →](https://lambda.ai/inference-models/nvidia/nemotron-3-ultra)

- [Current OpenRouter benchmark price →](https://openrouter.ai/nvidia/nemotron-3-ultra-550b-a55b/pricing)

✓ Current launch
MIT

### DeepSeek V4 Flash 0731

The official Flash release supersedes the Preview checkpoint. Its 166.9 GB mixed FP4/FP8 repository includes the DSpark speculative module, supports 1M context and up to 384K output, and exposes low, high and max reasoning effort. DeepSeek reports large agentic-benchmark gains—including results above V4 Pro Preview—but its code-agent scores use a model-specific harness and two listed tests are internal, so validate the gain on your own stack.

**First node** — Official reference: 4× GB300; profile other layouts
**Serve with** — vLLM or SGLang with DSpark enabled
**Licence fee** — €0 / $0

- [Official model card and MIT licence →](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731)

- [Official API features and pricing →](https://api-docs.deepseek.com/quick_start/pricing)

- [vLLM deployment recipe →](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Flash?hardware=b300&features=tool_calling,reasoning)

- [SGLang deployment cookbook →](https://docs.sglang.io/cookbook/autoregressive/DeepSeek/DeepSeek-V4)

<a id="kimi-k3"></a>

✓ Current launch
Kimi K3 License

### Kimi K3

Moonshot’s released 2.8T-parameter multimodal MoE has 104B active parameters, selects 16 of 896 experts, supports a 1M context and uses MXFP4 weights with MXFP8 activations. Its 96-shard weight index totals 1.56 TB before runtime and cache headroom.

**First node** — 8× B300 or 8× MI350X/MI355X
**Serve with** — SGLang, vLLM or TokenSpeed; validate recipes
**Licence trigger** — Separate deal for large Model-as-a-Service operators

- [Official weights, model card and licence →](https://huggingface.co/moonshotai/Kimi-K3)

- [Official technical report and repository →](https://github.com/MoonshotAI/Kimi-K3)

- [SGLang hardware recipes and validation notes →](https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K3)

- [Current Kimi K3 market-price reference →](https://openrouter.ai/moonshotai)

- [Video: Meet Kimi K3](https://www.youtube.com/watch?v=bn0atstgavo)

△ Current option
Apache 2.0

### Qwen3.5 397B‑A17B

A still-current alternative now ranked behind Qwen3.8. Multimodal MoE with 397B total / 17B active parameters. The calculator now uses NVIDIA's roughly 251 GB NVFP4 repository and a measured 4× B300 SGLang/MTP profile at 8K input / 1K output, replacing the much larger 807 GB BF16 planning route.

**First service cell** — 4× B300 NVFP4 for the measured profile
**Serve with** — SGLang with the pinned NVFP4/MTP path
**Licence fee** — €0 / $0

- [Official model card →](https://huggingface.co/Qwen/Qwen3.5-397B-A17B)

- [NVIDIA NVFP4 build →](https://huggingface.co/nvidia/Qwen3.5-397B-A17B-NVFP4)

- [Inspect the measured InferenceX profile →](https://inferencex.semianalysis.com/)

<a id="glm-53"></a>

! Hosted preview
Licence pending

### GLM‑5.3

A post-trained successor that uses the same base model as GLM‑5.2. Z.ai reports large gains on long-horizon coding and agent benchmarks, but the launch status is split: Coding Plan access is live, the general API is marked coming soon and weights are targeted for release within two weeks of 14 August. Keep GLM‑5.2 for self-hosted planning until the 5.3 checkpoint, licence and engine recipes can be inspected.

**Available now** — GLM Coding Plan and ZCode
**API change** — Thinking required; low, high or max effort
**Self-hosting** — Wait for weights, licence and measured recipes

- [Official launch, score table and evaluation methods →](https://z.ai/blog/glm-5.3)

- [Official API status and migration requirements →](https://docs.z.ai/guides/llm/glm-5.3)

- [Coding Plan and coding-agent setup →](https://docs.z.ai/devpack/overview)

- [Compare 5.3 capability with 5.2 deployment →](https://isaiuseful.com/local-models.html.md#glm)

↻ Current self-host
MIT

### GLM‑5.2

A 753B model with up to 1M context. The BF16 repository is about 1.51 TB and needs at least 2 TB aggregate HBM with useful headroom; the calculator instead uses the released FP8 build and a matched 8× B200 profile.

**First node** — 8× B200 for FP8; ≥2 TB HBM for BF16
**Serve with** — vLLM or SGLang
**Licence fee** — €0 / $0

- [Official model card, files and licence →](https://huggingface.co/zai-org/GLM-5.2)

- [Review the measured B200 FP8 profile →](https://lambda.ai/inference-models/zai-org/glm-5.2)

△ Current fallback
Apache 2.0

### Mistral Small 4

A current compact fallback rather than the catalogue default; validate its vendor comparisons on your own tasks. 119B total / 6.5B active, multimodal, 256K context and 24-language support. Its official 70.8 GB NVFP4 build is a clean Blackwell fit; SGLang documents TP1 on B200/B300, while FP8 needs TP2 on H100/H200. Native W4A4 acceleration requires Blackwell.

**First node** — 1× B200 with NVFP4; benchmark before scaling
**Serve with** — SGLang or vLLM; validate the exact NVFP4 kernel
**Licence fee** — €0 / $0

- [Official model card →](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603)

- [Official NVFP4 build →](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603-NVFP4)

- [SGLang deployment topology →](https://docs.sglang.io/cookbook/autoregressive/Mistral/Mistral-Small-4)

! Current · review
Modified MIT

### Kimi K2.7 Code

A downloadable trillion-parameter coding and agent model with 32B active parameters, native INT4 and 256K context. The calculator uses a matched 8× B200 vLLM run; the modified terms, nightly runtime and model-specific kernels still need review.

**First node** — 8× B200 at native INT4
**Serve with** — vLLM, SGLang or KTransformers
**Licence fee** — €0 / $0

- [Official model card and modified licence →](https://huggingface.co/moonshotai/Kimi-K2.7-Code)

- [Review the measured B200 profile →](https://lambda.ai/inference-models/moonshotai/kimi-k2.7-code)

↓ Older · validate
MIT

### DeepSeek V3.2

A prior-year reasoning baseline retained for comparison and existing deployments. A 685B FP8 reasoning and agent model with a 690 GB repository. The card documents vLLM and SGLang; production tool parsing still needs your own robustness tests.

**First node** — 8× H200-class; pilot other stacks
**Serve with** — vLLM or SGLang
**Licence fee** — €0 / $0

- [Official model card and MIT licence →](https://huggingface.co/deepseek-ai/DeepSeek-V3.2)

- [Review the measured H200 sweep →](https://docs.gpustack.ai/2.0/performance-lab/deepseek-v3.2/h200/)

↓ Older baseline
Apache 2.0

### Mistral Large 3

A prior-year Mistral baseline, not a first recommendation against current releases. 675B total / 41B active, multimodal and 256K context. Mistral documents FP8 on one 8× H200 or B200 node and an NVFP4 checkpoint on one 8× H100 or A100 node. H100 and A100 do not have native FP4 Tensor Cores, so that older-GPU route depends on model-specific fallback or dequantization—not native NVFP4 W4A4 acceleration.

**First node** — 8× B200 for native NVFP4; 8× H200 for FP8
**Serve with** — vLLM; pin the hardware-specific path
**Licence fee** — €0 / $0

- [Official model card and deployment notes →](https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512-BF16)

↓ Older baseline
Apache 2.0

### gpt-oss-120b

Retained as an older compatibility and footprint baseline; start with current candidates unless this exact package wins your acceptance tests. About 117B total / 5.1B active, tool-capable and packaged in MXFP4 to fit one 80 GB H100. Minimum fit is not the same as an economical service profile: the calculator uses a measured 2× B200, 8K/1K TensorRT‑LLM run so packed capacity is grounded in observed throughput.

**First service cell** — 2× B200 for the measured capacity profile; H100 fits the build
**Serve with** — TensorRT‑LLM, vLLM or supported reference stack
**Licence fee** — €0 / $0

- [Official model card and licence →](https://huggingface.co/openai/gpt-oss-120b)

- [Inspect the measured InferenceX profile →](https://inferencex.semianalysis.com/)

**Commercially usable does not mean restriction-free.** Before listing any model, retain the exact licence and notice files, scan model code, document provenance limits, red-team the served build and define an abuse policy. Qwen3.8’s custom licence requires the notice to travel with copies; commercial products or services above 100 million monthly active users or $20 million monthly revenue must display the model name, while Model-as-a-Service or AI Work Assistant businesses above $50 million aggregate revenue in any consecutive 12 months need a separate Qwen licence. MiniMax M3 is also open-weight under custom terms, not Apache 2.0 or MIT: commercial deployments require visible attribution and a one-time notice, and products or services above the licence's $20M annual-revenue threshold require prior written authorization. Architecture, quantization, context and concurrency determine the real memory requirement.

<a id="hardware"></a>

NVIDIA platform ladder

## Start with platform scale—then see what it takes to become a provider.

The ten current and roadmap tiers keep node and rack choices legible. The provider-scale view then connects those purchases to research, open-model labs, frontier fleets and the EU and US public routes that can help bridge the gaps.

The ten-name map

### Five node tiers. Four NVL72 racks. One NVL144 rack.

Read the node class first, then the rack class. Open the card for the NVIDIA platform map, videos, NVFP4 guidance, prices, Rubin analysis and OEM alternatives.

H100 through GB300 are shipping systems; **Vera Rubin NVL72** is in its production ramp with preliminary published specifications. **V300** uses the later Rubin Ultra roadmap's 576 GB-per-GPU assumption; **VB300 NVL72** is a 72‑GPU comparison domain derived from half a Kyber rack, not an announced NVIDIA SKU. Kyber NVL144 remains the actual Rubin Ultra roadmap rack label.

8-GPU node · Hopper
**H100 → H200**

The mature entry tier. H200 keeps the same node shape and raises accelerator memory from 640 GB to 1.13 TB.

8-GPU node · newer
**B200 → B300 → V300**

More memory and newer low-precision engines; V300 is a derived roadmap planning node rather than an orderable SKU.

72-GPU NVLink rack
**GB200 → GB300 → Vera Rubin → VB300**

Vera Rubin is the current production-ramp generation. VB300 is a later, derived Rubin Ultra comparison—not the name of Vera Rubin NVL72.

144-GPU roadmap rack
**Kyber NVL144**

The Rubin Ultra/V300 roadmap tier changes rack density, power delivery and price class again.

From number format to facility

### See the platform NVIDIA is describing—then translate the spectacle into a bill of materials.

These vendor videos connect NVFP4, Vera, Rubin, DSX facilities, a scientific workload and the full GTC keynote. Use them to understand the intended system shape; use the tables below for procurement questions, facility constraints and explicit planning caveats.

- [Video: What Is NVFP4? Faster LLM Inference Without Losing Quality](https://www.youtube.com/watch?v=UTfg-_EGurw)

Low precision · NVIDIA Developer

### NVFP4 can make large models smaller—but the GPU matters

NVFP4 stores most values in four bits and uses fine-grained scaling to preserve more accuracy than a crude four-bit conversion. The catch is simple: native W4A4 acceleration starts with Blackwell Tensor Cores. Hopper can use selected checkpoints through a W4A16 fallback, but that is not the same speed path.

- [Video: NVIDIA Vera Rubin Platform Ramping into Full Production | Built for the Era of Agents](https://www.youtube.com/watch?v=jMZgjAVR7bo)

Platform overview · NVIDIA

### Vera Rubin is a full system, not just a GPU

Use the overview to see how NVIDIA frames CPUs, GPUs, networking and software as one agent platform. The delivered configuration still needs an exact vendor quote and acceptance test.

- [Video: NVIDIA Vera—The CPU for Agents](https://www.youtube.com/watch?v=vLfrBembjsk)

Host architecture · NVIDIA

### The CPU still shapes the service

Vera is NVIDIA's host-side story for agent systems. Translate that promise into memory bandwidth, data movement, storage and software requirements for the workload you will actually serve.

- [Video: NVIDIA DSX Powers Gigawatt‑Scale AI Factories at Maximum Efficiency](https://www.youtube.com/watch?v=cf40vNN5_Js)

Facility blueprint · NVIDIA

### At rack scale, the building joins the stack

DSX makes the facility-level ambition visible. Power delivery, cooling, networking, operations and recovery are part of the product long before a gigawatt becomes relevant.

- [Video: Advancing Scientific Discovery in the Agentic AI Era](https://www.youtube.com/watch?v=Il4dhCv0Li0)

Workload context · NVIDIA

### Start with the scientific job, not the rack

The discovery story is a useful demand-side counterweight to hardware spectacle. Define the models, data, latency and evaluation first; only then choose the infrastructure tier.

- [Video: NVIDIA GTC Keynote 2026](https://www.youtube.com/watch?v=jw_o0xr8MWU)

Full keynote · NVIDIA

### See the complete platform story in one sitting

The full GTC 2026 keynote connects Vera Rubin, agents, networking and AI factories in NVIDIA’s own long-form narrative. Use it for context, then return to the evidence tables for procurement decisions.

<a id="nvfp4"></a>

NVFP4 in plain English

### A promising format with one hard boundary: native acceleration needs newer NVIDIA GPUs.

NVFP4 is not a universal “make any model four times faster” switch. It reduces stored weights and memory traffic, but the checkpoint, serving engine, kernels, context and GPU generation must all agree.

Promising · hardware-gated

### Smaller weights can mean a larger model, more replicas or more cache.

NVFP4 groups 16 four-bit values under a higher-precision scale. NVIDIA reports about 4.5 bits per quantized value including block-scale overhead, roughly 3.5× less storage than FP16 and 1.8× less than FP8. Real checkpoints stay mixed precision: Nemotron 3 Ultra is 352.3 GB rather than the 309.4 GB that pure 4.5-bit math would suggest.

#### Read the labels this way

- **Native NVFP4:** Blackwell or Blackwell Ultra can execute W4A4 through FP4 Tensor Cores, subject to a supported kernel.
- **No native NVFP4:** Hopper H100/H200 lacks FP4 Tensor Cores. A model-specific W4A16 fallback may preserve weight-memory savings, but activations use 16-bit math.
- **Roadmap assumption:** the capacity arithmetic is useful, but support is not a shipping-product promise.

Same B300 core · different integration boundary

### HGX vs DGX: choose who integrates, licenses and supports the node.

**HGX B300** is NVIDIA’s eight-GPU baseboard and interconnect; an OEM turns it into a server with the CPU, RAM, storage, chassis, cooling, firmware, BMC and hardware-support route. **DGX B300** is NVIDIA’s integrated system baseline with DGX OS preinstalled. Compare the delivered configuration, entitlement certificate, facility fit and support path—not just the GPU label.

Why OEM HGX

#### Fit the node into the fleet you already operate.

OEMs can supply liquid cooling and rack density, preferred CPU, memory, storage and networking choices, and standard fleet controls such as iLO, iDRAC or XClarity where applicable. They can also offer regional procurement, warranty, spares and existing vendor-contract terms—sometimes at a lower or differently structured price. NVIDIA certification validates systems, but hardware support normally comes through the OEM or channel.

Why DGX

#### Buy a more tightly integrated NVIDIA baseline.

DGX B300 arrives with Ubuntu, NVSM, DCGM, driver/CUDA, Docker, Container Toolkit and DOCA-OFED/MST in its DGX OS stack; partner or NVIDIA field installation and entitlement registration are part of the supported path. That can simplify accountability, updates and escalation, while an OEM build needs more configuration and entitlement diligence.

| Boundary | DGX B300 | OEM HGX B200/B300 vs OEM GB200/GB300 NVL72 |
| --- | --- | --- |
| Base software | DGX OS stack is preinstalled with the system. | OEM image and fleet tooling vary; NVIDIA AI Enterprise can be supported, but enterprise support is an optional purchase. |
| NVIDIA AI Enterprise | Blackwell DGX licenses are purchased separately. Hopper DGX systems include NVIDIA AI Enterprise in the DGX software bundle; this does not carry into Blackwell. | Purchase and metric are separate unless the quote explicitly includes them. The five-year exceptions are only eligible H100 PCIe/NVL and H200 NVL GPUs in NVIDIA-Certified systems (A800 40 GB Active: three years), not generic HGX SXM or this Blackwell hardware; activation ties the term to the selected GPU. |
| Mission Control | Separately purchased or entitled unless the exact order says otherwise. It is recommended for DGX B200/B300 and required for GB200/GB300 NVL72 management. | The current Mission Control 2.3.x support matrix excludes OEM HGX B200/B300. Its OEM/partner support applies to listed GB200/GB300 NVL72 deployments; verify the exact release and configuration. |
| Mission Control footprint | For supported B300 deployments, the current requirements/support matrix lists ten separate control-plane nodes; include them in facility and TCO planning. | The same ten-node control-plane requirement is listed for supported GB-series deployments—not for the excluded OEM HGX B200/B300 path. |
| Support accountability | NVIDIA support requires the applicable active entitlement; exact inclusions depend on the order. | Hardware support is through OEM/channel; software, integration and escalation can be split across vendors. For an included eligible-GPU term, start is the OEM board ship date plus 90 days. |

- [NVIDIA Certified Systems validation →](https://docs.nvidia.com/certification-programs/latest/nvidia-certified-systems.html)

- [NVIDIA HGX AI Factory reference architecture →](https://docs.nvidia.com/enterprise-reference-architectures/hgx-ai-factory/latest/overview.html)

- [DGX B300 base stack and installation →](https://docs.nvidia.com/dgx/dgxb300-user-guide/introduction-to-dgxb300.html)

- [NVIDIA AI Enterprise licensing guide →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/licensing.html)

- [NVIDIA Mission Control requirements and support matrix →](https://docs.nvidia.com/nvidia-mission-control/)

Commercial platform path · checked 10 August 2026

### Hardware is only one layer of the NVIDIA commercial stack.

Open the stack to see which layers are software products, which are included system software and which require a separate commercial entitlement.

Order-line warning

### Blackwell DGX hardware alone does not grant a lifetime NVIDIA AI Enterprise or Mission Control license.

The exact order and entitlement control the software, support level and term. Selected H100 PCIe/NVL and H200 NVL GPUs include a five-year NVIDIA AI Enterprise subscription, but it is time-limited—not lifetime—and does not establish an included entitlement for HGX SXM or current Blackwell DGX/HGX. Hopper-generation DGX systems include NVIDIA AI Enterprise in the DGX software bundle; Blackwell DGX B200/B300 and GB200/GB300 systems require separate purchase unless the exact order or EC says otherwise. NVIDIA AI Enterprise is per GPU: subscriptions include support during their term; perpetual use is indefinite with five years of Business Standard support, renewable thereafter. Current list pricing is $4,500 per GPU for one year, $18,000 per GPU for a discounted five-year subscription, or $22,500 per GPU perpetual with five years of Business Standard support: for eight GPUs, that is list-price arithmetic of $36,000/year, $144,000/5 years, or $180,000 perpetual plus five-year support.

There is no public free-forever entitlement for the complete supported production suite. The general production trial is 90 days, includes Omniverse, excludes Run:ai and has support governed by its offer and Entitlement Certificate (EC); it is an evaluation route, not a production purchase.

- [NVIDIA AI Enterprise licensing exceptions and terms →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/licensing.html)

- [NVIDIA AI Enterprise pricing and trial routes →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/pricing.html)

> Visual: Layered architecture · links open NVIDIA documentation. One possible commercial stack—not a bundled bill of materials. Products can be selected independently, and support or entitlement depends on the exact order and supported configuration.

Layered architecture · links open NVIDIA documentation
**One possible commercial stack—not a bundled bill of materials.**

Products can be selected independently, and support or entitlement depends on the exact order and supported configuration.

Outcomes and workloads
- Applications
- Agents
- Model APIs
- Analytics
- Training

Application development

#### Build, customize and prepare workloads

Developer-facing tools; availability may depend on the chosen software entitlement.

- [**NIM** Model-serving APIs Packages model-serving APIs for deploying supported models.](https://docs.nvidia.com/nim/)

- [**NeMo Framework** Model development tools Builds, customizes and evaluates models and related workflows.](https://docs.nvidia.com/nemo-framework/user-guide/latest/overview.html)

- [**RAPIDS** Accelerated data science Accelerates GPU-backed data science and analytics workflows.](https://docs.rapids.ai/)

- [**Omniverse** Development and simulation platform Free for development, production and redistribution; community support is free, while enterprise support needs the applicable subscription.](https://docs.omniverse.nvidia.com/dev-guide/latest/common/NVIDIA_Omniverse_License_Agreement.html)

Inference serving

#### Optimize, expose and coordinate model responses

Serving components are separate choices; a workload does not need all of them.

- [**TensorRT-LLM** Optimized LLM inference Optimizes large-language-model inference on NVIDIA GPUs.](https://nvidia.github.io/TensorRT-LLM/)

- [**Triton Inference Server** Model-server endpoints Exposes model servers through production-oriented inference endpoints.](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/)

- [**Dynamo** Distributed inference coordination Coordinates distributed inference across GPUs and serving components.](https://docs.nvidia.com/dynamo/)

Infrastructure management

#### Schedule workloads and operate supported infrastructure

Entitlement-sensitive layers; confirm the configuration, term and support route.

- [**Run:ai** GPU scheduling and governance Schedules and governs shared GPU workloads.](https://docs.nvidia.com/run-ai/index.html)

- [**Mission Control** Supported AI-factory operations Operates supported AI-factory deployments; entitlement and support are configuration-specific.](https://docs.nvidia.com/nvidia-mission-control/)

- [**Base Command Manager (BCM)** Cluster provisioning Provisions and manages supported compute clusters.](https://docs.nvidia.com/base-command-manager/index.html)

- [**UFM** InfiniBand fabric management Manages InfiniBand fabrics used by supported systems.](https://networking-docs.nvidia.com/ufmenterpriseum/6221)

- [**NetQ** Network observability Observes and troubleshoots network state and faults.](https://docs.nvidia.com/networking-ethernet-software/cumulus-netq/)

System software

#### Operating system, GPU platform and container access

DGX OS is a system foundation; it is not itself a production-software entitlement.

- [**DGX OS** Supported DGX OS stack The supported operating-system stack for DGX systems.](https://docs.nvidia.com/dgx/dgx-os-7-user-guide/)

- [**DCGM** GPU health and telemetry Supplies GPU telemetry, health checks and diagnostics.](https://docs.nvidia.com/datacenter/dcgm/latest/index.html)

- [**CUDA** GPU programming platform Provides the GPU programming platform and libraries.](https://docs.nvidia.com/cuda/)

- [**NVIDIA Container Toolkit** GPU-enabled containers Makes NVIDIA GPUs available to container workloads.](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/)

System and hardware foundation
- DGX or certified systems
- NVIDIA GPUs
- Networking
- Storage
- Active support where purchased

Deploy anywhere
- Cloud
- Data center
- Edge
- Local workstation

- [Licensing, support and Blackwell DGX distinction →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/licensing.html)

- [Current NVIDIA AI Enterprise list pricing →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/pricing.html)

- [NVIDIA DGX platform purchase boundary →](https://www.nvidia.com/en-us/data-center/dgx-platform/)

Operator-assembled alternative

### An open-source control plane can lower license fees—but moves operations to you.

Open the diagram for a realistic surrounding stack, including the vendor dependencies that remain at the hardware boundary.

> Visual: Operator-assembled architecture · examples, not a required bill of materials. Open components, explicit operator ownership. Each layer is replaceable. The engineering team owns the integration, upgrades, recovery path and production support boundary.

Operator-assembled architecture · examples, not a required bill of materials
**Open components, explicit operator ownership.**

Each layer is replaceable. The engineering team owns the integration, upgrades, recovery path and production support boundary.

Outcomes and workloads
- Applications
- Agents
- Model APIs
- Analytics
- Training

Applications and agents

#### Compose the user-facing workload and its stateful agent paths

These tools supply interfaces and orchestration; the operator still owns identity, permissions and durable business state.

- [**LibreChat** Self-hosted model interface Provides a multi-user interface for local and hosted models, agents and MCP tools.](https://www.librechat.ai/docs)

- [**LangGraph** Stateful agent orchestration Builds explicit stateful agent graphs, checkpointing and human-in-the-loop paths.](https://langchain-ai.github.io/langgraph/)

Model serving and gateways

#### Expose models through an API contract you own

Choose a serving engine and gateway deliberately; the workload does not need every option.

- [**vLLM** GPU model serving Serves supported models through a high-throughput, OpenAI-compatible API.](https://docs.vllm.ai/en/stable/)

- [**SGLang** Distributed model serving Runs language and multimodal models from one GPU to distributed clusters.](https://docs.sglang.io/)

- [**LiteLLM Proxy** OpenAI-compatible gateway Routes, authenticates and meters requests across local and hosted endpoints.](https://docs.litellm.ai/)

Training, evaluation and lifecycle

#### Train, evaluate and promote reproducible model artifacts

Model code, adapters, datasets, evaluation and registry evidence are separate operating responsibilities.

- [**PyTorch** Training framework Provides the core tensor and training framework for model and adapter development.](https://pytorch.org/docs/stable/index.html)

- [**Transformers + PEFT + TRL** Models, adapters and training loops Supplies model implementations, parameter-efficient adapters and supervised or preference-training routes.](https://huggingface.co/docs/transformers/index)

- [**MLflow** Experiment and artifact tracking Tracks run parameters, metrics and artifacts for comparison and promotion evidence.](https://mlflow.org/docs/latest/)

Workload control

#### Pick one primary scheduling and queueing model

Combining schedulers without a clear owner creates competing admission and recovery rules.

- [**Kubernetes** Container orchestration Schedules GPU-enabled container workloads through device plugins and resource requests.](https://kubernetes.io/docs/tasks/manage-gpus/scheduling-gpus/)

- [**Kueue** Kubernetes job queueing Adds quota-aware admission and queueing for batch workloads on Kubernetes.](https://kueue.sigs.k8s.io/docs/)

- [**Volcano** Batch scheduling Adds batch and high-performance workload scheduling to Kubernetes.](https://volcano.sh/docs/home/introduction/)

- [**Slurm** HPC workload manager Allocates accelerators and other generic resources for queued compute jobs.](https://slurm.schedmd.com/gres.html)

- [**NVIDIA GPU Operator + KAI Scheduler** GPU enablement and gang scheduling Installs and manages Kubernetes GPU software while KAI Scheduler adds AI workload and gang scheduling.](https://docs.nvidia.com/gpu-operator/latest/)

Provisioning and configuration

#### Rebuild nodes from documented state

Bare-metal lifecycle and desired configuration are separate responsibilities.

- [**MAAS** Bare-metal provisioning Discovers, commissions and provisions physical machines.](https://maas.io/docs)

- [**Foreman** Host lifecycle management Provisions and manages physical and virtual host lifecycles.](https://docs.theforeman.org/)

- [**Ansible** Configuration automation Applies repeatable configuration and operational automation across nodes.](https://docs.ansible.com/)

Delivery and supply chain

#### Build, promote and retain deployable artifacts

Workflow automation and a private registry still need access control, signing policy and retention rules.

- [**Argo Workflows** Kubernetes-native workflows Runs DAG and step-based workflows as Kubernetes resources.](https://argoproj.github.io/argo-workflows/)

- [**Harbor** Artifact registry Provides a private registry with project-level distribution and retention controls.](https://goharbor.io/docs/)

Observability

#### Own telemetry, dashboards, logs and alerts

Keep sensitive prompts and identifiers out of labels and unbounded log streams.

- [**OpenTelemetry** Telemetry instrumentation Standardizes collection and export of traces, metrics and logs.](https://opentelemetry.io/docs/)

- [**Prometheus** Metrics and alerting Scrapes time-series metrics and evaluates alerting rules.](https://prometheus.io/docs/introduction/overview/)

- [**Grafana** Dashboards and exploration Visualizes and explores operational data from configured sources.](https://grafana.com/docs/grafana/latest/)

- [**Loki** Log aggregation Indexes labels around log streams for storage and investigation.](https://grafana.com/docs/loki/latest/)

Data, network and security

#### Assemble infrastructure controls as separate services

Storage, policy and secrets still need backups, upgrades and tested recovery.

- [**Ceph** Distributed storage Provides object, block and file storage across a managed cluster.](https://docs.ceph.com/en/latest/)

- [**MinIO / AIStor** S3-compatible object storage The current vendor documentation is for AIStor; verify product terms and the archived OSS route.](https://docs.min.io/aistor/)

- [**Cilium** Networking and policy Provides eBPF-based networking, policy and observability for cloud-native workloads.](https://docs.cilium.io/en/stable/)

- [**Vault** Secrets management Controls access to secrets and dynamic credentials under HashiCorp's current terms.](https://developer.hashicorp.com/vault/docs)

Hardware and vendor boundary

#### Open control software does not make the accelerator stack open

Pin and validate the OEM, driver, CUDA, collectives and firmware compatibility chain.

- [**Redfish** Hardware management API Standardizes server-management APIs; OEM implementations and extensions still vary.](https://www.dmtf.org/standards/redfish)

- [**NVIDIA Driver** GPU host driver Connects the operating system and supported NVIDIA accelerator hardware.](https://docs.nvidia.com/datacenter/tesla/driver-installation-guide/)

- [**CUDA** GPU programming platform Provides the NVIDIA GPU programming platform and libraries.](https://docs.nvidia.com/cuda/)

- [**NCCL** Collective communication Coordinates collective communication across supported GPUs and networks.](https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/)

- [**Platform firmware** OEM-specific lifecycle Firmware updates remain platform-specific and inside the vendor support boundary.](https://docs.nvidia.com/dgx/dgxb300-fw-update-guide/)

**Ownership trade:** this route can avoid NVIDIA AI Enterprise and Mission Control fees, but it transfers integration, upgrade validation, incident response, security patching, recovery automation and support ownership to the operator. It is not a claim of feature parity with Mission Control.

- [Find every component in the tools catalogue →](https://isaiuseful.com/tools.html.md#tools-ai-infrastructure)

- [NVIDIA HGX AI Factory boundary →](https://docs.nvidia.com/enterprise-reference-architectures/hgx-ai-factory/latest/overview.html)

- [NVIDIA Certified Systems validation →](https://docs.nvidia.com/certification-programs/latest/nvidia-certified-systems.html)

Vendor explainers · not independent evidence

### PNY Pro explains the commercial software story; verify entitlements against the order.

These vendor videos are useful orientation for NVIDIA AI Enterprise, not proof that a particular DGX or OEM quote includes a license, support term or Mission Control.

- [Video: Introducing NVIDIA AI Enterprise](https://www.youtube.com/watch?v=duemyiNxl4M)

Vendor explainer · PNY Pro

### Introducing NVIDIA AI Enterprise

Use PNY Pro's overview to understand the vendor's platform framing, then confirm the license metric, term and support on the order.

- [Video: NVIDIA AI Enterprise | End-to-End Platform for Production AI](https://www.youtube.com/watch?v=gJ5OvKFStIs)

Vendor explainer · PNY Pro

### Production AI is an entitlement question

The explainer describes the commercial platform. It does not establish a paid production entitlement or its duration for any hardware purchase.

Strong recommendation

### Host the gear in a datacenter.

A 120–600 kW liquid-cooled rack is a facility project before it is a model project. Shortlist colocation providers that already support high-density direct-liquid cooling, redundant power, coolant distribution, carrier-neutral networking, remote hands, physical security and the client’s compliance regime. Get written facility acceptance for the exact NVIDIA/OEM rack before signing the hardware order.

- [NVIDIA provider facility requirements →](https://docs.nvidia.com/dsx/ncp/inference-provider-requirements/home)

- [NVIDIA NVL72 reference architecture →](https://docs.nvidia.com/enterprise-reference-architectures/nvl72-ai-factory/latest/overview.html)

| Platform family | Scale + accelerator memory | Facility envelope | Best planning use | Ballpark, ex VAT |
| --- | --- | --- | --- | --- |
| H100 Available 8‑GPU Hopper node   <br> [Official H100/H200 system specs →](https://docs.nvidia.com/dgx/dgxh100-user-guide/introduction-to-dgxh100.html) | 8× H100 · 640 GB HBM3 ✕ No native NVFP4   <br> Speculative fit: ≈0.74–0.82T parameters through a supported W4A16-style fallback. | **10.2 kW max · air** 8U rack server; ordinary high-density datacenter deployment. | Mature, lower-capex CUDA node for inference, adapters and evaluation. | **≈€263–403k / $300–460k** Broad 2026 complete-system quote band. |
| H200 Available 8‑GPU Hopper node   <br> [Official H100/H200 system specs →](https://docs.nvidia.com/dgx/dgxh100-user-guide/introduction-to-dgxh100.html) | 8× H200 · 1.13 TB HBM3e ✕ No native NVFP4   <br> Speculative fit: ≈1.30–1.45T parameters through a supported W4A16-style fallback. | **10.2 kW max · air** Same DGX node envelope as H100. | Memory-first Hopper choice for larger checkpoints without Blackwell migration. | **≈€350–438k / $400–500k** Broad 2026 complete-system quote band. |
| B200 Available 8‑GPU Blackwell node   <br> [Official B200 system specs →](https://docs.nvidia.com/dgx/dgxb200-user-guide/introduction-to-dgxb200.html) | 8× B200 · 1.44 TB HBM3e ✓ Native NVFP4   <br> Speculative single-checkpoint fit: ≈1.65–1.85T parameters. | **14.3 kW max · air** 1,550 CFM and 48,794 BTU/hr at the system ceiling. | General Blackwell serving and post-training node. | **€438–569k / $500–650k** Public-reseller and integrator planning range. |
| B300 Available 8‑GPU Blackwell Ultra node   <br> [Official B300 system specs →](https://docs.nvidia.com/dgx/dgxb300-user-guide/introduction-to-dgxb300.html) | 8× B300 · 2.3 TB HBM3e ✓ Native NVFP4   <br> Speculative single-checkpoint fit: ≈2.65–2.95T parameters. | **14.5 kW max · air** 49,476 BTU/hr published system ceiling. | Largest current single-node memory tier before NVL72. | **≈€569–744k / $650–850k** Editable 2026 planning band; require an exact OEM quote. |
| V300 Derived 8‑GPU Rubin Ultra roadmap node   <br> [Inspect the V300/Kyber roadmap estimate →](https://wccftech.com/nvidia-rubin-ultra-rack-estimated-to-cost-21-million-hbm4e-swelling-to-1-5m-per-unit/) | 8× V300 · 4.6 TB HBM4e ◇ Roadmap assumption   <br> ≈5.3–5.9T parameters if this derived node retains NVFP4-class support. | **≈33 kW · full DLC planning** One eighteenth of the Kyber power target; not a vendor system specification. | Roadmap node comparison before deciding whether NVL72 scale is justified. | **≈€0.9–1.3M / $1.0–1.5M** Derived planning band, not a list price or announced node. |
| GB200 NVL72 Available 72‑GPU Blackwell rack   <br> [Official GB200 NVL72 specs →](https://www.nvidia.com/en-us/data-center/gb200-nvl72/) | 72× Blackwell · 13.4 TB HBM3e ✓ Native NVFP4   <br> ≈15.5–17.2T parameter-equivalent capacity; real services normally use replicas. | **≈120 kW · direct liquid** Residual air remains for networking and storage. | First rack-scale tier for large distributed inference and training. | **€2.45–2.98M / $2.8–3.4M** Reported 2026 purchase-quote range. |
| GB300 NVL72 Available 72‑GPU Blackwell Ultra rack   <br> [Official NVL72 component design →](https://docs.nvidia.com/enterprise-reference-architectures/nvl72-ai-factory/latest/components.html) | 72× B300 · 20 TB HBM3e ✓ Native NVFP4   <br> ≈23–25.6T parameter-equivalent capacity; usually spent on replicas, cache and test-time compute. | **Up to 142 kW · full DLC** Facility CDU and rack leak-detection integration required. | High-memory NVL72 reasoning, inference and post-training. | **€5.25–5.69M / $6.0–6.5M** Reported purchase quotes; lower figures are often component/BOM estimates. |
| Vera Rubin NVL72 2026 production ramp · preliminary specs   <br> [Official Vera Rubin NVL72 specs →](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/) | 72× Rubin + 36× Vera · 20.7 TB HBM4 ✓ Native NVFP4 + 3-bit LUT path   <br> 1,580 TB/s aggregate HBM bandwidth and 260 TB/s NVLink 6 switch bandwidth. | **≈240 kW comparison input · full DLC** 72 × 3.337 kW all-in per-GPU power used by the SemiAnalysis model; not a published rack nameplate. | High-interactivity, long-context and multitrillion-parameter inference where decode bandwidth and scale-up latency dominate. | **Quote required** SemiAnalysis owner-TCO assumption: $3.57/GPU-hour, not a rental rate or purchase quote. |
| VB300 NVL72 Derived 72‑GPU V300 comparison domain   <br> [Inspect the V300/Kyber roadmap estimate →](https://wccftech.com/nvidia-rubin-ultra-rack-estimated-to-cost-21-million-hbm4e-swelling-to-1-5m-per-unit/) | 72× V300 · 41.5 TB HBM4e ◇ Roadmap assumption   <br> ≈48–53T parameter-equivalent if the derived rack retains NVFP4-class support. | **≈300 kW · 800VDC + full DLC** Half-Kyber planning envelope; not an announced NVIDIA rack. | Like-for-like NVL72 comparison before the jump to Kyber NVL144. | **≈€8.75–10.50M / $10–12M** Half-Kyber planning band with contingency; not a quote. |
| Kyber NVL144 Rubin Ultra/V300 analyst roadmap   <br> [NVIDIA 800VDC rack-power direction →](https://developer.nvidia.com/blog/?p=100571) | 144× V300 · 82.9 TB HBM4e ◇ Roadmap assumption   <br> ≈96–106T parameter-equivalent if the roadmap system retains NVFP4-class support. | **≈600 kW · 800VDC + full DLC** Megawatt-class facility design around the compute rack. | Frontier scale where power architecture becomes a first-order design. | **≈€18.38M / $21M** Bank of America analyst ASP estimate reported July 2026. |

**These are ballparks, not list prices or model-fit promises.** NVFP4 capacity bands reserve 20–25% of advertised accelerator memory and assume roughly 5.0–5.2 effective bits per stored parameter, reflecting mixed-precision checkpoints rather than pure four-bit arithmetic. They exclude unusually large embeddings, multimodal towers and cache. H100 through GB300 prices use broad complete-system or reported rack quote bands; V300, VB300 and Kyber use explicitly labelled roadmap arithmetic. Actual draw follows utilization and power caps; network, storage, CDU/pumps, PUE overhead, spares, deployment and support are additional. Every price excludes VAT.

- [NVIDIA NVFP4 format and Blackwell support →](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)

- [Nemotron mixed-precision checkpoint evidence →](https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/)

- [2026 H100/H200 system quote bands →](https://intuitionlabs.ai/articles/data-center-gpu-pricing-2026)

- [Reported GB200 and GB300 purchase quotes →](https://www.tomshardware.com/tech-industry/artificial-intelligence/price-of-nvidias-vera-rubin-nvl72-racks-skyrockets-to-as-much-as-usd8-8-million-apiece-but-server-makers-margins-will-be-tight-nvidia-is-moving-closer-to-shipping-entire-full-scale-systems)

- [Morgan Stanley GB300 component estimate →](https://finance.yahoo.com/sectors/technology/articles/morgan-stanley-estimate-says-single-114101104.html)

- [V300 and Kyber roadmap estimate →](https://wccftech.com/nvidia-rubin-ultra-rack-estimated-to-cost-21-million-hbm4e-swelling-to-1-5m-per-unit/)

Early silicon · evidence checked 28 July 2026

### Vera Rubin is clearly ahead—but “10×” describes one old baseline at one speed.

CoreWeave measured a pre-production Vera Rubin NVL72 rack; SemiAnalysis then normalized it against July 2026 GB200 and GB300 InferenceX results. The useful question is not whether Rubin wins, but which Blackwell software vintage, user interactivity and cost boundary you compare.

100 tok/s/user
**2.00× vs GB200**

Rubin delivers 1,115,625 output tok/s/MW versus 558,569 on the July 2026 GB200 recipe; the corresponding GB300 ratio is 2.15×.

150 tok/s/user
**10.18× vs 2025 GB200**

This is the headline region. Against current July 2026 recipes, the lead is 2.74× over GB200 and 2.58× over GB300—not tenfold.

300 tok/s/user
**5.39× vs GB300**

Rubin delivers 96,446 output tok/s/MW. GB300 reaches 17,892 at its last viable frontier point; GB200 cannot reach this target.

Owner TCO · 300 tok/s/user
**5.03× cheaper**

The model produces $3.076 per million Rubin output tokens versus $15.456 for GB300. This is modeled ownership cost—not cloud rental pricing.

! Early external result

#### What was actually tested

DeepSeek R1 0528 671B at FP4, single-turn 8K input / 1K output. CoreWeave says both sides used NVFP4, speculative decoding with MTP, disaggregated prefill/decode with Dynamo, wide expert parallelism and TensorRT‑LLM. “Interactivity” is output tokens per second per user.

#### How the comparison was normalized

Throughput counts output tokens only, while power and cost include all prefill and decode GPUs. SemiAnalysis uses all-in power inputs of 2.168 kW/GPU for GB200, 2.553 for GB300 and 3.337 for Rubin, with the same PUE across the direct-liquid-cooled systems.

#### What remains unproven

The Rubin result came from a Dell engineering-sample rack without scale-out fabric and was supplied by CoreWeave; SemiAnalysis says it did not independently verify it. The workload is one older model and one short, single-turn recipe—not a multi-turn coding-agent trace.

Memory system
**20.7 TB HBM4 · 1,580 TB/s**

Each Rubin GPU carries up to 288 GB and 22 TB/s. Capacity helps model and KV-cache residency; bandwidth attacks the token-by-token decode bottleneck.

Scale-up fabric
**260 TB/s NVLink 6**

The 72-GPU rack provides 3.6 TB/s per GPU of all-to-all bandwidth. Counted writes reduce synchronization traffic for device-initiated communication.

Tensor Core path
**2× work per clock**

Rubin doubles the K dimension handled by its matrix instruction. Fewer K-loop iterations reduce overhead in throughput-, memory- and latency-bound kernels.

MoE data movement
**Inline TMA overrides**

A shared tensor descriptor can override addresses and strides in the instruction, avoiding an in-memory descriptor rewrite when switching experts with the same layout.

Weight compression
**3.125 bits/weight raw**

The new 3-bit LUT-B path stores an index plus an eight-entry E4M3 codebook per 512-weight block and resolves it inside the matrix operation. Accuracy depends on fitting and calibration.

Software transition
**SM100 kernels can start on SM107**

Rubin can reuse important Blackwell-family kernels for bring-up, unlike the Hopper-to-Blackwell break. Peak performance still needs Rubin-specific tuning; CUDA 13.4 support is a developer preview.

Planning arithmetic

#### Kimi K3 shows why 3-bit LUT support matters

For a 2.8T-parameter model, a 4.25-bit raw payload is about 1.49 TB; 3.125 bits is about 1.09 TB, roughly 394 GB less. At 288 GB HBM4 per Rubin GPU, raw weights alone move from about six packages to four. This excludes KV cache, activations, runtime buffers, sensitive higher-precision layers and parallelism replication—and it is not a released Kimi K3 LUT checkpoint.

**Do not bank the saving yet.**

One codebook serves 512 weights, whereas NVFP4 adapts a scale over groups of 16. The programmable codebook may preserve more useful values, but calibration, model quality and production kernels decide whether the paper capacity saving becomes a service gain.

> Visual: Reconstructed in HTML from SemiAnalysis data. Output throughput per all-in utility MW. Bars are scaled within each interactivity target; longer is better. Labels retain the exact output-token rate.

Reconstructed in HTML from SemiAnalysis data

#### Output throughput per all-in utility MW

Bars are scaled within each interactivity target; longer is better. Labels retain the exact output-token rate.

Vera Rubin · Jul 2026
GB300 · Jul 2026
GB200 · Jul 2026
CoreWeave GB200 · 2025

##### 100 tok/s/user

Vera Rubin
1,115,625
**1,115,625**

GB300
517,848
**517,848**

GB200
558,569
**558,569**

GB200 · 2025
300,000
**300,000**

##### 150 tok/s/user

Vera Rubin
786,208
**786,208**

GB300
305,023
**305,023**

GB200
287,289
**287,289**

GB200 · 2025
77,224
**77,224**

##### 200 tok/s/user

Vera Rubin
300,000
**300,000**

GB300
66,799
**66,799**

GB200
77,218
**77,218**

GB200 · 2025
50,000
**50,000**

##### 250 tok/s/user

Vera Rubin
140,264
**140,264**

GB300
37,049
**37,049**

GB200
36,394
**36,394**

GB200 · 2025
39,123
**39,123**

##### 300 tok/s/user

Vera Rubin
96,446
**96,446**

GB300
17,892
**17,892**

GB200
*Not feasible*

GB200 · 2025
*Not feasible*

> Visual: Reconstructed in HTML from SemiAnalysis TCO data. Modeled owner cost per million output tokens. Bars are scaled within each target; shorter is better. Values include prefill and decode GPUs under the source's ownership assumptions.

Reconstructed in HTML from SemiAnalysis TCO data

#### Modeled owner cost per million output tokens

Bars are scaled within each target; shorter is better. Values include prefill and decode GPUs under the source's ownership assumptions.

Vera Rubin · Jul 2026
GB300 · Jul 2026
GB200 · Jul 2026
CoreWeave GB200 · 2025

##### 100 tok/s/user

Vera Rubin
$0.266
**$0.266**

GB300
$0.497
**$0.497**

GB200
$0.423
**$0.423**

GB200 · 2025
$0.786
**$0.786**

##### 150 tok/s/user

Vera Rubin
$0.378
**$0.378**

GB300
$0.842
**$0.842**

GB200
$0.799
**$0.799**

GB200 · 2025
$3.046
**$3.046**

##### 200 tok/s/user

Vera Rubin
$0.991
**$0.991**

GB300
$3.837
**$3.837**

GB200
$3.016
**$3.016**

GB200 · 2025
$4.715
**$4.715**

##### 250 tok/s/user

Vera Rubin
$2.099
**$2.099**

GB300
$6.859
**$6.859**

GB200
$6.471
**$6.471**

GB200 · 2025
$6.013
**$6.013**

##### 300 tok/s/user

Vera Rubin
$3.076
**$3.076**

GB300
$15.456
**$15.456**

GB200
*Not feasible*

GB200 · 2025
*Not feasible*

Inspect the complete reconstructed data tables
**Output tokens per second per all-in utility MW**

| Recipe | 50 tok/s | 100 tok/s | 150 tok/s | 200 tok/s | 250 tok/s | 300 tok/s | 350 tok/s |
| --- | --- | --- | --- | --- | --- | --- | --- |
| GB200 NVL72 · Jul 2026 | 705,543 | 558,569 | 287,289 | 77,218 | 36,394 | Not feasible | Not feasible |
| GB300 NVL72 · Jul 2026 | 707,037 | 517,848 | 305,023 | 66,799 | 37,049 | 17,892 | Not feasible |
| CoreWeave GB200 · 2025 | 500,000 | 300,000 | 77,224 | 50,000 | 39,123 | Not feasible | Not feasible |
| CoreWeave Vera Rubin · Jul 2026 | 1,330,000 | 1,115,625 | 786,208 | 300,000 | 140,264 | 96,446 | 70,703 |

**Modeled owner cost per million output tokens**

| Recipe | 50 tok/s | 100 tok/s | 150 tok/s | 200 tok/s | 250 tok/s | 300 tok/s | 350 tok/s |
| --- | --- | --- | --- | --- | --- | --- | --- |
| GB200 NVL72 · Jul 2026 | $0.333 | $0.423 | $0.799 | $3.016 | $6.471 | Not feasible | Not feasible |
| GB300 NVL72 · Jul 2026 | $0.364 | $0.497 | $0.842 | $3.837 | $6.859 | $15.456 | Not feasible |
| CoreWeave GB200 · 2025 | $0.472 | $0.786 | $3.046 | $4.715 | $6.013 | Not feasible | Not feasible |
| CoreWeave Vera Rubin · Jul 2026 | $0.223 | $0.266 | $0.378 | $0.991 | $2.099 | $3.076 | $4.179 |

**Decision rule:** use the July 2026 GB200 and GB300 recipes for a current buy/no-buy comparison; keep the 2025 GB200 line only as a software-maturity lesson. Re-run the exact model, context distribution, concurrency, service-level target and power boundary before procurement. The charts above are original HTML/CSS reconstructions of the published values, with unavailable frontier points shown as “not feasible.”

- [SemiAnalysis inference TCO and architecture analysis →](https://newsletter.semianalysis.com/p/vera-rubin-nvl72-vs-gb200-nvl72-inference)

- [CoreWeave original engineering-sample result →](https://www.coreweave.com/blog/nvidia-vera-rubin-nvl72-on-coreweave-10x-more-tokens-per-megawatt-than-blackwell)

- [NVIDIA Vera Rubin NVL72 preliminary specs →](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/)

- [NVIDIA Rubin architecture detail →](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/)

- [CUDA 13.4 developer-preview notes →](https://docs.nvidia.com/cuda/developer-preview/13.4/cuda-toolkit-release-notes/index.html)

- [InferenceX methods and results →](https://github.com/SemiAnalysisAI/InferenceX)

OEM alternatives

### AMD or Intel host CPUs can still feed native-CUDA NVIDIA GPU servers.

The CPU vendor affects host-memory bandwidth, data preparation, storage and operational standardization. CUDA compatibility follows the NVIDIA accelerator, so all four configurations below remain native CUDA systems.

| Orderable OEM server | Host + accelerator | Power + cooling | Good first use | Ballpark, ex VAT |
| --- | --- | --- | --- | --- |
| Dell PowerEdge XE9685L [Official Dell configuration →](https://www.dell.com/support/manuals/en-us/poweredge-xe9685l/pexe9685l_ism_pub/system-overview?guid=guid-997e27f8-549c-42e4-8d17-41c79b40b0c5&lang=en-us) | 2× AMD EPYC 9005 + 8× B200 · 1.44 TB HBM ✓ Native CUDA + NVFP4 | **6× 3 kW PSU capacity** 4U direct liquid cooling; reserve residual air. | B200 service node with an AMD-standardized CPU fleet. | **€438–569k / $500–650k** Accelerator-led planning range. |
| HPE ProLiant Compute XD685 [Official HPE configurations →](https://www.hpe.com/us/en/collaterals/collateral.a00073553enw.html) | AMD EPYC + 8× H200, B200 or B300 H200: ✕ no native NVFP4   <br> B200/B300: ✓ native | **Up to 12 high-watt PSUs** Direct-liquid-cooled B200/B300 configurations; quote actual max draw. | A supported AMD-host path from H200 through Blackwell Ultra. | **€438–744k / $500–850k** Wide range spans GPU generation and support. |
| Dell PowerEdge XE9680L [Official Dell configuration →](https://www.dell.com/support/manuals/en-us/poweredge-xe9680l/pexe9680l_ism_pub/system-overview?guid=guid-eef3f3fc-57c9-4939-958a-7128c53241b4) | 2× 5th-gen Intel Xeon + 8× H200 or B200 H200: ✕ no native NVFP4   <br> B200: ✓ native | **6× 3 kW PSU capacity** 4U direct liquid cooling; reserve residual air. | B200 service node for Intel-standardized estates. | **€438–569k / $500–650k** Accelerator-led planning range. |
| Lenovo ThinkSystem SR680a V3 [Official Lenovo B200 guide →](https://lenovopress.lenovo.com/lp2247-thinksystem-sr680a-v3-with-b200) | 2× 5th-gen Intel Xeon + 8× B200 · 1.44 TB HBM ✓ Native CUDA + NVFP4 | **3.2 kW PSU modules** 8U air-cooled design; validate rack airflow and redundancy. | Air-cooled B200 where liquid plumbing is unavailable. | **€438–569k / $500–650k** Accelerator-led planning range. |

Capacity, training and market scale

### What each hardware tier can actually do—and how it becomes a model provider.

A server is a useful product unit. A model lab is a fleet, software organization, power contract and customer pipeline. Open the card to connect hardware ownership with the organizations and services it can support.

> Visual: AI model organization scale ladder

**Visual reading order:**
1. 01 · 1–8 accelerators Research project **One workstation or node** Evaluate, fine-tune and serve existing weights; pretrain compact models; prove a bounded workflow. A failed experiment costs hours or days, not a datacenter programme. **People on the work** ≈2–30 direct **Value / market cap** N/A or <$50m org
2. 02 · 1k–20k+ H100e Independent model lab **Mistral · DeepSeek · Kimi class** Pretrain large models, run ablations and post-training, then operate an API. Mistral disclosed a 3,000-H200 run; DeepSeek’s reported H100/H800 pool is around 20,000. Moonshot’s fleet is not disclosed. **People on the work** ≈200–2,000 direct **Value / market cap** ≈$12–75bn lab value
3. 03 · 600k–2m H100e Frontier lab **xAI · Anthropic · OpenAI** Run several training and post-training programmes, generate synthetic data, absorb failed runs and serve global products. Capacity is usually spread across clouds and campuses. **People on the work** ≈2,000–10,000 direct **Value / market cap** ≈$0.85–1.0tn lab value
4. 04 · 2m–5m+ parent pool Hyperscaler-backed frontier **Amazon · Microsoft · Meta · Google · Alibaba** The parent pool feeds cloud tenants, first-party models, recommenders and other AI. Alibaba’s total is not separately disclosed; the whole pool is never one model run. **People on the work** ≈10k–100k+ cross-stack **Value / market cap** ≈$0.29–3.9tn parent

| Owned shape | What it can actually do | Service reality | Organization it supports |
| --- | --- | --- | --- |
| 1–8 accelerators Workstation to one node | Evaluation, LoRA/Q‑LoRA, small-model pretraining and one production replica of a model that fits. | **Pilot to ≈1,200 heavy users** Model and node dependent; benchmark the exact build. | Research group, internal platform team or specialist service with cloud overflow. |
| 72–144 linked accelerators One NVL72/NVL144-class domain | Large distributed inference, multiple replicas, long-context cache, RL and substantial post-training. | **≈2,500–80,000+ heavy users** A planning range, not a vendor benchmark. | Regional model cloud or a major enterprise AI platform; the facility is now part of the product. |
| 1k–20k+ H100e Many racks and training domains | Pretrain competitive large open models, maintain several model lines and operate a public API. | **Provider-scale, still capacity constrained** Training and inference compete for the same fleet. | Mistral/DeepSeek/Kimi-class lab with dedicated infra, capital and model operations. |
| 600k–2m H100e accessible Multi-campus and multi-cloud | Parallel frontier programmes, large synthetic-data and RL pipelines, global inference and rapid retraining. | **Global consumer + enterprise products** No single training job consumes the whole estate. | Frontier lab backed by hyperscaler contracts, very large financing and energy commitments. |
| 2m–5m+ H100e owned Parent pool across regions and products | Rent compute to other labs, train first-party models, run internal AI and absorb multi-year silicon supply. | **Cloud platform + model catalogue** Customer workloads and first-party teams compete for allocations. | Amazon/AWS, Microsoft/Azure, Google and Alibaba Cloud; Meta has the owner scale without a public cloud. |

**People and value are scale context, not qualification rules.** “People on the work” estimates direct model, product, platform, silicon and datacenter effort—not every employee in the parent company. Independent labs are valued by private funding rounds; only listed parents have a market cap. July 2026 reference points include Mistral at 900+ employees, DeepSeek at an implied ≈$52bn, OpenAI at ≈$852bn and Anthropic at ≈$965bn.

- [Mistral headcount →](https://mistral.ai/careers/)

- [Mistral valuation →](https://mistral.ai/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai/)

- [DeepSeek implied valuation →](https://www.investing.com/news/stock-market-news/chinese-filing-implies-deepseek-valuation-of-around-52-billion-4796314)

- [OpenAI valuation →](https://openai.com/index/accelerating-the-next-phase-ai/)

- [Anthropic valuation →](https://www.anthropic.com/news/series-h)

<a id="cloud-routes"></a>

Rent, customize or build

### AWS, Azure, Google Cloud and Alibaba Cloud are both model shelves and infrastructure routes.

A hyperscaler can sell a managed API, provision dedicated throughput, rent a training cluster and train its own model family. Open the card to compare the four routes.

AWS · AMAZON

#### Bedrock + SageMaker AI

Use Bedrock for managed access to hundreds of foundation models; move to SageMaker AI and HyperPod for customization, distributed training and controlled deployment.

**First-party models** — Amazon Nova 2, Nova Sonic and Nova multimodal embeddings
**Other providers** — OpenAI and open/model-partner catalogues, Anthropic (CSP accounts not supported)
**Silicon route** — NVIDIA GPU plus Amazon Trainium and Inferentia

**Owner pool**

≈2.45m H100e
**Parent market cap**

≈$2.54tn · 24 Jul 2026
- [Amazon Bedrock →](https://aws.amazon.com/bedrock/)

- [Amazon Nova models →](https://aws.amazon.com/nova/models/)

- [Nova on SageMaker →](https://docs.aws.amazon.com/nova/latest/nova2-userguide/nova-model.html)

AZURE · MICROSOFT

#### Microsoft Foundry + Azure OpenAI

Use Azure-hosted model APIs with standard or provisioned deployment; move to managed compute and Azure ML when the workload needs custom weights, capacity or training control.

**First-party models** — MAI, Phi, Model Router and healthcare AI models
**Other providers** — OpenAI, Cohere, DeepSeek, Meta, Mistral, xAI, Anthropic (CSP accounts not supported) and other models.
**Infrastructure route** — Azure GPU estates, Foundry managed compute and Azure ML

**Owner pool**

≈3.42m H100e
**Parent market cap**

≈$2.84tn · 24 Jul 2026
- [Foundry model catalogue →](https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)

- [Azure ML compute routes →](https://learn.microsoft.com/en-us/azure/machine-learning/concept-compute-target)

GOOGLE CLOUD · ALPHABET

#### Gemini Enterprise Agent Platform

Use the platform for managed Gemini and partner APIs, provisioned throughput, tuning and agent workflows; use Managed Training, GKE, Cloud TPU or NVIDIA GPU capacity when the workload needs infrastructure control.

**First-party models** — Gemini, open-weight Gemma, Imagen, Veo and Google embeddings
**Other providers** — Meta, Mistral, Qwen, Anthropic (CSP accounts not supported) and other Gemini Enterprise Agent Platform Model Garden options
**Silicon route** — Cloud TPU including Trillium/Ironwood, plus NVIDIA GPU estates

**Owner pool**

≈5.05m H100e
**Parent market cap**

≈$3.89tn · 24 Jul 2026
- [Gemini Enterprise Agent Platform overview →](https://docs.cloud.google.com/gemini-enterprise-agent-platform/overview)

- [Gemini Enterprise Agent Platform models →](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models)

- [Google Cloud TPU route →](https://docs.cloud.google.com/tpu/docs/intro-to-tpu)

ALIBABA CLOUD · QWEN

#### Model Studio + PAI

Use Model Studio for OpenAI-compatible Qwen and partner APIs; use Platform for AI for dedicated deployment, fine-tuning and distributed training through DLC, EAS and Lingjun resources.

**First-party models** — Qwen text, reasoning, coding and multimodal families; Wan media models
**Other providers** — DeepSeek, Kimi and GLM availability varies by region
**Open-weight route** — Deploy Qwen weights with vLLM/SGLang or buy a managed Qwen API

**Owner pool**

Not separately public
**Parent market cap**

≈$0.29tn · 24 Jul 2026
- [Alibaba Model Studio →](https://www.alibabacloud.com/help/en/model-studio/what-is-model-studio)

- [PAI training and deployment →](https://www.alibabacloud.com/help/en/pai/use-cases/llm1/)

Scaling effect

#### More compute improves the search, the model and the service—but not in a straight line.

When model size, training data and compute stay balanced, language-model loss tends to improve along a power law: every extra unit buys a smaller marginal gain. The practical leap is broader than one benchmark score. Larger fleets buy more parallel experiments, longer post-training, more inference-time reasoning, replicas, recovery capacity and faster iteration. Better algorithms can compress the gap; weak data, networking or utilization can waste it.

- [OpenAI scaling-laws research →](https://openai.com/index/scaling-laws-for-neural-language-models/)

- [Chinchilla compute-optimal training →](https://arxiv.org/abs/2203.15556)

Best public comparison · latest like-for-like baseline is end-2025

### Lab-access compute, parent-owner pools and a fleet-cost proxy.

Open the card for a logarithmic or linear comparison of accessible compute and parent-owned pools. “H100e” compares peak AI operations with an NVIDIA H100 and does not equal real model throughput.

Start with the logarithmic order-of-magnitude view, then switch to linear to make equal chart width mean equal H100e—and therefore an equal step in the simple fleet-cost proxy. This is not an inventory audit. Solid bars estimate operational compute a lab could access; striped bars mark a disclosed run, floor or bound; outlined bars are parent-owned pools.

Axis spacing
Log view: each equal step is roughly ten times more accessible compute.
| Provider | Category | Fleet-cost proxy | Accessible compute / disclosure |
| --- | --- | --- | --- |
| Mistral AI | OPEN WEIGHTS | Fleet proxy $0.08–0.15bn | ≥3k observed H200 run |
| Moonshot / Kimi | OPEN WEIGHTS | Fleet cost not public | Fleet not disclosed |
| DeepSeek | OPEN WEIGHTS | Fleet proxy $0.5–1.0bn | ≈20k reported H100 + H800 |
| Qwen / Alibaba | HYPERSCALER MODEL LAB | Fleet cost not isolated | Qwen allocation not disclosed |
| Meta AI / MSL | OPEN + FRONTIER LAB | Fleet cost not isolated | <1.7m lab bound; parent pool below |
| **Frontier API labs** |  |  |  |
| xAI | FRONTIER | Fleet proxy $15–35bn | ≈600–700k lab access |
| Anthropic | FRONTIER | Fleet proxy ≥$25–50bn | ≥1m lab access |
| OpenAI | FRONTIER | Fleet proxy $43–85bn | ≈1.7m lab access |
| Google DeepMind | FRONTIER LAB | Fleet cost not isolated | <2m lab bound; parent pool below |
| **Hyperscaler parent owner pools · not one lab allocation** |  |  |  |
| Alibaba Cloud | CLOUD OWNER | Fleet cost not public | Not split out · China total ≈1.16m ex-smuggling |
| Meta | PARENT OWNER | Fleet proxy $58–115bn | ≈2.30m Q4 2025 owner pool |
| Amazon / AWS | CLOUD OWNER | Fleet proxy $61–122bn | ≈2.45m Q4 2025 owner pool |
| Microsoft / Azure | CLOUD OWNER | Fleet proxy $85–171bn | ≈3.42m Q4 2025 owner pool |
| Google | CLOUD + MODEL OWNER | Fleet proxy $126–252bn | ≈5.05m Q4 2025 owner pool |

**Fleet-cost proxy**

Plotted quantities × $25k–$50k per H100e: a broad accelerator, server and fabric replacement band derived from current complete-system pricing. It is not an accounting estimate of what the company paid.

It excludes buildings, grid connection, cooling, storage, energy, spares, financing and operations. TPU and Trainium economics differ; leases shift capex to a cloud owner; lab bounds are not costed when the parent allocation is unknown.

**Read DeepSeek correctly.** The ≈20,000 figure is accelerator-class chips, not 20,000 eight-GPU servers. Its V3 paper documents one 2,048-H800 training cluster and 2.788M H800 GPU-hours; wider fleet figures and additional H20/A100 hardware are not converted into the plotted number.

**Read Qwen and Alibaba separately.** Qwen is Alibaba’s model family; Alibaba Cloud is the infrastructure and API parent. Neither Qwen’s allocation nor Alibaba’s owner total is public. Epoch estimates identified Chinese owners together at ≈1.16m H100e in Q4 2025 before its optional smuggling estimate, so assigning that number to Alibaba would be wrong.

**Ownership is not availability.** Epoch AI’s Q4 2025 owner estimates are Google 5.05m, Microsoft 3.42m, Amazon 2.45m and Meta 2.30m H100e, but those pools also serve cloud customers and internal products. Its July 2026 campus directory provides a newer site check without a like-for-like company total.

- [Epoch AI owner-pool data and method →](https://epoch.ai/data/ai-chip-owners)

- [Epoch AI lab-access estimates →](https://epoch.ai/gradient-updates/frontier-labs-dont-use-most-ai-compute)

- [Operational campus directory →](https://epoch.ai/data/ai-data-centers)

- [Mistral 3,000-H200 run →](https://mistral.ai/news/mistral-3/)

- [DeepSeek fleet estimate →](https://www.csis.org/analysis/deepseek-deep-dive)

- [DeepSeek V3 report →](https://arxiv.org/abs/2412.19437)

- [System-price basis for fleet proxy →](https://intuitionlabs.ai/articles/data-center-gpu-pricing-2026)

- [Kimi capacity constraint →](https://apnews.com/article/4c66a2e0f557ce79d3cc2d769c9a6226)

Public bridges to scale

### Use public access, capital and demand before carrying the full infrastructure risk.

The European Union and United States offer different ladders from research access to provider-scale infrastructure. Open the card to compare the regional routes and their limits.

European Union

##### Shared compute, scale-up finance and sovereign demand

01 · ACCESS

##### Borrow the first large run

EuroHPC’s open industrial call offers AI Factory allocations above 50,000 GPU-hours with a stated ten-working-day approval target. Nineteen AI Factories and thirteen antennas are being assembled to give startups and SMEs compute plus technical support.

- [Apply for large-scale AI Factory access →](https://www.eurohpc-ju.europa.eu/large-scale-access-ai-factories_en)

- [Check the factory and antenna network →](https://www.eurohpc-ju.europa.eu/eurohpc-ju-launches-call-proposals-strengthen-european-ai-ecosystem-2026-04-28_en)

02 · FINANCE

##### Bridge startup to infrastructure company

TechEU says it will deploy €70bn of EIB Group equity, loans and guarantees in 2025–2027 to mobilize €250bn with partners across the innovation lifecycle, including AI and digital infrastructure. It is a financing channel, not an automatic subsidy.

- [Review the TechEU finance route →](https://www.eib.org/en/press/all/2025-314-europe-s-innovative-companies-get-boost-as-eib-group-launches-techeu-platform-to-simplify-financing)

03 · BUILD

##### Finance the leap to gigafactory scale

InvestAI plans to mobilize €20bn for several AI Gigafactories, with the EIB exploring advisory support and loans. The target shape is about 100,000 advanced chips per site. The formal call is due in summer 2026 and first construction is scheduled for 2027: this is future capacity, not online supply.

- [Track the AI Gigafactory call →](https://commission.europa.eu/topics/competitiveness/competitiveness-coordination-tool-projects/ai-gigafactories_en)

- [Read the EIB financing mandate →](https://www.eib.org/en/press/all/2025-491-eib-group-and-european-commission-join-forces-to-finance-ai-gigafactories)

04 · DEMAND

##### Turn sovereignty into a buying criterion

The Commission’s proposed Cloud and AI Development Act would streamline sites, energy and finance, create an EU sovereignty framework and establish common public-sector procurement. That can make trusted European capacity easier to specify and buy; it does not guarantee a contract.

- [Read the proposed Cloud and AI Development Act →](https://digital-strategy.ec.europa.eu/en/policies/cloud-and-ai-development-act)

United States

##### Research access, competitive funding, sites and procurement

01 · ACCESS

##### Use the national research resource

The NSF-led National Artificial Intelligence Research Resource connects US researchers, educators, startups and small businesses to public and private compute, models, data and expertise. By March 2026 it had supported more than 600 research teams and 6,000 students.

- [Explore current NAIRR access →](https://www.nsf.gov/focus-areas/ai/nairr)

- [Read the two-year progress update →](https://www.nsf.gov/cise/updates/nairr-2-years-advancing-american-artificial-intelligence)

02 · FUND

##### Fund efficiency and regional capacity

NSF-backed STRIDE can invest up to $21m over two years in deployment-ready AI infrastructure efficiency, with awards up to $3.5m. NSF Regional Innovation Engines can fund regional technology coalitions up to $160m over ten years. Both are competitive programmes, not project debt.

- [Review the AI Efficiency Challenge →](https://www.nsf.gov/tip/updates/nsf-supported-stride-ventures-launches-ai-efficiency)

- [Explore NSF Regional Innovation Engines →](https://www.nsf.gov/funding/initiatives/regional-innovation-engines)

03 · BUILD

##### Shorten the path to 100 MW+ campuses

A federal order defines qualifying AI data center projects as requiring more than 100 MW of new load and directs expedited federal permitting. DOE has also selected four federal sites for private AI data center and energy partnerships. This lowers siting friction; it does not supply a fleet or guaranteed grid capacity.

- [Read the data center permitting order →](https://www.whitehouse.gov/presidential-actions/2025/07/accelerating-federal-permitting-of-data-center-infrastructure/)

- [Review the selected DOE sites →](https://www.energy.gov/articles/doe-announces-site-selection-ai-data-center-and-energy-infrastructure-development-federal)

04 · DEMAND

##### Use federal procurement as a first market

OMB M-25-22 directs more efficient federal AI acquisition, while GSA’s OneGov consolidates technology buying and reported 20 vendor agreements by April 2026. The route can accelerate agency adoption, but eligibility is limited and no programme guarantees a provider contract.

- [Open current OMB memoranda →](https://www.whitehouse.gov/omb/information-resources/guidance/memoranda/)

- [Review the current OneGov route →](https://www.gsa.gov/buy-through-us/purchasing-programs/multiple-award-schedule/onegov)

EU operator example
**Mistral moved from a disclosed 3,000-H200 model run to production GB200 service in February 2026 and says it is targeting 200 MW of sovereign EU capacity by 2027.** The route mixes private capital, NVIDIA supply, datacenter partners, cloud revenue and sovereign-enterprise customers. It is a useful pattern—not proof that every provider should own the same fleet.

- [Inspect Mistral Compute’s live timeline →](https://mistral.ai/products/compute/)

- [See the 200 MW Campus AI agreement →](https://presse.bpifrance.fr/mistral-securise-jusqua-200-mw-de-capacite-de-calcul-avec-campus-ai-en-france/?lang=fra)

Read the comparison carefully
**The US programmes are analogues, not one-for-one equivalents.** They split research access, innovation funding, federal land and permitting, and government demand across separate agencies. Large-fleet capital still depends mainly on commercial customers and private finance, and eligibility in one lane does not imply support in the others.

- [Read the US AI Action Plan →](https://www.whitehouse.gov/releases/2025/07/white-house-unveils-americas-ai-action-plan/)

- [Review NAIRR’s operating scope →](https://www.nsf.gov/focus-areas/ai/nairr)

**Inference fit is not training fit.** Full training holds weights, gradients, optimizer states and activations, often using several times the inference memory before checkpoint and data-pipeline overhead. LoRA and Q‑LoRA are much lighter. Use the hardware table to shortlist a tier, the scale comparison to understand the organization around it, and the simulator below for the workload, quote and payback target you actually expect.

Try before you buy

## Rent the exact workload before committing to private hardware.

Before a five- or six-figure private-cloud commitment, rent the candidate stack long enough to discover whether the model, service shape and operating cost actually fit. A rental is a test drive, not a GPU-availability, price or hardware-validation guarantee.

**Run the acceptance pack first.** Test the exact model, engine, quantization, context, concurrency and acceptance harness you would buy for; retain the container, settings, logs, latency, throughput, memory and task-quality results.

 **Referral disclosure:** this isaiuseful.com link is a Runpod referral link. As checked 10 August 2026, eligible new first-time users must sign up through it with Google SSO and load their first $10: European users receive $5 credit, while non-European users receive a weighted $5–$500 credit (Runpod says most are $10 or less). Terms can change. If eligible, isaiuseful.com receives its referral bonus and earns Runpod credits on actual usage for the first six months: 3% of Pod spend and 5% of Serverless spend. Using the link supports the site.

- [Rent a Runpod test environment through this referral link →](https://runpod.io?ref=l40ix174)

- [Read Runpod’s current referral terms →](https://docs.runpod.io/accounts-billing/referrals)

<a id="economics"></a>

Users + ROI simulator

## Turn a vendor quote into a price per user—or per token.

Start from the model’s evidence-matched serving preset or lock a platform you already own. The planner packs repeatable replica cells into that platform, scales the packed loadout toward a user target and keeps every generated assumption editable.

Fits selected system
90 GB service build in 1.44 TB HBM, with 15% reserved for runtime headroom.

Pricing view
Monthly operating cash
**—**

Selected private billing minus modeled monthly operating cost.

Simple capex payback
**—**

Capex divided by positive monthly operating cash; excludes financing and demand ramp.

Gap to capital-recovery target
**—**

Monthly billing minus operating cost and the selected capital-recovery allowance.

Selected private quote / user / month
**—**

Same-model OpenRouter cost × the selected private quote multiplier.

Same-model OpenRouter / user / month
**—**

Selected local model’s market rate × workload.

Operating break-even / served user / month
**—**

Covers the costs explicitly modeled here, without capital recovery.

Target-payback price / served user / month
**—**

Covers modeled operating cost plus capex recovery inside the selected target.

Per-user gap to target-payback price
**—**

Selected private quote minus the target-payback price, per served user per month.

Estimated supported users
**—**

Constrained by the tighter of token throughput, concurrency and per-request memory.

Auto-sized loadouts for target
**—**

Calculated from the capacity of one packed loadout.

Selected API alternative / user / month
**—**

API rate × uncached input, cached input and output.

Selected private quote multiplier
**—**

The commercial ask currently applied to the same-model OpenRouter anchor.

Quote multiplier required for target
**—**

In-range results are rounded upward to the next 0.25× slider step so the recommendation never understates the target.

Modeled monthly operating cost
**—**

Monthly private billing at selected price
**—**

Selected API cost for served target
**—**

Change any assumption to test the scenario.

- [Open selected model’s current OpenRouter pricing →](https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b/pricing)

- [Open selected API pricing →](https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b/pricing)

Planning note
**This is a workload simulation—not a benchmark, capacity promise or complete financial model.**

**The financial outputs use two different thresholds on purpose.** Operating cash is private billing minus the recurring costs entered here. Simple payback divides capex by positive operating cash. The capital-recovery target additionally asks the service to recover capex within the selected number of years. A plan can therefore make positive operating cash and have a finite payback while still missing a faster target. Financing costs, tax, working capital, depreciation, residual value and any staffing, software, insurance or network cost not entered above are excluded.

Observed coding workloads vary sharply: one 12,000-developer study reported a 51M-token monthly median and about 380M at P90, while an agent-task study found up to 30× variation between runs. The higher presets represent continuously operating agent-worker equivalents, not ordinary human seats. Start there, then replace them with your own exports. [Review the developer workload study →](https://jellyfish.co/blog/is-tokenmaxxing-cost-effective-new-data-from-jellyfish-explains/) [Review the agent-task variance study →](https://arxiv.org/abs/2604.22750)

Each model starts from a complete replica shape, not a model-size speed multiplier alone. gpt‑oss‑120b, MiniMax M3, Qwen3.5 and DeepSeek V4 Pro now use 8K/1K single-node InferenceX profiles; Nemotron 3 Super and Ultra, GLM‑5.2 and Kimi K2.7 use matched Lambda runs; DeepSeek V3.2 uses a GPUStack sweep; and Kimi K3 keeps its throughput–interactivity curve. The MiniMax and Qwen memory gates include both data-parallel checkpoint copies in their measured four-GPU cells. Mistral Small 4, Mistral Large 3 and DeepSeek V4 Flash remain explicitly labelled reference-topology proxies. A complete proxy cell is no longer divided by all eight GPUs merely because several small cells pack into one node. [Inspect the InferenceX profiles →](https://inferencex.semianalysis.com/) [Nemotron 3 Super measured profile →](https://lambda.ai/inference-models/nvidia/nvidia-nemotron-3-super-120b-a12b) [Nemotron 3 Ultra measured profile →](https://lambda.ai/inference-models/nvidia/nemotron-3-ultra) [DeepSeek V3.2 measured sweep →](https://docs.gpustack.ai/2.0/performance-lab/deepseek-v3.2/h200/) [GLM‑5.2 measured profile →](https://lambda.ai/inference-models/zai-org/glm-5.2) [Kimi K2.7 measured profile →](https://lambda.ai/inference-models/moonshotai/kimi-k2.7-code) [Kimi K3 throughput curve →](https://vllm.ai/blog/2026-07-27-k3)

Nemotron 3 Ultra needs a special translation: its public run uses 8,192 input / 65,536 output tokens. The published input rate is therefore the observed token share of a decode-heavy saturation test, not an independently saturated prefill ceiling. The preset retains that observed value for context but uses measured decode throughput, interactivity and memory as its hard capacity gates; replace the profile with a mixed-shape sweep before procurement.

The service targets separate user-visible output speed from total server throughput. Independent API measurements checked on 1 August 2026 observed about 32 output tok/s for Kimi K3, 66 for GPT-5.6 Sol at maximum effort and 74 for Claude Fable 5. The 30 and 70 targets are parity anchors, 100 is a fast interactive target and 15 is for background agents; none is a latency SLA. gpt‑oss‑120b, MiniMax M3, Qwen3.5 and DeepSeek V4 Pro load measured points from their InferenceX concurrency sweeps where available, while Kimi K3 uses its approximate published curve. High-volume workloads also reserve the parallel active requests implied by monthly output volume and measured per-request speed; the per-request KV assumption is multiplied by that request count. Every single-point benchmark and topology proxy changes aggregate throughput through a conservative cross-model envelope calibrated to those 8K/1K sweeps: 0.70×, 1.00×, 1.50× and 1.90× the normalized frontier point. Those values are planning estimates, not measurements for the selected model; whole-platform rounding or a tighter memory gate can still leave two adjacent targets at the same purchase count. [Review the measured Kimi K3 API speed →](https://artificialanalysis.ai/models/kimi-k3) [Review the measured GPT-5.6 Sol API speed →](https://artificialanalysis.ai/models/gpt-5-6-sol) [Review the measured Claude Fable 5 API speed →](https://artificialanalysis.ai/models/claude-fable-5)

Prefix reuse changes the self-hosted prefill path, not output decoding. The simulator counts uncached input at full prefill work and cached input at the editable cache-hit allowance. Moonshot reports above 90% cache hits for Kimi K3 coding traffic, so its coding presets start at 90%; that provider result is not a promise for a different scheduler or workload. Cache lookup, state transfer, retention and eviction still consume memory and network capacity. [Review Moonshot’s K3 cache claim and direct pricing →](https://www.kimi.com/blog/kimi-k3) [Review vLLM’s hybrid prefix-cache design →](https://vllm.ai/blog/2026-07-27-k3)

Every preset assigns at least 20% of tokens to output; coding profiles use a 50/50 split so billed reasoning is not hidden. Replace workload, prefix reuse, cache-read work, aggregate throughput, concurrency, interactivity and per-request memory assumptions with telemetry from the same ISL/OSL test shape.

The same-model OpenRouter rate remains the grounding benchmark even when another API alternative is selected. Rates are stored as planning inputs checked on 1 August 2026 rather than fetched live; recheck the linked model page before quoting. Direct API comparisons exclude tool calls, cache writes, cache storage, long-context or regional uplifts and negotiated discounts. **All prices exclude VAT.** [Review the platform evidence →](#hardware)

<a id="cuda"></a>

Compatibility

## CUDA compatibility is a binary boundary, not a performance grade.

Only NVIDIA hardware runs CUDA binaries. AMD and Intel can still be excellent inference platforms, but the application, kernels and quantization must support their native stack.

| Platform | Runs CUDA binaries? | Native stack | Portability route | Provider decision |
| --- | --- | --- | --- | --- |
| NVIDIA Ampere · A100 | ✓ Native CUDA | CUDA · NCCL · NVLink   <br> ✕ No native NVFP4 | Use BF16/FP16 or supported integer paths. Some model cards document loading an NVFP4 checkpoint through a fallback. | Treat NVFP4 as model-specific weight compatibility here—not native FP4 throughput. |
| NVIDIA Hopper · H100/H200 | ✓ Native CUDA | CUDA · NCCL · NVLink · TensorRT‑LLM   <br> ✕ No native NVFP4 | FP8 is the native low-precision path. Selected NVFP4 checkpoints can fall back to W4A16. | Buy for mature CUDA and FP8—not for native NVFP4 throughput. |
| NVIDIA Blackwell · GB10/B200/B300/GB200/GB300 and RTX PRO | ✓ Native CUDA | CUDA · NCCL · NVLink · TensorRT‑LLM   <br> ✓ Native NVFP4 | Native W4A4 path when the engine, kernel and model support it. | Current launch target for NVFP4; validate GB10 and workstation-specific kernels separately from datacenter Blackwell. |
| AMD MI300X/MI325X/MI355X and Radeon | ↻ No—port and rebuild | ROCm · HIP · RCCL · Infinity Fabric | HIP source can target AMD or NVIDIA, but one compiled binary cannot run on both. | Strong memory economics when the exact engine, model and kernels pass a pilot. |
| Intel Gaudi 3 | ✕ No—separate stack | SynapseAI · Habana libraries · RoCE | Use Gaudi-supported frameworks and Optimum Habana; CUDA binaries do not transfer. | Treat as its own product lane with an explicit supported-model catalogue. |
| Intel GPU / XPU | ↻ No—source migration | oneAPI · SYCL · OpenVINO | SYCLomatic can assist CUDA-source migration; manual work and validation remain. | Useful for selected Intel-XPU workloads, not a Gaudi substitute. |
| tinygrad runtime | ◇ Backend choice | Its own small compiler with CUDA, AMD, NV, Metal and other backends | A tinygrad program can target different backends; tinygrad does not make AMD or Intel CUDA-compatible. | Use for owned compiler research and ports, not assumed vLLM feature parity. |

**Feature support changes faster than the silicon.** As of 26 July 2026, vLLM documents NVIDIA CUDA, AMD ROCm and Intel XPU routes, but feature parity and quantization support vary. NVFP4 is an NVIDIA format; native W4A4 acceleration is a Blackwell-generation capability, while Hopper uses a model- and engine-specific fallback where supported. Pin the container, driver, firmware, engine commit and model revision used in acceptance testing.

- [NVIDIA NVFP4 hardware explanation →](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)

- [Blackwell W4A4 versus Hopper W4A16 →](https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/)

- [NVIDIA CUDA libraries →](https://docs.nvidia.com/cuda-libraries/index.html)

- [AMD HIP binary-compatibility FAQ →](https://rocm.docs.amd.com/projects/HIP/en/docs-6.2.4/faq.html)

- [Current vLLM GPU support →](https://docs.vllm.ai/en/latest/getting_started/installation/gpu/)

- [Intel Gaudi documentation →](https://docs.habana.ai/en/latest/)

- [Intel CUDA-to-SYCL migration →](https://www.intel.com/content/www/us/en/docs/oneapi/programming-guide/2025-1/migrating-from-cuda-to-sycl-for-the-dpc-compiler.html)

tinygrad route

## Own more of the stack, accept more integration work.

tinygrad’s useful idea is not “cheap CUDA.” It is a small, MIT-licensed compiler/runtime and reference hardware path that makes the stack inspectable, portable and open to operator modification.

Reference lab · AMD

### tinybox red v2

**€10,502 / $12,000 ex VAT**

4× Radeon 9070 XT, 64 GB aggregate GPU memory, 128 GB system memory and 2 TB NVMe. Best for compiler work, AMD ports and smaller sharded models—not a 400B+ production endpoint. Confirm store stock before planning delivery.

- [Official tinybox red v2 →](https://tinycorp.myshopify.com/products/tinybox-red-v2)

- [Read tinygrad’s hardware pitch →](https://tinygrad.org/pitch.pdf)

Reference lab · NVIDIA

### tinybox green v2 Blackwell

**€65,640 / $75,000 ex VAT**

4× RTX PRO 6000 Blackwell and 384 GB aggregate GPU memory. It runs CUDA, but PCIe-attached workstation GPUs are not an NVLink HBM supernode; test throughput, concurrency and current store stock.

- [Official tinybox green v2 →](https://tinycorp.myshopify.com/products/tinybox-green-v2-with-4x-rtx-pro-6000-blackwell)

- [Read tinygrad’s hardware pitch →](https://tinygrad.org/pitch.pdf)

Software ownership

### Port only where it earns control

**€0 / $0 software licence**

Use tinygrad to understand and own selected kernels or models. Keep vLLM or SGLang for production until the port meets the same correctness, batching, observability and recovery tests.

**Before promising hardware, ask what must stay compatible.** Examples include an OpenAI-compatible endpoint; vLLM, SGLang, NIM or Triton; PyTorch or JAX; tool calling and structured JSON; MCP or Hermes-style agent harnesses; embeddings, reranking, vision, PDF processing or fine-tuning; plus SSO, audit logs, retention, data residency, context, concurrency and latency requirements.

A client may not know it has a CUDA dependency. An NVIDIA container, TensorRT‑LLM, FlashAttention, bitsandbytes, NCCL or a PyTorch CUDA wheel is one—even if the business owner never says “CUDA.” Ask for the current container, lockfile and a representative acceptance test.

- [tinygrad repository and backends →](https://github.com/tinygrad/tinygrad)

- [Project status and tinybox specifications →](https://tinygrad.org/)

01

### Hardware pluralism

Keep the option to buy NVIDIA, AMD or commodity cards where the workload and kernels allow it.

02

### Inspectability

A smaller compiler makes generated kernels and backend behavior easier for an operator to understand and modify.

03

### Acceptance tests first

Port the client’s real model and critical software path before making a platform-wide compatibility promise.

04

### Honest maturity

tinygrad describes itself as alpha and says it is not faster than PyTorch for most use cases yet. Treat it as an engineering lane, not magic production parity.

Capacity ladder

## Match the checkpoint first, then reserve serving headroom.

These are procurement starting points, not throughput guarantees. Keep 15–25% aggregate memory for runtime overhead and KV cache, then load-test the actual context length and concurrency.

✕ No native NVFP4 · 640 GB–1.13 TB

### H100 / H200

Use native FP8, or a checkpoint-specific W4A16 fallback when memory capacity matters. Speculative mixed-NVFP4 weight capacity is ≈0.74–1.45T parameters across an eight-GPU node, but it does not gain native W4A4 compute.

✓ Native NVFP4 · 1.44–2.3 TB

### B200 / B300

About 1.65–2.95T parameters by conservative mixed-checkpoint memory math. Nemotron Ultra, Qwen, DeepSeek and Mistral Large are more realistic first services; leave the rest for cache, replicas and concurrency.

✓ Native NVFP4 · 13.4–20 TB

### GB200 / GB300 NVL72

Roughly 15.5–25.6T parameter-equivalent weight capacity. In practice, spend that headroom on Kimi K3-scale sharding, multiple replicas, long-context cache and test-time compute.

◇ Roadmap assumption · 4.6–82.9 TB

### V300 / VB300 / Kyber

The 5.3–106T NVFP4-equivalent range is compounded planning arithmetic across derived or roadmap systems—not a shipped support matrix, benchmark or useful single-model target.

Price method

## One exchange rate, visible caveats.

USD prices were converted at the ECB reference rate for 20 July 2026: **€1 = $1.1426** . EUR hardware figures are rounded to the nearest euro; token prices are rounded to the nearest euro cent.

### Excluded from every price

VAT, sales tax, shipping, import duties, rack, networking, storage, facility upgrades, energy, support and integration unless the row says otherwise.

### Ballpark means planning input

Each row says whether its number comes from reported purchase quotes, a component/BOM estimate or a roadmap ASP estimate. Replace it with your vendor quote before making a return case.

### Requote before buying

GPU allocations, support terms and discounts move quickly. Keep the accelerator, memory, interconnect, cooling and software acceptance criteria fixed while comparing quotes.

- [ECB reference exchange rates →](https://www.ecb.europa.eu/stats/policy_and_exchange_rates/euro_reference_exchange_rates/html/index.en.html)

- [Morgan Stanley BOM breakdown reported by Tom’s Hardware →](https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidias-memory-costs-soar-485-percent-latest-ai-systems-now-cost-usd7-8-million-to-build-memory-now-comprises-25-percent-of-the-total-cost-rubin-gpus-a-mere-usd50-000-apiece)

- [Return to server prices →](#hardware)

Provider default
Launch one permissive model on one validated NVIDIA reference system with one datacenter partner. Add a second accelerator stack only when its memory economics survive the cost of maintaining a second software product.

- [Model the first service](#economics)

- [Check the software boundary](#cuda)

- [Return to local models](https://isaiuseful.com/local-models.html.md)
