Provider build guide · checked 15 August 2026

How Do You Build a Private Cloud for Open-Weight AI Models?

Start with downloadable weights and a serving stack you can operate inside your own security boundary. Then buy for memory, interconnect, power, cooling and measured token throughput—not just peak FLOPS.

14 + 1open-weight lines + hosted previewGLM‑5.3 ACCESS LIVE; WEIGHTS PENDING
2 customfrontier licences need reviewQWEN3.8 + K3 LARGE-SCALE MAAS TRIGGERS
14–600 kWNVIDIA system planning range8-GPU SERVER TO RUBIN ULTRA ROADMAP RACK
Recommendation catalogue

Deploy now, validate next, or retain only as a baseline.

Cards are ordered by present recommendation, not chronology or parameter count. Current launch means a 2026 release with usable access and a credible deployment route. Current option marks a specialized, superseded or review-gated 2026 line. Older baseline marks a pre‑2026 release retained for compatibility or comparison—not as the default. GLM‑5.3 remains a hosted preview until its checkpoint and licence ship; Qwen3.8‑Max and Kimi K3 require their custom commercial terms to be reviewed.

current launch · preferred current release current option · secondary or specialized review gate or hosted preview · incomplete route current self-host · newer hosted model exists older baseline · retained, not default
△ Current optionApache 2.0

Qwen3.5 397B‑A17B

A still-current alternative now ranked behind Qwen3.8. Multimodal MoE with 397B total / 17B active parameters. The calculator now uses NVIDIA's roughly 251 GB NVFP4 repository and a measured 4× B300 SGLang/MTP profile at 8K input / 1K output, replacing the much larger 807 GB BF16 planning route.

First service cell
4× B300 NVFP4 for the measured profile
Serve with
SGLang with the pinned NVFP4/MTP path
Licence fee
€0 / $0
! Hosted previewLicence pending

GLM‑5.3

A post-trained successor that uses the same base model as GLM‑5.2. Z.ai reports large gains on long-horizon coding and agent benchmarks, but the launch status is split: Coding Plan access is live, the general API is marked coming soon and weights are targeted for release within two weeks of 14 August. Keep GLM‑5.2 for self-hosted planning until the 5.3 checkpoint, licence and engine recipes can be inspected.

Available now
GLM Coding Plan and ZCode
API change
Thinking required; low, high or max effort
Self-hosting
Wait for weights, licence and measured recipes
↻ Current self-hostMIT

GLM‑5.2

A 753B model with up to 1M context. The BF16 repository is about 1.51 TB and needs at least 2 TB aggregate HBM with useful headroom; the calculator instead uses the released FP8 build and a matched 8× B200 profile.

First node
8× B200 for FP8; ≥2 TB HBM for BF16
Serve with
vLLM or SGLang
Licence fee
€0 / $0
△ Current fallbackApache 2.0

Mistral Small 4

A current compact fallback rather than the catalogue default; validate its vendor comparisons on your own tasks. 119B total / 6.5B active, multimodal, 256K context and 24-language support. Its official 70.8 GB NVFP4 build is a clean Blackwell fit; SGLang documents TP1 on B200/B300, while FP8 needs TP2 on H100/H200. Native W4A4 acceleration requires Blackwell.

First node
1× B200 with NVFP4; benchmark before scaling
Serve with
SGLang or vLLM; validate the exact NVFP4 kernel
Licence fee
€0 / $0
! Current · reviewModified MIT

Kimi K2.7 Code

A downloadable trillion-parameter coding and agent model with 32B active parameters, native INT4 and 256K context. The calculator uses a matched 8× B200 vLLM run; the modified terms, nightly runtime and model-specific kernels still need review.

First node
8× B200 at native INT4
Serve with
vLLM, SGLang or KTransformers
Licence fee
€0 / $0
↓ Older · validateMIT

DeepSeek V3.2

A prior-year reasoning baseline retained for comparison and existing deployments. A 685B FP8 reasoning and agent model with a 690 GB repository. The card documents vLLM and SGLang; production tool parsing still needs your own robustness tests.

First node
8× H200-class; pilot other stacks
Serve with
vLLM or SGLang
Licence fee
€0 / $0
↓ Older baselineApache 2.0

Mistral Large 3

A prior-year Mistral baseline, not a first recommendation against current releases. 675B total / 41B active, multimodal and 256K context. Mistral documents FP8 on one 8× H200 or B200 node and an NVFP4 checkpoint on one 8× H100 or A100 node. H100 and A100 do not have native FP4 Tensor Cores, so that older-GPU route depends on model-specific fallback or dequantization—not native NVFP4 W4A4 acceleration.

First node
8× B200 for native NVFP4; 8× H200 for FP8
Serve with
vLLM; pin the hardware-specific path
Licence fee
€0 / $0
↓ Older baselineApache 2.0

gpt-oss-120b

Retained as an older compatibility and footprint baseline; start with current candidates unless this exact package wins your acceptance tests. About 117B total / 5.1B active, tool-capable and packaged in MXFP4 to fit one 80 GB H100. Minimum fit is not the same as an economical service profile: the calculator uses a measured 2× B200, 8K/1K TensorRT‑LLM run so packed capacity is grounded in observed throughput.

First service cell
2× B200 for the measured capacity profile; H100 fits the build
Serve with
TensorRT‑LLM, vLLM or supported reference stack
Licence fee
€0 / $0

Commercially usable does not mean restriction-free. Before listing any model, retain the exact licence and notice files, scan model code, document provenance limits, red-team the served build and define an abuse policy. Qwen3.8’s custom licence requires the notice to travel with copies; commercial products or services above 100 million monthly active users or $20 million monthly revenue must display the model name, while Model-as-a-Service or AI Work Assistant businesses above $50 million aggregate revenue in any consecutive 12 months need a separate Qwen licence. MiniMax M3 is also open-weight under custom terms, not Apache 2.0 or MIT: commercial deployments require visible attribution and a one-time notice, and products or services above the licence's $20M annual-revenue threshold require prior written authorization. Architecture, quantization, context and concurrency determine the real memory requirement.

NVIDIA platform ladder

Start with platform scale—then see what it takes to become a provider.

The ten current and roadmap tiers keep node and rack choices legible. The provider-scale view then connects those purchases to research, open-model labs, frontier fleets and the EU and US public routes that can help bridge the gaps.

The ten-name map

Five node tiers. Four NVL72 racks. One NVL144 rack.

Read the node class first, then the rack class. Open the card for the NVIDIA platform map, videos, NVFP4 guidance, prices, Rubin analysis and OEM alternatives.

H100 through GB300 are shipping systems; Vera Rubin NVL72 is in its production ramp with preliminary published specifications. V300 uses the later Rubin Ultra roadmap's 576 GB-per-GPU assumption; VB300 NVL72 is a 72‑GPU comparison domain derived from half a Kyber rack, not an announced NVIDIA SKU. Kyber NVL144 remains the actual Rubin Ultra roadmap rack label.

8-GPU node · HopperH100 → H200

The mature entry tier. H200 keeps the same node shape and raises accelerator memory from 640 GB to 1.13 TB.

8-GPU node · newerB200 → B300 → V300

More memory and newer low-precision engines; V300 is a derived roadmap planning node rather than an orderable SKU.

72-GPU NVLink rackGB200 → GB300 → Vera Rubin → VB300

Vera Rubin is the current production-ramp generation. VB300 is a later, derived Rubin Ultra comparison—not the name of Vera Rubin NVL72.

144-GPU roadmap rackKyber NVL144

The Rubin Ultra/V300 roadmap tier changes rack density, power delivery and price class again.

From number format to facility

See the platform NVIDIA is describing—then translate the spectacle into a bill of materials.

These vendor videos connect NVFP4, Vera, Rubin, DSX facilities, a scientific workload and the full GTC keynote. Use them to understand the intended system shape; use the tables below for procurement questions, facility constraints and explicit planning caveats.

Low precision · NVIDIA Developer

NVFP4 can make large models smaller—but the GPU matters

NVFP4 stores most values in four bits and uses fine-grained scaling to preserve more accuracy than a crude four-bit conversion. The catch is simple: native W4A4 acceleration starts with Blackwell Tensor Cores. Hopper can use selected checkpoints through a W4A16 fallback, but that is not the same speed path.

Platform overview · NVIDIA

Vera Rubin is a full system, not just a GPU

Use the overview to see how NVIDIA frames CPUs, GPUs, networking and software as one agent platform. The delivered configuration still needs an exact vendor quote and acceptance test.

Host architecture · NVIDIA

The CPU still shapes the service

Vera is NVIDIA's host-side story for agent systems. Translate that promise into memory bandwidth, data movement, storage and software requirements for the workload you will actually serve.

Facility blueprint · NVIDIA

At rack scale, the building joins the stack

DSX makes the facility-level ambition visible. Power delivery, cooling, networking, operations and recovery are part of the product long before a gigawatt becomes relevant.

Workload context · NVIDIA

Start with the scientific job, not the rack

The discovery story is a useful demand-side counterweight to hardware spectacle. Define the models, data, latency and evaluation first; only then choose the infrastructure tier.

Full keynote · NVIDIA

See the complete platform story in one sitting

The full GTC 2026 keynote connects Vera Rubin, agents, networking and AI factories in NVIDIA’s own long-form narrative. Use it for context, then return to the evidence tables for procurement decisions.

NVFP4 in plain English

A promising format with one hard boundary: native acceleration needs newer NVIDIA GPUs.

NVFP4 is not a universal “make any model four times faster” switch. It reduces stored weights and memory traffic, but the checkpoint, serving engine, kernels, context and GPU generation must all agree.

Commercial platform path · checked 10 August 2026

Hardware is only one layer of the NVIDIA commercial stack.

Open the stack to see which layers are software products, which are included system software and which require a separate commercial entitlement.

Layered architecture · links open NVIDIA documentationOne possible commercial stack—not a bundled bill of materials.

Products can be selected independently, and support or entitlement depends on the exact order and supported configuration.

Outcomes and workloads
ApplicationsAgentsModel APIsAnalyticsTraining
Application development

Build, customize and prepare workloads

Developer-facing tools; availability may depend on the chosen software entitlement.

Inference serving

Optimize, expose and coordinate model responses

Serving components are separate choices; a workload does not need all of them.

Infrastructure management

Schedule workloads and operate supported infrastructure

Entitlement-sensitive layers; confirm the configuration, term and support route.

System software

Operating system, GPU platform and container access

DGX OS is a system foundation; it is not itself a production-software entitlement.

System and hardware foundation
DGX or certified systemsNVIDIA GPUsNetworkingStorageActive support where purchased
Deploy anywhere
CloudData centerEdgeLocal workstation
Operator-assembled alternative

An open-source control plane can lower license fees—but moves operations to you.

Open the diagram for a realistic surrounding stack, including the vendor dependencies that remain at the hardware boundary.

Operator-assembled architecture · examples, not a required bill of materialsOpen components, explicit operator ownership.

Each layer is replaceable. The engineering team owns the integration, upgrades, recovery path and production support boundary.

Outcomes and workloads
ApplicationsAgentsModel APIsAnalyticsTraining
Applications and agents

Compose the user-facing workload and its stateful agent paths

These tools supply interfaces and orchestration; the operator still owns identity, permissions and durable business state.

Model serving and gateways

Expose models through an API contract you own

Choose a serving engine and gateway deliberately; the workload does not need every option.

Training, evaluation and lifecycle

Train, evaluate and promote reproducible model artifacts

Model code, adapters, datasets, evaluation and registry evidence are separate operating responsibilities.

Workload control

Pick one primary scheduling and queueing model

Combining schedulers without a clear owner creates competing admission and recovery rules.

Provisioning and configuration

Rebuild nodes from documented state

Bare-metal lifecycle and desired configuration are separate responsibilities.

Delivery and supply chain

Build, promote and retain deployable artifacts

Workflow automation and a private registry still need access control, signing policy and retention rules.

Observability

Own telemetry, dashboards, logs and alerts

Keep sensitive prompts and identifiers out of labels and unbounded log streams.

Data, network and security

Assemble infrastructure controls as separate services

Storage, policy and secrets still need backups, upgrades and tested recovery.

Hardware and vendor boundary

Open control software does not make the accelerator stack open

Pin and validate the OEM, driver, CUDA, collectives and firmware compatibility chain.

Ownership trade: this route can avoid NVIDIA AI Enterprise and Mission Control fees, but it transfers integration, upgrade validation, incident response, security patching, recovery automation and support ownership to the operator. It is not a claim of feature parity with Mission Control.

Vendor explainers · not independent evidence

PNY Pro explains the commercial software story; verify entitlements against the order.

These vendor videos are useful orientation for NVIDIA AI Enterprise, not proof that a particular DGX or OEM quote includes a license, support term or Mission Control.

Vendor explainer · PNY Pro

Introducing NVIDIA AI Enterprise

Use PNY Pro's overview to understand the vendor's platform framing, then confirm the license metric, term and support on the order.

Vendor explainer · PNY Pro

Production AI is an entitlement question

The explainer describes the commercial platform. It does not establish a paid production entitlement or its duration for any hardware purchase.

Platform familyScale + accelerator memoryFacility envelopeBest planning useBallpark, ex VAT
H100Available 8‑GPU Hopper node
Official H100/H200 system specs →
8× H100 · 640 GB HBM3✕ No native NVFP4
Speculative fit: ≈0.74–0.82T parameters through a supported W4A16-style fallback.
10.2 kW max · air8U rack server; ordinary high-density datacenter deployment.Mature, lower-capex CUDA node for inference, adapters and evaluation.≈€263–403k / $300–460kBroad 2026 complete-system quote band.
H200Available 8‑GPU Hopper node
Official H100/H200 system specs →
8× H200 · 1.13 TB HBM3e✕ No native NVFP4
Speculative fit: ≈1.30–1.45T parameters through a supported W4A16-style fallback.
10.2 kW max · airSame DGX node envelope as H100.Memory-first Hopper choice for larger checkpoints without Blackwell migration.≈€350–438k / $400–500kBroad 2026 complete-system quote band.
B200Available 8‑GPU Blackwell node
Official B200 system specs →
8× B200 · 1.44 TB HBM3e✓ Native NVFP4
Speculative single-checkpoint fit: ≈1.65–1.85T parameters.
14.3 kW max · air1,550 CFM and 48,794 BTU/hr at the system ceiling.General Blackwell serving and post-training node.€438–569k / $500–650kPublic-reseller and integrator planning range.
B300Available 8‑GPU Blackwell Ultra node
Official B300 system specs →
8× B300 · 2.3 TB HBM3e✓ Native NVFP4
Speculative single-checkpoint fit: ≈2.65–2.95T parameters.
14.5 kW max · air49,476 BTU/hr published system ceiling.Largest current single-node memory tier before NVL72.≈€569–744k / $650–850kEditable 2026 planning band; require an exact OEM quote.
V300Derived 8‑GPU Rubin Ultra roadmap node
Inspect the V300/Kyber roadmap estimate →
8× V300 · 4.6 TB HBM4e◇ Roadmap assumption
≈5.3–5.9T parameters if this derived node retains NVFP4-class support.
≈33 kW · full DLC planningOne eighteenth of the Kyber power target; not a vendor system specification.Roadmap node comparison before deciding whether NVL72 scale is justified.≈€0.9–1.3M / $1.0–1.5MDerived planning band, not a list price or announced node.
GB200 NVL72Available 72‑GPU Blackwell rack
Official GB200 NVL72 specs →
72× Blackwell · 13.4 TB HBM3e✓ Native NVFP4
≈15.5–17.2T parameter-equivalent capacity; real services normally use replicas.
≈120 kW · direct liquidResidual air remains for networking and storage.First rack-scale tier for large distributed inference and training.€2.45–2.98M / $2.8–3.4MReported 2026 purchase-quote range.
GB300 NVL72Available 72‑GPU Blackwell Ultra rack
Official NVL72 component design →
72× B300 · 20 TB HBM3e✓ Native NVFP4
≈23–25.6T parameter-equivalent capacity; usually spent on replicas, cache and test-time compute.
Up to 142 kW · full DLCFacility CDU and rack leak-detection integration required.High-memory NVL72 reasoning, inference and post-training.€5.25–5.69M / $6.0–6.5MReported purchase quotes; lower figures are often component/BOM estimates.
Vera Rubin NVL722026 production ramp · preliminary specs
Official Vera Rubin NVL72 specs →
72× Rubin + 36× Vera · 20.7 TB HBM4✓ Native NVFP4 + 3-bit LUT path
1,580 TB/s aggregate HBM bandwidth and 260 TB/s NVLink 6 switch bandwidth.
≈240 kW comparison input · full DLC72 × 3.337 kW all-in per-GPU power used by the SemiAnalysis model; not a published rack nameplate.High-interactivity, long-context and multitrillion-parameter inference where decode bandwidth and scale-up latency dominate.Quote requiredSemiAnalysis owner-TCO assumption: $3.57/GPU-hour, not a rental rate or purchase quote.
VB300 NVL72Derived 72‑GPU V300 comparison domain
Inspect the V300/Kyber roadmap estimate →
72× V300 · 41.5 TB HBM4e◇ Roadmap assumption
≈48–53T parameter-equivalent if the derived rack retains NVFP4-class support.
≈300 kW · 800VDC + full DLCHalf-Kyber planning envelope; not an announced NVIDIA rack.Like-for-like NVL72 comparison before the jump to Kyber NVL144.≈€8.75–10.50M / $10–12MHalf-Kyber planning band with contingency; not a quote.
Kyber NVL144Rubin Ultra/V300 analyst roadmap
NVIDIA 800VDC rack-power direction →
144× V300 · 82.9 TB HBM4e◇ Roadmap assumption
≈96–106T parameter-equivalent if the roadmap system retains NVFP4-class support.
≈600 kW · 800VDC + full DLCMegawatt-class facility design around the compute rack.Frontier scale where power architecture becomes a first-order design.≈€18.38M / $21MBank of America analyst ASP estimate reported July 2026.
Early silicon · evidence checked 28 July 2026

Vera Rubin is clearly ahead—but “10×” describes one old baseline at one speed.

CoreWeave measured a pre-production Vera Rubin NVL72 rack; SemiAnalysis then normalized it against July 2026 GB200 and GB300 InferenceX results. The useful question is not whether Rubin wins, but which Blackwell software vintage, user interactivity and cost boundary you compare.

100 tok/s/user2.00× vs GB200

Rubin delivers 1,115,625 output tok/s/MW versus 558,569 on the July 2026 GB200 recipe; the corresponding GB300 ratio is 2.15×.

150 tok/s/user10.18× vs 2025 GB200

This is the headline region. Against current July 2026 recipes, the lead is 2.74× over GB200 and 2.58× over GB300—not tenfold.

300 tok/s/user5.39× vs GB300

Rubin delivers 96,446 output tok/s/MW. GB300 reaches 17,892 at its last viable frontier point; GB200 cannot reach this target.

Owner TCO · 300 tok/s/user5.03× cheaper

The model produces $3.076 per million Rubin output tokens versus $15.456 for GB300. This is modeled ownership cost—not cloud rental pricing.

Memory system20.7 TB HBM4 · 1,580 TB/s

Each Rubin GPU carries up to 288 GB and 22 TB/s. Capacity helps model and KV-cache residency; bandwidth attacks the token-by-token decode bottleneck.

Scale-up fabric260 TB/s NVLink 6

The 72-GPU rack provides 3.6 TB/s per GPU of all-to-all bandwidth. Counted writes reduce synchronization traffic for device-initiated communication.

Tensor Core path2× work per clock

Rubin doubles the K dimension handled by its matrix instruction. Fewer K-loop iterations reduce overhead in throughput-, memory- and latency-bound kernels.

MoE data movementInline TMA overrides

A shared tensor descriptor can override addresses and strides in the instruction, avoiding an in-memory descriptor rewrite when switching experts with the same layout.

Weight compression3.125 bits/weight raw

The new 3-bit LUT-B path stores an index plus an eight-entry E4M3 codebook per 512-weight block and resolves it inside the matrix operation. Accuracy depends on fitting and calibration.

Software transitionSM100 kernels can start on SM107

Rubin can reuse important Blackwell-family kernels for bring-up, unlike the Hopper-to-Blackwell break. Peak performance still needs Rubin-specific tuning; CUDA 13.4 support is a developer preview.

Reconstructed in HTML from SemiAnalysis data

Output throughput per all-in utility MW

Bars are scaled within each interactivity target; longer is better. Labels retain the exact output-token rate.

Vera Rubin · Jul 2026GB300 · Jul 2026GB200 · Jul 2026CoreWeave GB200 · 2025
100 tok/s/user
Vera Rubin1,115,6251,115,625
GB300517,848517,848
GB200558,569558,569
GB200 · 2025300,000300,000
150 tok/s/user
Vera Rubin786,208786,208
GB300305,023305,023
GB200287,289287,289
GB200 · 202577,22477,224
200 tok/s/user
Vera Rubin300,000300,000
GB30066,79966,799
GB20077,21877,218
GB200 · 202550,00050,000
250 tok/s/user
Vera Rubin140,264140,264
GB30037,04937,049
GB20036,39436,394
GB200 · 202539,12339,123
300 tok/s/user
Vera Rubin96,44696,446
GB30017,89217,892
GB200Not feasible
GB200 · 2025Not feasible
Reconstructed in HTML from SemiAnalysis TCO data

Modeled owner cost per million output tokens

Bars are scaled within each target; shorter is better. Values include prefill and decode GPUs under the source's ownership assumptions.

Vera Rubin · Jul 2026GB300 · Jul 2026GB200 · Jul 2026CoreWeave GB200 · 2025
100 tok/s/user
Vera Rubin$0.266$0.266
GB300$0.497$0.497
GB200$0.423$0.423
GB200 · 2025$0.786$0.786
150 tok/s/user
Vera Rubin$0.378$0.378
GB300$0.842$0.842
GB200$0.799$0.799
GB200 · 2025$3.046$3.046
200 tok/s/user
Vera Rubin$0.991$0.991
GB300$3.837$3.837
GB200$3.016$3.016
GB200 · 2025$4.715$4.715
250 tok/s/user
Vera Rubin$2.099$2.099
GB300$6.859$6.859
GB200$6.471$6.471
GB200 · 2025$6.013$6.013
300 tok/s/user
Vera Rubin$3.076$3.076
GB300$15.456$15.456
GB200Not feasible
GB200 · 2025Not feasible
Inspect the complete reconstructed data tables
Output tokens per second per all-in utility MW
Recipe50 tok/s100 tok/s150 tok/s200 tok/s250 tok/s300 tok/s350 tok/s
GB200 NVL72 · Jul 2026705,543558,569287,28977,21836,394Not feasibleNot feasible
GB300 NVL72 · Jul 2026707,037517,848305,02366,79937,04917,892Not feasible
CoreWeave GB200 · 2025500,000300,00077,22450,00039,123Not feasibleNot feasible
CoreWeave Vera Rubin · Jul 20261,330,0001,115,625786,208300,000140,26496,44670,703
Modeled owner cost per million output tokens
Recipe50 tok/s100 tok/s150 tok/s200 tok/s250 tok/s300 tok/s350 tok/s
GB200 NVL72 · Jul 2026$0.333$0.423$0.799$3.016$6.471Not feasibleNot feasible
GB300 NVL72 · Jul 2026$0.364$0.497$0.842$3.837$6.859$15.456Not feasible
CoreWeave GB200 · 2025$0.472$0.786$3.046$4.715$6.013Not feasibleNot feasible
CoreWeave Vera Rubin · Jul 2026$0.223$0.266$0.378$0.991$2.099$3.076$4.179
OEM alternatives

AMD or Intel host CPUs can still feed native-CUDA NVIDIA GPU servers.

The CPU vendor affects host-memory bandwidth, data preparation, storage and operational standardization. CUDA compatibility follows the NVIDIA accelerator, so all four configurations below remain native CUDA systems.

Orderable OEM serverHost + acceleratorPower + coolingGood first useBallpark, ex VAT
Dell PowerEdge XE9685LOfficial Dell configuration →2× AMD EPYC 9005 + 8× B200 · 1.44 TB HBM✓ Native CUDA + NVFP46× 3 kW PSU capacity4U direct liquid cooling; reserve residual air.B200 service node with an AMD-standardized CPU fleet.€438–569k / $500–650kAccelerator-led planning range.
HPE ProLiant Compute XD685Official HPE configurations →AMD EPYC + 8× H200, B200 or B300H200: ✕ no native NVFP4
B200/B300: ✓ native
Up to 12 high-watt PSUsDirect-liquid-cooled B200/B300 configurations; quote actual max draw.A supported AMD-host path from H200 through Blackwell Ultra.€438–744k / $500–850kWide range spans GPU generation and support.
Dell PowerEdge XE9680LOfficial Dell configuration →2× 5th-gen Intel Xeon + 8× H200 or B200H200: ✕ no native NVFP4
B200: ✓ native
6× 3 kW PSU capacity4U direct liquid cooling; reserve residual air.B200 service node for Intel-standardized estates.€438–569k / $500–650kAccelerator-led planning range.
Lenovo ThinkSystem SR680a V3Official Lenovo B200 guide →2× 5th-gen Intel Xeon + 8× B200 · 1.44 TB HBM✓ Native CUDA + NVFP43.2 kW PSU modules8U air-cooled design; validate rack airflow and redundancy.Air-cooled B200 where liquid plumbing is unavailable.€438–569k / $500–650kAccelerator-led planning range.
Capacity, training and market scale

What each hardware tier can actually do—and how it becomes a model provider.

A server is a useful product unit. A model lab is a fleet, software organization, power contract and customer pipeline. Open the card to connect hardware ownership with the organizations and services it can support.

01 · 1–8 accelerators

Research project

One workstation or node

Evaluate, fine-tune and serve existing weights; pretrain compact models; prove a bounded workflow. A failed experiment costs hours or days, not a datacenter programme.

People on the work≈2–30 directValue / market capN/A or <$50m org
02 · 1k–20k+ H100e

Independent model lab

Mistral · DeepSeek · Kimi class

Pretrain large models, run ablations and post-training, then operate an API. Mistral disclosed a 3,000-H200 run; DeepSeek’s reported H100/H800 pool is around 20,000. Moonshot’s fleet is not disclosed.

People on the work≈200–2,000 directValue / market cap≈$12–75bn lab value
03 · 600k–2m H100e

Frontier lab

xAI · Anthropic · OpenAI

Run several training and post-training programmes, generate synthetic data, absorb failed runs and serve global products. Capacity is usually spread across clouds and campuses.

People on the work≈2,000–10,000 directValue / market cap≈$0.85–1.0tn lab value
04 · 2m–5m+ parent pool

Hyperscaler-backed frontier

Amazon · Microsoft · Meta · Google · Alibaba

The parent pool feeds cloud tenants, first-party models, recommenders and other AI. Alibaba’s total is not separately disclosed; the whole pool is never one model run.

People on the work≈10k–100k+ cross-stackValue / market cap≈$0.29–3.9tn parent
Owned shapeWhat it can actually doService realityOrganization it supports
1–8 acceleratorsWorkstation to one nodeEvaluation, LoRA/Q‑LoRA, small-model pretraining and one production replica of a model that fits.Pilot to ≈1,200 heavy usersModel and node dependent; benchmark the exact build.Research group, internal platform team or specialist service with cloud overflow.
72–144 linked acceleratorsOne NVL72/NVL144-class domainLarge distributed inference, multiple replicas, long-context cache, RL and substantial post-training.≈2,500–80,000+ heavy usersA planning range, not a vendor benchmark.Regional model cloud or a major enterprise AI platform; the facility is now part of the product.
1k–20k+ H100eMany racks and training domainsPretrain competitive large open models, maintain several model lines and operate a public API.Provider-scale, still capacity constrainedTraining and inference compete for the same fleet.Mistral/DeepSeek/Kimi-class lab with dedicated infra, capital and model operations.
600k–2m H100e accessibleMulti-campus and multi-cloudParallel frontier programmes, large synthetic-data and RL pipelines, global inference and rapid retraining.Global consumer + enterprise productsNo single training job consumes the whole estate.Frontier lab backed by hyperscaler contracts, very large financing and energy commitments.
2m–5m+ H100e ownedParent pool across regions and productsRent compute to other labs, train first-party models, run internal AI and absorb multi-year silicon supply.Cloud platform + model catalogueCustomer workloads and first-party teams compete for allocations.Amazon/AWS, Microsoft/Azure, Google and Alibaba Cloud; Meta has the owner scale without a public cloud.
Rent, customize or build

AWS, Azure, Google Cloud and Alibaba Cloud are both model shelves and infrastructure routes.

A hyperscaler can sell a managed API, provision dedicated throughput, rent a training cluster and train its own model family. Open the card to compare the four routes.

AWS · AMAZON

Bedrock + SageMaker AI

Use Bedrock for managed access to hundreds of foundation models; move to SageMaker AI and HyperPod for customization, distributed training and controlled deployment.

First-party models
Amazon Nova 2, Nova Sonic and Nova multimodal embeddings
Other providers
OpenAI and open/model-partner catalogues, Anthropic (CSP accounts not supported)
Silicon route
NVIDIA GPU plus Amazon Trainium and Inferentia
Owner pool≈2.45m H100eParent market cap≈$2.54tn · 24 Jul 2026
AZURE · MICROSOFT

Microsoft Foundry + Azure OpenAI

Use Azure-hosted model APIs with standard or provisioned deployment; move to managed compute and Azure ML when the workload needs custom weights, capacity or training control.

First-party models
MAI, Phi, Model Router and healthcare AI models
Other providers
OpenAI, Cohere, DeepSeek, Meta, Mistral, xAI, Anthropic (CSP accounts not supported) and other models.
Infrastructure route
Azure GPU estates, Foundry managed compute and Azure ML
Owner pool≈3.42m H100eParent market cap≈$2.84tn · 24 Jul 2026
GOOGLE CLOUD · ALPHABET

Gemini Enterprise Agent Platform

Use the platform for managed Gemini and partner APIs, provisioned throughput, tuning and agent workflows; use Managed Training, GKE, Cloud TPU or NVIDIA GPU capacity when the workload needs infrastructure control.

First-party models
Gemini, open-weight Gemma, Imagen, Veo and Google embeddings
Other providers
Meta, Mistral, Qwen, Anthropic (CSP accounts not supported) and other Gemini Enterprise Agent Platform Model Garden options
Silicon route
Cloud TPU including Trillium/Ironwood, plus NVIDIA GPU estates
Owner pool≈5.05m H100eParent market cap≈$3.89tn · 24 Jul 2026
ALIBABA CLOUD · QWEN

Model Studio + PAI

Use Model Studio for OpenAI-compatible Qwen and partner APIs; use Platform for AI for dedicated deployment, fine-tuning and distributed training through DLC, EAS and Lingjun resources.

First-party models
Qwen text, reasoning, coding and multimodal families; Wan media models
Other providers
DeepSeek, Kimi and GLM availability varies by region
Open-weight route
Deploy Qwen weights with vLLM/SGLang or buy a managed Qwen API
Owner poolNot separately publicParent market cap≈$0.29tn · 24 Jul 2026
Best public comparison · latest like-for-like baseline is end-2025

Lab-access compute, parent-owner pools and a fleet-cost proxy.

Open the card for a logarithmic or linear comparison of accessible compute and parent-owned pools. “H100e” compares peak AI operations with an NVIDIA H100 and does not equal real model throughput.

Start with the logarithmic order-of-magnitude view, then switch to linear to make equal chart width mean equal H100e—and therefore an equal step in the simple fleet-cost proxy. This is not an inventory audit. Solid bars estimate operational compute a lab could access; striped bars mark a disclosed run, floor or bound; outlined bars are parent-owned pools.

Axis spacing
Log view: each equal step is roughly ten times more accessible compute.
Mistral AIOPEN WEIGHTSFleet proxy $0.08–0.15bn
≥3kobserved H200 run
Moonshot / KimiOPEN WEIGHTSFleet cost not public
Fleet not disclosed
DeepSeekOPEN WEIGHTSFleet proxy $0.5–1.0bn
≈20kreported H100 + H800
Qwen / AlibabaHYPERSCALER MODEL LABFleet cost not isolated
Qwen allocation not disclosed
Meta AI / MSLOPEN + FRONTIER LABFleet cost not isolated
<1.7mlab bound; parent pool below
Frontier API labs
xAIFRONTIERFleet proxy $15–35bn
≈600–700klab access
AnthropicFRONTIERFleet proxy ≥$25–50bn
≥1mlab access
OpenAIFRONTIERFleet proxy $43–85bn
≈1.7mlab access
Google DeepMindFRONTIER LABFleet cost not isolated
<2mlab bound; parent pool below
Hyperscaler parent owner pools · not one lab allocation
Alibaba CloudCLOUD OWNERFleet cost not public
Not split out · China total ≈1.16m ex-smuggling
MetaPARENT OWNERFleet proxy $58–115bn
≈2.30mQ4 2025 owner pool
Amazon / AWSCLOUD OWNERFleet proxy $61–122bn
≈2.45mQ4 2025 owner pool
Microsoft / AzureCLOUD OWNERFleet proxy $85–171bn
≈3.42mQ4 2025 owner pool
GoogleCLOUD + MODEL OWNERFleet proxy $126–252bn
≈5.05mQ4 2025 owner pool

Read DeepSeek correctly. The ≈20,000 figure is accelerator-class chips, not 20,000 eight-GPU servers. Its V3 paper documents one 2,048-H800 training cluster and 2.788M H800 GPU-hours; wider fleet figures and additional H20/A100 hardware are not converted into the plotted number.

Read Qwen and Alibaba separately. Qwen is Alibaba’s model family; Alibaba Cloud is the infrastructure and API parent. Neither Qwen’s allocation nor Alibaba’s owner total is public. Epoch estimates identified Chinese owners together at ≈1.16m H100e in Q4 2025 before its optional smuggling estimate, so assigning that number to Alibaba would be wrong.

Public bridges to scale

Use public access, capital and demand before carrying the full infrastructure risk.

The European Union and United States offer different ladders from research access to provider-scale infrastructure. Open the card to compare the regional routes and their limits.

European Union
Shared compute, scale-up finance and sovereign demand
02 · FINANCE
Bridge startup to infrastructure company

TechEU says it will deploy €70bn of EIB Group equity, loans and guarantees in 2025–2027 to mobilize €250bn with partners across the innovation lifecycle, including AI and digital infrastructure. It is a financing channel, not an automatic subsidy.

03 · BUILD
Finance the leap to gigafactory scale

InvestAI plans to mobilize €20bn for several AI Gigafactories, with the EIB exploring advisory support and loans. The target shape is about 100,000 advanced chips per site. The formal call is due in summer 2026 and first construction is scheduled for 2027: this is future capacity, not online supply.

04 · DEMAND
Turn sovereignty into a buying criterion

The Commission’s proposed Cloud and AI Development Act would streamline sites, energy and finance, create an EU sovereignty framework and establish common public-sector procurement. That can make trusted European capacity easier to specify and buy; it does not guarantee a contract.

United States
Research access, competitive funding, sites and procurement
01 · ACCESS
Use the national research resource

The NSF-led National Artificial Intelligence Research Resource connects US researchers, educators, startups and small businesses to public and private compute, models, data and expertise. By March 2026 it had supported more than 600 research teams and 6,000 students.

02 · FUND
Fund efficiency and regional capacity

NSF-backed STRIDE can invest up to $21m over two years in deployment-ready AI infrastructure efficiency, with awards up to $3.5m. NSF Regional Innovation Engines can fund regional technology coalitions up to $160m over ten years. Both are competitive programmes, not project debt.

03 · BUILD
Shorten the path to 100 MW+ campuses

A federal order defines qualifying AI data center projects as requiring more than 100 MW of new load and directs expedited federal permitting. DOE has also selected four federal sites for private AI data center and energy partnerships. This lowers siting friction; it does not supply a fleet or guaranteed grid capacity.

04 · DEMAND
Use federal procurement as a first market

OMB M-25-22 directs more efficient federal AI acquisition, while GSA’s OneGov consolidates technology buying and reported 20 vendor agreements by April 2026. The route can accelerate agency adoption, but eligibility is limited and no programme guarantees a provider contract.

Inference fit is not training fit. Full training holds weights, gradients, optimizer states and activations, often using several times the inference memory before checkpoint and data-pipeline overhead. LoRA and Q‑LoRA are much lighter. Use the hardware table to shortlist a tier, the scale comparison to understand the organization around it, and the simulator below for the workload, quote and payback target you actually expect.

Try before you buy

Rent the exact workload before committing to private hardware.

Before a five- or six-figure private-cloud commitment, rent the candidate stack long enough to discover whether the model, service shape and operating cost actually fit. A rental is a test drive, not a GPU-availability, price or hardware-validation guarantee.

Users + ROI simulator

Turn a vendor quote into a price per user—or per token.

Start from the model’s evidence-matched serving preset or lock a platform you already own. The planner packs repeatable replica cells into that platform, scales the packed loadout toward a user target and keeps every generated assumption editable.

The model preset will choose and pack its evidence-matched platform.

Cost + sales preset

The platform choice loads quote, power and colocation starting points. Add costs omitted from those three fields below; financing is not modeled.

Fits selected system

90 GB service build in 1.44 TB HBM, with 15% reserved for runtime headroom.

Pricing view
Monthly operating cashSelected private billing minus modeled monthly operating cost.
Simple capex paybackCapex divided by positive monthly operating cash; excludes financing and demand ramp.
Gap to capital-recovery targetMonthly billing minus operating cost and the selected capital-recovery allowance.
Selected private quote / user / monthSame-model OpenRouter cost × the selected private quote multiplier.
Same-model OpenRouter / user / monthSelected local model’s market rate × workload.
Operating break-even / served user / monthCovers the costs explicitly modeled here, without capital recovery.
Target-payback price / served user / monthCovers modeled operating cost plus capex recovery inside the selected target.
Per-user gap to target-payback priceSelected private quote minus the target-payback price, per served user per month.
Estimated supported usersConstrained by the tighter of token throughput, concurrency and per-request memory.
Auto-sized loadouts for targetCalculated from the capacity of one packed loadout.
Selected API alternative / user / monthAPI rate × uncached input, cached input and output.
Selected private quote multiplierThe commercial ask currently applied to the same-model OpenRouter anchor.
Quote multiplier required for targetIn-range results are rounded upward to the next 0.25× slider step so the recommendation never understates the target.
Modeled monthly operating cost
Monthly private billing at selected price
Selected API cost for served target

Change any assumption to test the scenario.

Compatibility

CUDA compatibility is a binary boundary, not a performance grade.

Only NVIDIA hardware runs CUDA binaries. AMD and Intel can still be excellent inference platforms, but the application, kernels and quantization must support their native stack.

PlatformRuns CUDA binaries?Native stackPortability routeProvider decision
NVIDIA Ampere · A100✓ Native CUDACUDA · NCCL · NVLink
✕ No native NVFP4
Use BF16/FP16 or supported integer paths. Some model cards document loading an NVFP4 checkpoint through a fallback.Treat NVFP4 as model-specific weight compatibility here—not native FP4 throughput.
NVIDIA Hopper · H100/H200✓ Native CUDACUDA · NCCL · NVLink · TensorRT‑LLM
✕ No native NVFP4
FP8 is the native low-precision path. Selected NVFP4 checkpoints can fall back to W4A16.Buy for mature CUDA and FP8—not for native NVFP4 throughput.
NVIDIA Blackwell · GB10/B200/B300/GB200/GB300 and RTX PRO✓ Native CUDACUDA · NCCL · NVLink · TensorRT‑LLM
✓ Native NVFP4
Native W4A4 path when the engine, kernel and model support it.Current launch target for NVFP4; validate GB10 and workstation-specific kernels separately from datacenter Blackwell.
AMD MI300X/MI325X/MI355X and Radeon↻ No—port and rebuildROCm · HIP · RCCL · Infinity FabricHIP source can target AMD or NVIDIA, but one compiled binary cannot run on both.Strong memory economics when the exact engine, model and kernels pass a pilot.
Intel Gaudi 3✕ No—separate stackSynapseAI · Habana libraries · RoCEUse Gaudi-supported frameworks and Optimum Habana; CUDA binaries do not transfer.Treat as its own product lane with an explicit supported-model catalogue.
Intel GPU / XPU↻ No—source migrationoneAPI · SYCL · OpenVINOSYCLomatic can assist CUDA-source migration; manual work and validation remain.Useful for selected Intel-XPU workloads, not a Gaudi substitute.
tinygrad runtime◇ Backend choiceIts own small compiler with CUDA, AMD, NV, Metal and other backendsA tinygrad program can target different backends; tinygrad does not make AMD or Intel CUDA-compatible.Use for owned compiler research and ports, not assumed vLLM feature parity.
tinygrad route

Own more of the stack, accept more integration work.

tinygrad’s useful idea is not “cheap CUDA.” It is a small, MIT-licensed compiler/runtime and reference hardware path that makes the stack inspectable, portable and open to operator modification.

Reference lab · AMD

tinybox red v2

€10,502 / $12,000 ex VAT

4× Radeon 9070 XT, 64 GB aggregate GPU memory, 128 GB system memory and 2 TB NVMe. Best for compiler work, AMD ports and smaller sharded models—not a 400B+ production endpoint. Confirm store stock before planning delivery.

Reference lab · NVIDIA

tinybox green v2 Blackwell

€65,640 / $75,000 ex VAT

4× RTX PRO 6000 Blackwell and 384 GB aggregate GPU memory. It runs CUDA, but PCIe-attached workstation GPUs are not an NVLink HBM supernode; test throughput, concurrency and current store stock.

Software ownership

Port only where it earns control

€0 / $0 software licence

Use tinygrad to understand and own selected kernels or models. Keep vLLM or SGLang for production until the port meets the same correctness, batching, observability and recovery tests.

Before promising hardware, ask what must stay compatible. Examples include an OpenAI-compatible endpoint; vLLM, SGLang, NIM or Triton; PyTorch or JAX; tool calling and structured JSON; MCP or Hermes-style agent harnesses; embeddings, reranking, vision, PDF processing or fine-tuning; plus SSO, audit logs, retention, data residency, context, concurrency and latency requirements.

A client may not know it has a CUDA dependency. An NVIDIA container, TensorRT‑LLM, FlashAttention, bitsandbytes, NCCL or a PyTorch CUDA wheel is one—even if the business owner never says “CUDA.” Ask for the current container, lockfile and a representative acceptance test.

01

Hardware pluralism

Keep the option to buy NVIDIA, AMD or commodity cards where the workload and kernels allow it.

02

Inspectability

A smaller compiler makes generated kernels and backend behavior easier for an operator to understand and modify.

03

Acceptance tests first

Port the client’s real model and critical software path before making a platform-wide compatibility promise.

04

Honest maturity

tinygrad describes itself as alpha and says it is not faster than PyTorch for most use cases yet. Treat it as an engineering lane, not magic production parity.

Capacity ladder

Match the checkpoint first, then reserve serving headroom.

These are procurement starting points, not throughput guarantees. Keep 15–25% aggregate memory for runtime overhead and KV cache, then load-test the actual context length and concurrency.

✕ No native NVFP4 · 640 GB–1.13 TB

H100 / H200

Use native FP8, or a checkpoint-specific W4A16 fallback when memory capacity matters. Speculative mixed-NVFP4 weight capacity is ≈0.74–1.45T parameters across an eight-GPU node, but it does not gain native W4A4 compute.

✓ Native NVFP4 · 1.44–2.3 TB

B200 / B300

About 1.65–2.95T parameters by conservative mixed-checkpoint memory math. Nemotron Ultra, Qwen, DeepSeek and Mistral Large are more realistic first services; leave the rest for cache, replicas and concurrency.

✓ Native NVFP4 · 13.4–20 TB

GB200 / GB300 NVL72

Roughly 15.5–25.6T parameter-equivalent weight capacity. In practice, spend that headroom on Kimi K3-scale sharding, multiple replicas, long-context cache and test-time compute.

◇ Roadmap assumption · 4.6–82.9 TB

V300 / VB300 / Kyber

The 5.3–106T NVFP4-equivalent range is compounded planning arithmetic across derived or roadmap systems—not a shipped support matrix, benchmark or useful single-model target.

Price method

One exchange rate, visible caveats.

USD prices were converted at the ECB reference rate for 20 July 2026: €1 = $1.1426. EUR hardware figures are rounded to the nearest euro; token prices are rounded to the nearest euro cent.

Excluded from every price

VAT, sales tax, shipping, import duties, rack, networking, storage, facility upgrades, energy, support and integration unless the row says otherwise.

Ballpark means planning input

Each row says whether its number comes from reported purchase quotes, a component/BOM estimate or a roadmap ASP estimate. Replace it with your vendor quote before making a return case.

Requote before buying

GPU allocations, support terms and discounts move quickly. Keep the accelerator, memory, interconnect, cooling and software acceptance criteria fixed while comparing quotes.

Provider default

Launch one permissive model on one validated NVIDIA reference system with one datacenter partner. Add a second accelerator stack only when its memory economics survive the cost of maintaining a second software product.