19 + 3released model lines + pending weightsGLM-5.3 + ML4 PREVIEWS; BEAM ANNOUNCED
2 customfrontier licences need reviewQWEN3.8 + K3 LARGE-SCALE MAAS TRIGGERS
14-600 kWNVIDIA system planning range8-GPU SERVER TO RUBIN ULTRA ROADMAP RACK
Two-minute guide
Which private AI deployment route should you choose?
This page is collapsed into decision-sized sections. If the workload itself is not proven, start with one measured adoption workflow; if the gap may need retrieval or changed weights, use the model-intervention chooser before sizing infrastructure.
Deploy now, validate next, or retain only as a baseline.
Cards are ordered by present recommendation, not chronology or parameter count. Current launch means a 2026 release with usable access and a credible deployment route. Current option marks a specialized, superseded or review-gated line. Older baseline marks a pre-2026 release retained for compatibility or comparison, not as the default. Mistral Large 4 and GLM-5.3 have hosted access with weights pending; Reflection Beam is announced with early-access signup. Qwen3.8-Max and Kimi K3 require their custom commercial terms to be reviewed. The image and video cards are a short open-weight leader set, not a mirror of WaveSpeed's hosted catalogue.
✓ current launch · preferred current release△ current option · secondary or specialized! review gate or hosted preview · incomplete route↻ current self-host · newer hosted model exists↓ older baseline · retained, not default
✓ Current launchMiniMax Community License
MiniMax M3
A current launch choice for frontier coding and agents. A native-multimodal 428B-total / 23B-active MoE with a 1M-token model limit and downloadable weights. The repositories total about 795.5 GiB for BF16, 413.3 GiB for MXFP8 and 232.9 GiB for NVIDIA's NVFP4 build. The calculator uses a measured 4× B200 FP4/MTP profile; keep the official 8× B200 recipe as the compatibility fallback and validate the exact four-GPU build before quoting.
First node
Calculator: 4× B200 measured cell; official fallback: 8× B200
The published large-node route is mature enough to trial: vLLM verified M3 on H200, GB200, B300 and AMD MI300/MI350 systems, and InferenceX now publishes a current 4× B200 FP4/MTP throughput sweep. The official NVFP4 recipe is newer and does not yet prove one-Station comfort. Rent the reference shape, lock an acceptance set, then repeat it on the quoted workstation.
Qwen has released the self-hostable 2.4T-total / 95B-active MoE in full and block-scaled FP8 packages: about 4.89 TB and 2.50 TB of model files respectively. The open checkpoint is text-only, requires thinking mode, has a native 262,144-token context and can extend to 1,010,000. Do not treat it as a byte-for-byte copy of hosted Qwen3.8-Max, which adds vision input, non-thinking mode, built-in tools and a default 1M context.
A 30B-total / 3B-active hybrid Mamba-2, attention and MoE model for high-volume agent execution. The official NVFP4 weights total about 20.1 GiB before runtime and cache; NVIDIA documents one DGX Spark or one H100 as deployment paths, alongside Jetson, GeForce RTX 5090 and hosted/data-centre routes. Its one-million-token model limit is not a practical context promise for every device.
First node
1× DGX Spark for the documented local recipe; 1× H100 for data-centre serving
Serve with
vLLM; broader ecosystem support still needs exact-version tests
A fully open 120B total / 12B active hybrid MoE for multi-agent work, with weights, datasets, recipes and a 1M-token model limit. Its native NVFP4 build loads on one B200; a matched 8K/1K vLLM run provides a stronger starting profile than a generic model-size estimate.
NVIDIA’s 550B total / 55B active frontier orchestration model. The official mixed-precision NVFP4 checkpoint is about 352.3 GB and supports up to 1M context. Blackwell runs native W4A4; Hopper lacks native FP4 Tensor Cores and automatically uses a W4A16 fallback.
The official Flash release supersedes the Preview checkpoint. Its 166.9 GB mixed FP4/FP8 repository includes the DSpark speculative module, supports 1M context and up to 384K output, and exposes low, high and max reasoning effort. DeepSeek reports large agentic-benchmark gains, including results above V4 Pro Preview, but its code-agent scores use a model-specific harness and two listed tests are internal, so validate the gain on your own stack.
First node
Official reference: 4× GB300; profile other layouts
Moonshot’s released 2.8T-parameter multimodal MoE has 104B active parameters, selects 16 of 896 experts, supports a 1M context and uses MXFP4 weights with MXFP8 activations. Its 96-shard weight index totals 1.56 TB before runtime and cache headroom.
First node
8× B300 or 8× MI350X/MI355X
Serve with
SGLang, vLLM or TokenSpeed; validate recipes
Licence trigger
Separate deal for large Model-as-a-Service operators
A still-current alternative now ranked behind Qwen3.8. Multimodal MoE with 397B total / 17B active parameters. The calculator now uses NVIDIA's roughly 251 GB NVFP4 repository and a measured 4× B300 SGLang/MTP profile at 8K input / 1K output, replacing the much larger 807 GB BF16 planning route.
A post-trained successor that uses the same base model as GLM-5.2. Z.ai reports large gains on long-horizon coding and agent benchmarks, but the launch status is split: Coding Plan access is live, the general API is marked coming soon and weights are targeted for release within two weeks of 14 August. Keep GLM-5.2 for self-hosted planning until the 5.3 checkpoint, licence and engine recipes can be inspected.
A 753B model with up to 1M context. The BF16 repository is about 1.51 TB and needs at least 2 TB aggregate HBM with useful headroom; the calculator instead uses the released FP8 build and a matched 8× B200 profile.
! Public API previewWeights pending · checked 6 October 2026
Mistral Large 4 / Le Chonk
Mistral's 6 October preview is natively multimodal: roughly 1T total parameters (1.05T in the model documentation), 49B active and a 1M-token context. Mistral says it trained the model from scratch on 3,800 Grace Blackwell GPUs in its European datacentres and serves the preview on that infrastructure.
Vendor benchmark claims: Mistral reports leading US/European open-weight performance, strength in cybersecurity, finance, law and manufacturing, and visual grounding above closed frontier models. Inspect the named tasks and evaluation settings before treating those claims as workload results.
Available now
Public preview API via Mistral Studio
Weights
Announced for the end of October 2026
Self-hosting gate
Inspect the released licence, checkpoint and serving recipe first
! Announced · early-access signupApache 2.0 planned · checked 6 October 2026
Reflection Beam / 501B-A23B
Reflection's 5 October announcement describes a sparse MoE with 501B total and 23B active parameters for coding, reasoning and agentic work. Its score table is vendor-reported; claimed inference efficiency still needs reproduction on the released build.
Available now
Early-access signup; final red-teaming and evaluation underway
Weights + licence
Apache 2.0 release promised during October 2026
Deployment gate
Technical report, model card, runtime and hardware recipe pending
A current compact fallback rather than the catalogue default; validate its vendor comparisons on your own tasks. 119B total / 6.5B active, multimodal, 256K context and 24-language support. Its official 70.8 GB NVFP4 build is a clean Blackwell fit; SGLang documents TP1 on B200/B300, while FP8 needs TP2 on H100/H200. Native W4A4 acceleration requires Blackwell.
△ Current image specialistTencent Hunyuan Community License
HunyuanImage 3.0
A self-hostable 80B-total / 13B-active autoregressive image model, with base, instruction-following and eight-step distilled instruction checkpoints. The base and instruction repositories are each about 169 GB. Tencent recommends at least 3×80 GB GPUs for the base text-to-image model and at least 8×80 GB for the instruction model, so this is a cloud or multi-GPU server service rather than a personal-machine fit.
First node
Base: ≥3×80 GB; Instruct: ≥8×80 GB in the creator recipe
Serve with
Official custom Transformers code; validate the documented vLLM acceleration path
Licence gate
Custom Tencent terms; retain the licence and usage-policy review
Qwen's December image checkpoint is a credible server shortlist entry because the weights are released and its model card reports more than 10,000 blind AI Arena comparisons, where it ranked as the strongest open-source image model. The repository is about 57.7 GB. A validated vLLM-Omni recipe uses 2× B200 and measures about 99 GB peak VRAM in BF16 or 84-86 GB with tested low-precision paths, making server deployment the honest default rather than pretending this is an ordinary 24 GB local load.
Validated service cell
2× B200 in the published vLLM-Omni recipe
Serve with
Diffusers for reference; vLLM-Omni for measured serving
Black Forest Labs positions this 32B open-weight model as its maximum-quality route for text-to-image, single-reference editing and multi-reference editing. The full reference implementation needs an H100-equivalent GPU. Quantized execution can use a 24 GB RTX 4090 only by moving the text encoder to a remote service, so a fully self-hosted installation belongs in the server catalogue. The licence is non-commercial and requires the documented input or output filtering path.
First full local service
1× H100-equivalent GPU in the creator reference
Serve with
Official FLUX.2 code or Diffusers; retain required filters
The A14B text-to-video and image-to-video checkpoints are the server half of the Wan 2.2 family. They generate 480p or 720p video and use a diffusion mixture-of-experts design. Alibaba reports leading open and closed model results on its evaluation, but that remains a creator claim to reproduce. The official single-GPU route requires at least 80 GB VRAM; use the 5B 24 GB checkpoint on the local-model page when access matters more than the larger model.
First node
1× 80 GB GPU for the official single-GPU route
Serve with
Official Wan2.2 code, Diffusers or a validated ComfyUI graph
△ Current audio-video specialistMiniMax H3 Community License
MiniMax H3
A 33B dense model that generates synchronized video and 32 kHz stereo audio from text, images, video and audio references. The official H3-Base service produces 4-15-second clips at 768p and 24 fps; MiniMax's full 2K workflow still calls hosted Context-IR and Regenerate-2K APIs. The reference SGLang deployment uses four GPUs, so server or cloud infrastructure is the honest default.
First service cell
4× H200 in the verified resident SGLang recipe
Serve with
SGLang Diffusion; treat the one-Spark Sol Engine path as a specialist exception
Licence gate
Community terms exclude the EU, UK, United States and South Korea
A downloadable trillion-parameter coding and agent model with 32B active parameters, native INT4 and 256K context. The calculator uses a matched 8× B200 vLLM run; the modified terms, nightly runtime and model-specific kernels still need review.
A prior-year reasoning baseline retained for comparison and existing deployments. A 685B FP8 reasoning and agent model with a 690 GB repository. The card documents vLLM and SGLang; production tool parsing still needs your own robustness tests.
A prior-year Mistral baseline with downloadable weights; Large 4 is a newer hosted preview with weights pending. Large 3 has 675B total / 41B active parameters, multimodal input and 256K context. Mistral documents FP8 on one 8× H200 or B200 node and an NVFP4 checkpoint on one 8× H100 or A100 node. H100 and A100 do not have native FP4 Tensor Cores, so that older-GPU route depends on model-specific fallback or dequantization, not native NVFP4 W4A4 acceleration.
Retained as an older compatibility and footprint baseline; start with current candidates unless this exact package wins your acceptance tests. About 117B total / 5.1B active, tool-capable and packaged in MXFP4 to fit one 80 GB H100. Minimum fit is not the same as an economical service profile: the calculator uses a measured 2× B200, 8K/1K TensorRT-LLM run so packed capacity is grounded in observed throughput.
First service cell
2× B200 for the measured capacity profile; H100 fits the build
Commercially usable does not mean restriction-free. Before listing any model, retain the exact licence and notice files, scan model code, document provenance limits, red-team the served build and define an abuse policy. Qwen3.8’s custom licence requires the notice to travel with copies; commercial products or services above 100 million monthly active users or $20 million monthly revenue must display the model name, while Model-as-a-Service or AI Work Assistant businesses above $50 million aggregate revenue in any consecutive 12 months need a separate Qwen licence. MiniMax M3 is also open-weight under custom terms, not Apache 2.0 or MIT: commercial deployments require visible attribution and a one-time notice, and products or services above the licence's $20M annual-revenue threshold require prior written authorization. Architecture, quantization, context and concurrency determine the real memory requirement.
NVIDIA platform ladder
Start with platform scale, then see what it takes to become a provider.
The ten current and roadmap tiers keep node and rack choices legible. The provider-scale view then connects those purchases to research, open-model labs, frontier fleets and the EU and US public routes that can help bridge the gaps.
Infrastructure evidence · checked 5 October 2026
AI agents need CPUs as well as accelerators.
Agent tools add CPU, memory and storage work around model inference. Compare current server supply constraints, published platform capacity and workload measurements before sizing a service.
An agent repeatedly runs a model, calls tools and returns the results to the model. Large-model reasoning usually runs on a GPU or another accelerator; host software and many tools run on CPUs. The model can choose the next action while the host enforces permissions and executes it. More agent activity can therefore require more of both kinds of compute.
1 · Model inference
Choose the next step.
The model generates a response or tool request. Accelerator memory, bandwidth and compute shape large-model serving capacity.
2 · Host control
Schedule and authorize.
CPU software manages requests, state and tool permissions. This control program is distinct from the model's reasoning.
3 · Tool execution
Do the external work.
Browsers, compilers and tests can consume CPU and memory. Other tools use accelerators or wait on remote APIs and storage.
4 · Return and repeat
Feed results back.
The next model call may wait for a tool. Parallel agents add bursts, queues and idle intervals across the service.
Read supply data, platform capacity and market forecasts separately.
Signal
Published evidence
Planning limit
Server CPU supply pressure
Intel's Q2 2026 filing reports server volume up 9% year over year and demand exceeding available supply because of internal constraints. Intel quarterly filing →
Supplier-specific supply data. Confirm delivery and pricing for each product; the filing does not isolate how much demand comes from agents.
22,500+ environments per Vera CPU rack
NVIDIA describes a 256-CPU rack supporting more than 22,500 independent CPU environments for agent tools. This is a vendor capacity claim. NVIDIA Vera specifications →
Successful-task throughput and latency still need measurement. Workload, memory, storage and accelerator capacity constrain the complete service.
15-100 times more compute
On NVIDIA's 26 August 2026 earnings call, Jensen Huang estimates much greater compute demand for agents, depending on the problem. NVIDIA earnings transcript →
An executive estimate about compute broadly. CPU core counts and CPU-to-GPU ratios need workload-specific benchmarks.
A $220 billion server CPU market in 2030
AMD's 4 August 2026 earnings call forecasts this approximate market size and expects agent sandboxes to become its largest segment. AMD earnings transcript →
A company forecast with uncertain adoption and economics. Service sizing needs measured demand, throughput and costs.
Tool time is not the same as CPU utilization.
A CPU-focused preprint reports tool processing at 84.5-90.6% of runtime in selected Haystack retrieval experiments. That is a result for those workloads, not an average across all agents. A separate study of production traces and four agent frameworks finds bursty CPU demand and idle capacity between tool calls. Waiting for a remote API can lengthen a task while consuming little local CPU.
Size the complete loop before buying more hardware.
Measure model time, tool time, CPU active time, memory pressure, network waits and queueing at the concurrency you intend to serve. Track typical and slow-task latency, successful completions and cost per successful task. Test bounded tool concurrency, separate capacity for expensive builds, and safe caching before adding capacity to the measured bottleneck.
A simple speed limit: in a hypothetical sequential task with 20% model inference and 80% tool time, making inference twice as fast leaves 90% of the original runtime: only about 1.11× overall speedup. The tool portion might be CPU work, storage or a remote service; this arithmetic alone cannot identify the hardware to buy. Carry the measured service profile into the cost comparison →
The ten-name map
Five node tiers. Four NVL72 racks. One NVL144 rack.
Read the node class first, then the rack class. Open the card for the NVIDIA platform map, videos, NVFP4 guidance, prices, Rubin analysis and OEM alternatives.
H100 through GB300 are shipping systems; Vera Rubin NVL72 is in its production ramp with preliminary published specifications. V300 uses the later Rubin Ultra roadmap's 576 GB-per-GPU assumption; VB300 NVL72 is a 72-GPU comparison domain derived from half a Kyber rack, not an announced NVIDIA SKU. Kyber NVL144 remains the actual Rubin Ultra roadmap rack label.
8-GPU node · HopperH100 → H200
The mature entry tier. H200 keeps the same node shape and raises accelerator memory from 640 GB to 1.13 TB.
8-GPU node · newerB200 → B300 → V300
More memory and newer low-precision engines; V300 is a derived roadmap planning node rather than an orderable SKU.
72-GPU NVLink rackGB200 → GB300 → Vera Rubin → VB300
Vera Rubin is the current production-ramp generation. VB300 is a later, derived Rubin Ultra comparison, not the name of Vera Rubin NVL72.
144-GPU roadmap rackKyber NVL144
The Rubin Ultra/V300 roadmap tier changes rack density, power delivery and price class again.
From number format to facility
See the platform NVIDIA is describing, then translate the spectacle into a bill of materials.
These vendor videos connect NVFP4, Vera, Rubin, DSX facilities, a scientific workload and the full GTC keynote. Use them to understand the intended system shape; use the tables below for procurement questions, facility constraints and explicit planning caveats.
Low precision · NVIDIA Developer
NVFP4 can make large models smaller, but the GPU matters
NVFP4 stores most values in four bits and uses fine-grained scaling to preserve more accuracy than a crude four-bit conversion. The catch is simple: native W4A4 acceleration starts with Blackwell Tensor Cores. Hopper can use selected checkpoints through a W4A16 fallback, but that is not the same speed path.
Platform overview · NVIDIA
Vera Rubin is a full system, not just a GPU
Use the overview to see how NVIDIA frames CPUs, GPUs, networking and software as one agent platform. The delivered configuration still needs an exact vendor quote and acceptance test.
Host architecture · NVIDIA
The CPU still shapes the service
Vera is NVIDIA's host-side story for agent systems. Translate that promise into memory bandwidth, data movement, storage and software requirements for the workload you will actually serve.
Facility blueprint · NVIDIA
At rack scale, the building joins the stack
DSX makes the facility-level ambition visible. Power delivery, cooling, networking, operations and recovery are part of the product long before a gigawatt becomes relevant.
Workload context · NVIDIA
Start with the scientific job, not the rack
The discovery story is a useful demand-side counterweight to hardware spectacle. Define the models, data, latency and evaluation first; only then choose the infrastructure tier.
Full keynote · NVIDIA
See the complete platform story in one sitting
The full GTC 2026 keynote connects Vera Rubin, agents, networking and AI factories in NVIDIA’s own long-form narrative. Use it for context, then return to the evidence tables for procurement decisions.
NVFP4 in plain English
A promising format with one hard boundary: native acceleration needs newer NVIDIA GPUs.
NVFP4 is not a universal “make any model four times faster” switch. It reduces stored weights and memory traffic, but the checkpoint, serving engine, kernels, context and GPU generation must all agree.
Commercial platform path · checked 10 August 2026
Hardware is only one layer of the NVIDIA commercial stack.
Open the stack to see which layers are software products, which are included system software and which require a separate commercial entitlement.
Layered architecture · links open NVIDIA documentationOne possible commercial stack, not a bundled bill of materials.
Products can be selected independently, and support or entitlement depends on the exact order and supported configuration.
Outcomes and workloads
ApplicationsAgentsModel APIsAnalyticsTraining
Application development
Build, customize and prepare workloads
Developer-facing tools; availability may depend on the chosen software entitlement.
Ownership trade: this route can avoid NVIDIA AI Enterprise and Mission Control fees, but it transfers integration, upgrade validation, incident response, security patching, recovery automation and support ownership to the operator. It is not a claim of feature parity with Mission Control.
PNY Pro explains the commercial software story; verify entitlements against the order.
These vendor videos are useful orientation for NVIDIA AI Enterprise, not proof that a particular DGX or OEM quote includes a license, support term or Mission Control.
Vendor explainer · PNY Pro
Introducing NVIDIA AI Enterprise
Use PNY Pro's overview to understand the vendor's platform framing, then confirm the license metric, term and support on the order.
Vendor explainer · PNY Pro
Production AI is an entitlement question
The explainer describes the commercial platform. It does not establish a paid production entitlement or its duration for any hardware purchase.
144× V300 · 82.9 TB HBM4e◇ Roadmap assumption ≈96-106T parameter-equivalent if the roadmap system retains NVFP4-class support.
≈600 kW · 800VDC + full DLCMegawatt-class facility design around the compute rack.
Frontier scale where power architecture becomes a first-order design.
≈€18.38M / $21MBank of America analyst ASP estimate reported July 2026.
These are ballparks, not list prices or model-fit promises. NVFP4 capacity bands reserve 20-25% of advertised accelerator memory and assume roughly 5.0-5.2 effective bits per stored parameter, reflecting mixed-precision checkpoints rather than pure four-bit arithmetic. They exclude unusually large embeddings, multimodal towers and cache. Hopper H100 and H200 retain their cited complete-system quote bands. Available Blackwell B200 through GB300 prices start from broad complete-system or reported rack quote bands and add a 15% forward-order contingency. That contingency is a planning choice, not a claim that every signed order reprices: Bloomberg reported on 22 August 2026 that some of Nvidia's largest customers were told many early-2027 server shipments would rise by more than 15%, with the change varying by chip generation and memory configuration. The report specifically included Grace Blackwell and Vera Rubin systems; it does not support applying the same uplift to Hopper. V300, VB300 and Kyber keep their explicitly labelled roadmap arithmetic. Actual draw follows utilization and power caps; network, storage, CDU/pumps, PUE overhead, spares, deployment and support are additional. Every price excludes VAT.
Vera Rubin is clearly ahead, but “10×” describes one old baseline at one speed.
CoreWeave measured a pre-production Vera Rubin NVL72 rack; SemiAnalysis then normalized it against July 2026 GB200 and GB300 InferenceX results. The useful question is not whether Rubin wins, but which Blackwell software vintage, user interactivity and cost boundary you compare.
100 tok/s/user2.00× vs GB200
Rubin delivers 1,115,625 output tok/s/MW versus 558,569 on the July 2026 GB200 recipe; the corresponding GB300 ratio is 2.15×.
150 tok/s/user10.18× vs 2025 GB200
This is the headline region. Against current July 2026 recipes, the lead is 2.74× over GB200 and 2.58× over GB300, not tenfold.
300 tok/s/user5.39× vs GB300
Rubin delivers 96,446 output tok/s/MW. GB300 reaches 17,892 at its last viable frontier point; GB200 cannot reach this target.
Owner TCO · 300 tok/s/user5.03× cheaper
The model produces $3.076 per million Rubin output tokens versus $15.456 for GB300. This is modeled ownership cost, not cloud rental pricing.
Memory system20.7 TB HBM4 · 1,580 TB/s
Each Rubin GPU carries up to 288 GB and 22 TB/s. Capacity helps model and KV-cache residency; bandwidth attacks the token-by-token decode bottleneck.
Scale-up fabric260 TB/s NVLink 6
The 72-GPU rack provides 3.6 TB/s per GPU of all-to-all bandwidth. Counted writes reduce synchronization traffic for device-initiated communication.
Tensor Core path2× work per clock
Rubin doubles the K dimension handled by its matrix instruction. Fewer K-loop iterations reduce overhead in throughput-, memory- and latency-bound kernels.
MoE data movementInline TMA overrides
A shared tensor descriptor can override addresses and strides in the instruction, avoiding an in-memory descriptor rewrite when switching experts with the same layout.
Weight compression3.125 bits/weight raw
The new 3-bit LUT-B path stores an index plus an eight-entry E4M3 codebook per 512-weight block and resolves it inside the matrix operation. Accuracy depends on fitting and calibration.
Software transitionSM100 kernels can start on SM107
Rubin can reuse important Blackwell-family kernels for bring-up, unlike the Hopper-to-Blackwell break. Peak performance still needs Rubin-specific tuning; CUDA 13.4 support is a developer preview.
Reconstructed in HTML from SemiAnalysis data
Output throughput per all-in utility MW
Bars are scaled within each interactivity target; longer is better. Labels retain the exact output-token rate.
Decision rule: use the July 2026 GB200 and GB300 recipes for a current buy/no-buy comparison; keep the 2025 GB200 line only as a software-maturity lesson. Re-run the exact model, context distribution, concurrency, service-level target and power boundary before procurement. The charts above are original HTML/CSS reconstructions of the published values, with unavailable frontier points shown as “not feasible.”
AMD or Intel host CPUs can still feed native-CUDA NVIDIA GPU servers.
The CPU vendor affects host-memory bandwidth, data preparation, storage and operational standardization. CUDA compatibility follows the NVIDIA accelerator, so all four configurations below remain native CUDA systems.
Air-cooled B200 where liquid plumbing is unavailable.
€504-655k / $575-750kAccelerator-led range plus a 15% forward-order contingency.
Capacity, training and market scale
What each hardware tier can actually do, and how it becomes a model provider.
A server is a useful product unit. A model lab is a fleet, software organization, power contract and customer pipeline. Open the card to connect hardware ownership with the organizations and services it can support.
01 · 1-8 accelerators
Research project
One workstation or node
Evaluate, fine-tune and serve existing weights; pretrain compact models; prove a bounded workflow. A failed experiment costs hours or days, not a datacenter programme.
People on the work≈2-30 directValue / market capN/A or <$50m org
02 · 1k-20k+ H100e
Independent model lab
Mistral · DeepSeek · Kimi class
Pretrain large models, run ablations and post-training, then operate an API. Mistral disclosed a 3,000-H200 run; DeepSeek’s reported H100/H800 pool is around 20,000. Moonshot’s fleet is not disclosed.
People on the work≈200-2,000 directValue / market cap≈$12-75bn lab value
03 · 600k-2m H100e
Frontier lab
xAI · Anthropic · OpenAI
Run several training and post-training programmes, generate synthetic data, absorb failed runs and serve global products. Capacity is usually spread across clouds and campuses.
People on the work≈2,000-10,000 directValue / market cap≈$0.85-1.0tn lab value
04 · 2m-5m+ parent pool
Hyperscaler-backed frontier
Amazon · Microsoft · Meta · Google · Alibaba
The parent pool feeds cloud tenants, first-party models, recommenders and other AI. Alibaba’s total is not separately disclosed; the whole pool is never one model run.
People on the work≈10k-100k+ cross-stackValue / market cap≈$0.29-3.9tn parent
Owned shape
What it can actually do
Service reality
Organization it supports
1-8 acceleratorsWorkstation to one node
Evaluation, LoRA/Q-LoRA, small-model pretraining and one production replica of a model that fits.
Pilot to ≈1,200 heavy usersModel and node dependent; benchmark the exact build.
Research group, internal platform team or specialist service with cloud overflow.
Large distributed inference, multiple replicas, long-context cache, RL and substantial post-training.
≈2,500-80,000+ heavy usersA planning range, not a vendor benchmark.
Regional model cloud or a major enterprise AI platform; the facility is now part of the product.
1k-20k+ H100eMany racks and training domains
Pretrain competitive large open models, maintain several model lines and operate a public API.
Provider-scale, still capacity constrainedTraining and inference compete for the same fleet.
Mistral/DeepSeek/Kimi-class lab with dedicated infra, capital and model operations.
600k-2m H100e accessibleMulti-campus and multi-cloud
Parallel frontier programmes, large synthetic-data and RL pipelines, global inference and rapid retraining.
Global consumer + enterprise productsNo single training job consumes the whole estate.
Frontier lab backed by hyperscaler contracts, very large financing and energy commitments.
2m-5m+ H100e ownedParent pool across regions and products
Rent compute to other labs, train first-party models, run internal AI and absorb multi-year silicon supply.
Cloud platform + model catalogueCustomer workloads and first-party teams compete for allocations.
Amazon/AWS, Microsoft/Azure, Google and Alibaba Cloud; Meta has the owner scale without a public cloud.
People and value are scale context, not qualification rules. “People on the work” estimates direct model, product, platform, silicon and datacenter effort, not every employee in the parent company. Independent labs are valued by private funding rounds; only listed parents have a market cap. July 2026 reference points include Mistral at 900+ employees, DeepSeek at an implied ≈$52bn, OpenAI at ≈$852bn and Anthropic at ≈$965bn.
AWS, Azure, Google Cloud and Alibaba Cloud are both model shelves and infrastructure routes.
A hyperscaler can sell a managed API, provision dedicated throughput, rent a training cluster and train its own model family. Open the card to compare the four routes.
AWS · AMAZON
Bedrock + SageMaker AI
Use Bedrock for managed access to hundreds of foundation models; move to SageMaker AI and HyperPod for customization, distributed training and controlled deployment.
First-party models
Amazon Nova 2, Nova Sonic and Nova multimodal embeddings
Other providers
OpenAI and open/model-partner catalogues, Anthropic (CSP accounts not supported)
Use Azure-hosted model APIs with standard or provisioned deployment; move to managed compute and Azure ML when the workload needs custom weights, capacity or training control.
First-party models
MAI, Phi, Model Router and healthcare AI models
Other providers
OpenAI, Cohere, DeepSeek, Meta, Mistral, xAI, Anthropic (CSP accounts not supported) and other models.
Infrastructure route
Azure GPU estates, Foundry managed compute and Azure ML
Use the platform for managed Gemini and partner APIs, provisioned throughput, tuning and agent workflows; use Managed Training, GKE, Cloud TPU or NVIDIA GPU capacity when the workload needs infrastructure control.
First-party models
Gemini, open-weight Gemma, Imagen, Veo and Google embeddings
Other providers
Meta, Mistral, Qwen, Anthropic (CSP accounts not supported) and other Gemini Enterprise Agent Platform Model Garden options
Silicon route
Cloud TPU including Trillium/Ironwood, plus NVIDIA GPU estates
Use Model Studio for OpenAI-compatible Qwen and partner APIs; use Platform for AI for dedicated deployment, fine-tuning and distributed training through DLC, EAS and Lingjun resources.
First-party models
Qwen text, reasoning, coding and multimodal families; Wan media models
Other providers
DeepSeek, Kimi and GLM availability varies by region
Open-weight route
Deploy Qwen weights with vLLM/SGLang or buy a managed Qwen API
Best public comparison · latest like-for-like baseline is end-2025
Lab-access compute, parent-owner pools and a fleet-cost proxy.
Open the card for a logarithmic or linear comparison of accessible compute and parent-owned pools. “H100e” compares peak AI operations with an NVIDIA H100 and does not equal real model throughput.
Start with the logarithmic order-of-magnitude view, then switch to linear to make equal chart width mean equal H100e, and therefore an equal step in the simple fleet-cost proxy. This is not an inventory audit. Solid bars estimate operational compute a lab could access; striped bars mark a disclosed run, floor or bound; outlined bars are parent-owned pools.
Estimated lab accessObserved run / estimate boundParent owner poolNot disclosed
Axis spacing
Mistral AIOPEN WEIGHTSFleet proxy $0.08-0.15bn
≥3kobserved H200 run
Moonshot / KimiOPEN WEIGHTSFleet cost not public
Fleet not disclosed
DeepSeekOPEN WEIGHTSFleet proxy $0.5-1.0bn
≈20kreported H100 + H800
Qwen / AlibabaHYPERSCALER MODEL LABFleet cost not isolated
Qwen allocation not disclosed
Meta AI / MSLOPEN + FRONTIER LABFleet cost not isolated
<1.7mlab bound; parent pool below
Frontier API labs
xAIFRONTIERFleet proxy $15-35bn
≈600-700klab access
AnthropicFRONTIERFleet proxy ≥$25-50bn
≥1mlab access
OpenAIFRONTIERFleet proxy $43-85bn
≈1.7mlab access
Google DeepMindFRONTIER LABFleet cost not isolated
<2mlab bound; parent pool below
Hyperscaler parent owner pools · not one lab allocation
Alibaba CloudCLOUD OWNERFleet cost not public
Not split out · China total ≈1.16m ex-smuggling
MetaPARENT OWNERFleet proxy $58-115bn
≈2.30mQ4 2025 owner pool
Amazon / AWSCLOUD OWNERFleet proxy $61-122bn
≈2.45mQ4 2025 owner pool
Microsoft / AzureCLOUD OWNERFleet proxy $85-171bn
≈3.42mQ4 2025 owner pool
GoogleCLOUD + MODEL OWNERFleet proxy $126-252bn
≈5.05mQ4 2025 owner pool
1k10k100k1m6m H100e
Read DeepSeek correctly. The ≈20,000 figure is accelerator-class chips, not 20,000 eight-GPU servers. Its V3 paper documents one 2,048-H800 training cluster and 2.788M H800 GPU-hours; wider fleet figures and additional H20/A100 hardware are not converted into the plotted number.
Read Qwen and Alibaba separately. Qwen is Alibaba’s model family; Alibaba Cloud is the infrastructure and API parent. Neither Qwen’s allocation nor Alibaba’s owner total is public. Epoch estimates identified Chinese owners together at ≈1.16m H100e in Q4 2025 before its optional smuggling estimate, so assigning that number to Alibaba would be wrong.
Ownership is not availability. Epoch AI’s Q4 2025 owner estimates are Google 5.05m, Microsoft 3.42m, Amazon 2.45m and Meta 2.30m H100e, but those pools also serve cloud customers and internal products. Its July 2026 campus directory provides a newer site check without a like-for-like company total.
Use public access, capital and demand before carrying the full infrastructure risk.
The European Union and United States offer different ladders from research access to provider-scale infrastructure. Open the card to compare the regional routes and their limits.
European Union
Shared compute, scale-up finance and sovereign demand
01 · ACCESS
Borrow the first large run
EuroHPC’s open industrial call offers AI Factory allocations above 50,000 GPU-hours with a stated ten-working-day approval target. Nineteen AI Factories and thirteen antennas are being assembled to give startups and SMEs compute plus technical support.
TechEU says it will deploy €70bn of EIB Group equity, loans and guarantees in 2025-2027 to mobilize €250bn with partners across the innovation lifecycle, including AI and digital infrastructure. It is a financing channel, not an automatic subsidy.
InvestAI plans to mobilize €20bn for several AI Gigafactories, with the EIB exploring advisory support and loans. The target shape is about 100,000 advanced chips per site. The formal call is due in summer 2026 and first construction is scheduled for 2027: this is future capacity, not online supply.
The Commission’s proposed Cloud and AI Development Act would streamline sites, energy and finance, create an EU sovereignty framework and establish common public-sector procurement. That can make trusted European capacity easier to specify and buy; it does not guarantee a contract.
The standard grant ceiling is below €2.5M (€2,499,999), covering up to 70% of eligible innovation costs, usually over 24 months. Budget the remaining costs. Optional €1M-€10M EIC Fund investment uses equity or quasi-equity, including convertible loans; that component can dilute ownership. Coaching, mentors and investor connections are also available.
EU and Horizon Europe associated-country startups and SMEs can apply, as can individuals intending to establish an eligible company. Third-country applicants must establish or relocate the company before the full proposal. UK applicants qualify for grant-only support. Grant-only support is available once per legal entity during 2021-2027.
Accelerator Open accepts any technology field. It targets high-risk, scalable innovation: complete TRL 5 validation in a relevant environment before applying, then fund TRL 6-8 development. AI, hardware, robotics, climate and biotech can fit; an idea alone does not.
Start with a 12-page form, up to ten slides and a three-minute team video. Submit anytime; short proposals are batched on the first Tuesday each month at 17:00 Brussels time. Feedback normally takes 4-6 weeks. A GO lets you prepare the full proposal; successful applicants proceed to the EIC jury interview, followed by the grant agreement and investment due diligence where applicable.
Next short-proposal batch: 3 November 2026, 17:00 Brussels (18:00 Tallinn). The next 2026 full-proposal batch is 4 November, 17:00 Brussels. These are separate stages: a November short proposal cannot normally receive feedback in time for the next day's full-proposal batch. Check the next work programme and portal before scheduling a later full submission.
Research access, competitive funding, sites and procurement
01 · ACCESS
Use the national research resource
The NSF-led National Artificial Intelligence Research Resource connects US researchers, educators, startups and small businesses to public and private compute, models, data and expertise. By March 2026 it had supported more than 600 research teams and 6,000 students.
NSF-backed STRIDE can invest up to $21m over two years in deployment-ready AI infrastructure efficiency, with awards up to $3.5m. NSF Regional Innovation Engines can fund regional technology coalitions up to $160m over ten years. Both are competitive programmes, not project debt.
A federal order defines qualifying AI data center projects as requiring more than 100 MW of new load and directs expedited federal permitting. DOE has also selected four federal sites for private AI data center and energy partnerships. This lowers siting friction; it does not supply a fleet or guaranteed grid capacity.
OMB M-25-22 directs more efficient federal AI acquisition, while GSA’s OneGov consolidates technology buying and reported 20 vendor agreements by April 2026. The route can accelerate agency adoption, but eligibility is limited and no programme guarantees a provider contract.
Inference fit is not training fit. Full training holds weights, gradients, optimizer states and activations, often using several times the inference memory before checkpoint and data-pipeline overhead. LoRA and Q-LoRA are much lighter. Use the hardware table to shortlist a tier, the scale comparison to understand the organization around it, and the simulator below for the workload, quote and payback target you actually expect.
Try before you buy
Rent the exact workload before committing to private hardware.
Before a five- or six-figure private-cloud commitment, rent the candidate stack long enough to discover whether the model, service shape and operating cost actually fit. A rental is a test drive, not a GPU-availability, price or hardware-validation guarantee.
Run the acceptance pack first. Test the exact model, engine, quantization, context, concurrency and acceptance harness you would buy for; retain the container, settings, logs, latency, throughput, memory and task-quality results.
Referral disclosure: this isaiuseful.com link is a Runpod referral link. As checked 10 August 2026, eligible new first-time users must sign up through it with Google SSO and load their first $10: European users receive $5 credit, while non-European users receive a weighted $5-$500 credit (Runpod says most are $10 or less). Terms can change. If eligible, isaiuseful.com receives its referral bonus and earns Runpod credits on actual usage for the first six months: 3% of Pod spend and 5% of Serverless spend. Using the link supports the site.
Take the workload and ROI calculators into the physical world. Build a small pilot or a rack-scale factory, then explore the power, heat, networking, resilience and ownership cost around the GPUs.
Roadmap scenarios include Kyber Rubin Ultra NVL144, one 144-GPU rack with Vera host processing, and Rubin Ultra NVL576, eight 72-GPU racks linked by copper and optical NVLink. Supporting POD quantities, power, memory, prices and floor area are planning allowances. NVIDIA scale-up architecture, checked 8 September 2026 →
Two Sparks to a rack-scale factory
Build the floor. See the consequences.
Pack equipment into 48U racks with a visible power limit. Give each datacenter its own inventory. Storage, CPU workers and networking scale alongside the GPUs, and a facility-area estimate includes support racks, electrical rooms and cooling plant. Follow requests, responses, optical circuits, power and cooling between sites.
Published hardware specifications and editable planning allowances. No signup or external 3D engine. JavaScript enables the interactive workbench.
Explore 200, 400 and 800 Gb/s switches, KV cache tiers and CPU workers. Compare lit WAN capacity with fiber distance, gateway limits and shared-route failures. B300 AC and 54 VDC variants, high-voltage sidecars and facility 800 VDC conversion have distinct infrastructure needs.
Start with two Sparks.
racknex sells a 1.33U tray for two DGX Sparks and their power supplies. The planner reserves 2U and 480 W from two 240 W supply ratings; actual draw depends on the workload. The pair link and distributed memory need a supported software recipe.
NVIDIA's B300 user guide lists 14.5 kW; its facility guide uses 15 kW for the AC version and an estimated 19.7 kW peak. This planner starts at 15 kW, then adds supporting IT and design reserve. Airflow and coolant circulation use simple heat-balance equations.
A larger site does not automatically require 800 VDC. Check the ordered rack, its power shelves or sidecar and the facility distribution. NVIDIA's Vera Rubin POD connects five specialized rack roles; the planner's five-role preset is a teaching example, not the full POD specification.
Turn a vendor quote into a price per user, or per token.
Start from the model’s evidence-matched serving preset or lock a platform you already own. The planner packs repeatable replica cells into that platform, scales the packed loadout toward a user target and keeps every generated assumption editable.
Results paused, correct the highlighted assumptions.
Fits selected system
90 GB service build in 1.44 TB HBM, with 15% reserved for runtime headroom.
Pricing view
Monthly operating cash-Selected private billing minus modeled monthly operating cost.Simple capex payback-Capex divided by positive monthly operating cash; excludes financing and demand ramp.Gap to capital-recovery target-Monthly billing minus operating cost and the selected capital-recovery allowance.
Selected private quote / user / month-Same-model OpenRouter cost × the selected private quote multiplier.Same-model OpenRouter / user / month-Selected local model’s market rate × workload.Operating break-even / served user / month-Covers the costs explicitly modeled here, without capital recovery.Target-payback price / served user / month-Covers modeled operating cost plus capex recovery inside the selected target.Per-user gap to target-payback price-Selected private quote minus the target-payback price, per served user per month.Estimated supported users-Constrained by the tighter of token throughput, concurrency and per-request memory.Auto-sized loadouts for target-Calculated from the capacity of one packed loadout.Selected API alternative / user / month-API rate × uncached input, cached input and output.Selected private quote multiplier-The commercial ask currently applied to the same-model OpenRouter anchor.Quote multiplier required for target-In-range results are rounded upward to the next 0.25× slider step so the recommendation never understates the target.
Modeled target tokens / month-Input + output tokens for the smaller of target users and available capacity.Required blended price / 1M-Recovers operating cost and capex inside the selected target.Same-model OpenRouter / 1M-Blended using this workload’s input/output mix.Selected API alternative / 1M-Blended using the editable cache and output mix.Private price in this scenario / 1M-OpenRouter benchmark × selected private-deployment premium.
Results paused, correct the highlighted assumptions.
Calculated capacity by tier
Supported users within the available hardware, in allocation order.
Developers--
PM / PO research--
Chat--
Calculating fleet
The requested mix is checked against input and output throughput, active requests and per-request memory.
Hardware units required-The tighter of throughput and memory determines the fleet.Available input capacity-Prefill-equivalent input after the modeled prefix reuse allowance.Available output capacity-At selected serving utilization across available units.Available active-request capacity-Validated concurrency capped by the minimum per-request output rate.Target API cost / month-Selected input/output mix; tool and regional charges excluded.Required-fleet capex-Required units × quoted unit cost.Modeled target-recovery local / month-Three-year capex recovery plus energy at 1.25 PUE and €0.14/kWh, colocation, and 5% annual support. Financing, tax and unentered costs are excluded.
Change any workload or platform assumption to rebalance the fleet.
CUDA compatibility is a binary boundary, not a performance grade.
Only NVIDIA hardware runs CUDA binaries. AMD and Intel can still be excellent inference platforms, but the application, kernels and quantization must support their native stack.
Platform
Runs CUDA binaries?
Native stack
Portability route
Provider decision
NVIDIA Ampere · A100
✓ Native CUDA
CUDA · NCCL · NVLink ✕ No native NVFP4
Use BF16/FP16 or supported integer paths. Some model cards document loading an NVFP4 checkpoint through a fallback.
Treat NVFP4 as model-specific weight compatibility here, not native FP4 throughput.
NVIDIA Hopper · H100/H200
✓ Native CUDA
CUDA · NCCL · NVLink · TensorRT-LLM ✕ No native NVFP4
FP8 is the native low-precision path. Selected NVFP4 checkpoints can fall back to W4A16.
Buy for mature CUDA and FP8, not for native NVFP4 throughput.
NVIDIA Blackwell · GB10/B200/B300/GB200/GB300 and RTX PRO
✓ Native CUDA
CUDA · NCCL · NVLink · TensorRT-LLM ✓ Native NVFP4
Native W4A4 path when the engine, kernel and model support it.
Current launch target for NVFP4; validate GB10 and workstation-specific kernels separately from datacenter Blackwell.
AMD MI300X/MI325X/MI355X and Radeon
↻ No, port and rebuild
ROCm · HIP · RCCL · Infinity Fabric
HIP source can target AMD or NVIDIA, but one compiled binary cannot run on both.
Strong memory economics when the exact engine, model and kernels pass a pilot.
Intel Gaudi 3
✕ No, separate stack
SynapseAI · Habana libraries · RoCE
Use Gaudi-supported frameworks and Optimum Habana; CUDA binaries do not transfer.
Treat as its own product lane with an explicit supported-model catalogue.
Intel GPU / XPU
↻ No, source migration
oneAPI · SYCL · OpenVINO
SYCLomatic can assist CUDA-source migration; manual work and validation remain.
Useful for selected Intel-XPU workloads, not a Gaudi substitute.
tinygrad runtime
◇ Backend choice
Its own small compiler with CUDA, AMD, NV, Metal and other backends
A tinygrad program can target different backends; tinygrad does not make AMD or Intel CUDA-compatible.
Use for owned compiler research and ports, not assumed vLLM feature parity.
Feature support changes faster than the silicon. As of 26 July 2026, vLLM documents NVIDIA CUDA, AMD ROCm and Intel XPU routes, but feature parity and quantization support vary. NVFP4 is an NVIDIA format; native W4A4 acceleration is a Blackwell-generation capability, while Hopper uses a model- and engine-specific fallback where supported. Pin the container, driver, firmware, engine commit and model revision used in acceptance testing.
Own more of the stack, accept more integration work.
tinygrad’s useful idea is not “cheap CUDA.” It is a small, MIT-licensed compiler/runtime and reference hardware path that makes the stack inspectable, portable and open to operator modification.
Reference lab · AMD
tinybox red v2
€10,502 / $12,000 ex VAT
4× Radeon 9070 XT, 64 GB aggregate GPU memory, 128 GB system memory and 2 TB NVMe. Best for compiler work, AMD ports and smaller sharded models, not a 400B+ production endpoint. Confirm store stock before planning delivery.
4× RTX PRO 6000 Blackwell and 384 GB aggregate GPU memory. It runs CUDA, but PCIe-attached workstation GPUs are not an NVLink HBM supernode; test throughput, concurrency and current store stock.
Use tinygrad to understand and own selected kernels or models. Keep vLLM or SGLang for production until the port meets the same correctness, batching, observability and recovery tests.
Before promising hardware, ask what must stay compatible. Examples include an OpenAI-compatible endpoint; vLLM, SGLang, NIM or Triton; PyTorch or JAX; tool calling and structured JSON; MCP or Hermes-style agent harnesses; embeddings, reranking, vision, PDF processing or fine-tuning; plus SSO, audit logs, retention, data residency, context, concurrency and latency requirements.
A client may not know it has a CUDA dependency. An NVIDIA container, TensorRT-LLM, FlashAttention, bitsandbytes, NCCL or a PyTorch CUDA wheel is one, even if the business owner never says “CUDA.” Ask for the current container, lockfile and a representative acceptance test.
Keep the option to buy NVIDIA, AMD or commodity cards where the workload and kernels allow it.
02
Inspectability
A smaller compiler makes generated kernels and backend behavior easier for an operator to understand and modify.
03
Acceptance tests first
Port the client’s real model and critical software path before making a platform-wide compatibility promise.
04
Honest maturity
tinygrad describes itself as alpha and says it is not faster than PyTorch for most use cases yet. Treat it as an engineering lane, not magic production parity.
Capacity ladder
Match the checkpoint first, then reserve serving headroom.
These are procurement starting points, not throughput guarantees. Keep 15-25% aggregate memory for runtime overhead and KV cache, then load-test the actual context length and concurrency.
✕ No native NVFP4 · 640 GB-1.13 TB
H100 / H200
Use native FP8, or a checkpoint-specific W4A16 fallback when memory capacity matters. Speculative mixed-NVFP4 weight capacity is ≈0.74-1.45T parameters across an eight-GPU node, but it does not gain native W4A4 compute.
✓ Native NVFP4 · 1.44-2.3 TB
B200 / B300
About 1.65-2.95T parameters by conservative mixed-checkpoint memory math. Nemotron Ultra, Qwen, DeepSeek and Mistral Large are more realistic first services; leave the rest for cache, replicas and concurrency.
✓ Native NVFP4 · 13.4-20 TB
GB200 / GB300 NVL72
Roughly 15.5-25.6T parameter-equivalent weight capacity. In practice, spend that headroom on Kimi K3-scale sharding, multiple replicas, long-context cache and test-time compute.
◇ Roadmap assumption · 4.6-82.9 TB
V300 / VB300 / Kyber
The 5.3-106T NVFP4-equivalent range is compounded planning arithmetic across derived or roadmap systems, not a shipped support matrix, benchmark or useful single-model target.
Price method
One exchange rate, visible caveats.
USD prices were converted at the ECB reference rate for 20 July 2026: €1 = $1.1426. EUR hardware figures are rounded to the nearest euro; token prices are rounded to the nearest euro cent.
Excluded from every price
VAT, sales tax, shipping, import duties, rack, networking, storage, facility upgrades, energy, support and integration unless the row says otherwise.
Ballpark means planning input
Available Blackwell bands include a 15% forward-order contingency over the cited quote basis. Hopper and roadmap estimates remain unchanged. Replace every estimate with your vendor quote before making a return case.
Requote before buying
GPU allocations, support terms and discounts move quickly. Keep the accelerator, memory, interconnect, cooling and software acceptance criteria fixed while comparing quotes.
Launch one permissive model on one validated NVIDIA reference system with one datacenter partner. Add a second accelerator stack only when its memory economics survive the cost of maintaining a second software product.