NVIDIA H200 Systems
NVIDIA H200 Servers
The H200 is an H100 with substantially more memory and bandwidth per GPU — 141GB of HBM3e against 80GB. For training that changes little architecturally, but for inference it changes the economics: a model that needed two H100s to hold its weights and KV cache will often fit on one H200, which halves the node count for the same served throughput.
That makes H200 the platform to price against H100 when the workload is serving rather than training. We quote both so the comparison is on your numbers, not a vendor benchmark.
11 configurations below · quotes returned within 48 business hours · purchase orders accepted
NVIDIA H200 Servers we configure and supply
Every system below is quoted to your workload — accelerator count, CPU, memory, storage and fabric are specified together rather than sold as a fixed SKU.

H200 GPU Server
Expanded memory and bandwidth for the largest models and most demanding workloads.
- Run larger models per node
- Higher inference throughput
- Fewer nodes, less fabric

NVIDIA DGX H200
Eight H200 GPUs with 1,128GB of HBM3e — the memory-heavy Hopper DGX.
- More memory per GPU
- Same operational model
- Fewer nodes for inference

Dell PowerEdge XE9680 — 8x NVIDIA H200 SXM (HGX, 141GB HBM3e)
Dell's flagship 8-GPU HGX H200 node for frontier-scale AI training.
- Frontier-scale single node
- Vendor-backed reliability
- Cluster-ready fabric

Supermicro SYS-821GE-TNHR — 8x H200 SXM with BlueField-3 DPUs
8U HGX H200 SuperServer with 1:1 BlueField-3 networking for HPC at scale.
- True 1:1 GPU-to-NIC
- DPU-offloaded networking
- Serviceable density

HPE Cray XD670 — 8x H200 SXM, Direct Liquid Cooled
Liquid-cooled 5U HGX H200 node engineered for dense, efficient training rows.
- Higher density per rack
- Improved power efficiency
- Supercomputing fabric

Lenovo ThinkSystem SR685a V3 — 8x H200 SXM, AMD EPYC 9005
AMD EPYC Turin paired with 8x H200 for high-throughput generative AI.
- Massive CPU headroom
- Generative AI throughput
- Lenovo serviceability

Supermicro 4U — 4x H200 NVL PCIe (4-Way NVLink)
Four NVLink-bridged H200 NVL GPUs for high-memory inference in standard racks.
- 564GB pooled HBM3e
- Standard-rack friendly
- Right-sized for inference

Dell PowerEdge XE7745 — 4x H200 NVL PCIe Inference Node
Dell-supported 4x H200 NVL node tuned for enterprise inference deployment.
- Enterprise-managed inference
- Pooled high-memory serving
- Flexible deployment

Nexus H200 32-GPU Training Cluster (4× Dell PowerEdge XE9680)
141GB-per-GPU H200 nodes for training the largest models on fewer machines.
- More model per node
- Enterprise-platform reliability
- Integrated and accepted

Nexus H200 128-GPU SuperPOD (16× HGX H200 Nodes)
A full-scale 128-GPU H200 pod for frontier training and large-scale HPC.
- Reference-architecture certainty
- Frontier capacity, one machine
- End-to-end accountability

Nexus H200 32-GPU Inference Cluster with Parallel Storage
Low-latency 32-GPU H200 serving with a GPUDirect parallel storage backbone.
- Built for low latency
- Storage that keeps GPUs busy
- Highly available serving
What we need to quote accurately
A configuration quote takes minutes when these are known and days of back-and-forth when they are not. You do not need all of them to start — send what you have.
- The workload: model sizes, training or inference, and expected concurrency
- Node count, and whether this is a single system or a scaling cluster
- Rack power and cooling available per rack, and inlet temperature
- Existing estate — the vendor and management tooling you already run
- Network fabric: InfiniBand, Ethernet, and the speed you are standardised on
- Timeline, and whether the budget is approved or being built
How we quote
- 1
Send the requirement
An email, a bill of materials, or a rough description of the workload. All three work.
- 2
We validate the configuration
Accelerator, chassis, fabric, power and cooling checked against each other before anything is priced.
- 3
Itemised quote within 48 hours
Line-by-line pricing, current availability and lead time, with warranty terms stated.
H200 server — buyer questions
Is the H200 worth the premium over the H100?
For inference on large models, usually — the deciding factor is whether the extra memory removes a GPU from every node. If your model and KV cache already fit comfortably on an H100, the premium buys you headroom you may not use. If you are currently sharding across two GPUs to fit, an H200 can consolidate that. Send us the model and context length and we will quote both configurations side by side.
Which H200 platforms do you supply?
8-GPU SXM platforms from Dell, Supermicro, HPE and Lenovo, including direct liquid-cooled variants, and 4-GPU H200 NVL PCIe nodes for inference. The right one usually follows from what your data centre can cool and which vendor your operations team already runs.
Do H200 systems need liquid cooling?
Not necessarily, but an 8-GPU SXM node is a serious thermal load and many air-cooled rooms cannot take one per rack at full utilisation. We ask what your rack power and cooling budget is before recommending a chassis, because the honest answer is sometimes that air-cooled is fine and sometimes that it is not.
Compare with other platforms
NVIDIA H100 Systems
NVIDIA H100 Servers
View systemsNVIDIA Blackwell Systems
NVIDIA B200 & GB200 Blackwell Servers
View systemsDell PowerEdge
Dell PowerEdge XE9680 GPU Servers
View systemsBrowse the full range of GPU servers and AI clusters, or the component catalogue for drives, memory, optics and spares.
