
H100 vs B200: When Blackwell Is Worth It
A cross-generation comparison for buyers adding capacity. What Blackwell genuinely changes, why an existing H100 estate is often the cheaper answer, and the rack power and cooling work that has to be scoped first.
Blackwell has been shipping long enough that the useful question has changed. It is no longer whether a B200 outruns an H100 — it plainly does — but whether that margin is worth what it costs you in facility work, in schedule, and in the operational habits your team has already built around Hopper. If you are adding incremental capacity to an estate that already runs H100, the honest answer is frequently no. If you are starting a new build, or your serving tier is bandwidth-starved, it is usually yes. This is a guide to telling those two situations apart before the purchase order goes out.
What Blackwell actually changes
The H100 SXM5 is a single Hopper die carrying 80GB of HBM3 at 3.35 TB/s, rated up to 700W. The B200 is a different kind of object: two reticle-limit Blackwell dies joined by a high-bandwidth die-to-die link and presented to software as one GPU. In HGX B200 form it carries 180GB of HBM3e per GPU — 1.4TB across an eight-GPU baseboard, against 640GB on an HGX H100 — at close to 8 TB/s of memory bandwidth. NVLink doubles, from 900 GB/s per GPU on Hopper to 1.8 TB/s on Blackwell. Blackwell's tensor cores and its second-generation Transformer Engine add FP4 and FP6, which Hopper has no equivalent for.
Those are large numbers. What matters is which of them your workload can actually spend.
Where the specifications translate, and where they do not
Memory bandwidth is the gain that converts most reliably. Autoregressive decode in LLM inference is bandwidth-bound almost by definition — the weights are read out of HBM for every token generated — so the step from 3.35 TB/s to roughly 8 TB/s shows up directly in decode throughput, moderated by batch size and by how well your serving stack keeps the GPU busy. If your production bottleneck is serving throughput, this is the strongest single argument for Blackwell, and it does not require you to change anything about your model.
FP4 is a different matter. It can reduce the memory footprint of weights and raise throughput again on top of the bandwidth gain, but only if your serving stack supports it and your accuracy budget survives the quantisation. That is engineering work, not a hardware property. Treat FP4 as upside you have to earn rather than headroom you have bought.
Training gains are the least predictable of the three. A distributed training job that is communication-bound, dataloader-bound or checkpoint-bound will not deliver the headline speedup, because the GPU was never the constraint. If your H100s are already idling while they wait on storage or the network, Blackwell will let them idle faster. Measure utilisation on the estate you have before you price the estate you want.
Why an existing H100 estate is often the better buy
There is a bias in hardware procurement toward buying the newest thing, and it is worth naming. When the requirement is more of what you already run, homogeneity has real, unglamorous value that rarely appears in a comparison table.
- Identical nodes schedule as one pool. A scheduler that can treat every node as interchangeable is simpler to run and wastes less capacity than one juggling two hardware classes.
- Your driver, container and framework stack is already validated on Hopper. Adding matching nodes is a change with almost no software risk attached.
- Your spares kit, rack layout, cabling and power envelope are known quantities. Nothing about the building has to change.
- The platform itself is mature. Firmware behaviour, thermal limits and failure modes are well understood by now, and there is little left to discover about how these nodes behave under sustained load.
If your models fit in 80GB per GPU, your serving latency targets are being met, and the requirement is simply more concurrent jobs, more H100 is a defensible and often superior answer. The premium for Blackwell buys headroom you may not use.
The rack power and cooling step change
This is where cross-generation buyers get caught. An H100 SXM5 is rated up to 700W. A B200 in HGX configuration is specified at up to 1,000W per GPU, and the rest of the node grows with it. Air-cooled HGX B200 servers do exist — ours is a 10U chassis, against the 8U typical of an air-cooled HGX H100 node — but the extra height and airflow are the price of avoiding liquid, and they consume rack units you may not have. The direct liquid-cooled build of the same eight-GPU baseboard fits in 4U. At any real density, direct-to-chip liquid cooling stops being an option and becomes the assumption.
The questions that follow are building questions, not IT questions. What can a single rack draw before the PDU and the upstream breaker become the limit? Is there a coolant distribution unit, a secondary loop, and somewhere for the return heat to go? Will the floor take the weight? Who owns the leak-detection procedure? A Blackwell node you cannot power and cool is slower in practice than a Hopper node you can run today, and the facility work has its own schedule that runs in parallel with the hardware, not after it.
HGX B200 and GB200 NVL72 are two different purchases
These are often discussed as if they were tiers of the same product. They are not, and the distinction decides how much of your data centre is in scope.
An HGX B200 node is a generational replacement at the node level: eight GPUs on a baseboard, NVLink and NVSwitch inside the box, InfiniBand or high-speed Ethernet between boxes, in a rack you already understand. The operational model is the one you have from HGX H100. You are changing what is in the chassis, not how the cluster is put together.
GB200 NVL72 is not a server. It is a rack — 72 Blackwell GPUs and 36 Grace CPUs joined into a single NVLink domain, liquid-cooled throughout, with roughly 13.5TB of HBM3e addressable across the domain and a power density on the order of 120kW in one cabinet. The reason to buy it is the domain itself: workloads where a model, or an expert-parallel serving pattern, benefits from far more than eight GPUs behaving as one coherent unit. If your workloads run happily inside an eight-GPU domain today, NVL72 is not an upgrade for you at any price. For most enterprise buyers the real choice is HGX B200 or more Hopper, and rack-scale is a question for a later cycle.
Running both generations at once
You can own H100 and B200 in the same hall, but not usefully in the same distributed job. Collective operations run at the pace of the slowest participant, so a mixed training job donates the Blackwell advantage back. Keep them as separate scheduler pools with separate queues.
The pattern that works well is a split by role rather than by vintage. Hopper handles training and fine-tuning that fits within its memory, where the estate is already proven and utilisation is easy to keep high. Blackwell handles the production serving tier, where bandwidth converts directly into tokens per second and the case is clearest. That also gives you a contained place to do the FP4 and driver work, rather than doing it across the whole cluster at once.
The case for waiting, and the case against
Waiting is rational when your facility is not ready, when your models fit and your latency targets are met, or when the capacity is needed on a date that the electrical and cooling work cannot hit. It is not rational as a general posture. There is always a faster part coming, and an organisation that defers every cycle to catch the next one ends up buying nothing while paying for cloud in the meantime.
The practical test is a date and a bottleneck. Name the date the capacity has to be in production, and name the specific thing that is currently limiting you — memory capacity, memory bandwidth, GPU-hours, or something upstream that is not the GPU at all. If the bottleneck is bandwidth or per-GPU capacity and the date allows for facility work, Blackwell earns its premium. If the bottleneck is upstream, or the date is tight, more Hopper is the cheaper and faster answer, and it is not a compromise.
How Nexus Compute helps
As an independent procurement partner, we help you turn an H100 or B200 decision into a concrete, validated configuration — sourced through authorized channels and quoted within 48 business hours. Our specialists configure first and quote second, so what you receive actually works on day one. Tell us the workload, the need-by date and what your rack can genuinely draw, and we will tell you which generation the answer points to.
Systems covered in this article
B200 & GB200 Blackwell Servers
HGX B200 8-GPU nodes in air-cooled and direct liquid-cooled builds, plus GB200 NVL72 rack-scale and NVL2 MGX options.
H100 GPU Servers
SXM5 and PCIe H100 platforms for adding matched capacity to an estate you already run.
Request a Quote
Send the workload and your rack power envelope and we return a validated configuration.
Planning a hardware investment?
Tell us what you're trying to build. A procurement specialist will help you specify and quote the right configuration — within 48 business hours, no obligation.
