Skip to content

300,000+ products available to order|AI infrastructure + enterprise IT hardware, sourced through authorized channels.

Procurement hardware — How Much Does an AI Server Cost? What Actually Drives the Number
Procurement
Back to Resources
Procurement 11 min read August 28, 2026

How Much Does an AI Server Cost? What Actually Drives the Number

A line-by-line walk through what makes up the price of an AI server — accelerators, the host platform, the fabric, power and cooling, software and support, and commissioning — in the order budgets tend to forget them.

"How much does an AI server cost?" is a fair question with no honest single-figure answer. A four-GPU L40S inference node and an eight-way B200 training node are both AI servers. They are not in the same conversation. What is useful is not an average but the list of line items that make up the number, walked in the order teams tend to forget them. Once every line is on the page, the figure stops moving.

Why there is no single number

The spread in AI server cost is not vendor mischief. It is real configuration variance. The same chassis can be built with four PCIe accelerators or eight SXM modules on an HGX baseboard, with one network port or eight, air-cooled or plumbed into a liquid loop, on a three-year or a five-year support term. Each of those choices moves the total substantially, and several of them are decided by your building rather than by your workload. A quote is a function of a configuration, and a configuration is a function of answers you have to supply. A published headline figure is almost always a base chassis, not the system you will actually run.

The accelerators: the line everyone budgets for

GPUs are almost always the largest single line, and the one that gets the most attention. The current enterprise ladder runs roughly from the L40S at 48GB of GDDR6, through the A100 at 80GB of HBM2e, the H100 at 80GB of HBM3, the H200 at 141GB of HBM3e, to the B200 at 180GB of HBM3e. Memory capacity and memory bandwidth are what separate them for AI work, far more than any headline compute figure.

Form factor matters as much as the part number. SXM modules arrive mounted on an HGX baseboard with NVSwitch between them — you buy the assembly, commonly eight GPUs at a time, not individual cards. PCIe accelerators are bought and installed singly. Some support NVLink bridges between paired cards, but none give you the all-to-all bandwidth of a switched baseboard. If your workload spans many GPUs constantly, that difference is the whole point of paying for SXM. If it does not, you are buying interconnect you will not saturate.

This is the easiest line to over-buy. If your model already fits comfortably in 80GB and you are not memory-bandwidth bound, the newer part buys headroom you may never use. That can still be the right call — headroom has value when you cannot predict what you will be running in eighteen months — but it should be a decision, not a default.

The server around the GPUs

An accelerator does nothing on its own. Around it sits a dual-socket host, a memory configuration that should populate every channel rather than hit a capacity target cheaply, NVMe for staging and checkpoints, redundant power supplies sized for the full GPU draw, and enough PCIe lanes to feed both the GPUs and the network adapters without contention. A common design convention is to fit at least as much system memory as the aggregate GPU memory in the node.

None of this is glamorous, and all of it is where cheap quotes get cheap. A half-populated memory configuration, a single power supply, or a CPU chosen on core count without regard to PCIe lanes will produce a lower number and a node that underperforms its own GPUs. When comparing two quotes for the same accelerators, this is usually where the difference is hiding.

The fabric: the first line that goes missing

If you are buying one node, you can skip this section, and that is precisely why so many first budgets are wrong. Single-node budgets tend to survive contact with reality. Multi-node ones frequently do not, because the fabric between the nodes was never costed.

A multi-node training cluster needs a dedicated east-west network, and the standard design is rail-optimised: one high-speed adapter per GPU, so an eight-GPU node carries eight current-generation 400Gb/s-class ports for compute traffic alone. That is before the storage network and before out-of-band management. The line items are the adapters, the switches, the transceivers at both ends of every link, and the cabling itself — direct-attach copper for short runs, active optical or structured fibre for anything longer.

Optics scale with ports, and ports scale with GPUs. On a cluster of any size, the fabric is not a rounding error against the servers. A budget that lists compute nodes and no switch line is not a budget. It is a partial one that will need a second approval later, usually at the worst possible moment in the project.

Power and cooling: the second

An SXM-class data-centre GPU such as the H100 or H200 is rated up to around 700W under sustained load, and Blackwell parts are rated higher still. Eight of those, plus dual CPUs, a full memory configuration, eight network adapters and the fans to move air through the chassis, put a single node in the region of ten kilowatts. Blackwell-generation nodes draw more. That is one node. Many enterprise racks were provisioned for a fraction of that, spread across a dozen servers.

So the facility work is real capital spend: higher-amperage circuits and PDUs, possibly new busway, UPS and generator capacity to match, and a way to remove the heat — rear-door heat exchangers for air-cooled nodes, or a full liquid loop with a coolant distribution unit and manifolds for direct-liquid-cooled systems. Rack-scale Grace-Blackwell platforms such as GB200 NVL72 are liquid-cooled by design, not by preference. There is no air-cooled variant to fall back on.

This line goes missing more often than any other, for an organisational reason rather than a technical one. It sits with facilities, not IT, and is approved by a different person from a different budget. The electrical and mechanical contractors also have their own lead times, which run in parallel with the hardware lead time only if you start them in parallel.

Software, licensing and support

Enterprise AI software is a recurring line. NVIDIA AI Enterprise is licensed per GPU on a subscription basis, and cluster management, scheduling and monitoring tooling sit alongside it — whether you pay for a commercial stack or absorb the engineering time to run an open one. The engineering time is a cost even when it does not appear on a purchase order.

Hardware support is the line most often quoted at its cheapest tier to keep a bid competitive. Next-business-day parts and four-hour onsite response are different products, and on a cluster where a dead node stalls a training run, the difference is the point. Term length matters too. A five-year support wrap changes the number materially against three years. Compare quotes at the same support tier and the same term, or you are not comparing quotes.

Delivery, installation and commissioning

These systems are heavy. An eight-GPU node is a two-person lift at minimum, and a populated GB200 NVL72 rack weighs well over a ton and arrives on its own pallet. Freight, insurance, a loading dock, a lift and a floor rated for the load are preconditions, not details. Confirm the delivery path into the room before the order, not on the day the truck arrives.

Then someone has to rack it, cable it, level the firmware across every node, run burn-in, and validate that the fabric actually performs collective operations at the rate the design assumed. The day the pallet arrives is not the day the cluster works. Whether that labour is yours or your supplier's, it is real, and it belongs in the plan with a duration attached.

How to get a number you can defend

A quote is only as good as the brief behind it. To get one that will not move, put these on the page before you ask for pricing:

  • The workload: model sizes, training or inference, and expected concurrency.
  • Node count now, and whether this design has to scale later without being rebuilt.
  • The rack power and cooling you actually have available today, measured rather than assumed.
  • The vendor platforms and orchestration your operations team already runs.
  • The support tier and term you intend to buy, so quotes are comparable.
  • Who owns the facility work, and whether it sits inside this budget or a separate one.

Then price the whole thing in one pass. The most expensive mistake in AI infrastructure budgeting is not paying too much for GPUs. It is approving the compute, discovering the fabric and the facility work afterwards, and having to go back for a second approval on a project that has already been announced.

Nexus Compute configures first and quotes second. Tell us the workload, the node count, and the power and cooling you have, and we will return a complete bill of materials — accelerators, host platform, fabric, storage, software and support — sourced through authorised channels and quoted within 48 business hours, with the facility requirements stated plainly rather than left for you to discover.

Systems covered in this article

Planning a hardware investment?

Tell us what you're trying to build. A procurement specialist will help you specify and quote the right configuration — within 48 business hours, no obligation.

AI server costGPU server priceAI infrastructure budgetProcurementHGXCapital planning