Pricing & Procurement
AI server price: what actually drives the number
Most people searching for this are building a budget rather than placing an order, and the number they need is not a list price for one box — it is the total for a system that will actually run their workload.
We are not going to print a figure here. Accelerator pricing moves with allocation, and a number published today would mislead you next quarter. What follows is the cost structure instead: the six line items that decide the total, in the order they usually surprise people. Then send us the requirement and you get a real, itemised quote.
48 business hours · purchase orders accepted · 57 GPU server configurations quoted from
The six lines in an AI server budget
The accelerators
The largest single line, and the one with the least stable pricing. It is decided by the part (H100, H200, B200, A100, L40S, MI300X), the form factor — SXM modules on an NVLink baseboard cost more than PCIe cards and are not interchangeable — and how many you need per node.
The server around them
CPUs, DDR5 in the terabytes for an 8-GPU node, and NVMe fast enough to keep the GPUs fed. Under-specifying storage is the most common way a build ends up GPU-starved: expensive accelerators idling while data arrives.
The fabric
For a single node, network adapters. For a cluster, InfiniBand or high-speed Ethernet with switches, optics and cabling, laid out rail-optimised. On a multi-node training build this is a substantial share of the total and it is the item most often left out of a first budget.
Power and cooling
An 8-GPU SXM node at sustained load is a thermal problem before it is a budget problem. Rack PDUs, the power feed itself, and — above a certain density — direct liquid cooling and the facility water to serve it. Frequently the longest-lead part of the whole project.
Software and support
Enterprise licensing where it applies, plus the support contract. A cheaper chassis that needs a second set of management tooling and a second runbook is often more expensive to own than the one that matches your existing estate.
Delivery and commissioning
Racking, cabling, firmware levelling and burn-in. Small against the hardware, but real, and worth agreeing before the pallets arrive rather than after.
Price by platform
The accelerator sets the floor. These are the platforms we quote most often — each page lists the systems and the questions worth settling before you ask for a number.
NVIDIA H100 servers
8x SXM5 HGX and 4x NVL PCIe. The production training default.
View systemsNVIDIA H200 servers
141GB HBM3e per GPU — usually fewer nodes for the same inference throughput.
View systemsNVIDIA B200 & GB200
HGX B200 nodes and GB200 NVL72 rack-scale systems.
View systemsNVIDIA A100 servers
Extending an existing estate, or MIG-partitioned inference.
View systemsDell PowerEdge XE9680
8-GPU HGX inside a Dell-managed operational model.
View systemsSupermicro GPU servers
The widest configuration range, air and liquid-cooled.
View systemsHow to get an accurate number quickly
- Tell us the workload — model sizes, training or inference, expected concurrency
- Say how many nodes, and whether this scales later
- Give us the rack power and cooling you actually have available
- Name the vendor and tooling your operations team already runs
- Send an existing bill of materials if you have one — we will price it line by line
If you are pricing individual components rather than a system, the component catalogue carries indicative pricing across drives, memory, processors and optics.
Common questions
How much does an AI server cost?
There is no single figure, and any site quoting one is describing a configuration that is probably not yours. The accelerator is the largest line, but an 8-GPU node also needs CPUs, several terabytes of memory, NVMe capable of feeding the GPUs, network adapters, and — on a multi-node build — an InfiniBand or high-speed Ethernet fabric that can approach a meaningful share of the total. Rack power and cooling are frequently the item nobody budgeted for. We return an itemised quote within 48 business hours against your actual requirement.
Why don't you publish prices for GPU servers?
Because accelerator pricing moves with allocation, and a number published today misleads a buyer reading it next quarter. We do publish indicative pricing across the component catalogue — drives, memory, optics and spares — where pricing is stable enough to be useful. For complete systems we quote, and the quote states current availability and lead time alongside the price.
What is the cheapest way to get into AI infrastructure?
Usually a single well-specified node rather than a small cluster. A 4-GPU PCIe system with L40S or H100 NVL cards avoids the InfiniBand fabric, the rack power upgrade and the liquid cooling that make multi-node builds expensive, and it fits in a standard rack. If the workload later justifies scale, the fabric can be designed so that first node is not stranded. We will say when that is the right call.
Do you accept purchase orders and work with procurement?
Yes. We quote on company letterhead with line-item detail, accept purchase orders, and can supply the documentation procurement and finance teams typically require to raise one.
Can you quote against a bill of materials we already have?
Yes, and it is the fastest route to an accurate number. Send the BOM as a spreadsheet or paste the part numbers — we will price it line by line, flag anything that will not work together, and note where a different part is materially cheaper for the same outcome.
