
Server and Rack Cooling: When Air Cooling Stops Being Enough
Most racks still run on air, and that's the right call for most of them. The point where it stops working is more specific than a single power number: here's how to tell which side of it your next deployment falls on.
Air cooling still handles the majority of enterprise racks, and for most deployments it remains the correct, boring, reliable answer. The interesting question isn't whether liquid cooling is better in some abstract sense (at high enough density it clearly is) but at what specific point a given rack's air-cooling design actually stops working, because that point is more precise than "when the GPUs arrive."
What air cooling is actually good at
A well-designed air-cooled server (the right heatsink for the CPU's TDP, adequate fan count and airflow path, correctly managed hot-aisle/cold-aisle separation at the rack and row level) comfortably handles a wide range of conventional enterprise workloads. Modern high-core-count Xeon and EPYC CPUs at 300-400W each are still well within the range vendor-qualified passive and active heatsinks are designed for, in servers with sensible chassis airflow. The failure mode people worry about, thermal throttling, is much more often caused by a wrong or degraded heatsink, blocked airflow from cable management, or a hot aisle that isn't actually isolated from the cold aisle, than by air cooling being fundamentally inadequate for the load.
Component-level cooling matters as much as room-level design. A CPU heatsink rated for the wrong socket's TDP, or a chassis fan replaced with a lower-spec part during a repair, silently reduces the thermal headroom the rest of the design assumed. This is a more common cause of unexplained throttling than most teams initially suspect, and it's worth confirming heatsink and fan specifications against the actual installed CPU whenever a chassis has been serviced.
Where the numbers stop working
The one-line answer people want, "liquid cooling above X kW per rack," is directionally right but the real constraint is airflow and heat-rejection capacity at the room level, not a fixed number. A rack that a well-designed data hall can comfortably air-cool at 15-20kW might be genuinely unmanageable at half that density in a facility with poor hot-aisle containment or undersized CRAC capacity. That said, as a practical planning heuristic: single racks moving past roughly 20-30kW are the range where most conventional data center air-handling designs start to struggle, and where direct liquid cooling (or at minimum rear-door heat exchangers) moves from optional to the practical requirement.
GPU-dense compute is what actually forces this conversation for most teams today. A single 8-GPU HGX-class server with current-generation SXM accelerators, dual high-core CPUs and a full network complement can draw well into five figures of watts on its own. Several such nodes in one rack exceeds what air handling in a conventionally provisioned data hall was ever designed to remove, regardless of how good the individual server's internal cooling is.
Rear-door heat exchangers: the middle step
Before committing to a full liquid-cooling loop, a rear-door heat exchanger (a passive or fan-assisted coil mounted on the back of the rack, chilled water flowing through it) extends the useful range of air-cooled servers considerably by capturing heat at the rack before it reaches the room, rather than relying on room-level CRAC capacity to remove it after the fact. It's a real, well-proven middle step for facilities that have chilled water available but don't want to commit to direct-to-chip liquid cooling for every server. It doesn't change what's happening inside the server chassis: the CPUs and GPUs are still air-cooled internally, but it substantially raises how much heat a given rack footprint can support before the room itself becomes the bottleneck.
Direct liquid cooling: when the chassis itself needs it
Direct-to-chip liquid cooling, cold plates mounted on the CPU and GPU dies themselves with coolant circulated through a manifold to a coolant distribution unit, is the answer once a single component's heat output exceeds what any air-based heatsink can remove regardless of airflow. This is now standard, not optional, on the highest-power accelerator platforms; rack-scale systems built around the densest current GPU generations are liquid-cooled by design, with no air-cooled variant offered, because no air-based heatsink and fan combination can remove that much heat from that small a die area.
Committing to direct liquid cooling is a facilities decision with real prerequisites: a coolant distribution unit, a secondary fluid loop, leak detection, and (often the part most underestimated) a documented procedure for what happens when a leak is detected. It is not a line item you tick on a purchase order; it's a capital project with its own timeline that needs to run in parallel with, not after, the compute procurement.
A decision framework, not a threshold
- Conventional enterprise racks under roughly 10-15kW: air cooling, with attention to component-level heatsink and fan specification, is almost always sufficient and the right default.
- Racks in the 15-30kW range, or facilities with marginal air-handling capacity even below that: evaluate rear-door heat exchangers before assuming a full liquid loop is required.
- Any rack built around current-generation 8-GPU accelerator platforms at real density: plan for direct liquid cooling from the start, and treat the facilities work (coolant distribution, secondary loop, leak detection) as a parallel project with its own schedule, not a follow-on task after the servers arrive.
- Whatever the room-level answer, confirm every individual server's internal heatsink and fan configuration matches its installed CPU/GPU TDP: a correctly cooled room with one misconfigured chassis inside it still throttles.
How Nexus Compute helps
As an independent procurement partner, we help you turn a cooling decision into a concrete, validated configuration, genuine and fully warranted, quoted within 48 business hours. Tell us your rack power budget and what the room can actually remove today, confirmed by facilities rather than assumed, and we'll specify the servers, heatsinks and (where the numbers call for it) the liquid-cooling path to match.
Systems covered in this article
System Cooling
OEM chassis fans, Dynatron and Noctua CPU coolers, and switch/router fan trays across current server sockets.
Power Distribution & UPS
The PDU and UPS capacity a denser rack needs alongside the cooling upgrade.
GPU Servers
The dense compute that usually forces the air-versus-liquid decision in the first place.
Planning a hardware investment?
Tell us what you're trying to build. A procurement specialist will help you specify and quote the right configuration within 48 business hours, no obligation.
