A single Nvidia GB200 NVL72 rack pulls somewhere around 120kW. A decade ago, a "hot" rack pulled 10-15kW and datacenter operators considered that a design challenge. The gap between those two numbers is not incremental — it's an order of magnitude — and it explains why the industry that spent thirty years perfecting air conditioning is now retrofitting itself around pipes, coolant, and pumps.
Air cooling didn't fail because someone got lazy. It failed because physics stopped cooperating. Air is a poor conductor of heat compared to liquid, and moving enough of it fast enough to pull 100+kW out of a 42U cabinet without turning the datacenter floor into a wind tunnel is no longer practical. Liquid cooling isn't a nice-to-have upgrade anymore — for the highest-density AI racks, it's the only thing standing between a cluster and thermal throttling.
Below we explain why air cooling hits a wall, how liquid cooling data centers work (direct-to-chip, immersion, and rear-door or in-row heat exchangers), how the approaches compare, what the shift means for colocation customers and operators, the cost picture, and the limitations still being worked out.
Why Air Cooling Hits a Wall
Air cooling works by blowing chilled air across heat sinks attached to chips, carrying the heat away, and returning warmer air to a CRAC (computer room air conditioner) or CRAH (computer room air handler) unit to be cooled again. It's simple, cheap, and well understood, and it's fine for racks pulling single-digit or low-double-digit kilowatts.
The problem is heat capacity. Water carries roughly 3,500 times more heat per unit volume than air for the same temperature rise. That means moving the same amount of thermal energy with air requires vastly more volume and velocity than moving it with liquid. Datacenter engineers, following thermal guidelines like ASHRAE's, describe this in practical terms:
- Above roughly 15-20kW per rack, air-cooling fan power starts rising faster than useful cooling gained, because you need increasingly forceful airflow to compensate for air's poor thermal conductivity.
- Above roughly 30-40kW per rack, hot spots become unmanageable — GPUs at the front of a chassis get cooler air than GPUs at the back, creating thermal gradients that force conservative clock speeds across the board.
- Beyond about 50kW, most air-cooling configurations simply cannot keep junction temperatures in a safe operating range without unacceptable acoustic noise, energy draw, or physical airflow constraints (raised floors, hot/cold aisle containment, CRAC capacity).
GPU-dense AI training and inference racks now regularly exceed all three thresholds at once. A rack of eight to seventy-two high-end accelerators, packed tightly for interconnect bandwidth reasons, generates a heat load that air simply cannot evacuate fast enough — not without cooling infrastructure so large it defeats the purpose of dense compute in the first place.
How Liquid Cooling Data Centers Actually Work
"Liquid cooling" isn't one technology — it's a family of approaches that all substitute liquid for air at some stage of heat removal. The three that matter for datacenters today differ in how close the liquid gets to the chip.
Direct-to-Chip (Cold Plate) Cooling
This is the dominant approach in new AI-cluster builds. A metal cold plate, machined with internal microchannels, sits directly on top of the CPU, GPU, or other high-power silicon in place of a traditional heat sink. Coolant — typically a water-glycol mixture — is pumped through the cold plate, absorbs heat at the source, and carries it out of the server to a coolant distribution unit (CDU), which transfers that heat to a facility-level liquid loop or an outdoor heat rejection system.
Direct-to-chip cooling handles the hottest components (GPUs, CPUs) directly while leaving lower-power components like memory, NICs, and power delivery on conventional air cooling, often with rear-door heat exchangers or in-row cooling handling the residual air-cooled load. This hybrid approach is why most 100kW+ AI racks aren't purely liquid-cooled — they're liquid-cooled where it counts and air-cooled everywhere else.
Immersion Cooling
Here, entire servers — motherboard, GPUs, drives, everything except perhaps power supplies — are submerged in a bath of dielectric fluid, a liquid engineered not to conduct electricity. Heat moves directly from every component into the surrounding fluid, which is then circulated to a heat exchanger.
Immersion comes in two flavors:
- Single-phase immersion: the fluid stays liquid throughout the cycle, absorbing heat and being pumped to an external cooler, similar in principle to a cold plate loop but bathing the whole server rather than just the chip.
- Two-phase immersion: the fluid is engineered to boil at a low temperature directly on the hot components, and the vapor condenses on a cooled lid or coil above the tank, releasing its heat and dripping back down as liquid. This uses the latent heat of vaporization — a much larger energy transfer per unit of fluid than simple temperature-based heat absorption — making it extremely efficient, though the fluids involved are typically more expensive and the fluorochemical formulations have drawn environmental scrutiny (some are PFAS-related compounds).
Rear-Door and In-Row Heat Exchangers
A less invasive step up from pure air cooling: liquid-cooled panels are mounted on the back of a rack (rear-door heat exchangers) or between rows (in-row units), intercepting hot exhaust air and cooling it with a liquid coil before it re-enters the room. This doesn't touch the chip directly, so it's less effective at very high densities, but it's a lower-disruption retrofit for facilities not ready for full direct-to-chip infrastructure.
Comparing the Approaches
| Method | Density supported | Retrofit complexity | Efficiency (PUE impact) | Typical use case |
|---|---|---|---|---|
| Rear-door / in-row heat exchanger | Up to ~30-40kW/rack | Low — bolts onto existing racks | Moderate improvement | Mixed-density legacy datacenters |
| Direct-to-chip (cold plate) | 50-130kW+/rack | Medium — needs facility liquid loops, CDUs | High improvement | New AI training/inference clusters |
| Single-phase immersion | 50-100kW+/rack | High — new tank infrastructure, server redesign | High improvement | Purpose-built HPC/AI facilities |
| Two-phase immersion | 100kW+/rack | Very high — specialized fluids, sealed tanks | Highest improvement | Extreme-density, edge, specialized HPC |
Why It Matters Right Now
Average rack density hit 27kW in 2026, up 69% year over year. That single number captures the whole story: this isn't a niche problem affecting a handful of frontier AI labs running exotic superclusters. It's the average rack across surveyed datacenters climbing fast enough that infrastructure designed even three or four years ago is already being outpaced.
A few forces are converging to push density this hard:
- GPU power draw is climbing per chip, not just per rack. Successive generations of AI accelerators have increased thermal design power significantly, so even a rack with the same physical GPU count now dissipates far more heat than its predecessor.
- Interconnect physics favors tight packing. Training large models efficiently depends on high-bandwidth, low-latency links between accelerators (NVLink, InfiniBand, and similar fabrics). Spreading GPUs across more, less-dense racks to ease cooling would degrade training performance — so operators are choosing to solve the heat problem rather than sacrifice interconnect density.
- Colocation and cloud customers are demanding it. Enterprises renting AI capacity want it delivered in the densest, most performant configuration available, and providers that can't offer high-density liquid-cooled space are losing that business to those that can.
This creates a genuine bind for datacenter operators: the facilities that were profitable and fully depreciated under an air-cooling design are now the wrong shape for the workloads customers want to run in them. That's driving a wave of retrofits, new-build liquid-ready facilities, and a scramble among colocation providers to add CDU and liquid-loop capacity before their air-cooled floor space becomes a stranded asset — a dynamic playing out alongside the shift toward grid-interactive data centers more broadly.
Benefits of Liquid Cooling for AI Data Centers
Liquid cooling is often described as a necessity, which it is at the top end. It also brings advantages that make it attractive below the point where air physically fails.
Much Higher Rack Density
The headline benefit is simply being able to run 100kW-plus racks at all. Because liquid carries far more heat per unit volume than air, direct-to-chip and immersion systems can remove heat loads that would overwhelm any practical airflow design, even with aggressive containment and oversized air handlers. That lets operators pack accelerators as tightly as interconnect performance demands rather than spreading them out to stay cool.
Sustained Performance Without Throttling
When air cooling struggles, chips at the back of a chassis run hotter than those at the front, and the whole system settles at conservative clock speeds. Removing heat at the source keeps junction temperatures stable and even across components. For expensive accelerators, avoiding throttling means getting the performance you paid for on every training run.
Lower Cooling Energy per Unit of Compute
Pumps moving liquid use far less power than the fans and chillers needed to move the same heat with air. Facilities running warm-water loops can reduce mechanical chiller use substantially. The result is less of the electricity bill spent on cooling relative to compute delivered, which matters as energy becomes one of the largest operating costs of AI infrastructure.
More Compute in Less Floor Space
One liquid-cooled rack can replace several lower-density air-cooled racks. In supply-constrained real estate markets, delivering the same compute in a smaller footprint can be the difference between fitting a cluster into an existing hall and needing a new building. Shorter distances between accelerators also help interconnect performance.
Quieter, Calmer Data Halls
Pushing air hard enough to cool dense racks creates serious acoustic noise and turbulent hot spots. Liquid systems reduce the reliance on high-speed fans, making halls quieter and thermal behaviour more predictable. That improves working conditions for technicians and makes capacity planning less of a guessing game, since heat removal no longer depends on how air happens to flow around a crowded row.
Liquid Cooling Use Cases
Different liquid cooling methods fit different densities and facility situations. These are the main places each is being used.
Large AI Training Clusters
Training clusters built on racks like the GB200 NVL72, drawing around 120kW, are the clearest case. Tight packing is needed for high-bandwidth links between accelerators, and air cannot remove the resulting heat. New builds for these clusters are standardising on direct-to-chip cold plates for GPUs and CPUs, with rear-door or in-row units handling the residual air-cooled load from memory, networking and power components.
High-Density Inference Deployments
Inference has historically run at lower densities than training, but large models and rising per-chip power are pushing inference racks up as well. Operators deploying dense inference capacity use direct-to-chip cooling to keep performance consistent under sustained load, especially when the same facility also hosts training hardware. Sharing one liquid loop across both workloads simplifies the plant and lets capacity shift between training and inference as demand changes.
Colocation Retrofits for AI Tenants
Colocation providers with air-cooled halls face customers who want AI-ready space. A common route is to retrofit part of a facility, often with rear-door heat exchangers for moderate densities or a dedicated hall with a new liquid loop and CDUs for higher ones. This lets providers serve AI tenants without rebuilding the whole site or disrupting existing customers, and learn liquid operations on a contained footprint first.
High-Performance Computing
Scientific and engineering supercomputers adopted liquid cooling earlier than most enterprise datacenters, because dense CPU and accelerator nodes hit the same thermal limits. Purpose-built HPC facilities use direct-to-chip and single-phase immersion, and their experience with loop design and maintenance has informed AI deployments. Many of the practices now standard in AI halls, from negative-pressure loops to leak detection, were refined in these environments first.
Extreme-Density and Specialised Sites
Two-phase immersion appears in extreme-density, edge and specialised HPC deployments where its efficiency justifies the cost and handling of specialised fluids and sealed tanks. Environmental scrutiny of some fluorochemical fluids means these deployments are watching reformulation efforts closely, and some operators prefer single-phase immersion to avoid that exposure.
Common Liquid Cooling Mistakes
Buying Advertised Rack Power Without Checking the Loop
A provider can advertise 100kW racks while lacking the facility-level liquid distribution and heat rejection to deliver that at scale. Customers who sign for rack power alone can find their actual deployable density capped well below what they planned. The question to ask is how much heat the whole facility can reject, not what one rack is rated for.
Ignoring the Air-Cooled Remainder
Most direct-to-chip servers still send part of their heat to air from memory, network cards and power delivery. Plans that assume the rack is fully liquid-cooled underestimate the air handling still needed. The ratio of liquid to air heat load determines how much of the density gain is real, and an undersized air side can quietly throttle the very servers the liquid loop was meant to free up.
Skimping on CDU Redundancy
A coolant distribution unit failure can take an entire liquid-cooled row offline. Treating CDUs like ordinary plant equipment without N+1 or 2N redundancy turns one component failure into a major outage for very expensive hardware. Maintenance windows for CDUs also need planning, since servicing one should never mean shutting down the row it feeds.
Waiting Until Density Forces a Retrofit
Operators who delay liquid readiness until customers demand it often pay for rushed construction and long equipment lead times. Building liquid-ready capacity ahead of demand, or at least designing halls so loops can be added cleanly, is usually cheaper than an emergency conversion carried out around live customers.
Underestimating Serviceability and Skills
Swapping a component in a direct-to-chip system means disconnecting and reconnecting coolant lines, and immersion requires handling fluids and validated hardware. Teams that don't train technicians or specify reliable quick-disconnect fittings see longer repair times and more leak risk than the technology itself warrants.
Liquid Cooling Best Practices for Businesses and Builders
If you're renting or building datacenter capacity for AI workloads, rack density and cooling method aren't back-office details — they directly affect what you can deploy, how fast, and at what cost.
For colocation and cloud customers
- Ask about facility liquid-loop capacity, not just rack power. A datacenter can advertise "100kW racks" but if it lacks a facility-level liquid distribution system and adequate heat rejection (dry coolers, cooling towers, or chillers sized for the load), you may be sold theoretical capacity that isn't actually deliverable at scale.
- Understand the hybrid reality. Most GPU servers today still rely on some air cooling for non-GPU components even when direct-to-chip handles the accelerators. Ask what percentage of total rack heat load is liquid-cooled versus air-cooled, since that ratio determines how much of your density gain is real.
- Factor in water usage and sustainability commitments. Liquid cooling loops need water somewhere in the chain (even closed loops usually reject heat to water-cooled chillers or evaporative systems at some point), which matters for organizations with water-usage or sustainability reporting obligations.
For datacenter operators and builders
- Size CDUs with redundancy, not just capacity. A coolant distribution unit failure can take down an entire liquid-cooled row, so N+1 or 2N CDU redundancy is becoming standard practice for mission-critical AI infrastructure.
- Weigh retrofits against new builds site by site. Adding liquid loops, CDUs, and piping to an existing shell is disruptive but typically faster and less capital-intensive than a greenfield liquid-ready facility, especially in supply-constrained real estate markets.
- Plan for plumbing interoperability up front. Coolant quality standards, quick-disconnect fitting types, and facility-to-rack interface specs vary by vendor, and interoperability issues can lock an operator into a single hardware ecosystem if not planned for carefully.
A simplified decision framework
For teams evaluating whether a workload needs liquid cooling, the rough rule of thumb operators use is:
- Below ~20kW/rack — air cooling with good hot/cold aisle containment is usually sufficient and cheapest.
- 20-50kW/rack — rear-door or in-row heat exchangers can often bridge the gap without a full liquid-loop buildout.
- 50-100kW/rack — direct-to-chip cooling becomes close to mandatory for sustained, quiet, efficient operation.
- 100kW+/rack — direct-to-chip is standard, and immersion cooling starts becoming competitive, particularly for extreme-density or specialized HPC deployments.
Cost considerations
Liquid cooling isn't just a technical decision — it changes the capital and operating cost profile of a facility in ways worth planning for explicitly.
- Upfront capital costs are higher. CDUs, facility-level piping, coolant, and specialized racks all add cost compared to a conventional air-cooled build. For a retrofit, the disruption of running new plumbing through a live datacenter floor adds further expense and scheduling complexity.
- Operating costs typically fall. Once installed, liquid cooling generally reduces the energy spent on cooling relative to compute delivered, since pumps use far less power than the fans and chillers needed to move an equivalent amount of heat with air. Facilities that shift to warm-water loops can cut mechanical chiller use substantially, lowering the electricity bill tied to cooling specifically.
- Density gains offset per-rack cost. Because a single liquid-cooled rack can replace several air-cooled racks at lower density, the cost comparison isn't liquid-vs-air on a like-for-like rack basis — it's liquid-vs-air for the same total compute delivered in the same floor space, which usually favors liquid once density passes the crossover point discussed above.
- Total cost of ownership favors early adoption for AI-heavy workloads. Operators who wait to retrofit until density forces the issue tend to pay a premium for rushed construction and equipment lead times, compared to those who build liquid-ready capacity ahead of demand.
Limitations and Open Questions
Liquid cooling solves the heat-removal problem, but it introduces new ones that the industry hasn't fully settled.
- Leak risk is real, and the stakes are higher. A coolant leak near live electronics is a far more consequential failure mode than a fan dying in an air-cooled server. Cold-plate systems mitigate this with negative-pressure loops (so any leak draws air in rather than pushing coolant out) and leak-detection sensors, but the failure mode is qualitatively different from anything air-cooling operators had to plan for.
- Serviceability changes. Swapping a GPU in a direct-to-chip system means disconnecting and reconnecting coolant lines — a more involved procedure than pulling a card out of an air-cooled chassis, and one that requires trained technicians and quick-disconnect fittings rated for repeated use without degradation.
- Immersion cooling raises material and environmental questions. Two-phase dielectric fluids have historically included fluorochemicals facing regulatory scrutiny over persistence in the environment, and the industry is actively working on alternative formulations. Single-phase immersion fluids avoid some of these concerns but bring their own handling, disposal, and material-compatibility considerations (not every cable jacket or component coating tolerates long-term immersion equally well).
- Standardization is incomplete. Unlike air cooling, where a rack from one vendor works in essentially any datacenter with adequate CRAC capacity, liquid-cooled racks depend on facility-specific loop designs, coolant chemistry, and connector standards — efforts like the Open Compute Project aim to close that gap — that aren't yet fully interoperable across vendors. This raises switching costs and can create vendor lock-in at the facility level.
- Retrofitting older buildings has real limits. Not every existing datacenter has the floor loading, ceiling height, or piping runways to add a facility liquid loop economically. Some older facilities may simply not be viable candidates for high-density AI tenancy regardless of investment.
What to Watch Next
A few trends will shape how this plays out over the next several years:
- Facility-level liquid readiness becomes a standard spec sheet item, the way power density and network uplink speeds already are, making it easier for customers to compare providers on true delivered cooling capacity rather than marketing claims.
- Coolant and connector standardization efforts from industry bodies aiming to reduce vendor lock-in and make liquid-cooled hardware more interchangeable across facilities.
- Warm-water cooling adoption, where facilities cool with higher-temperature water (reducing or eliminating mechanical chiller use) to cut energy consumption — an approach some large-scale operators have already used to lower cooling-related power draw significantly.
- Immersion fluid reformulation as environmental regulation tightens around certain fluorochemical classes, pushing vendors toward alternative dielectric fluids with lower environmental persistence.
- Continued climb in per-rack density as accelerator power draw keeps increasing generation over generation, which will keep pushing the threshold at which liquid cooling becomes mandatory rather than optional further downward over time.
Teams evaluating high-density AI infrastructure decisions can get hands-on help scoping cooling and deployment tradeoffs from Woyce Technologies.
FAQ
What rack density requires liquid cooling?
There's no hard cutoff, but most operators consider air cooling impractical above roughly 30-50kW per rack. Below that, containment and airflow management can usually keep up; above it, direct-to-chip or immersion cooling becomes necessary to maintain safe operating temperatures without excessive fan power or noise. In practice, many facilities run a hybrid: liquid for the GPU racks, conventional air for storage, networking and general compute. The right threshold for you depends on your building, climate and existing chiller plant.
What's the difference between direct-to-chip and immersion cooling?
Direct-to-chip cooling routes liquid through a cold plate mounted on individual hot components like GPUs and CPUs, while other server parts stay air-cooled. Immersion cooling submerges the entire server in dielectric fluid, cooling every component simultaneously. Direct-to-chip is generally easier to retrofit into existing server designs; immersion typically requires purpose-built enclosures.
Is liquid cooling more energy-efficient than air cooling?
Generally yes. Liquid's higher heat capacity means less energy is spent moving the cooling medium around, and liquid-cooled facilities often report lower Power Usage Effectiveness (PUE) than comparable air-cooled ones because they need less mechanical chiller and fan power to remove the same amount of heat. The exact savings depend on the facility's design and climate, so treat any reported figure as site-specific.
Does liquid cooling increase the risk of hardware damage from leaks?
It introduces a new failure mode that didn't exist with air cooling, but modern direct-to-chip systems mitigate it with negative-pressure loops, leak sensors, and quick-disconnect fittings designed to seal automatically when detached. Leak risk is a real engineering consideration, not a solved problem, but it's actively managed rather than a reason to avoid the technology.
Can existing datacenters be retrofitted for liquid cooling?
Many can, particularly with rear-door heat exchangers or in-row units that don't require a full facility liquid loop. Full direct-to-chip or immersion retrofits are more involved, requiring piping, CDUs, and sometimes floor-loading or ceiling-height changes, so feasibility depends heavily on the building's original design. A common path is to dedicate one hall or a group of rows to liquid-cooled AI racks rather than converting the whole site, which limits disruption to existing customers and lets operators learn on a smaller footprint first.
Why are AI workloads specifically driving this shift?
AI training and inference clusters pack many high-power accelerators tightly together to maximize interconnect bandwidth between chips, which concentrates far more heat per rack than traditional enterprise computing. Combined with each new GPU generation drawing more power than the last, this pushes density well past what air cooling can handle. Spreading the same hardware across more racks would ease cooling, but it lengthens the connections between accelerators and wastes expensive floor space, so operators accept the cooling challenge rather than giving up density.
Is immersion cooling safe for standard server hardware?
Most modern servers aren't submersion-rated out of the box — immersion typically requires hardware validated for compatibility with the specific dielectric fluid used, since some materials, adhesives, and coatings can degrade over prolonged immersion. Vendors increasingly offer immersion-validated server configurations, but it's not a universal drop-in replacement for air-cooled chassis.
Conclusion
AI hardware has pushed rack power from the 10 to 15kW range into 100kW and beyond, and air simply can't carry that much heat out of a cabinet efficiently. Liquid cooling has gone from a niche option for supercomputers to a requirement for the densest AI clusters.
The main approaches solve different problems. Direct-to-chip cold plates target the hottest components and are the most common path for new GPU deployments. Immersion cools everything at once but needs validated hardware and purpose-built tanks. Rear-door and in-row heat exchangers are often the easiest retrofit for moderate densities. Most real facilities will mix these with conventional air.
The trade-offs are practical rather than theoretical: higher upfront cost, new plumbing and maintenance skills, leak management, floor loading limits, and supply chains for coolant distribution units that are still catching up with demand. Retrofitting an older building is possible, but feasibility depends heavily on its original design.
If you're planning AI workloads and need to decide between cloud GPUs, colocation and on-premises clusters, start with the density you actually need. For help mapping those choices, explore our cloud architecture services.
