Uniqcli

InsightsBuying Guides

High-Density GPU Rack Power and Liquid Cooling: A Facility Readiness Guide

Modern AI systems span very different facility classes. NVIDIA's DGX B200 documentation lists a maximum input power of 14.3 kW for one 10U appliance. A rack of multiple servers plus switches and storage can move far higher. NVIDIA's GB300 NVL72 reference architecture states that a complete liquid-cooled rack can reach 142 kW. “GPU rack” therefore is not a sufficient planning unit.

By Uniqcli Team · · 6 min read · Updated

Facilities engineer inspecting liquid-cooling manifolds and hoses on a high-density GPU rack
Facilities engineer inspecting liquid-cooling manifolds and hoses on a high-density GPU rack

Key takeaways

  • Use system- and rack-level vendor data; never estimate the facility from GPU TDP alone.
  • Capture typical, tested peak and maximum loads with the assumptions behind each.
  • Verify upstream electrical capacity, redundancy, connectors, protection and maintenance—not only rack PDU rating.
  • Direct liquid cooling requires a defined technology-water/facility-water boundary, water quality, controls and leak response.
  • Design for transients, partial failures, adjacent-rack effects and maintenance modes.
  • Commission with representative heat load before accepting the production workload.
On this page

Facility readiness starts with the exact configured system, realistic workload profile and deployment pattern. It ends with commissioned electrical and cooling interfaces, alarms, failure procedures and documented ownership.

Start with a load inventory

List every device in the rack: compute, switches, storage, management, CDU pumps/controllers and auxiliaries. For each, collect quantity, typical draw under the target workload, tested peak, maximum/nameplate, heat output, airflow or liquid requirement and feed arrangement.

The DGX B200 user guide lists 14.3 kW maximum input, 1,550 CFM airflow and 48,794 BTU per hour for that appliance. These are manufacturer planning values for a specific system. Apply the same discipline to every chosen OEM model.

For rack-scale GB300, use the current NVL72 reference-architecture component data. It describes a full-rack maximum of 142 kW and integrated liquid-cooling components. Do not transpose that value to an HGX B300 server rack or assume every workload constantly reaches it.

Create at least three scenarios: idle/degraded, expected sustained mission load and maximum credible load. Include workload transients and simultaneous restart. Note whether power caps are configured, how they are enforced and what performance loss occurs when a cap is reached.

Build the electrical path

Trace from utility or generator through switchgear, UPS, distribution, busway/RPP, branch circuit, rack PDU, cord and device. Record capacity, redundancy, protection, voltage, phase, connector and monitoring at every handoff. A 60 kW rack connected to two 60 kW feeds does not necessarily have 120 kW usable capacity; an A/B redundant design may require either feed to carry the full protected load.

Review upstream diversity. Two rack PDUs can land on the same panel or UPS and share a failure domain. Confirm generator/UPS behavior, maintenance bypass, battery runtime, harmonic/current characteristics and the site's policy for continuous loading of circuits.

Power-shelf architectures deserve special attention. The GB300 NVL72 documentation describes multiple power shelves. Map every shelf to feeds and protection, and understand system behavior after a feed or shelf failure.

Produce a one-line diagram and branch schedule. Label physical receptacles and cords to match. Configure metering and alerts before burn-in. Capacity that exists only in a spreadsheet is not operational capacity.

Decide whether air cooling is viable

Air cooling can remain appropriate for many H200 and B200 server configurations, but high airflow creates constraints. Check cold-aisle supply, containment, return path, static pressure, fan curves, floor/tile or duct delivery, allowable inlet temperature, altitude derating and neighboring loads.

Assess the room, row and rack. A room may have enough aggregate tons while one aisle recirculates hot exhaust. A rack may receive adequate average airflow while top servers ingest hotter air. Use temperature and pressure sensors at multiple heights and observe behavior with doors and containment in their final state.

Rear-door heat exchangers can remove heat near the rack but introduce water, door weight, hose routing and maintenance concerns. They do not reduce electrical load. Specify approach temperature, flow, pressure drop, condensate policy, leak detection and behavior when the door is open or exchanger unavailable.

If the air-cooled design requires extreme fan energy, noise or room changes, compare direct liquid cooling on lifecycle and schedule—not just equipment cost.

Design the liquid-cooling boundary

Direct liquid cooling commonly uses a technology cooling system that interfaces with a facility water system through a coolant distribution unit. Define which party owns:

  • Server cold plates and internal hoses.
  • Rack manifold and quick disconnects.
  • Technology coolant and water-quality control.
  • CDU heat exchanger, pumps, filters and controls.
  • Facility supply/return piping and heat rejection.
  • Leak detection, containment, drains and emergency response.

Obtain required supply temperature, allowable return, flow, pressure, pressure drop, coolant chemistry, particulate limits, materials compatibility and dew-point margin. Values vary by manufacturer and configuration. Use the exact site-preparation guide and approved fluids.

Design hose routing for service. Avoid bend, strain and trip hazards. Confirm dripless connector performance, labeling and keyed/mistake-proof connections. Account for expansion, purge/fill, trapped air and maintenance access.

Condensation risk depends on fluid temperature and room dew point. A colder loop is not automatically better. Instrument and control the system to remain within vendor and environmental limits.

Plan CDU, water and controls

Size the CDU for expected and maximum heat load, flow/pressure requirements, growth and redundancy. State whether N+1 means pumps, CDUs, heat exchangers or the full cooling path. Confirm power feeds for pumps and controls; a protected compute rack with an unprotected cooling controller has an incomplete availability design.

Place the CDU based on pipe length, floor loading, service clearance, containment and leak consequences. An in-row, in-rack or facility-level CDU changes responsibility and failure domains. Coordinate with fire protection, drainage and building controls.

Water quality is a lifecycle process. Specify sampling, filtration, treatment, corrosion monitoring, microbial control where applicable and recordkeeping. Materials from server cold plates through facility piping must be compatible. Do not introduce an unapproved additive because it is convenient locally.

Integrate telemetry: supply/return temperature, flow, pressure, pump state, filter differential, leak detection, conductivity/quality indicators and rack load. Define alert thresholds, recipients, retention and automated protective actions. Avoid an automatic shutdown that creates a worse mission outcome without approved sequencing.

Design failure and maintenance modes

Analyze loss of one utility feed, rack PDU, server power supply, pump, CDU, facility-water loop, sensor, controller, network management path and leak-detection zone. For each, state remaining capacity, system response, workload action, alarm and recovery.

Determine ride-through time. Thermal mass and coolant volume may provide seconds or minutes, not a maintenance window. Coordinate power and cooling shutdown sequences so pumps do not stop before compute load falls. Test emergency power-off and leak response under controlled conditions.

Plan maintenance with the system running and stopped. Can a filter, pump or hose be serviced without draining the full loop? Is isolation available at rack and device level? Are caps, absorbent materials, spill kits, replacement hoses and approved coolant on site? Who is authorized to disconnect liquid lines?

Consider workload placement. Orchestration may drain nodes before maintenance or enforce rack power caps. Document the control integration and ensure facilities does not assume the AI platform will shed load unless that behavior has been tested.

Commission and document

Before compute installation, pressure-test and flush/purge the loop under the approved procedure. Verify water quality, sensor calibration, valve position, controls communication and leak detection. Inspect electrical torque, grounding, phase balance, cord mapping and protective settings.

Use staged load. Start pumps/controls, energize infrastructure, bring up nodes in sequence and run a representative workload long enough to reach thermal equilibrium. Capture rack power, branch currents, supply/return temperatures, flow, pressure, component temperatures, fan/pump behavior and alarms.

Test failures: remove a feed, fail a pump or simulate an allowed sensor/alarm condition under manufacturer/site direction. Verify workload response and recovery. Complete the acceptance record with readings and thresholds.

Deliver one-line electrical, mechanical flow diagram, rack elevation, equipment schedules, set points, alarm matrix, water baseline, test results, maintenance plan and ownership/escalation contacts. Update them after every material configuration change.

Can an 80 kW GPU rack be cooled with air?

Do not answer from rack watts alone. Some specialized facilities can deliver very high airflow or use rear-door heat exchange, but feasibility depends on server airflow, inlet limits, containment, static pressure, room/row heat rejection, noise, altitude, redundancy and adjacent loads. Direct liquid cooling often becomes more practical at high density, yet it adds water and controls interfaces. Ask the server manufacturer and qualified facility engineer to compare complete air, rear-door and direct-liquid designs for the actual site and workload.

How Uniqcli can support readiness

Uniqcli can connect the selected AI server rack to a configuration-specific power and cooling assumption sheet, integrated BOM, rack elevation and factory test. Request a facility-readiness review before hardware is ordered.

Final building design, stamped engineering and authority approval remain with qualified site professionals. The integration goal is to give them verified equipment data and interfaces early enough to avoid a delivery-day surprise.

Engineering note: Power and liquid-cooling values are safety- and configuration-sensitive. Use current manufacturer site-preparation documents and qualified electrical/mechanical professionals.

Coolant distribution, rear-door and in-row units this catalog carries

Ask AI about Uniqcli

Why buyers use Uniqcli

Related reading

InsightsBuying Guides

NVIDIA H200 vs B200 vs B300 for Government AI Infrastructure

Choosing between NVIDIA H200, B200 and B300 is not a contest to buy the newest accelerator. It is a decision about the mission workload, the supported server platform, the network and storage data path, the facility envelope, the deployment date and the evidence an agency will need to accept the system. H200 remains a capable Hopper-generation option with mature server designs. B200 moves into the Blackwell generation and a different power-and-cooling class. B300, also called Blackwell Ultra at the GPU generation level, increases memory and arrives in several system forms—including HGX B300 servers and the much denser GB300 NVL72 rack.

· 7 min read

InsightsBuying Guides

Is NVIDIA H300 Real? H200, B300 and GB300 Explained

As of August 25, 2026, NVIDIA does not list a current product named “NVIDIA H300” in its official AI Enterprise support matrix or current data-center platform documentation. The search term usually reflects a mix-up between H200, the Hopper-generation GPU, and B300, the Blackwell Ultra GPU. It can also be a mistaken shorthand for GB300, the Grace Blackwell Ultra superchip and rack-scale systems built around it.

· 6 min read

InsightsBuying Guides

NVIDIA H200 Price: What a Government Buyer Actually Needs to Budget

There is no durable, universally valid “NVIDIA H200 price.” An agency does not deploy a bare headline price; it deploys a configured server or appliance with CPUs, memory, local storage, NICs, fabric, rack power, cooling, software, integration, support and a data path. Availability, warranty, OEM configuration, delivery location and acquisition path can change the quote materially.

· 6 min read

About the author

Uniqcli Team

Uniqcli's newsroom, buying guides and glossary are produced by our in-house team — seven procurement and technology professionals who source, screen and integrate IT and security hardware every day, working with two editors. Practitioners draft from live sourcing and integration work; editors review every piece for accuracy and plain language before it publishes.

More about the Uniqcli Team

Ready to scope your program?

Talk to a Uniqcli engineer, or send a bill of materials for a TAA-verified quote — no payment up front.