GPU Clusters Are a Power Problem Before They’re a Compute Problem

Want to know why brand new AI halls sit dark for years?

They seldom fail to arrive on schedule because of GPUs. The GPUs arrive. They are unpacked, racked and wired…. And then they sit there, because no one figured out how to feed them first.

Here’s the uncomfortable part:

Procuring compute is easy. Powering that compute effectively is the hard part. And most organisations figure that out well after signing the purchase order.

What you’ll walk away with:

  1. Why The Real Bottleneck Is Electrons, Not Silicon
  1. What 140 kW Racks Did To The Building
  1. The Long Lead Items That Decide Your Timeline
  1. How To Phase A Modular Switchgear Lineup
  1. Why Power Now Drives Site Selection

Why The Real Bottleneck Is Electrons, Not Silicon

A GPU order ships in months. A utility connection does not.

That’s where schedules fail, silently. Worldwide data centre electricity demand is forecast to double to approximately 950 TWh by 2030, with AI-intensive facilities expanding several times more than the overall average. Each megawatt needs to eventually end up somewhere material — on a feeder, past a breaker, into a busway, out to a rack.

You want clusters for machine learning? Okay. But first lesson: so the first question you ask on any cluster project is not “how many GPUs?” No no no… it’s MUCH more boring than that:

Can the building actually take the load — and when?

That answer resides in the electrical room. An MV switchgear solution takes your incoming medium voltage feed and conditions it into manageable, protected, switchable power for your entire hall. Modular switchgear lineups are designed to be built in modules allowing capacity to be added phase by phase as you need it instead of purchasing for maximum capacity on day one. In cases where space is at a premium, medium voltage gas insulated switchgear does the same thing but in a fraction of the space, important when cooling units are fighting for every square metre.

Get that layer wrong and the compute plan on the whiteboard means nothing.

What 140 kW Racks Did To The Building

Rack density is the thing that broke the old playbook.

An old school enterprise rack used to pull between 5kW and 10kW. Now a single NVIDIA GB200 NVL72 rack can draw 120kW to 140kW by itself. That’s 1 cabinet consuming more power than a whole row once did.

Think about what that changes:

  • Distribution: require more copper, larger breakers and upsized busway for higher currents than the old halls ever saw.
  • Cooling: when air stops working above ~100 kW/rack, liquid cooling becomes the default — and pumps/CDUs have their own power & protection requirements.
  • Redundancy: one fault on a dense row now destroys many times more compute than before.
  • Space: electrical and mechanical rooms expand, your available white space shrinks.

Neither does the load play nice. Training runs have thousands of GPUs spike to near-full draw simultaneously, and cut right back down. Those transients are tough on protection settings and tough on the utility supplying you.

Density is awesome when you look at dollars-per-square-metre. It’s murderous on 70’s era infrastructure.

The Long Lead Items That Decide Your Timeline

Here’s the part that catches people out.

Power mobility equipment is becoming more scarce than compute equipment. Lead times for substation transformers have increased from approximately 140 weeks in 2023 to 160 weeks in 2026, while switchgear lead times are also far above historical averages.

Then there’s capacity value. Four to seven years are being reported in interconnection queues in the busiest markets – Northern Virginia, Phoenix, Dallas. Four years is greater than the useful life of the GPUs under consideration for a site.

None of this is solved by throwing more money at the problem, either. Lines are being added at factories, but skilled labour, electrical steel and testing bays limit how fast production can ramp up.

So the sequence most teams use is backwards. It usually looks like this:

  1. Approve the compute budget
  1. Choose a site
  1. Design the building
  1. Order the electrical gear
  1. Discover the gear arrives after the lease starts

Power should come first. Equipment slots second. Compute third. Anything beyond that is speculation disguised as a project plan.

How To Phase A Modular Switchgear Lineup

Now for the practical bit.

The beauty of a modular switchgear product lineup is that GPU clusters don’t need AI load until they have AI load. Rather than coming all at once, that load will come in increments – a hall here, a row there, another tenant next quarter. Don’t build one huge distribution system for load that won’t be online for three years. That wastes money and locks in decisions too far in advance of anyone knowing what the racks will actually look like.

Phasing takes care of that. Finalise the cubicle design once, then replicate it. Batch your orders large enough to reserve factory slots early. Slot some empty spaces into the queue so downstream feeders can snap on later without downtime on live compute.

A few things worth locking down before anything is ordered:

  • Voltage class and fault rating — these dictate everything downstream and are very hard to change later
  • Footprint — air insulated gear needs more room; gas insulated saves it
  • Spare capacity — empty sections cost far less now than a retrofit later
  • Protection and metering — dense racks need visibility per feeder, not per building

Do that and expansion becomes a scheduled task instead of a rebuild.

Why Power Now Drives Site Selection

Site selection used to be about land, fibre and tax. Now it’s about megawatts.

US data centre growth is forecast to increase capacity from approximately 24 GW to around 100 GW from 2026 to 2030. The growth is fighting with utilities, manufacturers and renewables developers over transformers, breakers and switchgear.

The sites winning right now share a few traits:

  • An executed interconnection agreement, not a queue position
  • An energised substation nearby with genuine headroom
  • Room on site for the electrical build to expand
  • A utility relationship that started before the design did

Strip centers, other older industrial buildings with oversized electrical service are suddenly worth a lot of money for reasons completely unrelated to the building itself. It’s the feed that’s valuable. Speed to power is your only schedule advantage that money can consistently purchase.

Tying It All Together

GPU clusters look like a compute problem. They behave like a power problem.

Quick recap:

  • The chips are available — the megawatts often are not
  • 140 kW racks changed the electrical and cooling design completely
  • Transformers and switchgear now sit on multi-year lead times
  • A phased, modular switchgear lineup lets capacity grow with demand
  • Power availability should be settled before the compute budget is

Sort the power path first. The compute is the easy part.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *