AI Infrastructure: Power and Cooling Become a System-Level Challenge

September 28, 2026

The expansion of artificial intelligence is changing not only servers and processors, but the physical architecture of data centres. As the power density of modern AI and high-performance computing systems increases, so do the demands placed on grid connections, power distribution, switchgear and thermal management. Compute capacity alone is no longer the defining constraint. What increasingly matters is whether electrical power can be supplied reliably, heat can be removed efficiently and the supporting infrastructure can scale at the same pace as the IT load.

AI is becoming an energy and thermal-management question

The growth of generative AI is making the relationship between digital and physical infrastructure increasingly visible. As high-performance compute clusters expand, demand rises not only for processors and network equipment, but also for electrical capacity and heat rejection.

For operators, annual energy consumption is only one part of the problem. A more immediate question is whether a site can provide the required electrical power reliably and distribute it safely within the facility. Grid connections, transformers, switchgear, uninterruptible power supplies and backup systems all have to be matched to the planned IT load. Large AI clusters may also create more dynamic electrical behaviour than conventional computing environments, with rapid changes in load placing additional demands on power quality and control.

The thermal challenge develops in parallel. The more computing capacity is concentrated into a limited space, the more heat must be removed from that space. Conventional air cooling remains suitable for many IT environments, but as rack power densities increase, cooling by air alone becomes more difficult. High-density AI and HPC systems are therefore accelerating interest in liquid and hybrid cooling architectures.

There is no single technical solution. Direct-to-chip cooling removes heat close to processors and other high-load components. Rear-door heat exchangers capture heat at rack level, while immersion cooling may be considered for selected applications. Each approach changes the infrastructure surrounding the IT equipment by introducing additional elements such as Cooling Distribution Units, pumps, heat exchangers and liquid circuits.

That changes the engineering question. Cooling is no longer simply a building service added after the IT design has been completed. It becomes part of the computing architecture itself.

Power, cooling and building design are converging

The most important shift is therefore not one specific technology, but the closer coupling of disciplines that were previously planned more independently.

Higher rack density can reduce the amount of floor space required per unit of computing capacity, but it simultaneously increases demand on electrical distribution and heat removal. Liquid cooling systems require additional space, monitoring, maintenance access and redundancy concepts. Heavier racks, pipework and technical equipment can also alter the structural requirements of a facility.

Power systems, thermal management, IT architecture and building design increasingly need to be considered together from the earliest project stages.

For operators, this also raises the question of future density. A data centre designed around today’s rack loads may need to accommodate substantially higher loads only a few years later. If the supporting infrastructure cannot scale accordingly, later upgrades can become technically complex or economically unattractive.

This is particularly relevant for existing facilities. Not every data centre can absorb additional electrical capacity, new liquid-cooling circuits or heavier rack configurations without significant modification. Retrofit projects will therefore often require a mixed architecture in which existing infrastructure continues to serve lower-density areas while new high-performance zones are created for AI workloads.

Resilience extends beyond servers and networks

As AI becomes more deeply embedded in business-critical processes, the resilience of the underlying infrastructure becomes increasingly important.

A failed AI cluster does not necessarily indicate a problem with processors or software. Transformers, switchgear, UPS systems, pumps, heat exchangers and cooling systems can all become availability constraints.

This broadens the conventional concept of redundancy.

Redundant electrical supply offers limited protection if the cooling architecture contains a central single point of failure. Conversely, a highly resilient cooling system cannot compensate for dependence on a common switchboard or control path. Availability therefore has to be considered across the complete power and thermal chain.

Monitoring becomes part of this resilience model. Electrical loads, temperatures, flow rates and operating states need to be visible in combination rather than as isolated data sets. As IT and facility systems become more closely coupled, fault detection, capacity management and maintenance increasingly depend on a common operational view.

For security and critical-infrastructure operators, this convergence is particularly relevant. The operational continuity of digital systems depends increasingly on equipment that traditionally sat below the IT layer but is now just as critical to service availability.

Location determines what is technically possible

High-density computing also makes site selection more important.

Grid capacity is an obvious constraint, but it is only one of several. Climate, water availability, opportunities for heat reuse, available space and proximity to energy and data networks all influence the technical design.

In cooler regions, free cooling can reduce the need for mechanical refrigeration for parts of the year. In hotter and drier climates, the problem is different: high thermal loads have to be managed reliably without creating excessive water demand.

This makes regional differences particularly relevant for Europe and the Middle East.

A design that works well in Northern Europe may not be appropriate for Southern Europe or Gulf markets. Ambient temperature, humidity, water availability, electricity costs and local infrastructure conditions can alter the balance between air cooling, liquid cooling and other heat-rejection strategies.

The site therefore becomes part of the engineering architecture. The relevant question is not simply whether enough physical space exists for additional servers, but whether sufficient power can be provided and the resulting heat can be removed under the actual climatic and infrastructural conditions of the location.

Scalability becomes an industrial and supply-chain issue

The expansion of AI capacity is not only an engineering challenge. It is also a question of industrial availability.

Transformers, switchgear, UPS systems, power-distribution equipment, pumps, heat exchangers and Cooling Distribution Units all have to be manufactured, tested, transported and installed. As projects grow, the availability of this equipment becomes increasingly relevant alongside the supply of GPUs and servers.

This helps explain the growing use of modular and prefabricated infrastructure. Power and cooling assemblies that are integrated and tested before delivery can reduce on-site work and allow capacity to be added in phases.

The challenge lies in balancing standardisation with project-specific requirements. Data centres vary significantly in power demand, redundancy philosophy, site conditions and expansion plans. At the same time, operators face pressure to bring new capacity online faster.

The supply chain therefore becomes part of the data-centre strategy. The economic value of an AI cluster begins only when it enters productive operation. Delays in grid connections, switchgear, transformers or cooling infrastructure can therefore have the same commercial effect as delays in server delivery.

Efficiency becomes a design criterion, not a single metric

Growing energy demand also puts greater emphasis on efficiency.

It is not sufficient to consider IT consumption in isolation. Cooling, power conversion, pumps and other supporting systems contribute to total facility demand. At the same time, improvements in component efficiency can be outweighed by rapid growth in installed compute capacity.

Operators therefore have to manage two objectives at the same time: provide more power while limiting the infrastructure overhead associated with each unit of computing capacity.

Heat reuse can form part of that discussion. Where suitable users and temperature levels exist, waste heat from data centres can potentially be integrated into local heating networks or other thermal applications. Its usefulness, however, depends strongly on location, cooling architecture and the proximity of potential heat consumers.

Efficiency is therefore less a standalone figure than the result of decisions about site, cooling concept, electrical design and operational strategy.

The market is expanding supporting capacity

The shift in demand is beginning to influence industrial investment. One current example is Vertiv, which plans to expand its facility in Nové Mesto nad Váhom, Slovakia, by around 22,000 square metres over the next 18 to 24 months. The additional capacity is intended for power, switchgear and thermal-management technologies, including liquid-cooling solutions for high-density AI and HPC applications. The project illustrates the broader increase in demand for industrially available power and cooling infrastructure for high-density data centres.

AI infrastructure is becoming more physical

The next phase of AI deployment will not be determined by processor performance alone.

As compute clusters become larger and denser, the boundaries between IT infrastructure, electrical engineering and thermal management become increasingly difficult to maintain. Operators can no longer treat computing capacity as an isolated procurement decision. It has to be planned together with power availability, cooling, redundancy and the long-term operating model.

This also changes where bottlenecks can emerge. A shortage of GPUs may delay a project, but so can a missing grid connection, unavailable transformers, insufficient switchgear capacity or a cooling system that cannot support the planned rack density.

The consequences are strategic as well as technical. Data-centre operators will have to plan further ahead, manufacturers will have to increase industrial capacity and existing facilities will need clearer decisions about how far they can be upgraded before a new site becomes more viable.

As AI becomes more deeply embedded in business-critical processes, the physical foundations of digital infrastructure become harder to ignore.

Digital intelligence depends on real power, real cooling and real industrial capacity. The performance of future AI infrastructure will therefore be determined not only by models and processors, but by how effectively those physical systems are designed, supplied and operated as one integrated architecture.

Related Articles

DIN EN IEC 62676-4: Video Security Starts with the Security Objective

More megapixels do not automatically mean better security. The forthcoming DIN EN IEC 62676-4:2026-10, due for publication in October 2026, starts from a more fundamental question: what task is the video surveillance system actually intended to perform? The revised...

Share This