No items found.
No items found.

AI Performance Starts in the Data Centre

Designing AI Data Centers
8
Min Read
August 19, 2026
Share Article

Some of the most consequential performance decisions in an AI infrastructure project are made several months or years before the first GPU is installed. It could concern an electrical topology, the route of a cooling pipe, the position of an isolation valve or the telemetry exposed through the building management system. None of these decisions appears on a GPU specification sheet, yet each can determine whether the hardware delivers its intended performance consistently or spends part of its operating life constrained by the facility around it.

Hardware Sets the Ceiling. The Facility Determines What Gets Delivered.

Every generation of AI compute arrives with a specification prescribing what it can achieve under the conditions for which it was engineered: a defined power envelope, a stable thermal regime and a set of assumptions about how electricity reaches the rack and heat leaves the silicon. Those conditions are not incidental to the performance figures. They are the terms on which those figures hold.

The same hardware can therefore produce materially different outcomes in two different facilities. In one, the cooling system maintains stable temperatures through peak demand, electrical capacity is available at the required density and critical components can be maintained without interrupting live workloads. In another, thermal or electrical headroom runs short, capacity must be derated and some of the capability engineered into the hardware remains inaccessible.

GPU vendors have raised the performance ceiling dramatically, but in doing so they have also raised the standard required of the building beneath it. Denser racks, liquid-cooled systems and increasingly integrated power architectures transfer more responsibility to the data centre. As the hardware becomes more capable, the tolerance for infrastructure that cannot support it becomes smaller.

For a hyperscaler or enterprise buying AI capacity, the important question is therefore not simply which GPUs are installed. It is how much of their capability the provider can deliver, continuously and predictably, over the life of the contract.

The Infrastructure Customer Is Buying Continuity, Not Installed Inventory

An AI infrastructure agreement may appear to describe compute, storage and networking, but what ultimately buyers want is continuity: confidence that capacity will remain available and usable through planned maintenance, component failures, staffing changes and phased expansion, with service restored predictably when an interruption does occur. That continuity is a facility outcome. An availability commitment reflects the electrical and mechanical topology; an incident-response target depends on whether qualified engineers are on-site if needed; and timely customer notification requires monitoring that can connect a physical fault with the workloads it affects.

Continuity does not come from any single redundant component. It emerges from a chain of physical and operational controls, not just CapEx but OpEx too, which is why every external commitment must be translated into a facility-level obligation. A capable provider should be able to trace one availability clause backwards through the redundancy architecture, maintenance regime, staffing model, spares inventory, monitoring thresholds, escalation paths and change-control procedures that make it credible.

Where that traceability exists, continuity has been engineered into the facility rather than added to the contract as a promise. Without it, two providers may offer similar hardware and headline availability, but only one may be able to maintain, repair and expand its infrastructure without unnecessary disruption. The difference becomes clear when the facility is placed under pressure.

Continuity Is Decided Before the Building Is Finished

The difficult aspect of operational continuity is that it cannot be added neatly at the end of a project. By practical completion, many of the decisions that determine how safely and reliably a facility can be operated for the lease duration have already been fixed, often by teams whose responsibility ends at handover.

Does the design ensure concurrent maintainability of plant and equipment without compromising capital efficiency? Is there enough physical access for an engineer to reach, remove and service it safely? Does the building management system expose the information operators need to diagnose emerging faults through a planned, preventative maintenance approach that monitors asset performance? Can a later phase be energised and tested while an earlier phase carries customer workloads? 

These are operational questions expressed through design, each of them relatively inexpensive to resolve while the facility exists as a coordinated set of drawings and considerably more expensive once the answer has been cast in concrete, installed behind live plant or embedded in a control system. 

The operator therefore needs to enter the project early enough to change the design rather than simply learning how to work around it. In practice, that means appointing the facility management and security partner during concept design and spatial coordination, giving named operational leaders formal responsibility for reviewing operability and maintainability, and tracking their findings through to closure.

This reverses the conventional sequence in which an operator inherits a completed building and adapts its procedures to the limitations left behind. It is more demanding to procure because it requires operational expertise years before there is a facility to operate, but it produces an operating model designed into the building rather than retrofitted after handover.

Redundancy Is Not the Same as Resilience

Resilience is often reduced to a configuration such as N, N+1 or 2N, as though the central question were simply how much equipment to duplicate. Redundancy matters, but it does not describe the whole architecture of failure.

The earlier and more consequential decision is how capacity is divided. A single, very large facility and several smaller, independent buildings may present identical aggregate capacity and similar redundancy on paper, yet behave very differently when something goes wrong.

Independent buildings designed and commissioned as separate fault domains can contain the impact of a serious incident. They also allow capacity to be energised in stages, reducing dependence on a single construction critical path and bringing portions of the campus into service without waiting for the entire programme to complete.

The trade-off should be stated plainly. Independence requires duplicated plant, separate commissioning and greater coordination. It also creates a complex operating condition in which one building may carry live customer workloads while neighbouring phases remain under construction or commissioning.

That model only works if the campus operates as one controlled estate: one asset register, one governance structure, one set of operating procedures and one change-management system. Otherwise, apparently independent fault domains can become inconsistently managed facilities that share a postcode but gradually diverge in practice.

Higher Density Changes the Operating Model

The move towards liquid-cooled, high-density AI systems is often presented as a facilities-engineering challenge, but the consequences extend well beyond design and commissioning because the compute platform becomes much more closely tied to the mechanical plant that supports it. Once water-bearing infrastructure enters the white space, leak detection has to connect tray, rack and facility-level systems, while cooling distribution units, pumps, manifolds and controls all become part of the operational chain behind every live workload. A maintenance lapse that might once have reduced cooling efficiency can now affect the availability of an entire rack-scale system.

The market is compounding operational challenges; the denser the white space, the more GPUs per square meter. The more GPUs installed the more cabling required to interconnect, the more cabling the less room/flexibility for maintenance of fibre connections therefore more risk to availability SLAs, and so, designing for shorter and more efficient cable routes become a critical operational consideration something that historically was never front of mind during DC design in lower density scenarios. 

That closer relationship makes condition-based maintenance increasingly important, with operating thresholds agreed with equipment manufacturers during commissioning, when normal system behaviour can be established, rather than devised after faults begin to appear. It also requires a site-dedicated workforce with practical experience in high-density power and liquid cooling, capable of handling routine interventions and first-line diagnosis without relying on external specialists to travel to the facility.

The shortage of that expertise is becoming a programme risk in its own right. The industry is commissioning high-density infrastructure faster than it is producing engineers qualified to operate it, which means recruitment lead times, training pathways and competency assurance deserve the same scrutiny as long-lead electrical equipment.

Uptime Institute’s latest analysis found that failures to follow procedures had become an even greater contributor to outages, even as infrastructure resilience improved overall. However sophisticated the plant or extensive the redundancy, continuity still depends on whether the engineer responding under pressure has the right information, training and procedure to deal with the fault.

Audit Reveals How the Facility Really Operates

Enterprise customers do not have to accept operational maturity on trust. They can test it through the evidence a facility produces during normal operations.

Change management is often the most revealing place to begin. Who authorised a change? How was it classified? What systems and customers could it affect? Was there a tested rollback plan? If it crossed multiple buildings, was it reviewed as a campus-wide change? Does the record show what was planned, what happened and what was learned?

The same principle applies to maintenance histories, access logs, incident reports, training records and periodic reviews of access rights. In a mature operation, this material exists as a natural by-product of the way work is conducted. It can support customer audits and certification programmes without requiring the organisation to reconstruct its own history.

An immature operation behaves differently. Documentation is assembled in the weeks before an audit, inconsistencies appear between systems and important records begin suspiciously close to the date on which someone first asked to see them.

The useful question is not whether the provider has a policy. It is whether the evidence shows that the policy has governed the facility since mobilisation.

Five Questions Every AI Infrastructure Buyer Should Ask

For enterprises evaluating AI infrastructure providers, five questions can separate installed capacity from dependable capability:

  1. Trace one availability commitment. Select a clause in the service agreement and ask the provider to identify the facility-level obligations that support it, including topology, maintenance, staffing, monitoring and escalation.
  2. Ask when the operator was appointed. Did the operating partner join while the design could still be changed, or after practical completion? What specific design decisions did its involvement influence?
  3. Understand the fault domains. How is capacity divided, what can fail independently and what would a serious incident in one part of the campus mean for workloads elsewhere?
  4. Find out who responds at three in the morning. Are engineers with high-density power and liquid-cooling competence available on-site if needed, or does the first response depend on an external escalation path?
  5. Request operational records, not policy documents. Review actual change, maintenance and incident records extending back to mobilisation.

How Radiant Designs Continuity into the RIBA Process

At Radiant, we have structured our current campus procurement around this principle, using the RIBA Plan of Work as an operational control mechanism rather than simply a reporting convention. By dividing a project into defined stages, the framework makes clear when decisions remain relatively easy to change and when they are about to become permanent and expensive.

Our contractors and suppliers for MEP, CSA, facility management and security are appointed no later than Stage 2 and remain involved throughout Stages 3 and 4. This means (although not accountable for the design) the functional team responsible for delivery, commissioning and operations of the designed space has been formed early in the project, relationships formed and interfaces mitigated as early as possible. The sequence below shows how that involvement develops from early design review to an operating model ready for practical completion.

Radiant's RIBA based Data Center Design

Across these stages, our civil, structural and architectural partner, mechanical and electrical partner, and operator are procured in parallel and coordinated through a joint design-review forum. Each stage gate requires formal approval, allowing interfaces between structure, plant and operations to be resolved while they remain drawings rather than emerging later as change orders or interruptions to live capacity.

Before mobilisation, a single operational interface document defines the boundary between the responsibilities of Radiant appointed facility management and those of Radiant’s in-house team. By Stage 4, the maintenance schedules, spares strategy and monitoring thresholds are also ready to activate at practical completion.

This model requires earlier investment in operational expertise and more coordination across several connected procurements. Our judgement is that both costs are lower than discovering at handover that a facility performs as specified but is difficult to maintain, expand or operate with the continuity promised to customers.

Conclusion: Designing for Continuity Sustains Performance Over Time

AI performance is not determined by hardware alone. It depends on a facility designed to maintain stable power, cooling and operations through maintenance, component failures and phased expansion. That continuity must be established early, while plant topology, maintenance access, fault domains, telemetry and operational controls can still be shaped around live customer workloads.

Treating the data centre as an integral part of the AI system design creates the conditions latest-generation GPUs need to perform consistently throughout their operating life. As silicon vendors continue to raise the performance ceiling, reaching it consistently will depend on whether continuity has been designed into the data centre from the outset.

FAQs

No items found.

How To's

No items found.

Related Articles