The Boundary That Makes an AI Factory Repeatable

The boundary for a repeatable AI factory
13
Min Read
October 5, 2026
Share Article

As we noted in our initial post in the series, a first principles repeatable AI factory with an agnostic input power strategy should survive a change of power source with minimal impact to the data center design, buildability and operability.

Across a datacenter portfolio, there will be a real diversity of inputs: different utilities, generation technologies coupled with supply constraints and routes to expansion. Each of these options arrives with its own voltage level, fault behavior, protection scheme, power quality, and interconnection obligations. Left alone, those differences cascade into bespoke distribution, controls, and commissioning at every site.

In the opening post, we argued that those differences should leave as much of the factory design intact as possible. As we work our way in from the outside the first stop on that journey is the electrical architecture and in many ways it is the most defining.  

Which decisions survive a change of source? What evidence supports carrying them forward? And what does preserving that freedom cost?

Our position is that source flexibility belongs in the original architecture. A developer should be able to add supply options over time without repeatedly revisiting the factory’s electrical distribution, controls, commissioning and operating model. Furthermore, the work required to commission additional energy/power sources on a campus should become more predictable as the portfolio grows.

This is not easy by any means. It requires investment in interfaces, models and qualification before all the variables are fully established. It also requires restraint: preserving an option is worth paying for only when its commercial value exceeds the cost of keeping it available.

Building in versatility in the early design stages requires our team to be smart with expansion capability that is baked into the initial stages of our AI factory architecture It is true this comes with small capex offsets during the enablement phase of a project but the upside is a significant amount of opex and capex saved in the future, where changes during operations are impossible. This consideration is critical to future proofing a campus - especially when exponential silicon evolution puts the longevity of all development projects starting from scratch, at risk.

In this case, the return is broader than access to more power. 

A stable handoff lets a project commit capital to parts of the factory while specific upstream choices remain open. Done well, it changes the sequence in which the business can make decisions.

Decide what a source change is allowed to reopen

We should judge a source boundary by the downstream decisions it protects and the cost of doing so.

Putting more conversion, storage or reserve generation upstream may make the factory easier to standardize. It may also have the effect of putting the project underwater economically. At the other extreme, passing every source constraint through to the factory spreads adaptation across equipment, controls, workload commitments and operating procedures. A lower equipment price can leave the operator carrying tighter restrictions on usable compute.

The Adaptation Boundary

The adaptation boundary

Drag the boundary. Everything left of it is absorbed before the factory. Everything right of it is passed through to it. Both extremes have a price.

Conceptual model. Curve is illustrative, not measured data.
Power
Source
→
Absorb upstream
conversion · storage · reserve generation
Pass through
equipment · controls · workloads · procedures
Boundary
→
AI
Factory
All absorbed upstream All passed to factory
Total exposure vs. boundary position
workable range upstream in factory
Exposure is lowest through the middle band, where enough is absorbed to keep the factory reusable and enough is passed through to keep capital in line. It climbs toward either end.
Capital cost, upstream
Conversion, storage and reserve you build before the factory. Overbuild and the project can go underwater.
Compute risk, downstream
Adaptation spread across equipment, controls and workloads. Pass through too much and usable compute gets restricted.

‍

The design task is to locate the boundary where the total cost of adaptation and reuse makes sense across the sites the business actually expects to develop.

That means starting with a credible range of supply arrangements and a specified factory service, then testing the cost of making the two meet. Solar, wind, fuel cells, biomass, nuclear, gas and hydro all belong in the opportunity set. Whether they progress from opportunity to evaluations depends on the complete supply arrangement meeting the requirement, including whatever conversion, firming and controls it needs. The platform builds a library of validated power supply configurations over time. Each new project starts from that library. When a new power source is proposed, we check whether it fits an existing design/test-fit and its operating limits, or needs further investigation, engineering and testing.

There needs to be a clear distinction between a source that fits an existing service power specification and one that changes it.

Suppose a new arrangement requires tighter limits on factory load ramps or outages because of maintenance. If the factory already operates within those limits, the change may remain upstream. If meeting them requires a new scheduling policy, a different storage duty or a restriction on customer workloads (impacting SLAs that can be achieved), the change has crossed the boundary. When comparing power supply options, we need to include the cost of changing compute schedules and storage operation, along with any revenue lost to restrictions on customer workloads.

This is where an interface specification becomes commercially useful. It identifies the assumptions under which downstream commitments remain valid. A proposed supply arrangement either preserves those assumptions or comes with an explicit price for changing them.

Source proposals should include that scope of change alongside their energy economics. It gives the developer a basis for comparing the engineering, equipment, qualification and operating consequences before selecting the supply.

Specify the service the factory accepts

The power specification should state what each source must deliver and what the factory’s power system must guarantee to downstream equipment. We can then compare supply proposals on the design changes they require, alongside energy price and delivery date. Without that specification, we risk discovering those changes after procurement, when they become change orders. Without that specification, we risk discovering those changes too late in the project lifecycle, when LLE (long lead-time equipment) procurement is on order.

The power specification has three parts.

The first is electrical. It fixes the voltage and frequency range at the handoff, the capacity increments the factory draws in, the fault level and short-circuit contribution the downstream protection is designed for, the grounding arrangement, the harmonic and power-quality limits, the power factor, the ride-through behavior, and the ramp rates the source has to hold. For example, voltage THD for grid-tied generation at medium voltage should remain below 5%, a baseline for off-grid AI factories to reduce strain on downstream UPS and BESS equipment, preserve equipment life, and limit maintenance costs. The power specification also states the redundancy philosophy: what the factory assumes about availability and what it does when the source falls short.

The second is controls and data. A repeatable factory reads its supply through a common signal model. That means agreed telemetry points, timestamp conventions, a shared state and alarm taxonomy, power-quality measurement, and a defined set of dispatch and load-management commands. It also means a cybersecurity boundary that does not move every time the source changes. 

Two sites that report power in different vocabularies cannot share a commissioning script or an operating procedure. Beyond the cost of procuring and operating multiple tools across the portfolio for the same function, the commissioning effort is also a critical factor in achieving the speed to market when deploying a repeatable AI factory.

The third is commercial and operational. The power specification states the guaranteed capacity profile, the maintenance states the source can enter and what the factory does during them, the restoration sequence after a loss, the black-start assumptions where they apply, and the acceptance tests that prove a new source meets the power specification before it carries load.

These values are engineered per jurisdiction and per topology. A grid code in one country sets a different fault level from another. An onsite generation package has a different dynamic behavior from a utility feed. The power specification is a defined set of categories and a decision logic for filling them. Every site fills the same categories from its own source, and each presents the factory with the same kind of interface and the same acceptance evidence.

Each power proposal should show how the supply will meet the specification. If changes to the supply or the factory are required, the proposal should identify them and state their cost and effect on schedule and operation.

Reuse depends on the assumptions behind the design

Before we reuse a design, we need to know which studies and test results still hold. The equipment may be identical, but a different power supply or different operating conditions can change the real world performance.

Each design should record the assumptions behind its studies and tests. A protection study may assume a particular fault contribution. A recovery sequence may require a particular load ramp. If a new power supply changes either, we should know which work needs to be repeated and which results we can still use.

We want the qualification package to make that decision clear: what still applies, why it applies and what needs to be checked again. Otherwise, we risk paying to repeat work we shouldn’t, or pay for changes based on results that no longer support it.

Across a portfolio, this becomes a library of power supply configurations we have already qualified. Each entry records the factory version, operating conditions, assumptions, models and acceptance criteria used in that work. The next project can use this record to identify what carries over and where new engineering is needed.

That record also helps us act on problems found in operation. When we correct a control problem at one site, we should review other sites that rely on the same assumptions. The correction may require different settings at each location. Sharing the reasoning behind the fix lets each site check whether it applies. Copying settings without that context can spread a local problem across the fleet.

Price the transition out of bridge power

Let’s take our initial example of a campus that starts on generation and later connects to the grid. There are three things to price:

  1. The first phase of campus configuration
  2. The built-out campus configuration
  3. The work required to move between them

The transition (#3) can involve equipment, commissioning of new controls and renewed acceptance. Its cost is easy to understate when the initial generation package and the future utility connection sit in different contracts and budgets.

Bridge proposals should be evaluated against the intended grid-connected configuration from the outset. A proposal that reaches first power earlier may still be the better choice even with a substantial transition cost. But that cost has to be visible when the choice is made.

The same applies to the role of the original assets after connection. Generation and storage might remain in service, be redeployed or be retired. Each outcome implies a different initial specification and residual value. A future grid-services role is worth including when the operating rights, asset capabilities and commercial route support it. Until then, it is an option rather than dependable revenue.

NVIDIA and Emerald AI’s March 2026 announcement describes this lifecycle explicitly: colocated generation and storage would provide bridge power, then remain in use after the campus connects to the grid. Their proposal combines those resources with software that adjusts compute workloads, helping the factory reduce its draw from the grid when needed.

For procurement, this suggests a more useful form of optionality. Ask suppliers to identify the enabling work that must be done now, the equipment that can be added later, the outage required, and the evidence that must be renewed.

We recommend pricing the delayed-grid case as well. An arrangement that is attractive for a short bridge period may be a poor basis for several years of operation - and at Radiant what we build is built to last.

This will change which proposals look competitive but has the added benefit of giving the construction and operations teams a transition they can plan for.

The key lever Radiant can pull in these types of projects is to form the project team as early as possible; Design, Construction, Delivery, Operations so that these long term considerations are designed for at the genesis of the project – it is the only way to make sure all eventualities are planned, designed and costed for. The traditional model of handing over each stage sequentially is a tried and tested approach to reduce capex at each milestone but when you are developing and operating AI factories you cannot afford to start considering operational risk too late in the delivery of the project.

Flexibility changes the service being sold

A flexible power supply is worth more when the compute load can flex with it. A site that can pause or throttle jobs can run on a source that a constant load would break.The capability is real. But it only counts when the jobs running at that site can actually be delayed, and it cannot be sold to a customer as firm capacity.

Emerald AI and NVIDIA, tested the compute side of this approach in an earlier field demonstration published in Nature Energy. They coordinated representative AI workloads on a 256-GPU cluster in Phoenix, reducing cluster power consumption by 25 percent for three hours while maintaining the tested quality-of-service guarantees. The result was achieved without additional energy storage or hardware modifications.

Overview of the Emerald Conductor Architecture

This tells us it is possible to reduce power demand by adjusting workloads, the question becomes how much of that capability the business is prepared to commit, to whom, and on what terms.

If a supply arrangement depends on curtailment, someone has to own the corresponding right to defer or constrain compute. That right cannot exist only in the energy model while the capacity is sold under commitments that prevent its exercise. 

Equally, an operator should not give away valuable workload flexibility without recognizing what it saves in power infrastructure or earns in access to capacity. Conventional delivery models cannot provide this flexibility because contracts across providers do not fully align. Exercising it requires both the authority to act and direct ownership across power, infrastructure, and compute, giving the operator control to optimize AI factory and GPU delivery around the project’s full set of constraints.

There is a less obvious dependency here. The electrical assets can stay unchanged while the workload mix drifts. A site whose power plan assumed its jobs could be deferred may start serving work with tighter latency or completion requirements, often inference rather than training. That work cannot be curtailed, so the flexibility the power arrangement counted on is no longer there. The one-line looks the same. The operating freedom behind it has narrowed.

For that reason, the approved flexibility power specification should be part of capacity allocation and customer contracting. A change in workload commitments can require the same scrutiny as a change in generation. Finance, energy and compute operations need a common view of the capacity that is actually available under each condition.

Energy storage introduces a related allocation problem. If the energy case depends on dispatching a reserve that the resilience case assumes will remain available, the investment model contains incompatible commitments. The resolution may be additional storage, a narrower dispatch obligation, more generation or a different compute promise. It should be selected on the economics of the whole service. Adding energy storage is capital intensive, can trigger additional safety and fire requirements at certain capacity thresholds, and requires space that constrained sites may lack, particularly in Europe.

We need to evaluate power savings alongside their effect on the compute business. If cheaper power requires us to delay jobs, reserve capacity to catch up on deferred work, or limit what we can promise customers, those costs belong in the same investment case.

The migration choice behind 800 VDC

The move toward 800 VDC gives this argument a current procurement test.

OCP’s August 2026 update from Google, Microsoft and NVIDIA describes two deployment paths: local conversion beside the racks using existing 480 VAC infrastructure, and conversion from medium-voltage AC into facility-scale 800 VDC distribution. The first preserves upstream infrastructure where AC capacity and row space are sufficient. The second changes the distribution architecture. OCP distinguishes the published ±400 V sidecar specification from subsequent native 800 V development.

Those paths preserve different assets and create different commitments. A project should choose between them with a view of the expected compute generations, the condition of the existing infrastructure and the cost of the eventual migration.

The useful procurement question is: “what does an offered design preserve when the next change arrives. Which distribution assets remain? Where does conversion move? Which interfaces and protection studies reopen?”

The answer determines whether a later upgrade is contained within a defined part of the facility or spreads into an operational rebuild.

OCP’s solid-state-transformer specification takes an interface-based approach to the conversion equipment, covering expected behavior as well as electrical connections. Revision 0.3 still leaves parameters for further definition and testing. Those open items matter when deciding which procurement assumptions are firm enough to commit against.

The protection work also constrains how far standardization can go. OCP’s LVDC white paper addresses fault contributions from DC-link capacitors and the coordination required to isolate faults. A converter substitution can therefore change the validity of downstream protection evidence even when nominal voltage is unchanged. 

Our view is that source migration and distribution migration should be managed as separate, explicitly linked changes. A utility connection does not have to become the occasion for a complete rack-power redesign. A new rack generation does not automatically justify reopening the source strategy. Combining the changes may be sensible, but the choice should be supported by an asset-retention and qualification plan.

Measure whether the boundary is working

A boundary that cannot be measured will invariably drift. The claims in this post are testable, and a portfolio should track them so the power specification stays honest as sources and equipment change.

We look at six measures:

  1. Source-side engineering hours per MW. The engineering labor spent on the power source and its handoff to the factory, measured per MW of capacity. The goal is to reduce this across the portfolio as archetypes get reused.
  2. Downstream deviations caused by the source. Count the times a source difference forced a change below the handoff, into distribution, controls, or operating procedure. Ideally - times = zero.
  3. Study and test reuse coverage. The share of protection studies, power-quality analysis, and acceptance tests carried forward from an earlier site without rework. Low coverage means the power specification is failing on its assumptions.
  4. Power-block lead time. The interval from a confirmed source to accepted power at the factory interface. When the interface is standard, that equipment can be ordered and built from a known design instead of engineered from scratch, which shortens the wait.
  5. Commissioning defects at the interface. Faults found at the source-to-factory boundary during commissioning. A rising count points to gaps in the power specification that that manifest themselves in the field.
  6. Time to qualify a new input archetype. How long it takes to add a source type to the qualified library and prove it against the power specification. This sets the real limit on how quickly the portfolio can take on a new energy asset.

Together these show whether the boundary is doing its job or whether site variation is still reaching the factory. They also give the investment committee something to hold the architecture to after the first build.

Put a value on the decisions that remain open

The financial value of a stable boundary is that it changes when capital has to be committed and how much of that commitment survives a later choice.

Suppose a project is evaluating two credible supply arrangements. If both have demonstrated that they can meet the same factory requirements, selected downstream packages can proceed while the source decision remains open. The outstanding choice still carries risk, but its effect on committed work can be priced and contained.

If the two arrangements require materially different downstream designs, the claimed optionality is weaker. The business is carrying parallel designs, delaying procurement or accepting rework. That can be justified but it needs to appear in the investment case as the cost of preserving the choice.

The portfolio calculation adds another dimension. Qualifying a source arrangement that is slightly more expensive at the first site may reduce engineering and procurement uncertainty at several later sites. Conversely, forcing every site to carry equipment or headroom needed by only one configuration can destroy the value of commonality. The decision needs both a project case and an explicit account of what subsequent deployments are expected to inherit.

There is a time limit on that inheritance. A qualification programme earns its return on deployments that arrive while its assumptions remain applicable. The investment case should identify those deployments and allow for the cost of updating the evidence as equipment, controls and workloads change. A large hypothetical pipeline is a weak justification for paying today to qualify configurations the business may never use.

Since Radiant finances, designs, builds and operates AI factories the power selection belongs in the same investment case as the compute commitments and operating model. Any benefit claimed for an earlier start has to survive equipment procurement, construction, the planned source transition and the obligations the operator inherits.

For the next source proposal, I would ask for three things alongside the power price and delivery date:

  • The downstream commitments it preserves: designs, equipment orders and test evidence that remain valid, with the assumptions stated.
  • The full cost of admitting and later changing it: adaptation, qualification, operating restrictions and the intended transition path.
  • The work the next deployment inherits: reusable models, configurations and acceptance evidence, with site-specific work identified separately.

Those answers give the investment committee a way to compare power opportunities and give engineering a clear scope to deliver. They also expose where the proposed architecture still depends on unresolved assumptions before those assumptions become equipment orders or customer commitments.

The next article takes that discipline into the powered shell: how much integration and verification can be completed before the equipment reaches the site.

‍

Related Articles