Agentic AI May Finally Be Here. It Is Changing How We Build AI Factories

AI Factories for Agentic AI
5
Min Read
October 7, 2026
Share Article

Agentic AI has not gone to plan. Expectations were sky-high, and reality has been ordinary. Most enterprises still cannot run agents that are reliable, governed and worth the money. The few that made headlines did so for the wrong reasons. Plenty of people are tokenmaxxing their LLMs. Very few are doing anything truly agentic at work.

The first wave stalled on the basics. Enterprise work needs permissions, context, clean data, security, exception handling and someone accountable. A model can write a plausible plan. Running a process correctly inside a business is a different problem. Every integration turns the agent into a governance question, and every extra step makes evaluation harder. The same task can take five model calls or five hundred, so cost swings wildly. Latency hurts in any loop that matters. The technology got ahead of the operating model.

That is changing. Like everything else in AI, it is changing faster than expected, and the signals are stacking up.

Pricing. Salesforce does not price Agentforce like a SaaS module. Microsoft does not price Copilot Studio by the seat. Both are moving to credits, actions and consumption. An agent works around the clock. A seat licence charges as if it were one person working nine to five.

Connectors. MCP, Agent2Agent, Responses APIs, tool registries, browser agents and sandboxed execution are all maturing, and they all point the same way. Models get more useful with more data access. Agents now get that access through standard interfaces.

Consumers. Meta launched Muse in September. It is an agent that shops, books travel and handles email for you. It climbed to the top of the US App Store, and each Muse agent runs in its own virtual machine in Meta's cloud. Now multiply that by Meta's user base. Instinct runs entirely over text and is reportedly raising at $10B.

NVIDIA. This is the loudest signal. NVIDIA no longer talks only about training and inference. It now talks publicly about agentic inference, context memory, KV-cache routing, disaggregated serving and rack-scale systems as one workload. It talks privately to the big players too. NVIDIA has the best view of the ecosystem. It has every reason to move everyone to where the puck is going. Think of the old EF Hutton advert. When NVIDIA talks, the industry builds.

Software. Claude Code and Codex are not autonomous yet. They need direction, they get stuck and their work needs review. That is fine. Software is the first market where agents can work for real. The workflow already has the scaffolding. Repositories hold the context, terminals run the code, tests check it, pull requests get it reviewed and version control rolls it back. Software is showing the market what agentic demand needs: tool access, scoped tasks, places to run code, permission gates, observability and a way to check the work was right. The first agentic market is not "AI does the whole job". It is AI doing more of the loop inside a workflow that catches its mistakes.

Inference is a response. Agency is a loop.

Inference demand is what a model creates when it answers a prompt. Agentic demand is what a system creates when it tries to finish a task. The first phase of AI infrastructure was training. The second was inference. The third is still inference on paper, but the economics are agentic. That shift will drive the next redesign of AI infrastructure.

One prompt becomes a workload

The biggest forecasting mistake is counting visible prompts. An agent turns one prompt into a task graph. That graph holds planning tokens, retrieval queries, tool calls, interim summaries, validation steps, revised plans and hidden reasoning. The user sees one request. The infrastructure sees a chain of model calls.

Token demand can rise as token prices fall. Cheaper tokens make longer workflows worth running, so more tasks qualify. The metric that matters moves from cost per million tokens to cost per completed task.

Spend per task still has a ceiling. No CIO, CTO or CFO will sign off an open-ended agent bill. So agents will route work across model tiers. Large reasoning models plan. Cheap models extract and classify. Specialised models handle code and retrieval. Plain deterministic software does the validation. The winning design is orchestration within budget, latency and policy limits. Token volume will grow and get harder to forecast at the same time. "Tokens in, tokens out" will not cut it.

Agents need factories, not endpoints

Basic inference is a serving problem. You put a model behind an API and tune it for throughput and cost per token. Agentic demand is a systems problem. The endpoint model breaks once the output stops being a token and becomes finished work. That changes how the factory gets built, turned up and run.

Stop building one kind of rack. Prefill and decode make different demands on compute and memory. Much of retrieval and validation runs on CPUs. A uniform GPU cluster is only part of the machine. The factory becomes pools of different node types, with prefill and decode split where the workload justifies it. A scheduler places each step on the right tier. Mixed hardware is the design.

Treat KV-cache as a managed memory tier. Many agent steps reuse the same model and prompt prefix. Their KV blocks can sit in HBM, CPU memory, or local and shared NVMe. That raises the value of DRAM and fast NVMe across the cluster. Memory capacity and bandwidth become a first-order spec next to GPU count.

Design for east-west traffic. Disaggregation moves KV-cache between prefill and decode nodes. Traffic to retrieval and storage services adds to the load. In NVL72-class systems, NVLink turns the rack into one unit of compute. The scale-out fabric, InfiniBand or Spectrum-X, now needs sizing for cache movement on top of training collectives.

Plan for confidential compute. Agents run generated code and touch private and sovereign data. Sandboxes contain the code. Some tenants need protection from the infrastructure operator itself. For them, add GPU and CPU TEEs, encrypted memory and attestation. Build that overhead into the performance and capacity budget from day one.

Make telemetry the billing system. Once you sell completed tasks instead of tokens, you have to meter every step for every tenant. That means compute, memory and attributed power. You cannot bill or cap what you cannot see at step level. Fine-grained metering belongs in the operating layer from the start.

Size power and cooling for duty cycle. Agents keep inference capacity busy for longer, with fewer quiet periods. NVL72-class racks already draw well over 100 kW and need liquid cooling as standard. Size power and cooling for the rack's peak draw and the workload's real duty cycle.

None of this comes from one model getting smarter or one GPU running hotter. Work arrives, gets broken down and moves through stages. It uses shared resources, gets checked and gets finished. It has queues, bottlenecks and throughput limits. That is a factory.

The winners will not be the ones with the best models. They will turn models into finished work reliably, on budget, within policy and at predictable latency. Most of today's architectures were not built for that.

That is how we are designing our next builds. We plan memory, fabric, metering and confidential compute from day one, next to the GPUs. We size for the work agents create, and benchmarks rarely measure that. Bigger clusters will not win this. Coherent systems will.

Related Articles