With Blackwell NVL72, NVIDIA made the rack its primary scale-up compute unit, extending tightly coupled GPU communication beyond servers. Vera Rubin advances that architecture with 72 GPUs across 18 compute trays, each pairing four Rubin GPUs with two Vera CPUs. NVLink 6 delivers 3.6 TB/s of bidirectional bandwidth per GPU. For organizations investing in Vera Rubin, the question is not just performance, but how consistently the rack serves customers, teams and workloads with different capacity requirements.
A rack-scale system is the right architectural building block but may not be the right allocation for every team. Reserving an entire rack for a team that needs only part of it can leave expensive capacity unavailable to others. Sharing that capacity without the right isolation and placement controls creates a different problem: infrastructure that is easier to allocate than it is to operate securely and efficiently.
At Radiant, we treat allocation, isolation and placement as a single design challenge. Our FlightDeck architecture uses NVLink partitioning to place tenant boundaries inside the rack, helping you serve a broader mix of demand from the same investment.
Compute trays are the modular building blocks from which we assemble tenant capacity. Each tenant receives complete CPU-and-GPU units, with GPUs connected through the same rack-scale fabric. The rack remains intact; allocation boundaries follow tenant tray assignments and GPU membership in each NVLink partition.
NVLink partitioning makes the fabric tenant-aware
An NVLink domain describes GPUs connected by the physical scale-up fabric; a partition defines an isolated group within it. The NVSwitch control plane enforces the boundary, preventing multi-node NVLink communication between partitions. Each GPU belongs to one partition at a time, and partitions cannot extend beyond their physical domain. These hardware-enforced properties underpin the multi-tenant design.
Assigning nodes to different customers does not itself establish GPU-fabric isolation. NVLink partitioning supplies that boundary while preserving communication among GPUs within each partition.
Placement addresses where workloads run. Tightly communicating GPU groups need resources in the same suitable partition, not merely the same rack. Partitioning governs which GPUs may communicate over NVLink; topology-aware placement ensures workloads can use that connectivity.
In the example below, teams A and B each receive 32 GPUs, while team C receives 8 GPUs. Each team retains NVLink connectivity within its partition, with NVSwitch preventing NVLink traffic between teams. The rack’s power, cooling and switch infrastructure remain shared.

Real-world applications of NVLink partitioning
FlightDeck groups compute, network and storage into topology blocks that are reserved and pinned to specific accounts within NVIDIA’s tenancy. Each block is allocated as one unit, with consistent performance characteristics and enforced isolation.
1. Protecting capacity for production inference
A team running inference services for other teams reserves a topology block for its model endpoints. FlightDeck keeps the associated compute, network and storage committed to that account, preventing other teams’ workloads from consuming the reservation. The service retains its allocated capacity even when demand elsewhere in the cluster increases.
2. Reserving complete environments for distributed training
A team preparing a fine-tuning or reinforcement-learning run reserves its topology block as a single unit. Connected GPUs, network resources and storage are allocated together, avoiding a partial reservation that secures compute without the supporting resources. The team receives a complete allocation with uniform performance characteristics and security boundaries.
3. Isolating team-specific model development
Different teams develop their models within separate, account-bound resource blocks. Dedicated compute allocations and isolation across networking and storage keep each team within its authorized environment. This lets the administrator support concurrent projects on the same infrastructure while maintaining clear ownership and separation of reserved resources.
Where multi-tenancy creates the returns
- Less stranded capacity, more productive use. Sharing a single rack allows cloud providers to serve customers needing fewer than 72 GPUs, while enabling enterprise IT to support more internal teams on the same footprint. Spreading fixed infrastructure costs across more active workloads improves ROI, provided demand remains steady and operational costs are factored in.
- Better placement, lower cost per useful result. Keeping communication-intensive workloads within an NVLink partition can reduce time spent waiting for data exchanges, helping allocated GPU time produce more useful training and inference output. Partitioning provides isolation, not faster links; topology-aware placement makes the connectivity useful.
- Isolation that makes shared infrastructure usable. Hardware-enforced NVLink boundaries let tenants share a rack without cross-tenant NVLink access. Network, identity, storage and management controls complete the security design.
InfiniBand partitioning extends isolation beyond the rack
NVLink partitioning isolates GPU communication within a rack. InfiniBand partitioning complements it on the scale-out network, using partition keys (PKEYs) to define logical network partitions and restrict communication to authorized members.
FlightDeck aligns both layers with tenant allocations, so a team can connect its resources across multiple racks while keeping its compute traffic separate from other teams sharing the same physical network.
Plan your Vera Rubin deployment
Vera Rubin gives you considerable compute capacity in a single rack. The return depends on how much of that capacity your customers can use, how efficiently your workloads communicate and whether the isolation model allows different tenants to share it.
These are deployment decisions that Radiant addresses through FlightDeck. Planning team allocations around the rack’s actual topology gives businesses more ways to use their investment, while retaining the larger configurations that demanding workloads need.