managed slurm

Slurm on Top, Kubernetes Underneath.

Slurm scheduling on Kubernetes-operated infrastructure. Researchers keep the queues, partitions, and batch workflows they already use. Radiant operates provisioning, recovery, and lifecycle underneath.

The scheduler your team knows, the cluster they don't have to run

Slurm Interface

Familiar queues, partitions, priorities, and batch workflows, unchanged

Kubernetes Operations

Automated provisioning, scaling, and node recovery across the GPU fleet

One Control Plane

Compute, networking, storage, and lifecycle managed together

Pre-wired clusters

A cluster stands up with scheduler, node images, accounting, shared storage, GPU drivers, and recovery loops pre-configured.

Built for the Longest Runs

Direct-to-GPU scheduling

Jobs run on bare-metal capacity, no virtualization layer

Self-healing

Failed nodes are isolated, replaced, and reintegrated mid-run

Dynamic queues

Shift capacity across teams, projects, and training phases as priorities change

Unified GPU pool

Slurm and Kubernetes-native workloads draw from one capacity pool, no stranded compute between schedulers