Hardware

NVIDIA launches Cluster Readiness Engine for GPU validation

NVIDIA has released the Cluster Readiness Engine, an open-source Kubernetes controller designed to identify faulty GPU nodes before they disrupt massive AI training workloads.

NVIDIA Developer Blog2 days agoHardware
Image: NVIDIA Developer Blog

NVIDIA has introduced the NVIDIA Cluster Readiness Engine, an open-source Kubernetes controller designed to test and verify GPU clusters before production workloads begin. Licensed under Apache 2.0, the tool runs actual distributed workloads across topology-aware node groups to pinpoint hardware and configuration issues. This active testing approach helps operators catch silent performance degradations, such as a single slow GPU or a failing network link, which might otherwise go unnoticed by standard passive diagnostics until a massive training job fails.

The tool operates via a layered API consisting of Certification, Workflow, and Job custom resources. It features adaptive fault isolation, which automatically bisects failing node groups and reruns tests until it isolates the specific faulty nodes. Its built-in test catalog covers five NVIDIA Collective Communications Library variants, including all-reduce, all-gather, all-to-all, loopback, and loopback across NVIDIA NVSwitch. It also includes the NVIDIA Data Center GPU Manager level-4 diagnostic suite and NVIDIA NeMo pretraining workloads using NVIDIA Nemotron 5 models at 8B and 56B parameters.

For practitioners, the engine introduces a WorkloadRun API that automates repetitive multi-node GPU setups on Kubernetes. It supports torch, mpi, and exec frameworks, automatically generating Kubeflow TrainingRuntime configurations and injecting shared-memory volumes. To prevent deadlocks on busy clusters, the tool supports gang-aware schedulers like the KAI Scheduler. The engine integrates with the NVIDIA AI Cluster Runtime for configuration validation and NVSentinel for continuous telemetry monitoring, together forming the operating layer of the NVIDIA DSX AI Factory Platform.

To deploy the engine, operators need Kubernetes 1.29 or later, kubectl, Helm 3.x, and the NVIDIA GPU Operator. Systems using NVIDIA GB200 NVL72 or GB300 NVL72 architectures also require the NVIDIA DRA Driver to manage ComputeDomain resources. Setting up the controller requires running the nvcrectl setup init command, which installs the necessary custom resource definitions and the Kubeflow Trainer.

This is our own summary of reporting by NVIDIA Developer Blog

More in Hardware