
Nebius Rated Platinum in SemiAnalysis ClusterMAX™ 3.0
Nebius Rated Platinum in SemiAnalysis ClusterMAX™ 3.0
SemiAnalysis has rated Nebius Platinum in ClusterMAX™ 3.0, its highest tier for evaluating AI cloud providers.
The verdict in their own words? “Nebius is unquestionably an industry leader with strong offerings in every category.”
Platinum is proof of the kind of AI infrastructure we’ve been building: performant, resilient, observable, and engineered to hold up when real workloads hit real clusters.
SemiAnalysis didn’t stop at peak GPU performance. Over several months, its team ran training and inference workloads on live NVIDIA GB300 and NVIDIA HGX B300 clusters: burn-ins, storage under concurrent load, node and full-rack recovery.
Real racks, real conditions, the kind of load that breaks the type of infrastructure that only looks good on paper.
Good clusters are reliable.
When you think of a good cluster, you think of it being boring and reliable.
This is how SemiAnalysis described us: “A good cluster is generally boring. On Nebius, burn-ins run without errors; Slurm is topology-aware; the packages we need are on the cluster and generally recent; WAN is good; the orchestration layers are free of footguns.”
For infrastructure teams, that kind of boring is hard to build, and when you’re choosing where to run production AI, it means less time firefighting infrastructure and more time running the workloads you actually care about.
The report also singled out Soperator, our open-source Slurm-on-Kubernetes solution, as a technical point of distinction: it combines the scheduling model HPC teams expect from Slurm with Kubernetes underneath, using a /jail persistent volume backed by VirtioFS to abstract away the storage complexity that usually comes with running Slurm over Kubernetes.
Soperator can run on-premises or in other clouds, and its adoption already extends beyond Nebius. SemiAnalysis noted that other vendors supplied test clusters built on Soperator.
SemiAnalysis’s take: “Nebius’s approach makes this stress-free for the operator.”
Stress the storage. Find the problem. Fix it. Test again.
Testing surfaced real problems, and Nebius fixed them before the report closed.
Under high-concurrency storage workloads, SemiAnalysis found slow 4 KiB sequential allocating writes and metadata consistency issues during concurrent directory creation at 128+ ranks.
So we retuned the filesystem. Then they tested it again.
The errors were resolved. Per-client performance outperformed the median, and SemiAnalysis verified that the storage layer scaled to demanding workloads across two racks.
Break a node. See what happens.
SemiAnalysis also tested what happens when the hardware itself stops cooperating. After a synthetic error was injected into an NVIDIA GB300 NVL72 rack, Nebius detected the failure immediately and automatically returned the node to service. No manual intervention was required. That matters on NVIDIA GB300 NVL72, where node failures are considerably more complicated to handle than on traditional NVIDIA HGX systems.
SemiAnalysis notes that the other NVIDIA GB300 NVL72 racks it tested did not automatically reboot and return failed nodes.
“The fact that this process is self-driving counts in Nebius’ favor.”
Visibility is paramount.
Automation only matters if operators can see what’s happening underneath it. SemiAnalysis called the dashboards we provided “an excellent suite” for cluster health visibility.
Health checks, automated remediation, and observability are the same problem from three angles: keeping GPU infrastructure operational without making customers manage that complexity themselves.
The bar stays high
Platinum is significant because of what sits behind it: independent testing of the infrastructure, including the parts that are difficult to get right and easy to overlook.
It’s the standard we’ve been building for.
Learn more about the results here



