Skip to content
AI Primer
breaking

SemiAnalysis reports configuration failures in Vultr's ClusterMAX test

SemiAnalysis found vulnerable libraries, login-pod setup errors and storage configuration problems that left pods stuck in Vultr's ClusterMAX test. It also reported strong Lustre throughput alongside networking reliability problems.

3 min read
SemiAnalysis reports configuration failures in Vultr's ClusterMAX test
SemiAnalysis reports configuration failures in Vultr's ClusterMAX test

TL;DR

  • Vultr landed below Bronze, with setup errors recurring almost a year after its earlier evaluation, according to the ClusterMAX announcement.
  • Vulnerable components and broken configuration arrived in the test environment; the test notes describe an incomplete login pod and storage that left a pod stuck in ContainerCreating.
  • Lustre was a bright spot: the storage test measured 91.7 GB/s aggregate reads across 32 clients.

The failing default StorageClass pointed to a product Vultr explicitly documents as unsupported on bare-metal servers. Grafana also showed zero NVLink traffic because its DCGM profiling fields were unconfigured, according to the test notes.

Below Bronze

An ugly handover: Vultr received SemiAnalysis's below-Bronze participation rating, with basic configuration failures repeating almost a year after ClusterMAX 2.0.

SemiAnalysis also disputed Vultr's “world’s largest privately held” cloud infrastructure and hyperscaler claims, while noting its modern GB300 and MI355X offerings.

Vulnerable test images

Known vulnerabilities applied to five components in the images Vultr built for testing, according to SemiAnalysis's account:

  • CUDA
  • DCGM
  • Docker
  • runc
  • ConnectX firmware

Slinky login pod

The Slinky login pod arrived with two missing pieces:

  • No /shared mount.
  • No Pyxis plugstack configuration.

Researchers orchestrated from a worker pod until Vultr support reprovisioned the login pod mid-campaign, fixing both problems.

Kubernetes storage

Storage requests through vultr-block-storage, the default StorageClass in the tested Kubernetes layer, provisioned and bound without complaint. The consuming pod then remained in ContainerCreating, according to the storage findings.

Vultr's block-storage FAQ says those volumes cannot attach to bare-metal servers. They support virtual instances such as Cloud Compute, Optimized Cloud Compute and Cloud GPU.

Lustre throughput

The worker nodes already had a Lustre tier attached. It delivered 91.7 GB/s aggregate read throughput across 32 clients, which SemiAnalysis's report described as excellent.

Grafana was configured, but NVLink read zero because the DCGM_FI_PROF_NVLINK_* fields were unconfigured, according to the monitoring findings.

NVIDIA's DCGM profiling reference identifies the relevant bandwidth counters:

  • PROF_NVLINK_TX_BYTES (1011): transmitted bytes per second.
  • PROF_NVLINK_RX_BYTES (1012): received bytes per second.

Both report averages over a time interval and exclude protocol headers. The reported zero came from a telemetry configuration failure.

Scale-out networking

SemiAnalysis said Vultr had scale-out networking reliability problems that discouraged partners from starting deals or expanding existing ones.

Public cluster audit

ClusterMAX 3.0's published methodology includes a public configuration audit that runs before load testing:

  • Distribution: the clustermax package, with uv pip install clustermax as the listed install command.
  • Runtime: approximately 15 minutes.
  • Per-check results: pass, warn, fail or skipped.

The audit covers hardware inventory, software and firmware versions, GPU access, containers, scheduling, networking, storage, health monitoring and security. SemiAnalysis keeps the remainder of its testing suite internal.

Share on X