跳到正文
SemiAnalysis· Jordan Nanos·· 16 天前精选AI 评分76

SemiAnalysis 发布 ClusterMAX 3.0 GPU 云评级:Nebius 升入铂金级

ClusterMAX 3.0: The Industry Standard GPU Cloud Rating System Returns

AI 导读

SemiAnalysis 发布 ClusterMAX 3.0 GPU 云评级,覆盖 77 家受测供应商与 323 家市场观察对象,较 2.0 的 209 家继续扩大,仅 19 家获得奖牌级评级。Nebius 加入 CoreWeave 升入铂金级,Google Cloud 进入金级,Azure 降至银级,Fluidstack 降为不可用,Crusoe 降至铜级。

AI 生成摘要 · 以原文为准

关注理由

SemiAnalysis 第三版 GPU 云评测覆盖 323 家供应商并给出分层结果,同时梳理融资、Blackwell 部署与可靠性实践,为评估 neocloud 服务能力与采购谈判提供参考框架。

正文 · AI 翻译

译文尚不完整,完整内容请切换到原文。

本文包含付费订阅者专属的附加内容。升级以获取完整访问权限。

在 ClusterMAX 上一次大版本发布后的 8 个月里,垂涎欲滴的投资者们几乎已经把能塞支票的口袋都塞满了。GPU 供应已经归零。与此同时,我们一直在努力让集群经受严苛考验。

几周前,我们用一些来自我们探测 neocloud 安全实践经历的 R 级轶事对这份报告进行了预告,还引出了一份你可能听说过其作者的 neocloud 客户发布的 PSA(公益广告)。

今天,我们终于完整呈现测试的广度和深度,ClusterMAX 3.0 比以往任何时候都更加全面。这包括计算、网络、存储、编排、UI、监控、支持,以及你能想到的几乎所有 GPU 云上需要检查的内容。我们会解释谁在做交易、谁的工程师在辛勤工作,以及谁的集群你最好在谈判时要求附赠一瓶阿司匹林。

废话不多说,这就是 ClusterMAX 3.0 的领奖台:

来源:SemiAnalysis ClusterMAX 3.0,2026 年 9 月
来源:SemiAnalysis Neocloud Dashboard,供我们的 AI Cloud TCO Model 订阅者使用

结果(执行摘要)

播客摘要讨论现已在 YouTube 上推出!

  • ClusterMAX 3.0 首次发布即对 neocloud 行业进行全面评估,涵盖 77 家提供商。

  • 我们将市场覆盖范围扩大到 323 家提供商,高于 ClusterMAX 2.0 的 209 家、ClusterMAX 1.0 的 169 家,以及最初 AI Neocloud Playbook and Anatomy 文章中的 124 家。

  • 作为这项研究的一部分,我们目前已采访了 200 多位 neocloud 的终端用户。

  • 我们更新了 涵盖 10 个类目的逐项标准列表,并更新了我们对 Slurm、Kubernetes、独立机器、监控仪表盘和健康检查的期望的直接描述。所有这些内容都已在我们网站上发布。我们鼓励各提供商在开发其产品时使用这些列表。我们仍将这些列表视为我们采访终端用户经验的综合汇总,使其能够代表终端用户对其云提供商所期望的功能。

  • Nebius 与 CoreWeave 一同进入铂金级。虽然 CoreWeave 仍然为其他人树立着技术标杆,但 Nebius 如今已被确立为一家持续保持高于其他家溢价定价的提供商。Nebius 强有力的商业决策使他们能够以卖方价格服务整整一类新实验室。

  • Google Cloud 与 Oracle 一同进入黄金级。Azure 移至白银级,Fluidstack 移至不可用级,Crusoe 降至青铜级。Lambda、Firmus 和 TensorWave 保持在白银级,而 GMI 从青铜级升至白银级。

  • 许多公司从白银级(或黄金级)降至青铜级或更低。本轮我们提高了门槛,全球仅有 19 家 neocloud 获得了奖章级评级。

  • 我们在青铜级和表现不佳级之间设立了一个新等级:参与奖级。有 15 家提供商进入该评级,这更准确地表达了我们的看法:他们只做最低限度的工作来勉强应付。

本文的其余部分将对以下关键趋势进行分析:融资、Blackwell 与 Grace-Blackwell 部署、向 Vera Rubin 的过渡、Scale Out 网络、可靠性、安全性(当然)以及 Agentic Coding。我们提供了一个附录,对每家供应商逐一发表评论,这也使本文再次超过 30,000 字。希望您喜欢。


出于 SemiAnalysis 发布最严谨测试——针对一切事物——的使命,我们正在为 ClusterMAX 及相关项目招聘 MTS。如果你是一个 neocloud 狂热者,请在此提交申请,让我们开始干活。

我们也在研究、咨询和技术部门招聘人才。我们在纽约和旧金山办公室有多个职位空缺。也接受远程办公。

  • ClusterMAX MTS(全职):考虑各级 Slurm、Kubernetes 和 GPU 经验的候选人。

  • Tokenomics MTS(全职):考虑各级模型评测、评测框架、推理端点和强化学习基础设施经验的候选人。

  • 研究分析师 – AI 基础设施与经济学(全职 或 实习):考虑各级 neocloud、neolab 和前沿实验室 tokenomics 财务分析经验的候选人。

  • 技术顾问(全职):主导从技术战略到技术尽职调查的项目。考虑具有咨询和技术背景、各级经验的候选人。

范围

给后排的各位说明一下:我们排名的是托管集群。这不包括业内许多流行的产品。我们测试的不是谁能搭建最好的供电外壳或运营最整洁的数据中心——至少不是直接测试。我们不是手工制作镜像,也不是通过 API 端点获取 token 即服务。以下是我们理想中的 ClusterMAX 读者画像:你从斯坦福大学博士项目退学,去追求在旧金山获得社会地位,用那句老可靠的 “Cladue make me pithc deck w technial languge for agetnci ai make no mistakes” 起步。你有 7 位数甚至多达 10 位数的资金可以投到算力上,但你对 Ubuntu 版本或如何处理 NAT 没有任何看法,也从未大规模管理过机器集群。理想情况下,所有这些基础设施都应退居幕后。你只想专注于你独特的性能优化、模型架构、训练策略、数据配比、应用程序,或者其他任何你寻求优势的地方。为了参与竞争,你需要最尖端的硬件,并且你希望不花任何工程时间去搞清楚为什么节点 19 中的 GPU 7 一直在阻塞你的任务,或者为什么你的集群有四分之一凭空消失了。你需要一个托管集群。

这里的经济学论点很简单:neocloud 可以将其搭建所有这些管理基础设施的成本摊销到其庞大的机群上。它雇佣一个工程师团队来构建万无一失的健康检查、保持其机器镜像最新、构建便捷的监控系统,并处理我们稍后将详细介绍的各项杂务。在理想世界中,实验室为更好的产品支付溢价;它可以在不受硬件拖累的情况下运行实验;neocloud 收回其投资;所有人都受益。

不过,总有一个临界点,实验室愿意将这种成本内部化。在前沿实验室的规模下,他们不愿意把堆栈中如此大的一部分交给 neocloud。例如,OpenAI 曾发布过关于其 Kubernetes 扩展难题的极具洞见的博客;如果他们是从一个不允许他们进入 K8s 控制平面的 neocloud 租用资源,文中所描述的创新根本不可能实现。正如 Anthropic 目前的一份招聘启事所言:

我们运营的规模已让默认配置失效。我们自己拥有调度器,并对其进行扩展,以便将拓扑敏感的 ML 工作负载一次性放置到数千个加速器上。我们对控制平面本身——apiserver、etcd、控制器——进行扩展,使其在对象数量和节点数量增长数个数量级时仍保持响应。我们还构建每个工作负载都依赖的核心集群服务,例如服务发现,使其能在同样的压力下正常运转。

可以说,像 OpenAI 和 Anthropic 这样速度与规模的实验室需要同时协同设计其整个堆栈,一个现成的 Slinky 实现——即便是优秀的实现——也不够用。

前沿实验室在管理其基础设施上显而易见的用心,正说明了 ClusterMAX 为何重要。据我们最近统计,OpenAI 和 Anthropic 职位描述中包含“Kubernetes”的开放职位的薪资范围总和为 $27,769,274-$42,687,654,更不用说他们已经配备的庞大团队所支付的薪水了。如果你并不预期会出现万亿美元级的流动性事件,而你需要足够好的基础设施来给自己一个一搏的机会,你就需要一个 neocloud 为你承担重任。

至少在纸面上,较小的 neocloud 客户是大卫对抗歌利亚。OpenAI 和 Anthropic 在各自logistic指数曲线上拥有惊人的增长率,而我们的 数据中心行业模型目前预测,到 YE2027 它们将占所有实验室算力的 56.3%。

OpenAI 算力容量。来源:SemiAnalysis Tokenomics 模型。

正如我们稍后将更详细描述的那样,由于融资的原因,这里还存在马太效应:盈利的前沿实验室比较小的实验室更容易锁定容量,而后者则被巨额预付款、更差的价格以及提前数年规划的更大难度所困。

这是否意味着托管集群已死?cLUsTerMaX iS wAshED?实际上,尽管 Anthropic 和 OpenAI 占据的份额逐年增大,但托管集群作为一项业务仍在呈指数级增长。我们已帮助无数实验室找到算力,还有更多实验室正在寻找,愿意付费,并等待更多容量上线。Anthropic 和 OpenAI 是最大的客户,但还存在一个由数亿美元级交易组成的长尾,其总价值是巨大的。

托管集群在与前沿实验室裸机的竞争中占据优势的一个可能情景涉及开源的进展。服务开源模型已经比许多人意识到的更有利可图;我们最近测算过“每兆瓦每年出售开源代币可以赚取超过 1 亿美元”,即使按当前的租赁价格,也仍有充足的利润空间。可以设想开源模型的市场份额不断增长,无论边际采用者是对成本更敏感,还是算法进步缩小了与闭源模型的差距。无论具体动态如何,这都将提升托管集群的相对地位。几乎任何供应端的分裂都会让实验室略微更不愿意自己处理基础设施,而略微更愿意依赖新云厂商。

这些集群之上的各层正在快速成熟,但它们超出了 ClusterMAX 的范围。有几个细分领域值得考虑。有推理端点,可以是公开的或私有的,按代币计费。有后训练服务,为特定用例定制开源模型。还有沙箱服务,为智能体提供基础设施,尤其是在 RL rollout 期间,并简化和优化容器与 CPU 管理。这些都可以宽泛地归入“GPU 服务”这一类别,它们之间的边界可能比较模糊。例如,一家云厂商在售出托管集群后,可能用闲置容量启动一个推理端点,或者它可能提供沙箱基础设施,同时在销售 RLaaS 时内部使用这些基础设施。我们将在不久的将来就此发布更多内容,但这超出了本报告的范围。

测试方法

在 ClusterMAX 3.0 中,我们向每家供应商提出了以下要求:

  • 32 块 GPU(4 个 8 路 HGX 节点,或 8 个 4 路 NVL72 节点)

  • 高带宽网络(我们要求 800G RoCE 或 XDR InfiniBand,但许多供应商当时还没有)

  • 10TB+ 高性能文件存储(支持 NFS/POSIX 挂载以及通过 csi 驱动提供的 RWX StorageClass)

  • 10TB+ 兼容 S3 的对象存储

  • 一个监控仪表板(通常基于 Grafana)

  • 5 天时间测试 Slurm,5 天时间测试 K8s(可以并行进行,即在 SonK 集群上,或按顺序进行,如果供应商想收回节点并重新配置的话)

  • 可接受 NVIDIA 的 Blackwell 而非 Hopper(B200、B300、GB200、GB300),或 AMD 的 MI355X。H100 已经推出超过 4 年了!

我们通过如下所述的 3 个阶段对集群进行测试。当然,我们始终通过与这些供应商的所有客户联系并参考他们的反馈来补充测试。

阶段 1:审计

首先,审计涵盖是非题。集群设置是否正确?软件是否已安装?是否为最新版本?各种工具是否按预期工作?运行大约需要 15 分钟。可在 GitHub 上获取,或通过 {uv} pip install clustermax 安装。

审计会在系统承载负载之前检查配置。它涵盖硬件清单、软件与固件版本、GPU 访问、容器、调度器配置、网络、存储、健康监控和安全。我们会检查哪些组件适用于所测试的环境,并显示每项检查的通过、警告、失败和跳过状态。我们已免费发布这一部分,并将长期维护。其余部分目前仅内部保留。

阶段 2:性能

性能测试对集群各单个组件以及整个集群的性能特性设定通过/失败阈值。这些通过微基准测试和真实场景基准测试来完成。

具体来说,我们测试:

  • GPU 计算

  • 网络

  • 存储

  • 生命周期

  • 训练

  • 推理

下面我们详细解释这些内容,但请注意,我们测试的内容会随时间演变。

GPU 计算

我们首先检查 GPU 上的真实场景 GEMM 性能。GEMM 是现代 AI 工作负载中最关键的操作,这一点我们之前已多次解释过,而分组 GEMM(Grouped GEMM)对现代 MoE 模型尤为重要。

我们在 cuBLASLt(或 hipBLASLt)和 DeepGEMM 上测试多种精度类型的分组 GEMM:BF16、FP16、TF32、FP32、FP8 E4M3、MXFP8 和 NVFP4,具体取决于集群中芯片所支持的内容。我们使用 Kimi K2.5、K3 以及 DeepSeek V3、V4 Pro 模型中不同 gate_up 和 down 投影在不同批大小下的一组通用形状。

然后我们测试 GEMM、GEMV 带宽和 MAMF(最大可达矩阵乘 FLOPS,来自 Stas Bekman),后者使用 cuBLASLt(或 rocBLAS)在每块 GPU 上按精度运行一段时间的扫描。我们会报告所有 GEMM 测试的 FLOPs。

来源:SemiAnalysis ClusterMAX 结果仪表板
来源:SemiAnalysis ClusterMAX 结果仪表板

我们还通过 GEMV 测试和 `nvbandwidth` 来测试内存。GEMV 以细长向量测量矩阵向量乘法,报告 GB/s,作为低批解码流式行为的代理指标。NVIDIA 的 `nvbandwidth` 工具内置于 DCGM 健康检查中,测量 h2d(CPU 到 GPU)、d2d(GPU 到 GPU)和 d2h(GPU 到 CPU)带宽。最后,我们还有一个自定义 PyTorch 脚本,通过实际的张量拷贝测量相同路径,记录锁定内存 h2d、d2h 以及本地 HBM 拷贝的带宽。

最后,我们为老派用户测试 HPL-MxP。HPL-MxP 使用低精度分解求解大型稠密线性系统,这意味着我们同时压测计算、内存和通信。该测试在单一工作负载中覆盖计算、HBM、经 scale-out 网络的 MPI 广播、NVLink 以及数值正确性。我们根据芯片类型报告多种数字格式的 FLOPs。

尽管 GEMM 是现代 AI 工作负载的核心,但这些微基准测试极少能区分各家提供商——提供商很难把这项功能搞砸。重要的是在长时间满负荷运行(burn-in)期间检查原始 GEMM 性能,如后文所述,如果芯片冷却不当,其时钟频率可能会下降以避免过热。在 GEMM 突发负载下,你不太可能发现什么值得注意的问题。我们在每个提供商上都运行这些测试,因为速度快,而且一旦失败集群会立即不可用,但我们更担心其他测试。

生命周期

为了了解集群在整个生命周期中的易用性,我们准备了一系列专门的测试。我们首先使用多种方法检查下载/上传速度的网络速率,主要来自 Cloudflare。接着我们测试通过 pip 和 uv 安装 PyTorch 和 vLLM 等常用软件包所需的时间,以及从 Docker Hub、ghcr.io 和 nvcr.io 下载容器,并从 Hugging Face 和 ModelScope 拉取模型。

然后我们在每个可用的存储层(home、共享文件系统、本地 NVMe scratch 以及任何额外的挂载点,分别称为 `/home`、`/data`、`/scratch` 和其他)上使用全新的软件包缓存重做 pip 和 uv install 测试。随后我们从每个位置运行 `import torch`,看需要多长时间。在大多数提供商上,这只需 1-2 秒,但在一些存在 LOSF 性能问题的文件系统上,它可能会莫名其妙地耗时 10-20 秒。

最后,我们使用之前下载的模型,运行 vllm serve 命令,测试从启动到能够访问 /v1/models 所需的时间。同样,某些提供商可以在 10-40 秒内将一个小模型从共享存储加载到 GPU 内存中,而一些提供商则需要超过 2 分钟。

Source: SemiAnalysis ClusterMAX Results Dashboard

网络

我们对节点间和节点内传输的通信性能进行了一整套广泛的基准测试。这包括 MPI 集合通信,即通过 MPI 启动 NCCL 或 RCCL 测试,针对 all-to-all、all-reduce、all-gather 及其他集合通信操作,测量从 8 B 到 16 GB 各种消息大小下的集合通信带宽和延迟。同样的工作负载也通过 PyTorch 中的 torch.distributed 运行。此外,我们使用 RDMA perftest,具体是 ib_write_bw 和 ib_read_bw,测量每条网络轨道上的点对点带宽,以及主机内存回退(通过 PCIe)的情况,以发现任何问题。

我们以 2 的倍数扩大 world size,直到用满整个集群,同时记录性能的扩展情况。我们还在禁用 NVLink 的情况下运行测试,尤其是在 GB200 或 GB300 NVL72 集群上,以便隔离评估 scale-out 网络的性能。在这里进行核算时务必小心,因为 world size、算法以及任务相对于 scale-out 边界的放置位置都会显著影响结果。

来源:快速检查我们集群的网络栈开箱即支持哪些功能
来源:SemiAnalysis ClusterMAX 结果仪表盘
来源:SemiAnalysis ClusterMAX 结果仪表盘

存储

我们运行 fio、ior、Elbencho、一个用 Torch 和 Torch DCP 编写的自定义 checkpoint 保存/加载基准测试,以及一个由 Skild AI 分享给我们的自定义 datagen 基准测试。Datagen 可能是最有意思的,因为它近似于他们在机器人数据流水线中的真实工作负载:生成视频(MP4)和 parquet 传感器数据集的聚合写入吞吐量。

不过,能发现最多问题的测试还是老掉牙的 fio。我们从 c1 一直到 c32(或者,如果我们的节点超过 4 个,就采用集群所能容忍的尽可能多的客户端),以 1MiB 顺序块和 4 KiB 缓冲 / 64 KiB 直接随机块,遍历缓冲/直接 I/O 的顺序/随机读写。如果存储有任何配置不当,在这个 fio 遍历中几乎总能找到某个设置表现出糟糕的性能。我们跟踪吞吐量、IOPS 和延迟,当然,如果测试在过多客户端下失败,我们还会发现元数据中的不一致之处。

来源:SemiAnalysis ClusterMAX 结果仪表盘

当然,这在很大程度上取决于分配给我们的卷有多大以及 SLO 是如何定义的。请对这些结果持保留态度——上图实际上只表明 Azure 给了我们一个 PB 级的挂载点。这种保留态度也说明了在大量集群上拥有数据的重要性。例如,可以想见,缓冲测试的延迟更高,而且如果测试运行时间不够长,结果也可能非常嘈杂。但延迟会高多少,结果又有多嘈杂?要弄清什么是“足够好”、什么需要向提供商反馈,与庞大的同类集群集合进行比较至关重要。

对于对象存储,我们针对一个 S3 存储桶再次运行这套测试。我们还运行一个自定义基准测试,测量顺序读/写时大对象的吞吐量、小对象的 PUT 和 GET 速率,以及读取延迟(p50、p90、p99)。在这项测试中,我们保持所有条件一致。不过一如既往,如果我们注意到某个集群在某个工作负载上表现吃力,我们会驻留下来,运行更多测试以更精确地隔离问题。

训练

进入真实世界后,我们开始观察网络或存储上的问题如何在真实测试中显现。我们通过 torchtitan 运行两个作业:使用 FSDP 在 C4 数据集上进行 llama 3.1 8B 预训练,以及通过 torchtitan 并采用专家并行进行 GPT-OSS MoE 训练。前者受 FLOPs 限制,后者受集合通信限制(除非你用的是 NVL72)。因此,当网络出现问题时,后者能让你一目了然。

来源:ClusterMAX 结果仪表盘
来源:ClusterMAX 结果仪表盘

推理

我们在单节点和多节点场景下、使用不同模型运行 InferenceX AgentX 基准测试,在开启性能分析追踪的情况下找出受计算限制、受内存限制和受通信限制的运行状态,并将 InferenceX 结果视为用于比较的基准测试性能结果。与我们的训练测试一样,这凸显了微基准测试中涵盖的各种集群组件的重要性。如果一个集群连 NCCL 测试都过不了,它肯定无法生成很多 token,而我们的 InferenceX 测试则展示了有效吞吐量(goodput)是如何流失的。

阶段 3:可靠性

基准测试中的出色表现并不意味着该提供商能驱动良好的 有效吞吐量,当然也不意味着他们会提供良好的支持体验。因此,我们通过多种方式测试可靠性、监控与自动修复。

老化测试

我们首先对 GPU 和网络进行 8 小时的老化测试,方法是在张量核心上运行大型 GEMM 的同时进行 all-to-all 通信,并用监控脚本跟踪温度、功耗、时钟频率、FLOPs、网络连通性、延迟、带宽,当然还包括跟踪内核环形缓冲区中的任何错误。你会惊讶于仅凭这个简单的测试就能引发多少硬件问题。我们必须强调,同时压测 GPU 和网络是极其重要的。老化测试需要这两种组件同时发生热胀冷缩,以逼近这些系统在真实工作负载的真实压力下的实际行为。很多提供商仍在使用分别压测 GPU 和网络的脚本。

网络互连

接下来,我们通过循环执行大消息尺寸的 all-to-all、all-gather 和 reduce-scatter,测试 scale-out 和 scale-up 网络互连的持续性能。我们报告平均、最小和最大带宽,以及集合通信错误。在前端网络上,我们还测试跨四个节点的每对有向节点之间的管理网络丢包率、RTT 和 TCP 带宽。这个额外的测试并不常发现老化测试遗漏的错误,但我们喜欢这些数据,而且它曾揭示出某些以太网网络在负载下的巨大性能差异。

存储

继续进行,我们通过在每个客户端数量下使用 fio(如前所述)运行 7.5 分钟的顺序混合读写,随后再运行 7.5 分钟的随机混合读写,来测试文件系统的耐久性。我们在整个分配范围内增加客户端数量,并报告随时间变化的带宽、IOPS、尾延迟和性能衰减。有趣的是,我们发现某些提供商在负载下与初始基准相比带宽衰减接近 40%。

当对象存储可用时,我们还会运行 15 分钟的混合 PUT、GET、LIST 和 DELETE 工作负载,测量吞吐量衰减、p99 GET 首字节时间漂移、限流和错误。在这些测试中我们观察到了大约 10% 的差异。

编排

在 Kubernetes 上,我们通过反复创建/删除 GPU pod 并检查 API 中的 GPU 分配来测试“抖动”。我们报告调度和启动延迟,并在 kubectl delete po 卡住时报告失败。我们还会增加仅 CPU 的沙箱 pod 批次,测量有多少达到 Ready 状态、耗时多久,以及其余的为何保持 Pending。

最后,我们通过反复启动一个挂载卷、写入并 fsync 一个文件、然后被删除的 pod 来测试 PVC 生命周期。我们在每轮之间等待清理完成,并使用一轮预热来从测量运行中排除初始镜像拉取的影响。

破坏性测试

最后,我们来到有趣的部分:故意搞坏东西。

首先,我们重启集群中的所有节点。你或许希望这不是破坏性测试,但它确实是。在 Kubernetes 上,我们会封锁(cordon)并排空(drain)节点,发出重启命令,并要求 boot ID 发生变化,然后等待节点重新回到集群。计时器在新 pod 中能在该节点上运行 nvidia-smi 时停止。在 Slurm 上,我们只需获取分配或通过 SSH 运行 sudo reboot。Slurm-on-Kubernetes 略有不同,因此我们会尝试从 Kubernetes 层面操作,以便正确测试。我们的脚本会全程计时。

在注入故障时,我们首先使用 DCGM 注入方法。如果它不能正常工作,我们会接着向内核日志写入一条合成的 NVIDIA XID 或 SXID 消息,并观察供应商现有的健康检查和调度器如何响应。我们不会动他们的健康代理和排空自动化,因此检测错误并采取行动要靠他们自己的工具。大多数健康检查都配置为从内核环形缓冲区读取,但也有例外,因此我们会与供应商沟通,确保我们的触发方式有效,并且我们有权限读取它。最后,作为一种可靠的方法,我们会重置 GPU 的上游 PCIe 桥,这会产生一个真实的 XID 79。

在所有这些情况下,我们都会记录检测故障所需的时间、节点停留在 `DRAIN` 状态的时间,以及节点回到集群后运行一条基本命令所需的时间。在 AMD 系统上,我们注入一个 UMC RAS 错误,并以同样的方式跟踪一切。

这是我们最关键的测试。那些我们听到大量客户抱怨可靠性的供应商,通常没有配备健康检查、监控仪表板和自动修复机制。可靠性是世界各地许多最大客户的首要标准,正如我们在文章“GPU 集群的真实成本到底有多高”中详细讨论的那样。

来源:ClusterMAX 结果仪表板

值得注意的是,我们的实际测试只是整体评分的一部分。有许多事情我们无法通过实际测试来了解,需要使用其他研究方法深入理解,比如大规模下的性能、长期可靠性、支持体验、定价、GPU 可用性和交付时间。

即将推出:智能体工作负载与 RL

我们还有两个测试套件,将在后续工作中详细介绍。

CPU 计算

CPU 是现代智能体工作负载的关键性能考量因素,我们之前已经讨论过这一点。我们开发了一套全面的基准测试,通过常用的 Linux 工具调用以及更多微基准测试,在即将推出的名为 SandboX 的项目中,将 CPU 性能精确地在这些智能体工作负载上推向极限。

SandboX 运行一系列具有代表性的 CPU 密集型工作负载,包括环形洗牌生产者-消费者队列、Redpanda(测量本地 Kafka 兼容 broker 及其吞吐量),以及动态负载基准测试。

我们还会对 Kubernetes 集群的沙箱生命周期进行基准测试,以了解将 RL 沙箱共置在 GPU 集群中剩余的 CPU/内存资源上的可行效果。我们会报告冷启动延迟、调度吞吐量、每节点密度、容量和基线表现。

强化学习

一次 RL 运行包含许多环节:生成器负责推出轨迹,环境在沙箱中执行动作,训练器接收 rollout、更新策略并将更新后的权重推送回推理引擎。扩展 RL 意味着在最大化训练器、推理引擎、环境沙箱执行、rollout 传输、权重同步等各方面的利用率之间取得平衡,任何一环成为瓶颈都可能使整个 RL 栈停滞。

我们还看到 Post-Training-as-a-Service(后训练即服务,即 RLaaS)的兴起:客户自带环境,或由驻场工程师帮忙构建,托管的训练服务商则搭建好对开源权重模型进行后训练的基础设施。卖点是定制化同时保护客户的 IP,这显然成了 Satya 的口头禅。

在我们即将推出的 PostTrainingX 基准测试中,我们会扫描训练器与推理之间的 GPU 分配、批次大小、rollout 并发度以及策略陈旧度上限,以研究它们对性能和训练稳定性的影响。

我们通过每个推理 GPU 每秒的输入/输出 token 数、rollout 延迟、KV-cache 占用率和前缀缓存命中率来跟踪生成器的性能;通过权重同步延迟、传输带宽和网络遥测来衡量通信效率;通过训练吞吐量、优化器步进时间、MFU、train/rollout KL、观察到的策略陈旧度、裁剪比例、梯度范数和被拒绝的 rollout 来衡量训练器的性能与稳定性。

对于环境,我们测量设置时间、工具执行时间、超时次数和基础设施故障。GPU 利用率、内存使用量和功耗提供系统层面的背景信息,而存档的轨迹、日志和确切的环境定义则使结果可复现。

PostTrainingX 的目标是对每一家提供商进行基准测试。这包括托管训练平台,如 Applied Compute、Azure Foundry、Fireworks、Thinky、Baseten、Trajectory、Engram、AI21 的平台等,以及开源框架,如 Miles、slime、prime-rl 和 verl。最终结果应能让客户获得数据,以决定是在开源框架上运行自己的后训练栈,还是把它交给托管的 RL 提供商。

我们已经在开源框架上开始运行,包括 Miles、prime-rl 和 verl,接下来将扩展到托管提供商。敬请期待……

行业趋势

融资

过去几个月里,neocloud 领域最热门的话题当属融资。人人都想“跟着钱走”,这正是我们在 AI Compute, Capital, and Markets Model 中所做的,该模型在我们一篇近期文章中拆解 NVIDIA 的 Backstop Universe 时发布。下面我们再讨论几个热门话题。

债务与获取债务

如今,neocloud 们已经意识到,一份信用良好的客户合同可以帮助它们以低于自身信用评级所能支持的利率获得债务融资。这意味着,没有长期承诺客户(或称“IG offtaker”)的提供商可能需要更昂贵的债务或类股权资本才能购买 GPU,从而启动或扩展其业务。与此同时,贷款方也需要证据证明提供商能够按时交付其集群并履行客户合同中的 SLA。举债不仅取决于可用于偿债的现金和 TCV,也同样取决于租约终止权。这促使多家保险提供商进入市场,推出参数化产品,为 neocloud 及其贷款方提供合同终止保护,通常在发生终止时提供一个月的过桥期以寻找新的 offtaker。我们非常看好这种结构,并认为它是贷款方为这些交易去风险的可靠方法。

交易对手风险与 Nvidia 的“资产负债表即服务”

在一个大型 neocloud 项目中,GPU 出租方和数据中心出租方可能在同一项目中面对不同的交易对手。例如,在 Anthropic 的 TPU neocloud 结构中,Broadcom 支持设备融资,而 Google 支持数据中心租金。由于供应链中的多笔贷款可能依赖于同一家 AI 实验室,即使每笔贷款的借款方是不同的 neocloud 或数据中心,也可能存在系统性、相关性风险。正如我们刚才讨论的,建设延期可能触发取消并损害现金流。

于是有了 Nvidia 的兜底支持。由于 Nvidia 通过收入下限、房东担保以及计划转让给第三方的租约来支撑 GPU 需求,它们能够帮助 neocloud 和 AI 实验室在无需超大规模云厂商参与的情况下获得投资级融资。Nvidia 在初始销售时赚取 GPU 利润,但现在还可以从超出其 “AICP floor” 的收入中获得分成。因此,它们的财务支持推动了其硬件的需求。资产负债表就是护城河。

折旧时间表

自我们去年十一月在其文章中抨击 Dr. Burry以来,这个话题已经稍稍冷却。也许是因为大家签署的都是那些 6 年期合同,又或者是那些恼人的 4 年高龄 H100 就是降价!谁知道呢。

无论如何,我们在建模折旧时将设备购置日期与投用日期分开。6 年折旧假设需要针对具体提供商的支撑依据,而在 NVIDIA 的 AICP 计划中,6 年收入下限实际上并不能确立 GPU 的使用寿命。我们的折旧敏感性测试对建模利润率的影响不小,但贷款方优先考虑的是 IRR,所以我们不会在此分享这些内容。当然,财务分析并不能确立买家未来愿意为 GPU 支付的实际价格。那取决于供需动态,以及终极问题:残值。由于贷款方无法将贷款还款计划与我们用于设备折旧的会计假设分开设定,我们被困在一个今天没有人愿意承担为 4、5 或 6 年旧 GB300 评估残值风险的世界里。

Blackwell 与 Grace Blackwell

现在让我们回到本文剩余部分的技术内容。好戏来了。

正如我们在 ClusterMAX 2.0 中详细讨论的那样,运营 8 路 H100 HGX 服务器与 GB300 NVL72 机架级系统之间的差异是巨大的。三四年前从安装 H100 起家的供应商,如今正面临以下问题:

  1. 基于 ARM 的 Grace CPU(而非 x86 的 Intel 或 AMD)

  2. 采用 sm100/sm103 的 Blackwell GPU(需 CUDA 13.0+)

  3. 强制性的直接液冷(DLC)

  4. 高压供电(每机架 130kW 以上,通过三相 408 或 480V 供电)

  5. 横向扩展网络,配备 800G CX-8 网卡和 51.2T 或 102.4T 交换机(800GbE RoCE Spectrum-X)或 115.2T 交换机(800Gb XDR InfiniBand)

  6. 纵向扩展网络,在机架级采用 NVLink 交换机和背板,72 路

  7. 用于前端(或融合前端/存储)网络的 BF-3 或 BF-4 DPU

  8. 配备 shuffle 板/shuffle 线缆的多平面网络

所有这些变化都对配置软件、监控系统、技术人员培训、设施级电气、制冷和机械系统、OEM 支持合同与合作关系等产生影响。

供应商在 Hopper 世代取得成功,并不意味着他们在 Blackwell 世代也会成功。

在我们的排名中,我们格外看重任何向我们交付了 Grace Blackwell 的供应商,坦率地说,在配置、监控以及可靠性/热备方面遇到困难时,我们给了他们大得多的宽容度。

这意味着,如果客户租用仅通过 RoCE 或 InfiniBand 连接的标准 HGX 节点,我们期望看到配备热备的自动修复:端到端一小时内完成替换,约 2 分钟内完成检测。然而,对于 NVL72 系统,在另一套纵向扩展网络上热插拔托盘是不合理的。

因此,在我们的测试中,只要健康检查能发现不健康节点(仍然要求在 2 分钟或更短时间内),并阻止工作负载被调度到该节点上,它就完成了自己的任务。如上所述,我们为许多实验室和云服务商提供 GB300 NVL72 SLA 方面的建议,而我们看到的最常见模式可以称为“NVL64+”。也就是说,每个节点都有 SLA,但一旦机架中 18 个节点中有 3 个或以上在某一时刻发生故障,整个机架就被视为宕机。这是由简单的 2 的幂次决定的:许多作业可被 64 整除,运行在 64 个 GPU 上,而填满 72 个 GPU 的最大配置是 world_size=8 * num_jobs=9。

迈向 Vera Rubin

有趣的是,从 Grace Blackwell 迁移到 Vera Rubin 所需的改动,远少于从 Hopper 迁移到 Blackwell。供应商仍然管理着基于 Arm 的 CPU、液冷机架级架构、72 路 NVLink 纵向扩展网络、多平面横向扩展网络以及一些 Bluefield DPU。想当年 GB200 刚推出时,这一切都是全新的!

主要变化在于横向扩展网络从 800G CX-8 网卡迁移到 1.6T CX-9 网卡,以及功耗从每机架约 130kW 增加到约 200kW,包括 800V 直流选项。

早期的需求比 Blackwell 转型期更为强劲——这部分是因为 将 GB 软件移植到 VR 比当年从 Hopper 移植到 GB 要容易——而且 我们从许多供应商那里听到,他们有望比预期更早交付 VR。系统启动调试出奇地顺利,物理设计看起来也成熟得多。

Scale Out(横向扩展)网络

横向扩展网络是供应商仍然掌握的最后一项重大设计决策。Nvidia 和 OEM/ODM 锁定了 NVL72 机柜的大部分内容。在机柜之外,供应商选择拓扑结构、布线方式,以及 Kubernetes 或 Slurm 如何将网络架构分配给作业。我们在每一层都发现了问题。

从物理层开始,800G CX-8 网卡是多平面的:每块网卡将其带宽分配到多个独立的交换架构中,这些架构称为平面(plane)。每个 GPU 的总带宽保持在 800G。更多的平面为每台交换机提供更多有效端口,也为每块网卡提供更多上行链路路径,而这需要一次乱序重排(shuffle)。每条网卡通道都必须落在正确的平面上,因此 shuffle 盒子将光纤映射放在一个配线模块中,而 shuffle 线缆则将其内置到线缆组件中。两种方式都可行,但在可靠性(需要维护设备的频率)和可维护性(修复或更换设备的难易程度)方面各有取舍。

Oracle 是这一物理设计的先驱,这在你拿到一台 GB300 节点时就体现出来了。我们收到的每个 4x GPU 托盘都有 16 个物理 200G 端口,即每个 GPU 4 个平面。Linux 暴露出 4 个可用的 800G RDMA VF(rdma_vf_rail0 到 rdma_vf_rail3),每个 GPU 一个,各自位于独立的 VRF 中。聚合的 PF 因没有可用 GID 而无法工作,所以我们不得不依赖他们的 SPCX NCCL 插件,每个连接使用 16 个队列对,并跨平面进行自适应路由。

为什么要费这个劲?答案是规模。Oracle 在这一物理设计之上构建了 Acceleron。他们的 MRC 拓扑论文展示了每块网卡如何拆分为 8x100G,从而将一台 51.2T 交换机变成 512 个端口,并在仅 2 层、最多 3 跳交换机的情况下,以全带宽支持 131,072 个 GPU。OpenAI 于 2026 年 5 月通过 OPC 发布了 MRC,Oracle 在 Stargate Abilene 运行该技术。软件层允许 MRC 将一个连接分散到多条路径和多个平面上,乱序放置数据,通过选择性 ack 恢复丢包,并使用 SRv6 源路由绕开故障或拥塞的链路。这一逻辑层在任何规模下都可能变得复杂。作业出问题的地方就在这里。

要分配一个 RDMA 设备并正确使用它,Kubernetes 必须向每个 pod 提供正确的设备和接口。换句话说,配置和使用 NVIDIA NetworkOperator 的方式有很多种。一旦配置错误,对所有人来说都是一场噩梦。

我们的脚本识别以下与 NetworkOperator 相关的部署模式(资源名称是我们所检查集群中的示例):

  • rdma_shared_device:rdma/rdma_shared_device_*,配有基于资源的 NAD,或主机网络回退方案

  • sriov_host_device:nvidia.com/hostdev 或 nvidia.com/rdma_host_dev,配有 HostDeviceNetwork NAD,用于直接设备访问和 GPUDirect RDMA

  • sriov_legacy:通过绑定 NAD 的 nvidia.com/<resource> 提供 VF,以及一个 SriovNetwork NAD,并在需要时加上 RDMA CNI

  • sriov_ib:一个 InfiniBand VF,例如 nvidia.com/mlnxics,配有 SriovIBNetwork NAD。分区网络架构还需要正确的 PKey 和 UFM 配置。

  • ovs_offload:使用 OVSNetwork NAD 的 nvidia.com/switchdev,用于硬件卸载的 Open vSwitch。

  • rdma_ib / rdma_roce:rdma/ib、rdma/roce 或 nvidia.com/rdma_*,每个 rail 一个 NAD,或回退到主机网络。

  • 提供商特定方案:AWS EFA(vpc.amazonaws.com/efa 和 libfabric)、DOKS(rdma/fabricN,配合 roce-net-fabricN@fabricN),以及 GKE DRANET(DRA claim 模板和 pod resourceClaims)。

设备设置以及其上的 NCCL 插件可能会影响作业性能。在 Google 的 GB300 上,我们测试时内置的 NCCL 库选择了一个不可路由的链路本地 GID 并挂起。设置 NCCL_IB_GID_INDEX=3 解决了这个问题,而 Google 的 gIB 插件在 16 节点、16 GiB all-to-all 测试中达到 99.5 GB/s,其他提供商为 95.3 GB/s,即满性能。

话虽如此,多节点 NVLink 本身就是另一个问题。NVIDIA 的 DRA 驱动中的 ComputeDomain CRD 为每个作业提供自己的 IMEX 守护进程和 channel claims,与 RDMA 无关。Google、GMI 和 Firmus 的 Kubernetes 使用了 ComputeDomain,但 RDMA 路径各不相同:Google 使用 DRANET,GMI 在主机网络上使用共享设备,Firmus 使用每 rail 的 VF。与此同时,Oracle、Azure Slurm、GMI Slurm、Firmus Slurm 和 Nebius Soperator 直接从主机提供 IMEX,没有租户 claim。Azure AKS 在没有任何租户可见的 ComputeDomain 的情况下运行了多节点 NVLink,尽管存在 GPU DRA ResourceSlices,而且 channel 来源和租户隔离边界都没有文档说明。

每个 ComputeDomain 必须保持在同一个 NVLink 域内,因为跨域需要 RDMA。换句话说,这是在分层式 scale-up NVLink + scale-out RDMA 网络上的拓扑感知网络。当我们错误地启动作业、不感知 ComputeDomain 时,只要放置跨机架,它们就会在 clique 设置阶段直接挂起。关于 EFA 不支持 DeepEP 和 MoonEP 的吐槽,我们留到附录中的 AWS 提供商评测再说。

可靠性、健康检查、自动修复与 SLA

在我们开始为 GPU 集群到底要花多少钱? 做研究的前后,我们开始向测试的每个集群注入故障。提供商之间的差异很大。处理一次故障需要两件事:

  1. 识别故障已经发生

  2. 正确地修复它

要识别故障,我们推荐使用监控仪表盘。一个好的仪表盘会显示故障组件、受影响的作业、当前调度器状态,以及每项检查上次运行的时间。一个过期的绿色结果并不能证明健康。当然,监控仪表盘依赖于某些遥测数据,所以你必须正确地(且安全地)设置 DCGM。

要修复故障,我们期望自动修复:重启并用热备节点替换。自动修复不会覆盖所有情况。有些故障需要人工干预、RMA 以及提供商必须直接与客户沟通的其他工作。我们在 我们关于这个主题的文章 中以有效吞吐(goodput)的形式量化了确切的成本。

NVIDIA 最近开源了 NVSentinel,这是一个面向 Kubernetes 的故障检测与修复系统。它读取 DCGM、syslog 和云端维护事件,然后可以对节点执行 cordon 和 drain、重置 GPU 或重启。默认安装仅用于监控。软件是简单的部分。重要的部分是把每个 XID 映射到提供商采取的动作。已被隔离的 XID 94 只需要重启应用。未被隔离的 XID 95 则需要恢复 GPU。显然,如果你天真地针对每个 XID 都重启节点,你会杀掉健康的工作负载;而如果在严重 XID 之后仍让 GPU 保持可调度,你就有工作负载崩溃的风险。

这也是为什么很多客户只想要裸金属。他们只是想让提供商关闭支持工单。关闭支持工单比听起来更难。一张工单可能来自 100 多种 XID 中的任意一种,每种都有不同的含义和解决方法。SOP 涵盖了从更换硬盘或电源、清洁线缆,到重新插拔内存、GPU 或网卡,一直到对单个部件或整个服务器托盘做 RMA 的所有操作。

自动修复有很多不同的选择,正确的选择取决于系统。此外,对 GB 与对 HGX 而言,修复是不同的问题。在 HGX 上,你可以换入一个带 8 块 GPU 的热备节点。在 NVL72 上,你无法向 NVLink 域中热插拔 8 块 GPU。一个故障托盘会导致 4 块 GPU 下线,可选方案是以降级状态运行整个机架,或者更换整个机架。好消息是 GB300 的故障率往往低于 GB200。无论如何,从错误到数据中心内动作的流程才是关键。

一些提供商告诉我们,他们在采取行动之前会等待一段时间,给客户留出把东西从节点上搬走的时间。这是自我安慰。我们模拟的是硬故障。节点已经死了,上面没有什么可搬的了。(一个普遍的事实是:如果提供商说他们选择不做某件事是为了“给客户留出可选性”,那就是自我安慰,而提供商认识到这一点只是时间问题。)

这就是自己拥有并运营数据中心与租用托管机房的区别。自己运营设施的提供商可以掌控现场的技术人员、备件和 SOP。而在托管机房里的提供商则只能依赖房东的远程 hands 及其排期。两者都体现在容量恢复的速度上,而这正是 SLA 必须衡量的东西。

在 SLA 方面,我们为 3 个层级的提供商制定了一套标准化的 SLA

在这些文档中我们提供以下内容:

  • 在两种场景下对“节点”、“机架”、“集群”和“站点”的技术定义(HGX 8 路系统,即 H100 或 B300;以及 MGX 4 路系统,即 GB200 或 VR NVL72)

  • 针对所有系统的“停机”的技术定义

  • 三个层级的对应 SLA(“青铜”、“白银”和“黄金”,包含相应的阈值以及与停机相关的赔偿条款)

  • 关于测量停机和提供赔偿的双向承诺

  • 建议不计入停机测量的排除项,例如软件升级、安全补丁和物理维护

  • 建议的提供商 SLA 达标情况月度审查节奏

  • 与低于建议阈值的停机相关的买方合同终止权利

  • 验收测试的定义,涵盖 GPU 计算、网络、存储和软件各方面应执行的测试类型的描述,以验证集群可正常运行并可被验收,并建议联系 SemiAnalysis 进行 ClusterMAX 技术评估

  • 买方在错过约定的验收日期时拥有的合同终止权利

  • “不可抗力”条款的定义

  • ……等等更多内容

我们鼓励算力买家(Neolabs)、云服务商(Neoclouds)、债务/融资合作伙伴以及保险提供商以这些条款作为其合同的基础,并在验收测试和月度 SLA 性能审查中加入 SemiAnalysis 技术评估。我们使用我们的 cmax 工具以及一套全面的专有流程来执行此技术评估。

如需更多信息,请通过 clustermax@semianalysis.com 联系我们。

安全

自从我们发布了“Most Neoclouds Suck at Security”一文以来,反响令我们备受鼓舞。我们的朋友们一直在他们的集群上运行 我们的 CLI,以找出哪些陈旧的软件包需要替换,而且总体上人们更加重视以应有的严肃态度对待这一问题。

不过,我们认为这值得更深入的关注。当关于 AI 是否会毁灭所有人的讨论达到白热化时,却几乎没有人分析它实际上可能摆脱控制并造成危害的具体途径。OpenAI 和 Anthropic 正在花费巨额资金进行功能增强研究,其中一些基础设施已被揭示存在严重疑点,然而很少有人评论说,GPU 集群——这些模型所依赖的物理主机——往往缺乏你所期望的免疫系统。

无论你对 AI 能力持何种看法,良好安全实践的必要性都同样成立。即使这些集群不会因为 AI 工具而变得更加脆弱,我们仍需面对这样一个事实:我们正花费数十亿美元建设的基础设施并不像它应有的那样安全。有很多令人不快的情景可以想象,它们所需要的不过是一个会用 Vim 并精通操作系统的人。读一读 Hugging Face 的报告,数一数那些让火势蔓延的普通网络安全失误有多少。

同样值得强调的是,这在某种程度上是一种结构性风险。当整个行业正在应对严苛监管和未经审视的社会抵制的种种危害时,LLM 工厂遭遇安全漏洞入侵是它最不应该招上头条的事情。这不仅出于网络安全方面的原因——每个被攻陷的节点都会增加坏行为者可探索的邻居数量——也出于普通的社会政治原因——无论好坏,批评者会用一竿子打翻整个 neocloud 行业。

做一个好邻居,别被黑了。

Agentic 编程

智能体编程已经在给托管 GPU 集群带来利润率压力,因为使用 AI 工具的客户更愿意选择裸金属或轻度托管的集群。我们在测试中不断看到这方面的证据。如果我们需要添加 Linux 用户,而集群上没有脚本,Codex /goal make useradd script and add my whole team 就省去了跑腿的功夫。尤其是在文档很差的集群上,智能体帮助我们攻克了各种可能的配置,让我们摆脱了阻塞。对应的 Slack 对话往往是这样的:

“嘿各位,XYZ 应该怎么用?”

“没事了,Codex 告诉我是 ABC。”

“对,就是这样。我们把它加到文档里。”

在有明确成功标准和充分上下文的环境中,智能体对我们的帮助最大;否则,它们处理 GPU 系统的方法就很幼稚。它们具备相关知识,但相对于看似合理却错误的方法,这些知识的权重并不高。例如,Claude 有一个习惯,就是用循环而不是作业数组来作业,从而把 Slurm 控制器压垮。我们的一个智能体认为配置 NVLink fabric 太麻烦,改为通过 InfiniBand 启动作业并宣布胜利;另一个在存储测试中犯了类似的错误,在 NVMe 而不是 NFS 上运行;还有一个在试图把 GPU 作业调度到 CPU 节点上之后,把性能差归咎于供应商。我们通常让智能体整夜运行,要么用 `/goal fix this part of the cluster so tests can run`,要么用 `/goal try to break this part of the cluster with a realistic workload`。总体而言,这使我们的产出得以倍增,但有时验证结果所花的时间比我们自己动手还要长。如果你对如何搭建集群有清晰的思路,智能体会让你更强大,你甚至可能愿意接受一个比平时更差的供应商。如果你不知道自己在做什么,你的 AGI 顾问会兴高采烈地教你把自己的脚打穿。在这个半透明的行业里,智能体放大的是技能差距,而不是抹平它。

CoreWeave 的上线流程是一个有启发性的例子,它本已是业内最严格的之一。如今,除了严格的经典流程之外,他们还让 LLM 对整个机群运行统计,寻找能更早发现故障模式并改进流程的规律。CoreWeave 还有一个 NodeBot,在自动化告警之上提供轻量级推理,总结情况并建议下一步操作,为参与其中的人类简化诊断。例如,团队描述过这样一个事件:一条 热界面材料泵出告警 本应把节点移出集群进行检查,但一条同时出现的 PMU halt 告警 却抢着重启节点而不是先做分诊。NodeBot 在冲突中识别出了正确的操作。CoreWeave 拥有大规模运行 GPU 的经验,掌握无数训练数据中不存在的经验法则,并让专家不断闭合反馈回路以改进对这些模型的使用;即使 OpenAI 用 AI 生成的内核造出了一颗了不起的芯片,我们也不认为 CoreWeave 会在短期内受到那些用 vibe coding 写数据中心自动化的新贵 neocloud 的威胁。

尽管说实话——行业里每个人至少在某些工作流中已经在使用 AI——但我们看到的唯一 AI 优先的产品是 AWS 提供的一个不错的 MCP 服务器。除此之外,所有文档名义上都是为人类编写的。供应商为迎合智能体所能做的最好的事情,就是提供详尽的文档——理想情况下,比人类所能读到的还要详尽得多——并在集群上留下“黄金配方”作为爬山优化的起点。除了那些尚未 AI 优先的传统流程之外,还有许多系统因为不是为智能体构建而明显表现更差。例如,尽管让 AI 管理你的作业可能有利于集群利用率,但它并不尊重为辅助研究而设立的旧式系统。CoreWeave 和 Nebius 有许多 Slurm 层的优化——带有 RBAC 和针对用户与组的优先级的真实作业记账系统、精心管理的分区、作业数组、日志和指标——但当一切都以智能体方式运行时,这些都被浪费了。CoreWeave 还告诉我们,他们见过头节点内存耗尽,因为远程操控的 Claude 在糟糕地编排集群。看着供应商改进其 AI 智能体可用性,将是一件有趣的事。

展望未来

如本文通篇所述,我们历来是在测试集群。但市场显然已经朝着两个相反的方向发展。

First, the biggest labs are buying bare metal at massive scale, measured in the 100’s of MWs, and have the talent in house to manage their own training clusters, orchestration software, monitoring and reliability. In essence, their neocloud providers (when they use them, and don’t do self build) are datacenter technicians who close tickets. This is a complicated business, and not to be taken for granted: there are 172 unique, documented ways an NVIDIA GPU can fail (XIDs). And many of these failure modes can overlap, resulting in an incredible surface area of different failure scenarios that require a combination of simple troubleshooting logic, human experience, and increasingly, agentic research by AI to recover from these failures.

Meanwhile, many startups are raising 10s or 100s of millions of $$ and spending effectively all of that money on compute. But since so much of this research is driven by RL, via continual learning on real production traces for long horizon agents, the needs for both offline, throughput-optimized inference, and online, latency-sensitive inference is very strong. At the same time, more and more customers are looking for the simplest way to purchase vanilla open source tokens for their workflows.

This leads to our next article, coming soon, where we will describe the Anatomy of an Inference Endpoint, moving towards the inclusion of inference endpoint testing in future versions of ClusterMAX. We believe that all neoclouds need to have an endpoints business going forward.

We will also be evaluating RL infrastructure beyond inference: sandboxes, hosted training, evals, and more. We continue to see lots of new providers enter (and exit) the market every day. It is an exciting time to be in AI, and neoclouds are at the centre of it.

If you got to this point in the article and are still eager to learn more about our individual experience on every provider, you should consider coming to work at SemiAnalysis!

We are actively hiring across Research, Consulting, and Technical Staff. We have multiple openings in our New York and San Francisco offices. Remote work is also acceptable.

ClusterMAX Engineer: all levels of experience with Slurm, Kubernetes, and GPUs considered (full time)

Tokenomics Engineer: all levels of experience with model evals, harnesses, inference endpoints, and RL infrastructure considered (full time)

Research Analyst - AI Infrastructure and Economics: all levels of experience covering the financials of neoclouds, neolabs, and frontier labs tokenomics considered (full time, or internship)

Technical Consultant: lead engagements from technical strategy to technical due diligence. all levels of experience from consulting and technical backgrounds considered (full time)

Thanks for reading, and we look forward to seeing you for future versions of ClusterMAX.

Appendix: Provider Reviews

Platinum

CoreWeave

CoreWeave remains the industry’s sharpest, sturdiest, most proactive neocloud.

Our CoreWeave clusters are exemplary in all crucial categories. The health checks work as intended, the reliability is excellent, and nearly all tests reach expected values out of the box. Along with Nebius, CoreWeave serves as a measuring stick for others in the industry—not only in the abstract sense of “excellent performer,” but also in the literal sense that we compare their clusters to lower-rated ones as a means of precisely describing others’ shortcomings. Whereas other neoclouds have to be nagged to refresh their moldy GPU drivers, CoreWeave has not only taken care of the basics, but stood up a significant independent stack for marginal performance and reliability improvements.

The ClusterMAX 2.0 report went into great detail on CoreWeave’s design decisions, including its bare metal provisioning, Slurm fork, use of DPUs, and health checks. Since then, one smart feature that the nerds at CoreWeave have added is called “GPU straggler detection,” which is integrated into their best-in-class dashboard. Rather than leaving their customers to dig through logs or do binary search over their fleet to isolate a bad rank, CoreWeave applies handrolled algorithms to NCCL telemetry and recommends a remediation strategy based on its findings. The CoreWeave team tells us that they use this on calls with their customers just about every day. There are all kinds of soft failures that don’t write an XID or leave any other obvious signature, but waste GPU hours as the job waits on a straggler.

Better than spotting bad GPUs is not scheduling them in the first place; CoreWeave also has systems to make sure every node clears a high bar before it’s trusted with a customer’s workloads. When a node is idle in CKS, CoreWeave runs burn-ins for 20-30 minutes of every hour, validating that scale-up and scale-out fabrics are healthy, GEMM numbers are up to spec, D2H and H2D bandwidth are high, and so on. These tests are pre-empted by customer workloads; invisible to the operator, they help make sure that it’s not production workloads that catch faults. CoreWeave is certainly not the only provider with active health checks, but these are unusually robust.

This diligence extends to CoreWeave’s bring-up process. Provisioning around 10,000 GPUs per week, CoreWeave believes that they have more data than God about what proves a Blackwell rack healthy. In CoreWeave’s view, Nvidia diags are good at catching many failures, but their uniquely deep dataset has allowed them to roll out a custom suite of supplementary tests, including for NVLink. The heart of this process is the Fleet LifeCycle Controller (FLCC), which “automates node provisioning, testing, and monitoring” and owns the path from power-on to handover. Importantly, as others operating fleets at scale can corroborate, there’s a long tail of obscure failure modes not well covered by Nvidia documentation. CoreWeave keeps statistics on these failures to better predict future faults and bakes its response patterns into FLCC. This is one reason that experience with operations at scale is a crucial criterion in our ratings. There is a huge difference between a newbie and a provider like CoreWeave that has rigorous runbooks, automation, and burn-ins. As described in the “Agentic Coding” section above, CoreWeave uses LLMs on top of this system to glean insights that would not be spotted by a human engineer. CoreWeave also works with manufacturers before delivery. Involvement varies by OEM or ODM and by production stage, but in general, the goal is to work backward through the supply chain to catch issues as early as possible: never catch in prod what you could have spotted in bring-up, and never catch in L11 diags what you could have spotted on the factory floor.

Their experience with Blackwell NVL72 bring-up and validation, along with their close relationships with Nvidia and Dell, have allowed CoreWeave to be the first provider to announce a VR200 NVL72 system passing L11 diags. CoreWeave has described VR as “the generation where we bring our own IP.” As the physical demands of compute continue increasing, in CoreWeave’s view, “how the building operates is now part of how the computer operates—they are not two individual things anymore.” Thus, CoreWeave introduced its custom Racky, Valvey, and RLCC in a recent blog post. Particularly interesting is Valvey, a programmable per-rack liquid cooling valve assembly. This gives CoreWeave fine-grained control over each rack’s cooling loop and also, in case of a leak or other emergency, the ability to trigger a shutdown. Despite beginning to adopt a hyperscaler mentality in building its own hardware, CoreWeave frames its approach as anti-hyper: whereas traditional infra emphasizes redundancy and is built to prevent failures, CoreWeave builds to fail gracefully and recover fast, limiting failure domains and allowing nimbler response to the inevitable failures.

Amidst the current trillion-dollar buildout, however, CoreWeave is responding to the massive increase in customer demand and dealing with increasing cost of capital. Of course, there is plenty of good news to mention. Its relationships with Meta, Microsoft, OpenAI, and NVIDIA—its largest customers—appear strong, and it recently struck a multi-year deal with Anthropic. It has announced splashy deals with HRT and Jane Street this summer. It is also aggressively expanding its capacity in APAC, indicating that it “expect[s] international markets to become a major driver of growth.” Yet it sits with $35B in debt and as the cost of capital increases, it takes longer to get projects off the ground. As a result, CoreWeave has become focused on long-term bare metal contracts, which are easy to fund with an investment-grade counterparty, rather than managed clusters, which have higher margins but which the market is less keen to fund. Moreover, our last writeup mentioned an impressive lineup of recent acquisitions, but since then its Core Scientific merger fell through and nothing else has been announced.

In terms of feedback, some of which is a bit nitpicky:

  1. We wish that enabling local GPUDirect Storage could be an 1 click button in the UX console

  2. Onboarding new users and where they put the the ssh pub key input in the UX console is extremely confusing to many users that tried

CoreWeave has begun to diversify from bare metal, first by offering self-service SUNK clusters, which gives them the option to sell into the on-demand market when they have capacity for it. It is also worth noting that not all “bare metal” contracts are the same. In all of CoreWeave’s bare metal contracts, they still manage the scale-out and frontend network, do burn-in, handle repair/replace, monitoring, and other general day-two operations. By comparison, other providers claim to be a neocloud, but are locked out of their own data hall. We will dig into more of these details in an upcoming article where we compare the different bare metal offerings from different providers.

CoreWeave has also launched a managed inference platform, which recently disclosed $100M ARR. We are actively testing this for an upcoming article on serverless inference endpoints, though were disappointed to find no support for Kimi K3 and GLM 5.3, and the only way to consume the public endpoints by the token is via the Weights & Biases portal, or via OpenRouter. For now, let’s just say that CoreWeave has some work left to do on their inference endpoints.

Nebius

Nebius was rated the 2nd-best neocloud in the world in ClusterMAX 2.0, yet it remained in Gold. This time, Nebius is unquestionably an industry leader with strong offerings in every category and the ability to command a significant price premium. Nebius moves into Platinum.

Nebius’s relationship with Nvidia remains strong, including a $2B investment in March, as Nvidia disclosed a 9.3% overall stake in July. Nebius was the first provider to deploy HGX B300s back in December of 2025, and has been among the first neoclouds in VR NVL72 bring-up, expecting to begin deployment late this year or early next.

Nebius’s buildout continues across a number of sites in the US and Europe as they recently raised their 2027 year-end contracted power target to 5 GW. In terms of offtakers, Nebius has signed large deals with hyperscalers Meta and Microsoft and smaller ones with Reflection AI and Palantir, making it one of a few neoclouds with multi-hundred-MW hyperscaler agreements that still regularly competes for much smaller startup contracts. They are the most active of all neoclouds in the short-term cluster market, serving many happy customers. The days of Nebius spooking prospects with their Russian accents are behind them. This is mirrored by its capacity procurement: unlike many of its peers, Nebius has opportunistically snatched up sites in the 5-20 MW range through a broad network of datacenter and software partners.

In terms of the clusters themselves, well, a good cluster is generally boring. On Nebius, burn-ins run without errors; Slurm is topology-aware; the packages we need are on the cluster and generally recent; WAN is good; the orchestration layers are free of footguns. One point of distinction technically between Nebius and everyone else is that Nebius builds and open-sources its preferred SonK flavor, Soperator. (CoreWeave, of course, builds SUNK in house, but it is not open-source.) This round, a Gcore cluster we tested used Soperator, and as did Voltage Park and FPT clusters from previous rounds. Nebius writes that “the real magic of Soperator lies in how we use ‘jail’ Persistent Volume”: Nebius provides a VirtioFS volume at `jail`, then bind-mounts it at `/`, so you don’t need to worry about pointing your cache to the right directory, or think much about taking your files with you when you `salloc` a node. You’ll see in our writeups of other providers a handful of gnarly issues with managing the storage system on SonK: Slurm and Kubernetes have different philosophies of persistence, and it’s not always straightforward to provide the illusion of bare metal when in fact you’re several layers up, standing on an ephemeral Kubernetes pod. Nebius’s approach makes this stress-free for the operator.

In particular, Nebius provided us a handful of storage tiers on every worker: `/jail` VirtioFS, `/home` NFS, `/mnt/data/` VirtioFS, `/mnt/local-nvme` ext4, and `/mnt/memory` tmpfs, as well as first-class S3. For I/O-heavy jobs, `/mnt/data` is recommended, whereas `/home` suffices for shared code and other data with lower-concurrency access. When we first tested `/mnt/data`, we saw very slow 4 KiB sequential allocating writes, while concurrent directory creation from 128+ ranks sometimes returned `EEXIST` due to inconsistent metadata visibility, killing our jobs. With our feedback on the pathological workloads, Nebius retuned this filesystem, and it proved sturdy and fast. We pushed it around and verified that the errors were resolved; per-client, it outperformed the median, and we could further verify that it scaled to handle demanding workloads across two racks.

Nebius has put a great deal of work into its health checks, and in our testing, they performed well. Nebius automatically returned a node to service after a synthetic error injection on our GB300 rack; the process is not yet optimized for speed, taking 8h40m from start to finish, but detection is immediate, and the process ultimately works. Moreover, the other GB300 racks we tested did not automatically reboot and return failed nodes, so the fact that this process is self-driving counts in Nebius’s favor. (As we describe in the “Blackwell and Grace Blackwell” section, it’s common not to autoremediate GB300 nodes at all, since there is no real way to hot-swap a node into an NVL72 rack, and troubleshooting is generally way more complicated than with HGX servers.) Nebius also provided us with an excellent suite of dashboards shedding light on, among other things, cluster health.

This all leads to Nebius watching their revenue per MW climb compared to CoreWeave’s (and their market cap along with it). Since they have so much more capacity to sell at the current prices, we expect this trend to continue. Anecdotally, we have seen Nebius get very aggressive during recent negotiations, including a case where they offered a tranche of capacity as a straight-up auction. In general, demanding high prices and sizable prepayments that reach 100% on 1 year commits seems like quite a nice business, especially when those prepayments can cover the entire capex of the servers at the current TCV. Infinite Project IRR anyone?

At their largest bare metal sites, Nebius continues to face delays in construction and permitting, including both Béthune and Vineland, which we have covered extensively for clients of our industry-leading Datacenter Model. But as more chips of all types come online, Nebius’s track record puts them in a strong position to continue to grow.

Overall, we are consistently impressed with Nebius. While we do believe they still trail CoreWeave on many technical aspects and relationships with both NVIDIA and the frontier labs, solid business decisions have established them as the default neocloud to serve the neolabs.

Gold

Google Cloud

Amidst the strange retreat of Google Brain and DeepMind from the frontier, GCP has been dealmaking like few others. Notable partners include Palo Alto Networks for $10B and Thinking Machines Lab for a “multibillion-dollar deal,” along with expansions of its lucrative relationship with AI safety org Anthropic. With acquisitions of cybersecurity startup Wiz for $32B and energy startup Intersect for $4.75B in cash, Google has been aggressive in expanding its technical capabilities. It also exposes itself to new revenue streams by ramping TPU sales instead of only renting them through GCP, a development we are very keen to track. Its own TPUaaS is still growing, though, thanks to a $5B deal with Blackstone that is targeting 500 MW in capacity. Google has gigawatts in the pipeline, and we look forward to continuing to track its continued competition with the other hyperscalers and labs for access to power and permitting.

With a unique position cutting across so many layers of the stack, Google may have more total engineering talent than any company in the world, yet it has a notably imperfect history in supporting innovation. In ClusterMAX 1.0, we noted a myriad of issues with its platform, yet we predicted GCP would reach Gold or Platinum promptly. It has taken until ClusterMAX 3.0 for this to come true, but Google is finally and comfortably among the best managed cluster providers.

The GCP GPU experience is less refined than that of our Platinum providers, CoreWeave and Nebius. The console feels like the DMV compared to the streamlined UIs of the younger, more focused neoclouds. Access required a Google Cloud CLI that was occasionally uncooperative, forcing us sometimes to switch to the browser-based connection in Google’s console, which was also unreliable. There was a small misconfiguration on our GB200 GKE cluster, where default StorageClass refused to attach to the `a4x-highgpu-4g` nodes, forcing us to try again with a non-default class to get block storage. Nitpicks aside, GKE is very well set up, and we enjoy using it. Google’s managed Slurm is officially generally available and quite solid, with good defaults and health checks configured. A lot of the testing that we did for ClusterMAX 2.0 on their managed Slurm offering is now valid, since the bureaucrats in GCP Product Management are allowing real customers to use the offering, not sticking it with labels like “Beta” or “Pre Release” or “not yet GA.” (By the way, many of the neoclouds we’ll discuss after Google should consider using some of these labels more often—the balance is somewhere in the middle here.)

The biggest point of distinction between our GCP cluster and the other clusters we tested was networking. Our first NCCL test hung because automatic GID selection chose the wrong address, but pinning `NCCL_IB_GID_INDEX=3` completed the run. Even then, our NCCL tests were unsatisfactory: you would like to see a roughly logistic curve as performance increases monotonically in message size, but instead we had the following jagged shape:

This could conceivably have been an issue with NCCL itself rather than a problem with our GCP configuration. NCCL’s heuristics don’t always pick the right protocol and algorithm for the message size and topology, and some hand-tuning is always expected. However, we have gigabytes of networking data gathered at this point, and we knew to expect better perf. The issue turned out to be simply that we did not have the gIB plugin enabled. This is a set of NCCL plugins Google that ships for improved perf on Google’s RoCE network. As of 26.07, gIB is baked into the NGC PyTorch image, so Google’s customers get the custom stack by default. With gIB installed, we got a better chart with 4.4% higher peak throughput for 16-node all-to-all. (Note that these tests are both with the multi-node NVLink fabric shut off to isolate performance of the scale-out network, which is RoCE in this case.)

While not perfect, this looks much better, and the topline number is solid. At this point, it was clear that the fabric was healthy, the configuration was good enough, and an engineer looking to squeeze more juice out of the system would have a solid baseline to start from.

This is amazing work by Nvidia and Google to set up auto activation for Google’s ConnectX-7/8 NCCL plugin. Previously, users would need to mess around with the correct library load paths and env vars to get it set up correctly and optimize performance on ConnectX NICs on GCP Nvidia GPU machines. We gave feedback a couple of quarters ago, and now, it is fully automated!

Source: Nvidia

GKE’s health checks worked well. As with other providers who have experience with NVL72 systems, Google’s health checks intervened but left it up to the operator to bring the sick node back to the fleet. In particular, after our XID was injected, the health check marked `GPUUnhealthy=True` and `cloud.google.com/health-check-status=warning`, but left the node schedulable. This configuration was easy to work with, it’s what GCP’s customers prefer as a default, and it can be configured to the user’s liking to respond to errors of various levels of severity.

To continue improving, GCP could take a look at all the quality-of-life changes that Nebius and CoreWeave have made in the recent past. It should ship a better console that’s easier to navigate, simplify IAM and RBAC, and ensure cluster access is painless, either through standard SSH or `kubectl` or by making its existing CLI bulletproof. Especially now that gIB is available in the default Nvidia images, we hope that GCP’s operators have no trouble getting good performance out of their GCP scale-out. They could also continue to pursue neolabs in the mid-market, improve the hands-on support experience, and improve their relationship with Nvidia. GCP consistently loses so much business to smaller, less capable neoclouds, because they either lack the GPU capacity or refuse to serve the market at a certain price point.

We look forward to testing GCP’s VR NVL72 systems when they are available and seeing GCP’s continued improvements on usability and performance.

Oracle

As we mentioned in ClusterMAX 2.0, Oracle is in a unique position among hyperscalers in that its growth must come from large contracts. Unlike AWS, Google, and Azure, Oracle does not have any agreements to sell frontier tokens, and unlike Meta and Google, it does not have a non-frontier lab it can hang its hat on. Instead, it continues to cut huge deals involving OpenAI, Meta, and Nvidia, reporting “the delivery of 850MW additional datacenter capacity” from June through August, an eye-watering number. We project its relative growth in 2027 to be the largest of the hyperscalers, though from the smallest base.

We tested two Oracle clusters, which we will describe in order. The first was a GB300 NVL72 cluster on Slurm, set on a multi-planar network. This was vanilla Slurm, but due to customer demand, Oracle has Slurm-on-OKE on its roadmap. Once our allocation began, we tried to find our way in through the console, but it proved, like all hyperscaler consoles, a forbidding place.

Importantly, Oracle’s console did not support RBAC—nothing on the console is connected to SSH users—so all our SSH keys had to be managed on the machine. At one point, attempting to append to `authorized_keys`, our intern forgot a `-a` flag and removed all our previous keys. Classic user error, obviously, but this goes to show why we prefer to have a hardened path on a console. We’d rather not trust our interns—especially that one, no offense—to type the magic words on the command line when a mistake can lock us out of the cluster. Thankfully, Oracle’s engineers caught the mistake immediately and fixed it for us.

We hope they fix this and add enterprise grade RBAC.

As expected, Oracle’s cluster performed well from day 1. They provided us with scripts for InfiniBand and NCCL tests, which was helpful because vanilla scripts do not properly use all 4 rails. Burn-in was flawless and Oracle’s managed Lustre saw strong throughput. Oracle’s health checks worked as intended, draining the node on our synthetic error but leaving up to the user to decide whether or not to terminate and replace the GB300 node. Oracle’s managed Prometheus and Grafana were good, though not perfect. For example, they are not integrated with the end user experience for job accounting. However, on the positive side, health metrics are now integrated into Cluster Manager in the console, making it easier to get a view of the crucial statuses at a glance.

Next, we tested OKE with MI355X. This was the only provider other than TensorWave who let us test-drive AMD GPUs. The ride was rough.

As we complained about last time, at first, we needed to SSH into an operator node to steer Kubernetes. This time, the team quickly solved our Oracle CLI issues and we were promptly able to download our credentials locally. By far the greater problem on this cluster, though, was with the NICs. Peak performance was near line rate, so there was no issue with the hardware that we could detect. Rather, the issue was with the ionic drivers. Our per-rail `ib_write`/`ib_read` sweep across all eight NICs hit 391 Gb/s on every rail, then wedged the node the moment it finished: when the pods were deleted, the perftest processes never completed their `rdma_cm` teardown, the forced `kill` timed out the NIC’s destroy-CQ command, and the RDMA admin queue on `ionic_0` and `ionic_4` went dead. Worse, both ports still showed `ACTIVE` at 400 Gb/s and the node stayed `Ready`. The passive checks did not check out the RDMA control path, and no active check had run in eight days. We found the issue by grepping `dmesg` after RCCL kept failing. Oracle diagnosed it quickly on a weekend and rebooted: their own `ibwrite` health check had failed on this cluster in the same way days before. After a call where we explained the pathological workloads, they spoke with AMD and pushed different drivers and firmware to the cluster, which solved the problem.

Separately, we had a GPU intermittently dropping off the PCIe bus, neither caught by the health check nor properly surfaced in the dashboard. Oracle promptly swapped out the node and responded to our feedback on the dashboard.

This was a mixed episode, but Oracle’s responsiveness was good. It’s especially undesirable that no health checks alerted, even though these drivers are known to be flaky. Still, Oracle’s close relationship with AMD is clearly an asset. They were able to speak to the relevant teams, diagnose the issue, and put a patch in place quickly. This type of support is not guaranteed at organizations of Oracle’s scale. Interestingly, Oracle is also the only provider ranked Silver or above, and one of two ranked Bronze or above, with zero installed libraries with applicable CVEs. This is more of an expectation than a benefit for us, but it was a clear demonstration that their automation works for deploying and upgrading clusters.

If we were a neolab, we would come away from this testing experience with confidence in working with Oracle. Yet this is a moot point given Oracle’s increasing commitment to bare metal deployments for frontier labs such as OpenAI in Project Stargate. Success in that business depends more on getting government approval for gas pipeline construction than it does on RCCL tests.

Still, Oracle has made progress since last we tested them, and the experience is very good overall. We look forward to continuing to keep tabs on them going forward.

Silver

Lambda

Lambda was most recently in the news for participating in a $35B Anthropic deal for 350 MW in Nueces County, alongside Hut8 and Nvidia. Before that, they were rumored to be preparing for a 2027 IPO, and they separately raised $1B, $926M, and $1B in debt for their buildouts. Amidst all their growth, after judging them Silver in ClusterMAX 2.0, we were interested in evaluating Lambda’s clusters this time around.

Onboarding began with a nice deck and PDF explaining what to expect on the cluster, and access was sorted out promptly.

We criticized Lambda in ClusterMAX 2.0 for their reliability; this round of testing, it was one of Lambda’s points of emphasis. Their passive health checks cover the relevant conditions and ultimately autoremediated three errors we simulated during testing. Two synthetic XIDs were autoremediated in less than 15 minutes, and there were connected dashboards to monitor node health. Our test that triggered a genuine XID 79 by resetting the PCIe secondary bus reset brought the node into `NotReady` with no `GpuXid`, no cordon, and no monitoring visibility on the Kubernetes layer, but it ultimately rejoined the fleet after 2 hours. Our node-reboot tests on the Kubernetes layer also completed without incident.

The orchestration layer has improved significantly, though it still has a handful of issues. Slurm configuration left CPUs per task unset, so Slurm’s default allocation gave one logical core unless the launch requested more. Passwordless sudo is not enabled on Slurm, as we are using managed slurm rather than unmanaged Slurm. Lambda is the only provider we’ve encountered that makes this distinction, but clearly some of their customers want root access to their own cluster, while others do not (?). We enjoy having root on our own cluster.

Also missing was SSH to the compute nodes without an allocation. This creates a headache during debugging if someone else is running a job and you can’t attach to their allocation, or a job is stuck. Another reason why we enjoy “yolo mode” on our clusters.

More importantly, Slurm had software out of date, including CUDA Toolkit 12.0 at the default location /usr/bin/nvcc (which is so old it’s incompatible with the B300 SM100 architecture), and unacceptably old drivers. On the bright side, Slurm did have topology settings configured, and the NCCL recipes that they gave us worked well. We had no authentic errors during testing, and all our tests, including burn-ins and storage, showed good performance.

Our preoccupation with Lambda’s 1-Click Clusters seems quaint in today’s world of capacity reservations backlogged for months. Indeed, as mentioned in the opening, Lambda is increasingly driving its bottom line via bare-metal builds, apart from managed clusters altogether. Regardless, Lambda’s orchestration continues to improve, and customers are happy, which lands them in the upper ranks of ClusterMAX 3.0.

Microsoft

Using Azure clusters can be very difficult due to the Microsoft bureaucracy. Their engineers do the best they can, but this is not an organization set up to be flexible in accommodating its customers. Microsoft customers do not come first in matters of taste.

During testing, we were not allowed to have our own Azure accounts, so we had to use the cluster under the user of one of Azure’s engineers, which prevented us from poking around the console and experiencing the authentic onboarding experience. Our GPUs were set to be in Virginia, but thanks to some obscure internal decree, we were not allowed to have a control plane there. Instead, our control plane ran in San Antonio, Texas. Access required setting up a VPN, which worked after a bit (see: literally 3 weeks) of back and forth. Pasting was disabled on the bastion, so to run our tests, we had to manually type in our GitHub and HuggingFace tokens. Like, literally type them in one character at a time. We got them right first try though, nbd 😮‍💨.

Once everything finally ran, we found a handful of packages out of date, and bad default NCCL configs kept our first run from going through. On AKS, the `managed-csi-premium-v2` `StorageClass` could not attach, and we got `EAGAIN` trying to run `fio` at anything more than 4 clients on Azure Files Premium. No parallel system attached to our Slurm cluster meant that one of our tests, which loaded DeepSeek-V4 from NFS, timed out.

Source: a typical Azure experience

When we went to run burn-in, we had a genuine failure on 10 nodes of our Slurm rack, killing our job. Microsoft caught the errors, traced them back to an unhealthy host, and opened up a Guest Health Report for repair. An `scontrol` resume command brought them back, and after that, both our Kubernetes and Slurm racks completed their burn-in without error.

CycleCloud drained the node immediately after we tested Azure’s health checks by injecting a synthetic XID, but it did not resume or replace it. Because we had the whole GB300 rack, this was satisfactory, as described in the “Blackwell and Grace Blackwell” section above.

Due to the account issue, we could not see our Kubernetes dash by default. Instead, Azure’s engineering team cooked up a custom dashboard setup for us, and it was extremely detailed—a solid triumph over the bureaucratic straitjacket. We want to shoutout our friend Xu Xue for this one, as we are quite certain that he went way above and beyond the call of duty to build this out for us, but wouldn’t admit it. His curated dashboard is open sourced here, a great resource for anyone building Grafana dashboards to monitor Nvidia GPU clusters from scratch.

Satya once said “we want to build out Azure as being fantastic for the long tail of the workloads,” and that Microsoft is “not in the business of just doing five contracts with five customers being their bare-metal service.” Yet it is exceedingly rare to come across a VC-backed lab that uses Azure for training and inference. “Five customers for bare-metal service” is overstating it. Really, it’s two: OpenAI and MAI. Soon to be three, with Anthropic diversifying.

Azure’s most interesting asset is that it can serve OpenAI’s models and keep 100% of the revenue. They still have full access to all of OpenAI’s IP. Yet Azure can separately boast of large bare-metal deals with OpenAI, as well as an increasingly lucrative relationship with Anthropic. This is an incredibly strong position to be in. If only for the pause.

As Azure mostly competes on a different plane with enormous deployments and hundreds of billions of capex, its ability to manage clusters only makes it more flexible. Like Google and Oracle, Azure gets its cashflow from offtakers who need no help setting up a `ComputeDomain`. But based on the quality of testing, these three could still cut smaller, higher-margin deals with labs who want a higher-touch experience.

This does not seem likely to happen any time soon, though. Support and the console is a headache, even if Azure is the only hyperscaler where you can possibly get the Nvidia Reference Architecture on your scale-out networking. As much as anything, we like to verify that these hyperscalers are light on their feet.

Firmus

One of Nvidia’s favorite up-and-comers, and our favorite in Asia-Pacific, Firmus announced in August a $2B funding round at $10.5B post-money for expansion throughout Australia, Malaysia, and Indonesia, following a $505M round at $5.5B just 4 months before. Firmus also raised a $10B debt facility in March for its Project Southgate buildout in Melbourne and Tasmania. Just last week, it disclosed 900 MW in total contracted capacity and announced OpenAI as an anchor tenant in its Malaysian expansions. All in all, Firmus has disclosed 18,400 GB300 GPUs in the deployment process today and 36,800 later this year in Tasmania, with multiple GWs and tens of thousands of VR in the pipeline.

We had a rough takeoff but a fairly soft landing while testing one of Firmus’s dev clusters. Accessing our chips required using a Microsoft Entra account, Microsoft SSO, then downloading and configuring a custom pre-release version of the vCluster CLI. We’re suckers for new tech, but where cluster access is concerned, we’d much rather just SSH. After a few rounds of back and forth, with the help of Firmus’s team, we were on.

…on the cluster, but not quite off and running. The default setups on both Kubernetes and Slinky left a lot to be desired. Slurm login had no `sudo`, no `vim` or `nano`, HPC-X installed but not on `PATH`, and no NVCC. More importantly, instead of receiving 4 Slurm nodes on one rack and 4 Kubernetes on another, we had 2 of each on each. This was not necessarily a problem, but Kubernetes had `nvidia.com/gpu.clique` configured with nothing consuming it, while Slurm used `topology/flat`. Thus, the nodes visible to each orchestrator spanned multiple racks, and there was nothing to avoid a multinode job needlessly crossing the rack boundary. We had to infer node nomenclature to make sure our tests were well placed.

On top of this, the Kubernetes control plane’s `etcd` visibility flickered several times during testing. which killed our jobs and caused both the Slurm and Kubernetes layers to restart. This was ultimately attributed to an issue upgrading switch firmware and did not recur for the last 10 days. We also lost a NIC to a wedge during reboot, which was not caught by Firmus’s health checks as they were not active on the dev cluster we tested. Firmus was the only GB300 cluster we had that exposed GPU HBM as NUMA nodes, an interesting configuration that we mark as a failure because it lets you accidentally swamp device HBM if you overflow on the host. WAN was inconsistent but often quite poor, maintaining 0.25 Gb/s to NGC. The last of our problems was that one of the racks saw significantly degraded performance on its NVLink fabric, which was explained by a power setting that Firmus had been experimenting with on their dev cluster; fixing this brought our numbers into the healthy range.

Thus, by the time testing ended, we had satisfactory though underfurnished Slinky and Kubernetes layers, we knew how to access them, and we could get strong performance across our test suite. We understand that managed clusters are not Firmus’s priority right now, but we believe there is plenty of low-hanging fruit here for their technical team. As Firmus works to bring enormous capacity online, we look forward to seeing them continue to make life easier for users of their GPU software stack.

TensorWave

TensorWave is an AMD-only neocloud whose MI355X offering we tested via K8s and Slinky. In the last writeup, we emphasized the difficulty of TensorWave’s onboarding process, which required back and forth just to get on the cluster and had a large number speedbumps even still. The improvement this time round is significant. Onboarding was smooth, the `kubeconfig` sent directly to us worked perfectly, the console allowed us easily to add team members with SSH access, and the TensorWave support team was attentive throughout our testing.

Being an AMD-only cloud, TensorWave is fighting an uphill battle dealing with AMD’s software stack, which, despite significant recent progress, still lies well behind Nvidia’s. We saw this on our TensorWave cluster on RCCL microbenchmarks, with all-gather and all-to-all hanging at 32 KiB even using TensorWave’s binaries and recipes. Other collectives showed uneven scaling and poor throughput at certain message sizes. [Editorial comment: RCCL? More like Rick L!] In general, it can be said that the AMD networking stack lacks the community support that Nvidia tooling offers. The same goes, of course, on the hardware side. More important than microbenchmarks, we see in practice that the GB300 NVL72 systems have the most demand across the industry from labs: their performance is so superior that, even accounting for their higher prices, they often offer the best perf per dollar as well. AMD’s rackscale answer, the MI455X Helios, is still ramping up production, and is certainly not yet on offer by TensorWave. Thus, TensorWave’s place in the rankings depends not only on its internal improvements, but on AMD’s ongoing work to close the gap relative to Nvidia.

For what it advertised, our TensorWave cluster was good across the board. North-south bandwidth was strong. Storage performance was exceptional and burn-ins completed without a problem. When we pushed around the Kubernetes layer, its reliability was good, and the Slinky layer was easy to use.

The health check immediately recognized the error that we simulated, bringing the node into `drng`.

The dashboard oddly marked the node as “Allocated,” rather than being broken out into its own category. However, shortly afterward, we were presented with a nice juicy button to click to approve replacing our “sick” node with a fresh one.

The replacement completed in approximately 43m. We would prefer that the autoremediation occur by default—we’d rather not even have to click a nice juicy button—but otherwise consider TensorWave’s monitoring and autoremediation exemplary.

Overall, TensorWave is the leading AMD-exclusive neocloud, providing a comparable (and sometimes better) experience on AMD GPUs than hyperscalers like Oracle and Microsoft. And they are not stopping. TensorWave has entered into an agreement with Fermi for 222 MW in the Texas panhandle with expansion rights up to 650 MW, contingent upon Fermi securing project financing. It has a tailwind with a $350M June Series B, leaving it at $1.55B post-money, and is currently working on large deals with hyperscalers. With both OpenAI and Anthropic announcing strategic partnerships with AMD for the MI450 generation, we expect that TensorWave is set to benefit, and AMD-exclusive neoclouds may be able to break into the Gold tier. We look forward to testing TensorWave’s MI455X Helios once it is available.

GMI

GMI is a neocloud headquartered in Mountain View with roots in Taiwan. With a $500M Nvidia deal announced last November and $12B more this March, hundreds of GB300 racks in the pipeline, and plans for VR down the line, they are keeping on the frontier in terms of hardware. Recently, they have also been growing their fleet by making deals with smaller neoclouds and reselling the compute to their established customer base, which offers an attractive cashflow profile because it does not require up-front capital outlay. After judging them the “top Bronze neocloud” in ClusterMAX 2.0, we were interested to test the progress of their managed clusters.

Onboarding was a bit bumpy, including an empty Kubeconfig downloaded from the console and an SSH key that we uploaded but was never synced to the cluster. After some back and forth with the team, we got in. The problem, it turned out, was that the console was view-only for our testing purposes, so we could not verify its functioning. Last round, we did not get access to a self-service console at all, so we count this as progress.

As last time, the compute was up to spec, burn-in was clean, and the InfiniBand fabric was sturdy. One notable improvement since ClusterMAX 2.0 was in GMI’s storage. GMI provided a performant RWX NFS default StorageClass, which was also available on the Slurm layer at `/home`. We had difficulty trying GMI’s S3, which turned out to be due to a typo in their doc. But once the Brothers Karamazov Karasev hopped on a call and fixed it, we verified that the object storage perf was also good. The biggest blocker in terms of cluster configuration was that GMI expected the IMEX domain to be configured on the Kubernetes cluster by a ComputeDomain CRD. This is perfectly viable—many other highly rated providers do the same thing—but GMI’s Kueue would not tolerate the claim, so our jobs sat in `Suspended`. To work around, we had to submit our jobs in a namespace with no `LocalQueue` named `default`. We recommend that GMI document that MNNVL jobs need to skip the queue, or, much better, that GMI set up its Kueue to tolerate DRA resource claims.

GMI’s dashboards, which did not exist last time we tested, were useful, including hardware information from DCGM and NVLink and InfiniBand monitoring.

We could not get the full GMI health-check experience due to a lack of spare capacity, but what we could test performed well. Nodes were quickly cordoned at the Kubernetes layer, then rebooted and auto-uncordoned. The Slurm layer also brought a node into `DRAIN` soon after a synthetic error was injected. GMI had alerts set up in Slack, so we could see their system detecting the error and thread to discuss next steps with support, which was convenient.

Our GMI experience was solid in every crucial way, and we look forward to continuing to test them as their capacity ramps quickly.

Bronze

Amazon Web Services (AWS)

Our last report summarized the AWS experience as a “headache.” This time around, we had to break out some Amazon Basic Care Ibuprofen Tablets, Fever Reducer and Pain Relief from Body Aches, Headache, Arthritis Pain and More, Brown 200 Count to get through our testing.

The difficulty with AWS begins before you get on the cluster—indeed, before you even try to provision the cluster: it begins when you contemplate AWS’s zoo of managed offerings. To get this all straight, we had to make a Venn diagram:

Source: Intern (I forget his name)

(Caveat lector: anyone not interested in the minutiae of AWS’s GPU offerings should skip this ¶.) SageMaker is AWS’s fully managed platform meant for minimal maintenance. If you don’t want to handle any infrastructure, it offers serverless model customization and data science environments that abstract away the underlying compute. At the next level up is the SageMaker HyperPod line, which supports both Slurm and EKS. The HyperPod line offers many conveniences like health checks, autoremediation, dashboards, and workload-level features like checkpointless training out of the box for users who want to preserve their `ssh` or `kubectl` access. If you want infrastructure abstracted away but want to handle the ML stack yourself, you can leave the SageMaker umbrella and choose EKS Auto Mode, which takes care of node provisioning, autoscaling, OS patching, and node repair. EKS Auto Mode is a wrapper on top of “managed EC2 instances,” but these are not actual EC2 instances in the sense that they do not show up in the EC2 console and do not provide the same level of customizability as ordinary EC2 instances. In EKS Auto Mode, users do not have access to the control plane, and they must allow AWS to recycle their workers every 21 days for security. Finally, for even greater customizability, there is EKS + Karpenter, which allows the user to manage autoscaling themselves with the open-source tool Karpenter.

Despite Amazon’s reputation as “customer-obessed,” it is clear that this novel-length GPU menu is an example of the adage “you ship your org chart” rather than a reflection of genuine demand segmentation. And the bureaucratic frustrations hardly end once you’ve ordered your cluster. This is a criticism of the organization rather than of any team in particular—AWS has shades of the 3rd-Century Roman Empire. To improve its standing in the next round of ClusterMAX testing, we recommend AWS choose another empire and/or time period for reference.

Nevertheless, there is a unifying trait of all these offerings: they fail to provision the first time you try. Thus, we made another clarifying Venn diagram:

Source: Ibid.

This finally brings us to the AWS operator experience. One SageMaker cluster we provisioned in the console failed due to insufficient IAM permissions. Some of the EKS clusters used Terraform to simplify provisioning, but that path is not battle-tested and requires being in the right subdirectory and on the right commit of the right branch of the right (in-flight) repo and knowing exactly which knobs to turn without making an invalid request. One cluster came up without Lustre and would not let us attach any, despite extended back-and-forth with AWS engineers. For the purposes of testing, we provisioned a new cluster in a new region to gather data on filesystem perf, but if we were a real-life lab, this would have been crippling, at least temporarily. The AWS team we worked with was consistently helpful, but it is clear that they don’t have the liberty to go solve problems or even insight into how the system fits together, especially compared to the flat, nimble neoclouds we’re judging them against. It should not require six engineers on multiple calls to provision a Slurm cluster in the first place. (Edit: AWS would like to clarify that it was five of their engineers and one of our engineers to make six.)

Once the goods were delivered, there were no blockers that kept us from getting reasonable perf out of the gate. The drivers, firmware, and most utilities were up to date. This is the bare minimum from a reputable provider like AWS, however, and there was no shortage of rough edges elsewhere.

A basic point of distinction between the AWS clusters we tested and their peer B200 clusters is that AWS uses Elastic Fabric Adapter (EFA), its proprietary scale-out fabric. In our testing, peak EFA throughput was strong. On a four-node 32-GPU all-to-all NCCL test over EFA on B200 we measured 59.44 GB/s busbw at 16 GiB message size, which works out to about 92% of scale-out line rate given the proportion of traffic that crosses the node boundary. On other networking tests with other traffic patterns and launching mechanisms, as long as message size was large enough, throughput was among the best B200 clusters. However, EFA consistently demonstrated worse latency than well tuned InfiniBand comps. For example, when the message size was reduced to 64 KiB, the same four-node 32-GPU all-to-all NCCL test took 85.26µs over EFA and under 50µs on InfiniBand peers; a delta of 30-45% between EFA and InfiniBand was consistent across small-message tests, which are latency-dominated. Although this may seem like a nitpick, scale-out latency is a critical metric for modern expert-parallel workloads, which frequently send small token vectors across the node boundary. This 64 KiB NCCL test, for example, is similar to the expert routing step of models like Kimi K3 and DeepSeek V4. Our findings are corroborated by an excellent Perplexity technical blog post about optimizing EFA, which reported similar results: EFA demonstrates satisfactory peak throughput but pays a latency penalty of 20µs or so on “the message sizes exchanged during MoE dispatch and combine.” (Note that this Perplexity post compares EFA to ConnectX-7 NICs whereas we tested ConnectX-8, but the nameplate throughput of the two generations is identical.) It’s not a good sign when Perplexity’s cracked engineers are writing painstakingly detailed blog posts about how to “enable” basic functionality on your custom network stack.

All-reduce tests on EFA. Note that message size must double many times before time to completion changes much; this is because at small sizes, it is latency, rather than line rate, that defines performance..

EFA still requires extra setup for expert-parallel inference in upstream vLLM and SGLang. AWS reports successful vLLM deployments, but users must install EFA userspace libraries and build the relevant communication components with EFA support. Depending on the workload, these include DeepEP, NIXL, or Mooncake Store. AWS’s DeepEP fork supports EFA through NCCL GIN. DeepEP V2 uses NCCL GIN, while its NVSHMEM-based V1 path is now documented as legacy. Upstream integration and default container packaging still require work. NCCL EP is another route under development; the linked vLLM integration remains a draft PR. In general, EFA is not a priority for frontier communications libraries, and users must wait weeks or months for what support they do get.

AWS is the only cloud distinguished in this way. Every other provider we tested uses InfiniBand or RoCE. Of course, AWS’s customers with GW deals have the resources to make sure their EFA is well tuned, but we still think it a shame that the largest provider’s scale-out stack is uniquely troublesome.

In a different way, storage was another basic cluster service that often proved uncooperative. The B200 cluster we provisioned in `ap-south-1a` failed every request for storage with an ambiguous `Insufficient capacity` error, and the `Failed` filesystems took 5h46m55s to delete.

To test Lustre performance required another cluster in another region, which meant another protracted bringup—another evening of terminal-watching trying to discern whether Terraform was making progress or not.

Source: Provisioning purgatory.

Claiming storage on AWS via Terraform is like raising a toddler: you can only watch and hope that once it’s done destroying it will start creating. And once the storage is created, it is not necessarily straightforward to access. On the Kubernetes layer, one cluster provided no StorageClass by default, and its only provided StorageClass left our PVC indefinitely in `Pending` because it had the wrong driver. This required us to define our own StorageClass with the correct provisioner.

Powered on and wired up, storage perf was mixed. The provided FSx stack struggled significantly with ordinary buffered I/O, reaching 9.2s in p99 sequential read latency on fio testing, compared to 312ms on the same metric with buffering disabled. Greater latency on the buffered setting is expected, but this is a pathological delta.

AWS’s health checks’ performance was, again, mixed. Our EKS Auto Mode cluster caught the XID we injected, set `AcceleratedHardwareReady=False`, and gave us a new and healthy node only 24m after the original injection. This is good. Meanwhile, our SageMaker HyperPod Slurm cluster immediately detected the XID, brought the node into `DRAIN`, and completed most of the reprovisioning process before tripping over its own feet when it required a deep health check before it could rejoin the fleet. This deep health check was submitted as a Slurm job and couldn’t be scheduled until, well, the node was added back to the fleet. Thus, an operator had to manually push through a `scontrol ... State=RESUME` command, at which point the fleet returned to full health. With our feedback, AWS promptly fixed this circular dependency, and a follow-up test on a single H100 node succeeded without intervention. Better.

Our EKS Karpenter cluster’s health checks were even more self-defeating. The `dcgm-server` DaemonSet did not tolerate the `nvidia.com/gpu:NoSchedule` taint, so it did not attach to the worker node, causing the monitoring agent to report `error connecting to nv-hostengine` and flip `AcceleratedHardwareReady=False` with `Reason=DCGMError`. Karpenter then marked the `NodeClaim` for termination, the replacement node was marked unhealthy for the same reason, and the system spun in circles until we manually modified the DaemonSet to tolerate the GPU taint. The health check had another allergic reaction when it noticed the IMEX daemon did not schedule. IMEX configures the NVLink fabric across nodes, but we had ordinary B200 nodes speaking via EFA, so there was nothing for IMEX to do: this daemon’s failure to land should have been harmless. Instead, the monitoring log emitted the schizophrenic message

ignoring IMEX health code on non-NVLink multi-node system, code=122

sending condition to exporter,

Reason=NvidiaFabricError,

Severity=Fatal

This would again have recycled a healthy node if we did not manually pin
NodeRepair=false to keep overzealous garbage collection from blocking our work. (Edit: AWS informs us that this was not overzealous garbage collection; rather, it was a genuine hardware error with an unclear error message. This is worse. In any case, an `nvidia-smi -q` of ours reported “Fabric Health Summary: Healthy” after this supposed failure.) Finally, as our testing came to a close, Karpenter tore down our B200 cluster 39 minutes before the EndDate, force-evicting a 4-node MPIJob in progress. We have gone into such detail here because these bugs are almost unbelievably elementary for the world’s largest cloud. Perhaps the single painfree part of the AWS lifecycle was Karpenter scale-up and scale-down, which completed in minutes with no errors. However, GPU capacity must be secured in large blocks and far in advance, so snappy autoscaling with Karpenter is a meaningless feature.

Despite the many problems with their product, AWS’s specialists and engineers were intelligent and responsive in our interactions with them. They were eager to show what they had improved, sought out feedback on documentation, circled back to rectify errors in testing, and even built a nice, forward-looking MCP server whose skills we plundered for our own codebase. AWS’s issues lie in its organizational structure. The number of basic issues that have slipped through the cracks is astonishing in an organization with AWS’s resources.

In February, AWS announced the expansion of its deal with OpenAI by $100B; in 2Q26, it posted 37% growth in net sales YoY; its Trainium business has hit $25B in ARR with triple-digit YoY growth; its Anthropic investment is appreciating to the tune of tens of billions per year; Bedrock continues to make bed-rocking margins on bed-rocking income.

Managed clusters are not AWS’s priority.

Gcore

Gcore is a Luxembourg-ish neocloud building out in ambiguous region called « Europe, » where it relies on other providers for capacity. They have made no flashy announcements recently, though they are adding Blackwell to their Hopper fleet. In ClusterMAX 2.0, we gave Gcore a Silver rating after testing their Kubernetes and SonK offering via Soperator. This time round, we only tried the SonK again.

Our insight into Gcore’s capabilities was limited by their lack of capacity: most providers granted us at least 4 B300 nodes for at least a week, but Gcore was so constrained that they could only part ways with 3 H200s. Nevertheless, there was plenty to keep us on our toes on this cluster.

Our testing on the Soperator layer required several rounds of back-and-forth with the Gcore team to fix some rough edges, but the team consistently proved helpful. When the majority of our first sweep failed due to a misconfiguration that prevented non-root users from running their workloads, the team responded promptly, confirming `kernel.apparmor_restrict_unprivileged_userns=0` would get us unblocked. Gcore’s team was again responsive in correcting another poor default when each Slurm worker reported `RealMemory=2048 MiB`, apparently having erroneously inherited CPU configurations. This kept us from even running `nvidia-smi`, and we were glad to see it quickly repaired.

Debugging was made slightly more challenging on the Soperator layer because inter-node SSH was disabled by default, but the team pointed us to a `soperator-createuser` convenience script to make life easier. The team also troubleshot an error with their dashboard as we tested.

On the Kubernetes layer, there was no default StorageClass, so users had to select NFS explicitly. The storage on the Slurm layer was properly configured, with the easy-to-use Soperator default of NFS at `/`, providing the illusion of a single shared volume wired to all the nodes. This storage performed well across workloads. The N/S network was fine, and the E/W network ultimately performed up to spec, but the RDMA devices had non-standard names and no `topology.conf`, so configuration took extra steps.

The XID that we injected was quickly caught and displayed in Grafana; interestingly, the node went `Down` on the Slurm layer, but on the Kubernetes layer nothing changed, even though it was injected via K8s. We could not test autoremediation because Gcore had no spare capacity.

We were impressed by Gcore’s attentiveness and their Soperator offering that, after some massaging, did the basics right. We look forward to testing Gcore’s B300/GB300 management, including its handling of XDR ConnectX-8 NICs, and seeing the improvements of their Soperator and Kubernetes products.

GMO

GMO’s managed Slurm cluster delivered strong networking and storage performance on two B300 nodes. Its main gaps were restricted profiling, limited enterprise controls, and a lack of autoremediation.

Slurm arrived with a configured head node, partitions, and topology.conf. Drivers and CUDA/NCCL/HPC-X modules were current and consistent across nodes, and Pyxis was available. Onboarding was straightforward, though the Japanese-only console required some manual ctrl+c, ctrl+v translation for us to understand.

Source: PoV you are a SemiAnalysis intern who dropped Japanese after two classes trying to locate a Grafana dashboard

Grafana access took us five days of troubleshooting. We initially missed the macOS login-keychain instruction, then entered the certificate password where the system password was required. GMO investigated and repeatedly followed up. The final mistake was ours, but clearer certificate instructions would have saved time. Or, no unnecessary custom certificates to access the website. Japanese security theater strikes again.

Source: Attempting to access our Grafana dashboard

The sixteen 400 Gb/s rails per node delivered strong NCCL results, though two nodes provide less scaling evidence than our usual four-node minimum. MPI launches needed explicit NCCL_IB_HCA settings, which GMO said it already supplied in nccl.conf, but which didn’t work for us initially. Burn-in showed no network errors, retries, or ECC errors, with normal temperatures and power.

Shared home/data/work directories worked out of the box. FIO sequential reads and writes, random reads, and torch.save were solid, built on a well-tuned DDN Lustre filesystem. However, our contrived tests such as `import torch` and vLLM Serve from shared storage performed poorly; GMO offered to investigate but we chalked it up to classic Lustre metadata performance tuning issues and moved on. Like many other Lustre implementations, the GMO storage offering lacked snapshots, automated backups, and disaster recovery options.

The cluster also lacked `sudo`, Docker on the workers, and required the use of snodes, a convenience script that they built to show metadata about the cluster, since they prevent users from running any scontrol commands on their own cluster. Security theater once again.

As a final step, SSH to worker nodes on the cluster was also blocked. And inside Slurm jobs, `RmProfilingAdminOnly=1` blocked Nsight Compute’s GPU counters, and perf was not usable for our profiling needs. We believe that managed clusters should allow users to profile their own workloads if they choose to go for the yolo-mode option. We wanted it, but GMO could not provide.

Finally, Grafana exposed useful DCGM metrics but lacked XID detection, estimated TFLOPS, SM Active, and SM Occupied views that we enjoy for performance and reliability monitoring. The Slurm job summary view was also pretty basic. Our synthetic XID 79 log injection produced no observed alert or node drain. GMO said it runs dcgmi health in the Slurm prolog, confirmed XID 79 was not whitelisted, and agreed to investigate for future. We finished our testing by determining that GMO lacked any autoremediation system in place, or at least we couldn’t verify it.

Source: Our Grafana Dashboard on GMO!

Overall, GMO’s lacking Kubernetes and enterprise features such as RBAC, SSO, and customer-accessible audit logs put it in the camp of niche, local player serving the Japanese market effectively. Our support experience with their team was helpful, but followed standard Japanese business hours. Going forward, we expect GMO to continue to be a leader in Japan, but lack the ambition required to expand their business beyond their local geography.

Verda

Helsinki-based Verda (formerly DataCrunch) now has $450M in total funding as of September 2026. They maintain a close relationship with Nvidia and continue to aggressively pursue financing. On the technical side, they have implemented a number of features since we tested them in November, and we are optimistic about their engineering team and roadmap.

Last round of testing, Verda’s Slurm was officially in beta, and they did not offer Kubernetes. This time, they offer a managed Kubernetes and have moved their Slurm offering to Slurm-on-Kubernetes to simplify deployment. We tested their cluster on both the Slurm and K8s layers; it had its fair share of rough edges, but represents significant progress and a strong starting point for continued development. We also found that security was solid.

The purchase and onboarding process was smooth, and the cluster provisioned without hiccups in 30m. The Grafana dashboard worked out of the box, and we SSHed onto the cluster without incident.

Once we got on the cluster, though, there were some bumps in the road. Our Slurm testing began with an interesting footgun: when we SSHed in, we assumed we were on the Slinky layer, but we actually had landed on the outer VM with wrappers that ran `salloc`, `sinfo`, and similar commands via `kubectl exec` into the `LoginSet`. The behavior was unusual—all Slurm commands escalated us to `root`, for example, and our Python environment disappeared—and it took an evening of debugging before we puzzled out that our Slurm access was illusory and a different SSH command would be required to access the Slinky layer properly.

Verda’s engineers were responsive, quickly hopping on a support call in which they reasonably explained the rationale behind these convenience functions and pointed us to the supported path. Although this particular feature caused us a headache, it’s bullish when engineers have gone out of their way to try to make their cluster easier to use.

The Slinky layer proved serviceable but imperfect. Verda ran a NCCL-test prolog that exceeded Slurm’s 10-second `MessageTimeout`, causing a worrisome though harmless `Prolog hung` alert every allocation. More significantly, Verda did not support Pyxis or Enroot, so we had to stand up our environment via a massive `.sif` file. Once our workaround was finished, we completed an 8-hour burn-in without incident and measured satisfactory numbers on our microbenchmarks.

Meanwhile, the Kubernetes layer could only be steered through the bastion, offering no way to download a `kubeconfig` through the console. This is a relatively simple fix to align Verda’s offering with users’ standard workflows.

The biggest shortcomings that surfaced during our performance sweep were of the shared filesystem. Gross throughput numbers were excellent, but cross-node file locks seemed simply not to be enforced. We found 100/100 exclusive-lock violations for both `flock()` and `fcntl()` across nodes, while same-node controls refused every attempt. This strongly suggests the issue is with `virtiofs` failing to publicize locking behavior. We also saw corrupted data when downloading with multiple clients from Hugging Face and `ENOENT` failures on Elbencho with 32 clients. This is more low-hanging fruit, and crucial given the fundamental importance of a reliable shared filesystem when running Slinky.

Lastly, Verda’s health checks programs are not yet fully configured. They are literal “health checks” in the sense that they evaluate whether a node is healthy, but there is nothing to consume their output—they exit `1`—“yup, this one’s dead”—and move on. A user must therefore bring their own reliability suite if they don’t want bad nodes silently killing their workflows. There is of course also no autoremediation provided by Verda.

Verda consistently communicated promptly and clearly and seemed to have engineers with the liberty and motivation to improve their product. On September 10, after our testing completed, they pushed a number of changes to Instant Clusters that appear meaningful. We look forward to seeing them continue to sand the edges of their Slinky and Kubernetes offerings, tune their storage system, and implement robust health checks and autoremediation.

Moonlite

Moonlite, a neocloud headquartered in Chicago, is a new inclusion as of ClusterMAX 3.0. With an ex-Crusoe founding team, Moonlite currently operates a Hopper fleet and has recently handed over its first B300s and GB300s, which it will be continuing to ramp up in the coming months. According to LinkedIn, Moonlite has the fewest employees of any company we worked with this round and oriented around engineering.

We tested Moonlite’s Kubernetes and Slinky. The handover process was smooth, with a nice onboarding document defining validated recipes and a Kubeconfig that gave us everything we needed. By default, everyone shared the same root user on Slinky; when we asked for non-root users, Moonlite made a `moonlite-adduser` convenience script in half an hour and baked it into the login-pod image same-day. This shows Moonlite’s responsiveness to feedback, but also their greenness. We loved their response. We would rather have a battle-tested system, ideally through the console, to handle basics like this.

On the hardware side, everything we tested worked well out of the box. NCCL tests hit ConnectX-7 specs, NFS was good, compute benchmarks were solid, and burn-in completed without a problem. However, it bears repeating that we were testing H100s, which by now are quite familiar to the industry and are therefore easier to get right than the GB300 NVL72s we had with other providers. Nonetheless, we were happy with our Hopper perf.

One component of the system that we could not test was the NVMe, which existed but could not be accessed. Nothing exposes it: no local-path `StorageClass`, `hostPath` unavailable, and Slurm workers are Slinky pods with only VAST `/home`. On the Kubernetes layer, there was no default `StorageClass` at all, and unqualified PVCs hung `Pending`.

We tested Moonlite’s health and monitoring system via a PCIe secondary bus reset, and Moonlite alerted us in 11 minutes. With our permission, they cordoned the node, ran diags, and returned it to our fleet after 57 minutes total, including a few minutes where they waited for our instruction on how to proceed. In practice, with actual customers, they define policies ahead of time so that their team can follow a runbook depending on the particular error. Moonlite’s typical SLA is to respond within 15 minutes as their team receives alerts in Slack and PagerDuty. We hear from their customers that so far that this is exactly what they do.

In our view, any autoremediation process that brings us back to a full fleet within an hour is good enough. But while manual processes are fine, at scale, we want to see this automated. We would also prefer greater visibility on the cluster in the event of an error. The only change visible to the user was that the node, while remaining `Ready` on the Kubernetes layer, advertised 7 GPUs instead of 8 in the time between the reset and Moonlite’s manual intervention.

The Slinky layer lacked any container runtime on the workers but otherwise got the job done, and had most of what we ask for on the login pod. Moonlite’s Grafana was fine, but it lacked detail at the scheduling layer, and does not display the most crucial information, which is error status.

What Moonlite attempted to do, it did well. We look forward to testing them again as their products mature and they get some GB300 calluses on their hands.

Together

In ClusterMAX 2.0, we wrote the following about Together:

Together is a strong provider with a robust cluster offering for both Slurm and kubernetes, but it is held back from the gold category due to reliability issues.

This round, Together’s issues—generally but not exclusively having to do with reliability—are cause for downgrading it to Bronze.

Two nodes of ours failed during a compute test; one came back, but the other was sent to RMA without notice, so we were left wondering where it went. Later in testing, Together’s Slurm `HealthCheckProgram` tried to run `dcgmi diag -r 1` under a 45s timeout on nodes that were already scheduled; the diag did not land, the timeout was misread as a failure, and the nodes were drained mid-run. Together quickly fixed this with our feedback. Separately, one of our nodes could not access VAST because its storage NICs were in `operstate=down`, which Together explained was due to a race condition in NIC initialization in provisioning their new lightweight VM stack. There was nothing to catch this error, and manual intervention was required to reprovision a healthy node and bring our test fleet back up to 4. While we like the sound of “reduced cluster provisioning time,” we obviously don’t like to see it affecting reliability.

Together’s health checks did work correctly in some instances, showing progress since ClusterMAX 2.0. When we triggered a PCIe secondary-bus reset on the Slurm cluster, the node was drained in 67s, and an automatic node replacement completed in 35m7s without manual intervention. A synthetic error also cordoned a node on the Kubernetes layer and correctly did not trigger autoremediation because we had `Auto Repair` turned off.

The Together Kubernetes and Slinky layers were both functional, though with a handful of challenges. Interestingly, Together advises to avoid using the Kubernetes layer of the Slinky cluster: to convert Slinky nodes to K8s, we were instructed to tear them down and reprovision as vanilla Kubernetes. The Slinky offering should therefore be evaluated as plain Slurm. Summarizing its pain points as briefly as possible: OCI image extraction failed because of default settings on `/tmp`, so we had to redirect extraction to `/scratch`; when this happened, Slurm exited `0` anyway; NFS rejected any filename containing `:`; and non-login worker shells have neither `nvcc` nor `mpirun` on `PATH`, so ordinary batch scripts behave differently from interactive setups and lack the tooling that they need. Together also offered S3, but it was slow and unreliable. None of these is fatal, but adding them up, along with other more minor errors, meant a few dev days before we were able to run our stuff unimpeded.

A public IP check confirmed that this flaky B300 cluster had org `AS53735 IREN`. This N+0 Prince George datacenter is one of two datacenters in the industry notorious for unreliability, along with Crusoe’s cursed Reykjavik site. We strongly suspect that one or both of these datacenters was not inaugurated with a Land Acknowledgement. We know of multiple exceptionally unhappy customers of Together struggling with link flaps, power issues, cluster access being shutoff, random upgrades that aren’t approved by the users going wrong and not being rolled back, and weekend-long outages that eventually get blamed on a single ISP (yes, a single ISP at this site—no redundancy in Northern Canada). Maybe their other sites are better but the TogetherAI site at IREN Prince George is completely bad.

After playing the blame game for some time, it is clear that reliability issues cannot be blamed entirely on underlying providers. Unforced errors from customer support engineers have created a lack of trust between customer and provider.

Of course, managed clusters are not Together’s bread and butter, nor are they its primary source of revenue. When, in July, they announced an $800M Series C, the post made passing reference (at best) to their managed clusters business, emphasizing, on the contrary, that “Together AI is a research-driven company.” The offtakers of their 250MW in Saudi are as yet undisclosed, but we are aware that Together would prefer to get paid per token, not per GPU-hour. Despite this, we hope that Together manages to Get It Together, and address its cluster reliability issues to make its Slurm and Kubernetes products something to be proud of.

Crusoe

The Crusoe section of ClusterMAX 2.0 ended on a cautious note:

Crusoe is at risk of being downgraded to ClusterMAX Silver due to many of their top individual contributor engineers quitting, leaving the culture in their cloud division beginning to resemble big tech. There are too many middle managers across the organization, especially in engineering. This has caused incredibly slow moving releases, such as their AutoClusters feature, leaving us with concerns about the future of Crusoe’s public cloud offerings. Chase needs to do a rapid course correction if he doesn’t want to lose all of his 10x engineers and eventually lose their Neocloud business with it.

In the eight months and change since we published that paragraph, Crusoe has had its fair share of wins. Its valuation has grown from around $10B to $30.9B, it has announced 900MW and 1.0GW campuses in Texas and many more elsewhere, and it remains an innovator in the wild wild west of datacenter construction. They have the largest real pipeline of any datacenter builder.

Its inference endpoints business is also extremely highly regarded, following the acquisition of some cracked Israeli engineers from Atero. They have also landed awesome deals with firms like Jane Street.

Plenty to keep their hands full and business churning, some clusters are amazing, but others from them are terrible.

Yet this round of ClusterMAX testing finds that Crusoe’s neocloud business has not course-corrected rapidly enough. The problems go back to the fundamentals: reliability and hygiene.

We emphasize health checks in our expectations and our write-ups, but genuine hardware failures are rare in testing because our ClusterMAX allocations are too small and short-lived. Crusoe bucks this trend. On an H100 cluster we had in May to test AutoClusters, we saw an impressive barrage of XIDs—more real errors on one cluster over five days than we saw in the rest of testing combined.

In a pattern that would repeat throughout testing, exceptionally poor performance in some areas was paired with mature services in others: this storm of XIDs was accompanied by convenient emails, attractively styled, politely notifying us of the errors.

In addition to being responsible for the majority of this round’s hardware failures, Crusoe bears the distinction of having had the oldest Linux kernel and the oldest GPU drivers we tested. Crusoe has since rectified these shortcomings, but at the time of testing there were several software versions that were well below our minimum expectations: the aforementioned NVIDIA drivers failed our `cmax audit security`, along with Docker, ConnectX firmware, and `runc`. These issues speak to a lax security posture and poor engineering hygiene.

Nor are these software-version criticisms merely academic. The Linux kernel that our B300 cluster ran, `5.15.0-185-generic`, was earlier than a Linux patch series called “writeback: Avoid lockups when switching inodes.” The particular problem that our kernel experienced is well explained by the description of commit `e1b849c`, which fixed it for upstream kernels:
```
There can be multiple inode switch works that are trying to switch inodes to / from the same wb. This can happen in particular if some cgroup exits which owns many (thousands) inodes and we need to switch them all. In this case several inode_switch_wbs_work_fn() instances will be just spinning on the same wb->list_lock while only one of them makes forward progress. This wastes CPU cycles and quickly leads to softlockup reports and unusable system.

```

When might a `cgroup` that owns thousands of `inode`s exit? One answer: when a Slurm job is torn down. In our case, as an ordinary workload wrapped up, every CPU on the NUMA node ended up stuck contending for the same `list_lock`, making the processor unusable. The local NFS client, which happened to be on the frozen NUMA node, couldn’t get CPU time, and a liveness probe for that client timed out and reported the node unhealthy. Kubelet tried to control the damage with a `SIGTERM` and a `StopContainer` operation, but the kernel was taking its sweet time doing the laundry—reassigning dirty `inodes`—and took over two hours before it could honor the termination requests. The final result was another `NotReady` email from Crusoe—this one, uniquely, downstream of an old CPU kernel.

By the end of our testing, this kernel was not merely old and unperformant: it was officially marked a security vulnerability by CVE-2026-64378, which exploited the buggy writeback behavior, published on July 25 with severity 7.8; a separate 7.8 for this kernel was posted on July 27. Thus, we added the Linux kernel to our collection of Crusoe software exposed to CVEs. Crusoe has informed us that they will be upgrading to a 6-major Linux kernel, a necessary patch.

Source: A Crusoe job posting on August 13, 2026

When we ran our XID injection on Crusoe’s Kubernetes layer, the error was quickly detected and a `Node Replacement` workflow was triggered, but the node could not be put in `drain` due to a mismatch in Slurm naming convention that slipped through Crusoe’s CI. The health monitoring system spun its wheels, repeatedly catching the error and trying but failing to mark the node unhealthy. Crusoe’s engineering team diagnosed the error impressively quickly, rolling out a replacement within hours. However, our replacement node had faulty storage NICs, immediately bringing it into `drain`. This issue persisted after the node’s VM was manually reset. Crusoe’s engineering team has informed us that the failure of AutoClusters to remediate in this latter event was “a known gap,” and that they will be “rolling out auto-remediation actions for this in a few weeks.” This is a surprisingly lethargic response to a basic failure of reliability.

Last and briefly, it bears mentioning that this cluster’s WAN was among the slowest we tested, clocking in between 0.47 Gbps and 0.78 Gbps depending on the context. Iceland has some skinny cables getting off the island.

All these problems add up to a cluster notorious for its flakiness, as confirmed by our friends at frontier labs who complain about both the number of failures and the apparent lack of interest from Crusoe engineers to fix the issues promptly. To be clear, we at one point viewed Crusoe as the #2 neocloud in the world, and as a result the #2 place in the world to rent GPUs. We made recommendations to this effect. Oh, how the mighty have fallen.

But while Crusoe’s machine images are badly due for a refresh, and their engineers can’t get Claude Tags or Codex approved, they do have automations like this set up in their internal Slack:

Does a CVE from November of 2025 count as “historical baggage”?

In all seriousness, this is a problem of priorities, a problem of bureaucracy run amok. We are not sure what to make of job postings like the above one for a senior Linux kernel engineer. Literally, we’re just asking everyone to keep things up to date. You don’t need to hire Linus to understand why this makes sense. So while it’s cool to see Crusoe flexing its muscles and paying top dollar for a good person, it’s not a lack of people or funding or technical skill that has brought Crusoe here in the first place. Our complaints focus on the most basic issues possible, not the Slinky configuration or console or CLI that no doubt take up much more engineering time and that we enjoy using.

We enjoy working with Crusoe’s engineering team and believe they have plenty of talent to improve their standing. Frankly, in the last few months, things seem to be getting back on track. Engineers are locked in, pushing code, and excited about upcoming releases. The Atero acquisition has been worth its weight in gold, skyrocketing Crusoe to the top of many buyer’s lists when it comes to inference endpoint quality. We love to see it.

Still, in the managed clusters business, TCO comes down to goodput, and goodput comes down to reliability, and Crusoe’s clusters are exceptionally unreliable.

Crusoe comes in towards the bottom of our Bronze category for this round.

Prime Intellect

Prime Intellect closed a $130M Series A this July at a $1B valuation, disclosing $100M of “annualized revenue run rate across compute, RL and post-training, sandboxes, inference, environments, and evaluations.”

We tested Slurm and Kubernetes B300 clusters and mostly enjoyed both of Prime Intellect’s orchestration layers, although each was missing a handful of packages we would’ve liked to have. The huge Slurm control plane was nice and snappy, having machines with 2TB of RAM for some reason. Plenty of swap space to run salloc.

After a quick fix to standardize the NCCL bootstrap interface, our networking tests ran up to 13 nodes on the ConnectX-8 NICs with excellent throughput. Storage was no issue, with Weka /data volume performing very well at 4 nodes and scaling steadily beyond. PrimeIntellect’s managed Grafana was one of the nicer ones we encountered. However, we were on a cluster that was still going through burn-in, so we saw a steady stream of errors. An easy way to test the monitoring and health checks on the system!

First, a node on our Kubernetes cluster was marked NotReady, apparently due to on-site techs monkeying with adjacent servers; it returned within 5 minutes. Then, during a PyTorch networking test, a different Kubernetes node stopped posting kubelet status and the distributed job could not complete; the node came back after a reboot and the job completed. The next day, 2 Slurm nodes went down, one with NVLink issues and the other stuck after a reboot. Both returned promptly after intervention, but the latter was soon back to its shenanigans with another surprise reboot, causing a storage test to fail. An MPI test aborted in its last phase, and the corresponding Slurm job remained in COMPLETING as all workers remained unreachable in the brief period before we handed the cluster back.

All of these reliability issues can be blamed on the fact that we got on this cluster in a small window while they were still doing burn-in for a paying customer. Fair enough. But the lack of compute at Prime’s disposal that would allow them hold back hot spares points to a bigger problem of growing capacity constraints. Prime is a reseller of other’s compute, and depends on their underlying providers to guarantee customers with a high quality experience. This comes with challenges that high quality training/RL frameworks, dashboards, and maniacal Slack engagement from a cracked team of engineers can’t always cover up.

With that said, Prime’s health checks performed impressively throughout, quickly intervening and limiting the damage of each error. We thought the system’s default responses were sensible and even liked the Slack alerts, which the team promptly followed up on. Given that Prime gets its capacity from a number of providers, the quality of their software layer is important for labs to feel comfortable with what they’re getting. Of the marketplaces we’ve tested, Prime seems to be the best.

This leaves Prime Intellect with a straightforward assignment, and one that we cannot have direct insight into: improve reliability, build a base-load of compute capacity, and rise in the rankings as a true neocloud. We have no doubts about the team’s technical ability; they have done a lot of hard things correctly on this cluster, and are the only provider on this list that injects real experience training models into their approach to software and support. For better or worse, their work elsewhere is much more impressive than the bread-and-butter managed compute offering evaluated here.

We got a preview of their hosted training product during this period and are very impressed. RL infrastructure is decidedly more complex than a simple managed Slurm or Kubernetes cluster, since its built on top of the same foundation but with three components that need to stay in sync:

  1. Training

  2. Inference

  3. Sandboxes/environments

Coming soon, we will be evaluating hosted training infrastructure providers and RLaaS companies for hire through a project provisionally called PostTrainingX. If we had to put out an initial cut at those rankings, Prime would be in the hunt for Platinum! Stay tuned.

DigitalOcean

This marks the first round of testing that we have been able to get the full DigitalOcean developer experience. Marketing itself as “the first cloud built end-to-end for the inference and agentic era,” DigitalOcean distinguishes itself from “bare-metal focused neo-clouds and inference wrappers that lack cloud platforms.” In fewer words: DigitalOcean sells, among other things, managed clusters.

DigitalOcean’s managed clusters are orchestrated by Kubernetes, which proved solid but slightly immature. Their onboarding doc asked us to install Multus ourselves, then create network attachments, install MPI Operator, and create the MPIJob manifest. There’s no reason not to set this up for the customer in advance. Once configured, our cluster performed well, hitting all the standard benchmarks for compute and XDR networking, and completing burn-in without incident. The cluster also lacked a RWX StorageClass out of the box, adding another basic setup step that the better K8s providers handle for you. Once the storage was configured, its performance was interesting: sequential writes were approximately 6x faster than sequential reads at several different client counts, the inverse of what we see on most clusters. We expect that DigitalOcean’s sequential read performance could be significantly improved with retuning.

When we went to test the health checks, our first XID injection went unnoticed: the monitoring system did not detect anything, and the node remained available for scheduling. When we alerted the DigitalOcean team, they responded promptly, found the bug in their system, and asked us to reperform the test later in the week. The second try went smoothly, and we had a fresh node 50 minutes after injection.

When a neocloud is “built end-to-end for [...] inference,” that means it doesn’t offer Slurm. The DigitalOcean team helpfully offered us a runbook to stand up Slinky, and it got the basics right, but would require significant work before it could be considered production-ready. No SSH access, nor, as mentioned previously, any RWX volume by default, which means that there is nowhere to keep a `/shared`. Anyone who plans to run Slurm on DigitalOcean should expect to stand up their orchestration themselves.

DigitalOcean will be adding 60 MW of capacity throughout 2027, and, aside from its managed K8s, it has an interesting selection of GPU services like Droplets and bare-metal Hopper and MI300X that fall outside the scope of ClusterMAX testing. But no plans for GB’s or VR’s from what we can see at this point. We look forward to the team continuing to improve its managed Kubernetes service the next time we test.

Hyperstack

Hyperstack, under parent NexGen Cloud, is a UK-based neocloud that advertises Kubernetes clusters in the US, Canada, and Norway. It recently raised $45M at a $354M valuation with plans for many AI new products for “full-cycle development.” As it announced a $34M debt facility to build out a B200 fleet, it also disclosed plans to deploy 4,500 B300s later this year and for an additional 56 MW online in 2027.

The managed Kubernetes cluster that we tested marks significant progress since ClusterMAX 2.0, and a solid foundation for its continued growth. Compute and networking tests were up to spec, and burn-in sustained good numbers with no hard errors. Our Kubernetes node-reboot recovery was quick, and we had no trouble mounting and unmounting PersistentVolumes. The N/S network was fine, and storage was good enough, although we were not able to get Weka and VAST, which Hyperstack usually offers, on our timeline.

Our synthetic XID was detected in under a second with a `NvidiaFatalXid=True` flag. The test was awkward as Hyperstack had no extra capacity, so we had to take a node out of our fleet to simulate a spare. The support team notified us that they had detected the error 12m after we injected it, and our spare rejoined the fleet after a total of 52m. Interestingly, at this point, the “sick” node was still not cordoned, which took until 1h30m to register on the Kubernetes layer. This was all driven by manual intervention by Hyperstack’s SRE team after our synthetic error triggered Grafana alerts.

Overall, this test was a success, and we appreciate Hyperstack’s commitment to helping it go smoothly. However, as the team is aware, the amount of human support required makes us worry about its replicability. Hyperstack offers 24/7 support and targets intervention in under 60m in the event of single-node failure, but it’s better not to have to count on another engineering team pushing the right buttons before your cluster gets back to full health.

The only outright error during our testing was that two pods, `cilium-envoy` and `csi-hyperstack-node`, crash-looped with `Too many open files`. This was a minor annoyance and did not affect our workflows.

We look forward to continuing to work with Hyperstack as they grow their team, gain experience managing larger clusters, and continue to automate and harden their infrastructure.

Participation Ribbon

Vultr

Vultr advertises itself as “the world’s largest privately-held cloud infrastructure company,” which they are not, and “the world’s largest privately held hyperscaler,” which they wouldn’t be even if they were. Still, they are a relevant player in neocloud land, with modern GB300s and MI355X including a claimed 50 MW AMD site in Ohio. In ClusterMAX 2.0, Vultr’s cluster was delivered with a handful of basic errors. Unfortunately, the pattern is the same almost a year later.

The images that Vultr built for our testing had many libraries with applicable CVEs, including CUDA, DCGM, Docker, runc, and ConnectX firmware. Our Slinky login pod did not mount `/shared` and lacked the Pyxis plugstack config, so we had to orchestrate from a worker pod until the Vultr support team reprovisioned the login pod mid-campaign, fixing both issues. There was another basic problem with the storage visible from the K8s layer: the default storage configuration was incompatible with our bare-metal cluster. The default StorageClass, `vultr-block-storage`, provisioned and bound without complaint, then the pod sat in `ContainerCreating` forever. However, the Lustre tier, which came attached to the worker nodes, was excellent, achieving 91.7 GB/s aggregate read across 32 clients. Grafana was configured but only partially correct, as NVLink read 0 due to unconfigured `DCGM_FI_PROF_NVLINK_*` fields.

Vultr is known to have been experiencing issues with the reliability in its scale-out networking, which has discouraged partners from beginning or expanding existing deals. We recommend that Vultr continue to improve the sturdiness of its clusters and work to create golden images with all relevant utilities and orchestration reasonably configured and up to date.

Vessl

Vessl is a new Korean neocloud who fancies itself a “multi cloud orchestrator” rather than a broker, buying bare metal long-term and providing a management layer with SLAs on top. They claim 5,000 GPUs currently in operation and have an ambitious target of 100 MW by the end of 2027. An old Vessl BizOps job posting says to expect “Speed. We MUST move fast, learn fast, and iterate constantly,” demanding “a minimum of 60 hours per week.”

We love the ambition; however, the immaturity of Vessl’s cluster, which we sampled in its Kubernetes and Slinky flavors, gives them plenty to work on. Our onboarding began with a notice to pin three settings to render the RDMA fabric usable. A heads-up is better than nothing, but we much prefer to be given a cluster that performs well out of the box.

The third bullet was in response to a particularly cumbersome configuration: by default, due to Vessl’s Kubernetes `LimitRange`, each container would hard-fail if it used more than 2GB of host memory. Although a defensive default might make sense in certain cases, this is a poor setting. Nodes rarely are called on to run more than one job at a time, so containers can expect to have all of the host DRAM to play with—hundreds of GB—and OOM-killing them for going above 2GB is stifling. We recommend that Vessl relax this limit, bringing it in line with the physical capacity of the host while also making failure more graceful.

In terms of management, Vessl provided no way to handle users on the Slinky layer except via traditional Linux tools on the machine. The SSH key we added from the console did not show up on the cluster, nor did the one we added via the Vessl CLI; we later learned that these utilities are independent of Vessl’s Blackwell clusters. We could therefore use only the private key and `kubeconfig` DMed to us on Slack.

The cluster provided a reasonable (though strange) `add-user` convenience script to make new Linux users.

It did not provide, however, basic packages like `pip`, `git`, or even `sudo`.

The cluster therefore took some work to get to the starting line. But our intern was still not done putting in reps as Linux sysadmin. Vessl ran the Slurm login node with an ephemeral OverlayFS root, so runtime changes to `/etc/passwd` and `/usr/bin` existed only in the container’s writable layer and were not preserved when Kubernetes recycled the pod. This caused our users to disappear midway through testing.

Another Slinky misconfiguration brought down the Slurm layer for part of testing. Our workloads wrote to an overlay that they did not clean up, accumulating 863 GiB in the root volume, causing Kubernetes to mark the node `DiskPressure` and restart it. This not only brought down the local worker, but the Slurm controller daemon as well, rendering the whole Slurm layer intermittently inaccessible. We infer that `slurmctld` was scheduled on to a node also running `slurmd`. Co-scheduling a `slurmd-pyxis` pod with a `slurmctld` is generally a bad practice for reasons well illustrated by this anecdote: the data plane should not be bogged down by the administrative duties of the control plane, and the control plane should not be within the blast radius of the data plane. Vessl quickly stopped the bleeding by having Enroot write to a larger volume, but to our knowledge the root cause—the co-scheduled worker and controller—remains unaddressed. It is clear that Vessl does not have experience running Slinky at scale in production.

A separate error made Kubernetes access inconsistent. The URL in our `kubeconfig`’s `server` field did not always resolve to the same IP address: Vessl had erroneously published the private `10.x` IP to the record set in addition to the public one. Sometimes our `kubectl` command resolved to the right (public) IP, and we could access, but other times it would resolve to the wrong one and our request would hang until a retry went through. We recommend that Vessl remove the `10.x` IP from the public record set.

The `NetworkPolicy` we set on the Kubernetes layer of our Vessl cluster was accepted but not enforced, as HTTP requests between all tenant pods were accepted without regard to defined rules. Vessl did not have a health check program configured and should prioritize this along with all the mundane “quality of life” features indicated above.

The hardware itself that Vessl offered performed well throughout testing, including on NCCL tests over InfiniBand. Given that it has such a large fleet advertised as coming online soon, Vessl has every motivation to harden its offering, providing superior services to its users and enabling it to command superior prices. We will be interested to observe Vessl’s progress when we next test.

Runpod

Runpod is a neocloud headquartered in Moorestown, New Jersey with a fleet including RTX PRO 6000, H100, H200, B200, and B300. This June, they raised $100M at a $1B valuation, announcing on Twitter that their revenue doubled to $240M ARR from February to June.

The cluster we were able to test was in Seattle, and was so new that it did not even have storage configured yet. Our “Instant Cluster” was just about instant, spinning up in under 2 minutes.

Runpod’s user model is unusual. In an ideal world, we manage permissions via a web console with RBAC, where we can add our team members by email, increase or decrease privileges, make groups, and drop in our SSH keys to be automatically synced to the cluster. Instead, taking advantage of its snappy provisioning, Runpod recommends minting Instant Clusters as needed. Engineers then operate their cluster with a single tenant and return them to the pool when finished. This means that budgeting, monitoring, and other such policies are handled above the cluster level, rather than via traditional tools like Linux user management or Slurm accounting. Runpod’s design choice is viable, but it is revealing we did not end up using the cluster in the suggested manner, preferring to fall back to the more familiar pattern of making Linux accounts tied to our SSH keys.

Our first burn-in failed as one of our nodes had port 1 links go down on two NICs, suspected to be due to an issue with a leaf switch. There was no health check to catch the error, but there was a Grafana dashboard that highlighted the error if you know where to look. With our feedback, the Runpod team quickly improved the dashboard to make InfiniBand state easier to spot.

On the second burn-in attempt, the networking performed well and the test completed without error.

The cluster was so new that there was no shared filesystem to test. On the software side, though, our initial audit found many things to be improved. The CUDA Toolkit, GPU driver, and NIC firmware were deprecated for security reasons, and the CUDA Toolkit was additionally a liability because 12.8 does not support SM103, the B300 SM architecture. This caused some of our benchmarks to fail on the first try. The cluster was also missing lots of expected tooling, with no container runtime and many basic HPC packages absent. Slurm also occasionally marked waiting jobs as `InvalidAccount`, which prevented jobs from running, even though `AccountingStorageEnforce=none` and `AllowAccounts=ALL`. The workaround was to spam `scontrol update JobId=<jobid> Account=me` until Slurm accepted.

We also saw a node drained when an image extraction filled our 50 GB root overlay, which was shared with `/var/spool/slurmd`, leaving `slurmd` unable to write. Our tenant lacked credentials to `scontrol update`, so our testing continued shorthanded. Separately, Linux rejected `unshare -Ur`, so we had to run with `udocker` instead. Our standard XID injection route was blocked because we lacked permission to write to `/dev/kmsg`; we had no mechanism to issue node reboots; passwordless `sudo` was not enabled by default. The theme here is clear: within the Runpod environment, users lack many permissions they need to get their work done.

We find the Runpod team unusually receptive to feedback and eager to explain their decisions. We like that the majority of their job postings are in serious engineering roles, and we hope they continue to work to strike a better balance between ease of access and developer control.

Bitdeer

Bitdeer was a new inclusion in ClusterMAX testing as of ClusterMAX 2.1. Yet another crypto miner pivoting to AI, Bitdeer already has an impressive 1.75GW in total electrical capacity according to its most recent announcement. However, this mostly feeds cryptocurrency mining rigs: our SemiAnalysis Datacenter Industry Model estimates they only have a handful of AI MW online. Still, Bitdeer is steering hard toward the higher-margin GPU business, with hundreds of MW already in the pipeline over the next couple of years as it works on sites, either converted or now, in Knoxville, Wenatchee, Fox Creek, and Rockdale, as well as in Norway and Malaysia. Hoping to grow its business vertically as well as horizontally, Bitdeer is not content with mere colo deals, with fledgling offerings from token-as-a-service to managed agents to, of course, managed clusters.

Source: bitdeer.ai

We spent a lot of time trying to figure out what exactly Bitdeer expected us to test when they gave us credits on their public console. Health checks were not set up, the monitoring dashboard did not work, many basic libraries like PyTorch and Docker were not installed, networking settings were empty, IMEX was not configured, and the console did not support ordinary conveniences like RBAC. Bitdeer’s monitoring installer downloaded a binary for the wrong CPU architecture. SSH worked eventually, but oddly required an RSA key rather than ed25519. Kubernetes queries repeatedly lost their connection. In short, the management of this cluster was strictly nominal. The most that can be said is that it had GPU drivers, NCCL, CUDA, and an OS and kernel, and that all were sufficiently modern.

Once we had the time to set things up, performance was fine. GPUs hit expected numbers on burn-in and NVLink supported the expected traffic.

However, the cluster’s WAN performance was exceptionally bad, averaging around 0.1 GB/s across several tests. The cause of this poor performance is unknown to us; needless to say, it would be an impediment to real-world usage.

Most of all, getting support was a massive pain. The entire technical team was based in Asia (Malaysia, or Singapore in our case), which meant 8-12h TAT when we were trying to troubleshoot issues. It was impossible to consider Bitdeer for anything other than this tier.

We look forward to seeing Bitdeer’s improvements in future rounds of testing as its managed cloud offering matures and it implements features like health checks and managed Kubernetes. It will be interesting to see where Bitdeer allocates its time and attention, as management has unambiguously announced that its neocloud business is not its main focus: “Executing on colocation lease agreements is THE top priority.”

Shadeform

Shadeform continues to focus on brokering and small scale marketplace reselling of individual GPU VMs with a small team. We enjoy the interface and working with the Shadeform team, but without more ambition they will stay in our lowest tier.

Radiant

Radiant was formed after Brookfield acquired Ori, a UK-based neocloud that we discussed in ClusterMAX 2.1. Unfortunately we have not been able to test a working, secure cluster. Despite almost a year of planning, along with quite a bit of marketing effort, and announcements that mention Billions and Gigawatts, Radiant is yet to deploy a Blackwell GPU.

FPT

FPT suffered from a number of security issues during our testing, which we described in ClusterMAX 2.1. We have yet to re-test and are not aware of any Blackwell capacity.

Core42

Core42 is a dark horse in this category. They have a solid team, allocation of GPUs, political backing, and an infinite money glitch in the form of Mubadala. We expect to see them put the pieces together and rocket up the rankings as they enter the US market.

Latitude

Latitude is making progress to keep their modest cloud business afloat, being recently acquired by Megaport and as a result going public. We had the chance to test some RTX Pro 6000 Blackwell servers, but as long as Latitude is missing the latest and greatest GPUs, we expect them to be stuck in this category.

IBM Cloud

We have not had the opportunity to re-try IBM Cloud since our horrendous experience in ClusterMAX 2.0.

BuzzHPC

Buzz is making waves in Canada, latching onto the Sovereign AI wave associated with the Carney vs Trump showdown, to the tune of 1.2GW and $50B in Saskatchewan, in partnership with Bell and others. This is quite the complement for their 300MW+ elsewhere in the country. BuzzHPC is one of the many public companies on the ClusterMAX list, via parent company HIVE Digital Technologies Ltd. Unfortunately, all the GPUs that we have tested from Buzz were in clusters of questionable quality, relying on questionable software vendor partner choices. We have not had the opportunity to re-test since ClusterMAX 2.0, but we hope this changes in the near future.

Neysa

Indian neocloud Neysa announced a $1.2B in financing this February, including $600M equity and $600M debt, in a round led by Blackstone that left them at $1.4B post-money. They claim “9,216 air-cooled NVIDIA B300 staged between December 2026 and March 2027,” with an equal number of AMD MI350X in the pipeline as well. Impressive growth, and clearly the leader in India in this regard.

This round, our testing unfortunately revealed a cluster that was riddled with errors. Many software packages that came pre-installed during cluster provisioning were months or even years out of date, with several serious CVEs applicable. CUDA Toolkit 12.0, for example, was released in December 2022, making it older than Neysa itself. Somehow this was installed on our cluster at /usr/bin/nvcc.

We enjoyed working with the Neysa team and encourage them to continue to improve their security posture and shore up their fundamentals as their capacity spikes and the bring on large customers from abroad. While current Neysa customers are mostly stuck on the Hopper generation, with a few B300s available, they are no GB’s (and, as a result, no direct liquid cooling) that have made it to production on the Indian subcontinent.

Vast.ai

A reseller/marketplace that keeps coming up in customer conversations for random GPUs on-demand, representing one of the better partnerships teams in the industry, finding unused capacity all over the world. Unfortunately, the UI is just terrible, so hard to use.

Hyperbolic

We haven’t had the chance to re-test Hyperbolic since the last round, but they got a basic security compliance attestation done finally! Hyperbolic moves onto the list.

STN

STN has been very confusing. One of the earliest clouds to deploy B300s, and one of the only providers with an option for “private cloud”, where customers take care of procuring the cluster themselves, and keep the chips on their balance sheet, but STN takes care of cluster operations. We hear from multiple customers how they enjoy this arrangement and prefer it over price gouging neoclouds.

We’ve always enjoyed working with the STN team, but unfortunately each of the last three times we’ve engaged with the team, we have run into some performance, reliability, or configuration issue, asked for a fix, been promised a fix, not gotten the fix, and then time has run out.

We regrettably keep STN in this tier until we can see a complete end to end testing experience as the customer would experience it.

Not Recommended - Underperforming

Based on hands-on testing, these providers can quickly rise to Bronze by fixing one or more critical issues, such as offering only older GPUs, missing basic security attestation (SOC 2, ISO 27001), misconfiguring key server features (leaving PCIe ACS enabled, or failing to enable GPUDirect RDMA), or charging for GPU hours during cluster creation or hardware downtime.

SharonAI

Sharon AI is an Australian neocloud that claims 132MW in its pipeline, with 116MW already contracted out, against a current footprint that we estimate around 4 MW. They do not have SOC 2 or ISO 27001. We were able to test 2 bare metal H200 nodes to verify basic functionality. We look forward to testing Sharon AI down the line as their offerings mature.

IREN

IREN continues to come up in our conversations with neocloud customers, and usually for the wrong reasons. They’ve got rock bottom prices on HGX B200 and B300 GPUs in Prince George and Mackenzie in Northern BC, Canada, and we’ve heard the stories. Multi-day power outages, network upgrades, storage failures, air quality controls leading to link flaps, XID’s everywhere. The #1 worst site in the industry according to users. And we are some of their users, as we’ve gotten to test IREN GPUs from multiple providers that resell their capacity. A lot of our users that aren’t having an good time that we are hearing from aren’t renting from IREN resellers but IREN themselves

With that said, the new builds in Childress and Sweetwater look much better (read: not N+0 power, cooling, and ISPs!). We’ve covered plenty about these sites in our Datacenter Model, suffice it to say that despite all the technical shortcomings in their cloud services, we expect the $2.1B investment from NVIDIA will help them cut less corners this time around.

We recommend that IREN stop pretending to offer managed clusters and inference endpoints in public marketing and investment materials, and focus on their bare metal offerings, where there is plenty of market demand for their services.

Hydra Host

Hydra Host raised a $100M Series A in June 2026, with NVIDIA among its backers. They are focusing on brokering deals, and our earlier testing is all we have to go on for managed cluster experience. We need to test a real managed cluster before considering an upgrade.

FarmGPU

We couldn’t make much progress in our testing of FarmGPU. The Slurm layer did not advertise GPU resources properly, while Kubernetes did not expose any RDMA devices for scale-out networking.

With that said, we very much appreciate FarmGPU’s open development culture, including their solid Grafana monitoring experience and detailed notes on provisioning. We find the technical team to be trustworthy and solid to work with, even if they are stretched too thin just trying to get a few small clusters deployed correctly.

WhiteFiber

WhiteFiber raised $159.4M gross at its initial IPO closing in August 2025. Our prior testing found good network performance but unusable Slurm/Kubernetes integration and weak job monitoring. More recent customer feedback suggested improvement, though the ephemeral Slurm login filesystem remained a concern. We need to verify that these operational problems are fixed.

PaleBlueDot

PaleBlueDot raised a $150M Series B led by B Capital in January 2026. We previously found its marketplace easy to use for individual VMs, but it lacked a complete managed-cluster experience. Unfortunately, they have shut down their on-demand console in favor of serving some big customers in dedicated datacenters.

Akamai

Not much has changed in the last year at Akamai/Linode, with single-node GPUs being the focus and no managed Slurm or Kubernetes available. We’ll keep watching to see if this sleeping giant wakes up anytime soon.

Hetzner

Hetzner’s low-cost hosting model still provides some cheap PCIe GPUs, specifically the RTX 4000 and 6000 Pro Blackwell Edition. We’re waiting to see if they take the plunge with anything bigger, building on all their datacenter operations experience to really serve the AI market.

Mithril

Mithril, formerly Foundry, announced $80M in funding back in 2024 to “restore the promise of public cloud to AI”. Today, it’s a marketplace wrapping 3 Nebius availability zones with H100 or H200, though only the H200’s have an option with a scale-out network.

Source: Mithril website

There were some announcements about TPU support that we were quite excited about, but that partnership seems to have been scrubbed away.

OVHCloud

OVHcloud has substantial infrastructure globally, operating in tons of datacenters globally, some of which are shared with top neoclouds on this list. But our concerns remain that general purpose IaaS is not the growth engine of the AI age with companies doing crazy things to get access to the latest and greatest. Another sleeping giant.

Massed Compute

Massed Compute last announced capital raise involved up to $300M in equity and revenue-share financing from Digital Alpha in August 2025. They have used this to drive reasonable growth for a bare metal business, but that alone does not cut it on the ClusterMAX rating system, especially since their slop cannon chatbot is still getting indexed by web crawlers.

Not Recommended - Unavailable

Our “Not Recommended - Unavailable” tier is the same as ever, but to clarify, for this section we include four types of companies:

  1. Those that we have previously tested who now claim to have zero spare capacity and/or refuse to test with us. This includes Fluidstack, Cirrascale, Lightning AI (recently merged with Voltage Park), Scaleway, CUDO Compute, Denvr Dataworks, and Atlas Cloud.

  2. Those that are yet to launch, yet to be tested, and are generally interesting to us, though we believe some unnamed of this bunch are misleading investors (i.e., lying) by claiming in public marketing materials that they do managed clusters when they really just do bare metal. This includes SpaceXAI, Mistral, Poolside Infrastructure Company, Nscale, Highrise, Corvex, Andromeda, Volta, Firebird, Tatra, Sesterce, Yotta, Boostrun, GlobalAI, Argentum and Qumulus.

  3. Those who are big and important, but who we cannot test properly due to geographic or regulatory concerns. This includes Alibaba Cloud, MegaSpeed, BytePlus, RunSun, SK Telecom, Naver Cloud and Indosat/Zankore/Lintasarta.

  4. Those who are too small to be relevant. These are still covered in our market view, but not part of the rating system.

Fluidstack

Fluidstack has (currently) exited the managed clusters market, as far as we are concerned, because their focus has shifted from managed Nvidia clusters to bare-metal TPU deployments at 100K+ chip scale. We look forward to testing with them again in the future, with whatever chips they may offer at that time.

Cirrascale

Unfortunately, after receiving a legal letter to modify our previous article’s description of our experience working with Cirrascale, we have not been able to establish a productive working relationship.

Lightning (merged with Voltage Park)

Following the merger of Lightning and Voltage Park, the combined company has unfortunately run out of GPUs, and we haven’t been able to test the reconciled products. This is despite the company consistently advertising that they are the 3rd biggest neocloud in the world in terms of GPUs deployed. They are not in the top 10.

Scaleway

Recently, we hear that Scaleway has de-prioritized a number of customer engagements and decided to voluntarily break ties with NVIDIA, in favour of AMD, for seemingly no good reason. We think this is a shame.

CUDO

CUDO did not make a representative managed Slurm or Kubernetes environment available for this cycle as it focuses on other business priorities (namely, managing the build out of a bunch of bare metal datacenters). We look forward to reassessing CUDO when an environment becomes available.

Denvr Dataworks

Despite changes in leadership and random bouts of activity in the community, we have yet to make any progress testing Denvr and have not seen any meaningful growth in the business that would allow them to dig themselves out of the hole their founders put them in.

Atlas Cloud

We haven’t been able to get GPUs from Atlas recently, despite big announcements about bare metal datacenters from their crypto parent company and random appearances on endpoint provider lists with their managed inference offering.

SpaceXAI

We love Colossus and can’t wait to test the SpaceXAI neocloud offering!

Mistral

Mistral just raised a €3B Series D, and has yet to confirm the validity of a hacker from TeamPCP claiming to put its entire codebase up for sale on the dark web. Anyway, we are very excited to test their neocloud offering. The team told us the platform was still coming up and onboarding its first customers in August, so we very much look forward to testing with them in the future. It seems obvious to us that companies with experience building training clusters for their own models (see above) will be successful neoclouds if they choose to pivot.

Poolside Infrastructure Company

Following a $6B licensiquihire by Nvidia that took an incredible group and buried them in the Nemotron bureaucracy, Poolside’s legacy lives on through the Poolside Infrastructure Company, which is developing their Project Horizon campus in West Texas. We have not yet tested anything from them, but their experience and ambition makes them an immediate contender if the skeleton crew decides to go for it on the neocloud business.

Nscale

Nscale announced a $2B Series C in March, announced the acquisition of AnyScale in May, and just filed an S1 to go public on NYSE a few days ago. It is unbelievable to us that they can have such a big business, and continue to claim in public marketing materials that they offer managed clusters and inference endpoints, when this is clearly not true. Obviously, though, they’ll continue to be successful. They’ve got access to lots of land and power, some construction and datacenter operations experience, and people really want that bare metal at a nice price.

Highrise

Highrise, Hut 8’s AI Cloud business, is constantly in the news. We hope we can test some managed clusters from them soon.

Corvex

Corvex continues to run scared from ClusterMAX, though they announced a $33M private placement in August to expand their datacenter capacity. Their choice of merger partner was Movano, maker of the Evie smart ring. Quite a change in product roadmap. Privately, they claim some secure government customers in the USG. We look forward to testing.

Andromeda

Andromeda takes capacity from other providers and puts a managed cluster service on top. We have discussed testing multiple times, and enjoyed working with the technical team, but find the business side to be a little shady. Resellers have a perverse incentive to keep the customer in the dark about who the underlying provider is, out of fear that the customer can just go around them and contract with them directly. Unfortunately, this can lead to bad games of broken telephone, and lots of finger pointing in the worst case. Specifically, we have a problem with non-circumvention clauses and hope that Andromeda can build a durable business model that allows them to be honest about their solid technical team and the value it can provide.

Specifically, we are aware of three different options from Andromeda:

  1. Rent compute - typical 1:1 neocloud relationship, where Andromeda aggregates capacity from underlying providers and resells it, but depends on them to deliver support

  2. Manage compute - some providers have a growing list of 3 or 4 different providers and want one throat to choke, so Andromeda becomes that throat

  3. Bring your own cluster - also called private cloud, this is when customers want to own their own hardware (and control their own supply chain) but still need help managing it all, effectively SRE-for-hire

We are seeing big growth in Option #3 on that list. And we look forward to testing some of Andromeda’s Blackwell capacity soon, wherever it may be.

Volta

Volta launched with $300M in venture funding and a separate $5B financing pool for customers. Its first large announced contract uses Bitdeer’s Norway site. We have not tested the cloud software or customer operations.

Firebird

Firebird is building GPU infrastructure in Armenia, with $60M in construction financing announced by Ameriabank, and expansion plans in Kazakhstan. Recently, they became the subject of a story about Nvidia chip access and the Armenia-Azerbaijan peace process; co-founder Razmig Hovaghimian responded that the project had roots going back years before those diplomatic developments. GPU diplomacy. Pressure’s on to deliver for the people of Armenia! We still need to see the managed cloud in action.

Tatra

Tatra is developing B300 and GB300 capacity in Slovakia, using the country’s nuclear and hydro power base. We have some great feedback about the technical chops of the father/son team—hard working, impressive family. For now, though, we need them to get their first tranche fully deployed so we can put them to the test.

Sesterce

We spun up an on-demand H100 through Sesterce in August. The console showed an out-of-date Ubuntu option, provisioning took its time, and eventually we got a working machine. Then we found Shadeform underneath. A broker on top of a broker. We would quite like to know how many companies need a margin before the customer gets to run a workload! With that said, Sesterce is pursuing some big bare metal Blackwell clusters with big offtakers. We want to test those.

Groq

Groq set out to challenge Nvidia, but after a licensiquihire took Jonathan Ross and most of the engineering team, the remaining company has pivoted… to renting out Nvidia GPUs! Becoming a neocloud is the move if you want another $350M to live out your dreams (as announced in August). We look forward to testing more than just endpoints one day. Jensen gets the chip team and another customer. Not bad.

Yotta

The four-node Yotta H100 Kubernetes cluster we tested had many problems. Within a week, there was genuine XID 94 on one node and a bad rail on another. Yotta’s bring-up process appears brittle, as our small cluster had important discrepancies from node to node. It also had many out-of-date packages. One of our nodes failed to come back online after we rebooted it. Yotta’s health checks did not intervene when we manually reset the secondary PCIe bus to trigger an XID 79, leaving us with one GPU unreachable until manual intervention, although the Kubernetes layer did nothing to distinguish the shorthanded node.

Yotta must successfully deploy Blackwell GPUs and improve its reliability to keep pace with the industry.

Darya

Darya is bringing GPU cloud infrastructure to Tajikistan, right on the border of Afghanistan, which certainly expands the map of places we need to test. Its H200 cluster launched in June 2025, and it subsequently signed an agreement with Yotta to develop a hydropower-fed AI datacenter in Darvoz. Interesting stuff, but after our experience navigating Yotta’s portals, we are particularly curious which parts of the operating model make the trip to Tajikistan. We have not yet had access to Darya’s cluster, though we would love to visit the site.

Source: Darya’s Datacenter on Google Maps

Boostrun

Boostrun has some solid bare metal, but bought its way past some of the Kubernetes engineering: vCluster says the company launched managed Kubernetes in under 45 days with no new platform engineering hires. They went with isolated tenant control planes, dedicated private nodes, and Netris for network provisioning. Sensible. We would rather see a small provider use working software than spend a year rediscovering Kubernetes. In customer discussions, we have still considered adding another operator for a fully managed training experience. We have tested a Boostrun cluster indirectly, through another provider reselling their capacity, but await the opportunity to engage with their support teams directly. We want to see how much of the day two operations Boostrun handles for themselves in a real engagement before moving them onto the list.

Global AI

Global AI takes the sovereign-cloud pitch quite literally: dedicated, single-tenant, air-gapped clusters sold by the data hall (or so they say). Customers decide when Global AI can access their systems, a solid offering for bare metal. However, when advertising an option for managed clusters, we need to see more. Just to assess bare metal we want to see things like how provisioning and security patches get handled, how comprehensive the monitoring stack is, how reliability and health checks work, and generally how competent the onsite team is at closing tickets quickly and correctly. We have not yet gotten to see this level of detail.

Argentum

Argentum’s website says “200,000+ GPUs available now.” Had you heard of Argentum before you got to this section of the article?

QumulusAI

QumulusAI found a different source of GPU debt: a $500M non-recourse facility through USD.AI, using GPU Warehouse Receipt Tokens as collateral for stablecoin borrowing, with financing for up to 70% of approved deployments. GPU Warehouse Receipt Tokens is the actual name. The announcement describes a facility, so we are not counting $500M of installed hardware. In our customer discussions, Qumulus has come up as a capacity supplier that may need another operator on top for managed training. We still need to inspect its own software and support, but initial customer feedback has been surprisingly solid.

Alibaba Cloud

Alibaba has considerably more software to show us than the usual neocloud startup. One example is its ACK Slurm operator: a SlurmCopilot component coordinates resource allocations between Slurm and Kubernetes so idle resources do not remain stranded in one scheduler. This is the kind of orchestration plumbing we like to poke at, and obviously this is a company that is not ready for a poking. We have an account, have been whitelisted for a few model endpoints, and have engaged directly with the sales team, but have yet to get access to any modern GPUs or this ACK software for testing.

Megaspeed

Plenty of hardware to investigate in their Malaysian and Indonesian sites, but no login yet.

BytePlus

BytePlus’s GPU service advertises switch-affinity placement on a multi-rail network, direct RDMA access to its vePFS parallel filesystem, and integration with our favourite model ever: Seedance 2.5. We have not yet completed an evaluation of their clusters, but there is clearly stuff to watch out for. Tik, tok.

Humain

Humain has given itself plenty to do. Alongside the Saudi GPU buildout, Tareq Amin announced a roughly 6 GW datacenter ambition, as well as Humain One, a real-time voice interface for any computer. We would be happy to start with a login and a functioning cluster, but we have yet to get a response from the team, leaving us wondering what secrets they’re hiding next to their Groq systems.

SK Telecom

SK Telecom runs over 1,000 B200s for Korea’s sovereign foundation-model initiative. As we covered in a recent article, the government also rented ~3,000 H100 equivalents from SK Telecom and Naver combined for the competition’s first round. SK Telecom has announced a much larger 2 GW NVIDIA DSX AI factory, expected to deploy Vera Rubin systems with SK Hynix HBM4, within SK Group’s planned 5 GW first-phase buildout. We have tested SK Telecom GPUs through Vessl, and they were solid, but nothing with SK directly. We need to see more from them on the international market to consider them a serious player, but the technical foundation is clearly there.

Naver

Naver, much like SK Telecom, is highly credible. It already operates hyperscaler-class datacenters, and its July plan with NVIDIA and Brookfield calls for expanding the GAK Sejong AI factory from an initial 55 MW buildout to 200 MW by 2028. We are interested in how much of that operational experience reaches external customers, vs the internal research teams at Naver. We have yet to test that experience.

Indosat (Zankore)

Zankore has announced the first phase of roughly 200 MW of GB300 NVL72 capacity is to be delivered in the first half of 2027. Even if they are a little behind schedule on deploying Blackwell, they are still well on their way to a long-term 1 GW ambition, primarily to serve offtakers (rumoured to be) from mainland China. And they’ve landed a $3.1B loan to go do it!

Where ClusterMAX Goes From Here

Read more

来源:SemiAnalysis · newsletter.semianalysis.com