CoreWeave announced on September 30 that Nvidia’s Vera Rubin NVL72 systems are now available on its cloud, with Cognition, the lab behind the Devin AI software engineer, as the first customer running production workloads. Cognition’s benchmarks showed up to 4.8x higher total token throughput per GPU compared with the previous Blackwell generation.
The announcement came at CoreWeave’s Fully Connected conference in San Francisco and moves Rubin from infrastructure bring-up into customer production. CoreWeave was the first cloud provider to stand up and validate a full Vera Rubin NVL72 rack back in June. The racks now run real customer work, not just validation tests.
What Cognition measured
Cognition worked with CoreWeave to bring up a Vera Rubin NVL72 cluster in early September, then ran its own inference benchmarks against a GB200 NVL72 baseline, the previous generation rack. The workloads were real software engineering tasks rather than synthetic loops: engineers sampled tasks from FrontierCode, Cognition’s benchmark suite, and deployed AI agents to solve them.
The results, published in CoreWeave’s press release:
| Metric | Result vs GB200 NVL72 |
|---|---|
| Total token throughput, SWE-2 inference | Up to 4.8x per GPU |
| Output token throughput, reinforcement learning | Up to 3.8x per GPU |
| Rack configuration | 72 Rubin GPUs, 36 Vera CPUs |
| Memory | 20.7 TB HBM4, about 1,400 TB/s aggregate |
The measurements came with a caveat CoreWeave itself flagged: these are Cognition workload measurements on CoreWeave Cloud, not standardized industry benchmarks. The comparisons were run at matched interactivity, meaning the agents were not given more wait time to inflate the numbers.
Why token throughput matters here is specific to agentic work. Devin works in a loop, reading large codebases, reasoning about changes, executing code and waiting on each step’s result before the next. When every step waits on the last one, throughput gains compound across the whole task. Cognition’s engineers framed it that way in the release: the gains mean more concurrent Devin sessions per GPU, faster research loops and lower cost per session, with no loss in generation speed.
“For an agentic workload where every step waits on the last one, that compounds into real work Devin gets done,” Cognition said in the announcement.
Cognition scaled to thousands of GPUs on CoreWeave in nine months and runs the full Devin lifecycle, training, reinforcement learning and production inference, on the same bare-metal CoreWeave Cloud platform. That single-vendor setup is what made the comparison clean: the same workloads, the same software stack, the only variable being the rack generation underneath.
The rack itself
Vera Rubin NVL72 is Nvidia’s successor to the GB200 NVL72 generation. Each rack pairs 72 Rubin GPUs with 36 Vera CPUs, Nvidia’s first processor designed specifically for agentic AI workloads, connected by sixth-generation NVLink fabric rated at 260 TB/s. Nvidia claims the platform delivers up to 10x better inference per watt than Blackwell, roughly one-fourth fewer GPUs for the same work and about one-tenth the cost per million tokens.
Those vendor numbers deserve the usual skepticism, but the CoreWeave deployment gives them a first real-world test bed. The systems run with Spectrum-X 102.4T Ethernet networking, and BlueField-4 DPUs handle infrastructure offload, the same silicon Nvidia is using in its agent-safety platform announced the same week. CoreWeave has added its own orchestration layer on top, with purpose-built software the company says is needed to make rack-scale systems perform at production scale rather than only in lab conditions.
Why the deployment order matters
There is a competitive angle beyond the benchmarks. CoreWeave has built its business on being first to each Nvidia generation, a position it held with Hopper and again with Blackwell. Being first to Rubin production keeps that streak intact and gives CoreWeave a window, however short, where it sells capacity its larger rivals do not have yet.
The customer choice is also deliberate. An agentic AI lab that publicly publishes its own throughput numbers is the most credible reference a cloud vendor can ask for right now, given how much enterprise interest is focused on coding agents as the first real commercial AI workload. Devin competes with agents from OpenAI, Anthropic and Google, and every one of those labs buys compute somewhere. The press release functions as both benchmark and sales pitch.
According to Converge Digest, CoreWeave also cited earlier tests showing Vera Rubin delivering up to 10x more DeepSeek R1 inference tokens per megawatt than GB200, a different metric aimed at energy-constrained data centers. Power, not chips, is increasingly the binding constraint on AI expansion, so efficiency per megawatt is the number data center operators care about most.
Context: an expensive quarter for AI compute
The Rubin deployment lands into a market rethinking compute economics. OpenAI is raising a reported $30 billion round at a $1.4 trillion valuation, the FTC has opened an industry-wide probe into frontier labs over agent incidents, and Nvidia itself just shipped a safety platform for rogue agents after the Hugging Face breach. Every dollar of that spending eventually lands on an infrastructure bill, and per-token cost is the number CFOs watch.
CoreWeave’s own stock has been volatile this year on debt-fueled expansion concerns, which makes each production milestone a proof point for the financing model. If Rubin racks genuinely deliver an order-of-magnitude cost improvement per token, the payback math for AI data center debt improves with it. If the gains stay at the benchmark level, 4.8x on one customer’s workload, the story is thinner.
There is also the question of supply. Nvidia’s production ramp for each new generation has been the industry’s recurring bottleneck, and first-mover capacity advantages tend to erode quickly once volume shipments reach the hyperscalers. CoreWeave’s window is real but it is a window, not a moat.
What comes next
Vera Rubin NVL72 is in limited availability on CoreWeave Cloud now. The practical questions over the next quarter are which customers follow Cognition into production, whether the 4.8x figure holds across workloads other than SWE-2 coding inference, and when rival clouds, which are also taking Rubin deliveries, publish their own numbers to compare against.
For now the industry has its first public, customer-run data point on the new generation, and it is a strong one.
