Cognition's 4.8x benchmark proves throughput, not cheaper agent work

CoreWeave's first Vera Rubin customer published a 4.8x throughput benchmark but no cost figure, while CoreWeave has hiked GPU prices 35% since July. The number CRWV holders need, cost per completed Devin session, exists nowhere in the releases.

Vincent JiangVincent Jiang · 3 min read
Share
An Nvidia DGX GB200 rack photographed full height, the previous-generation NVL72 system class that Cognition's Vera Rubin benchmark used as its baseline
1 / 6Slide 1 of 6
An Nvidia DGX GB200 rack, the previous-generation NVL72 system whose throughput is the 1.0x baseline in Cognition's Vera Rubin benchmark. The new Vera Rubin NVL72 racks running Devin at CoreWeave have not been shown publicly.

Cognition has a habit of announcing its milestones with numbers attached. On 8 September it was a Series E: more than $2 billion raised at a $48 billion valuation, with company-reported run-rate revenue climbing from $492 million in May to almost $900 million 1. On 29 September it was a MongoDB partnership selling Devin as the engine of enterprise migrations 2. And on 30 September, at CoreWeave's Fully Connected conference, it became the first customer anywhere running production workloads on Nvidia's Vera Rubin NVL72, reporting up to 4.8x total token throughput for its SWE-2 inference workloads against a GB200 NVL72 baseline, and 3.8x on reinforcement learning 134.

A benchmark without a bill

The releases are bare in the same places. No absolute throughput figures, no pricing, no energy use, no configuration detail sufficient to reproduce the comparison, and no count of completed tasks 1. CoreWeave's release calls the results "independent benchmarks," but Cognition's own engineers ran them against a baseline on CoreWeave's own cloud 4. The only translation into money is verbal: Chen Goldberg, CoreWeave's EVP of product and engineering, said the gains mean more concurrent Devin sessions per GPU and "lower cost per session," and no release published that cost 34.

Two reported gains, one baseline, no cost anywhere

0x2x4x6xGB200 baselineSWE-2 inference on Vera Rubin (Cognition)RL output tokens on Vera Rubin (Cognition)3.8x4.8x
Data
Token throughput, multiple of GB200 baseline
GB200 baseline1x
SWE-2 inference on Vera Rubin (Cognition)4.8x
RL output tokens on Vera Rubin (Cognition)3.8x
All values are token throughput expressed as multiples of the GB200 NVL72 baseline, the comparison the benchmark's stated methodology defines: the baseline sits at 1.0x by definition, with Cognition's reported gains on Vera Rubin NVL72 measured against it [1][4]. Both gains are vendor-associated benchmarks with no published absolute figures, prices or energy use. Sources: RuntimeWire [1]; Business Wire via Yahoo Finance [4].1,4

The heavier ledger is CoreWeave's

The stakes sit on two ledgers. Cognition needs to show its $2 billion raise buys compute that makes long agent runs cheaper while it prices Devin against Claude Code and Cursor 1. CoreWeave's ledger is heavier: it says it has 10x demand for every megawatt of capacity and is holding to an 8 GW buildout target by 2030, and Cantor Fitzgerald kept an Overweight rating and $176 target on the stock after the conference 5. That ramp is being financed against demand evidence whose first public data point is a benchmark the customer and the vendor co-authored. CRWV holders are the ones underwriting it.

What a session costs is the missing number

Throughput is not economics. If a Devin session consumes more tokens, a 4.8x gain in tokens per second shrinks to something far less impressive in cost per finished task, and CoreWeave has raised prices 25% since July plus another 10% on top, per the Cantor note 5. CoreWeave's own silicon-level claim, 10x token throughput per megawatt over GB200 on DeepSeek R1, is an efficiency number, not a bill, and it appears in the same release that omits every cost figure 4. Meanwhile Cognition sold the same agent into MongoDB's migration funnel the day before, promising work that took five to six hours now takes just over one 2. The week's real question is whether per-task price cuts in AI coding are being funded by efficiency or by capital. Until either company publishes cost per session on Vera Rubin versus GB200, the market is pricing the multiple, not the cost. Throughput is not economics.

Deepdive

AI-generated from this story and its cited sources. Not investment advice.

Reader comments

0 comments

    Sign up

    Get your curated digest

    After email confirmation, you will receive a daily digest of the most relevant news that matter to your portfolio