Cerebras Cuts Its Own Moat in Half and Calls It a 5x Upgrade
SemiAnalysis says Nvidia GPUs, not Cerebras wafers, serve OpenAI's GPT-6.1 Sol Ultrafast. Cerebras answered with no denial but a redesign that keeps only the decode half of inference on its silicon, in the week the stock fell below its $185 IPO price.
Vincent Jiang · 2 min read
The charge nobody denied
SemiAnalysis posted on September 30 that OpenAI's GPT-6.1 Sol Ultrafast runs on Nvidia GPUs at low batch size, not on Cerebras silicon 12. Low batch is the regime Cerebras was built for: few requests in flight, each answered fast 2. OpenAI's DevDay materials, published a day earlier, made Astra Ultrafast live, called the 6.1 Sol version "coming soon," set the tier at six times standard API rates, and named no chip 3.
CBRS slid to about $181, below the $185 price of its May IPO 45. OpenAI is Cerebras' largest customer by committed revenue 15. The tier's August preview, on GPT-5.6 Sol, ran on Cerebras hardware at up to 750 tokens per second 3; the version now promised tops out at 300 3. Neither company has confirmed or denied the claim 14.
Prefill moves out
Cerebras answered on October 1 with a redesign, not a denial. Its disaggregation post hands prefill, the compute-heavy first half of inference, to partner accelerators, with AWS Trainium and AMD Helios systems shown feeding a Cerebras decode pool, and reports a 5x capacity gain in early tests from the same wafer footprint 6. That figure is company-claimed and unverified. What the wafers keep, they still win: decode throughput on GPT-oss-120B measures 1,669 output tokens per second against 708 for the nearest rival, in Artificial Analysis figures the post reproduces 6.
The half Cerebras kept: decode still runs more than twice as fast as any rival
Data
| Output tokens per second | |
|---|---|
| Cerebras | 1,669 t/s |
| SambaNova | 708 t/s |
| Groq | 475 t/s |
| Microsoft Azure | 319 t/s |
| Nebius | 294 t/s |
| Baseten | 293 t/s |
Half a moat, sold on debt
The narrower architecture is already moving through resellers. Gimlet Labs plans 100 megawatts of Cerebras capacity and, with Cerebras, targets up to 3,000 tokens per second 78. General Compute's order, the first draw on an up-to-$400 million Upper90 debt facility, opens in Q1 2027 910. Both queue behind OpenAI, G42 and AWS for supply, against $25.4 billion in second-quarter remaining performance obligations, while Cerebras' own cloud and services revenue, up 281% year over year to $126.0 million, sells beside them 78.
The clock runs with the chips. Form 144 notices covering a combined $84.3 million of proposed sales by the COO and CFO landed September 29 under trading plans adopted June 30, and about 19.4 million shares became sale-eligible the next morning 411. Proposed is not sold, and the plans predate the drop 4.
The verdict watch
Two facts will price this stock: which silicon OpenAI names for 6.1 Sol Ultrafast, and whether 5x survives production concurrency. Nvidia is already productizing disaggregation as a single-vendor feature on its Rubin roadmap 8. Cerebras still owns the fastest half of inference; this week the market began pricing what the other half is worth.
More about NVIDIA
NVDA · Fiscal Q2 2027Revenue rose 106% to $96.2B, 92.5% of it from Data Center, “driven by the ramp of our Blackwell Ultra infrastructure”. Gains on equity stakes of $7.8B lifted net income to $59.7B.
Show as a table
| Line | Value |
|---|---|
| Revenue | $96.2B |
| Gross margin | 75.0% |
| Operating margin | 66.2% |
| Net income | $59.7B |
| Hyperscale | $48.7B |
| AI clouds & enterprise | $40.3B |
| Data Center | $89.0B |
| Edge Computing | $7.2B |
| Gross profit | $72.1B |
| Cost of revenue | $24.1B |
| Other income | $7.8B |
| Operating income | $63.7B |
| Operating expenses | $8.4B |
| Tax | $11.8B |
| R&D | $7.1B |
| SG&A | $1.4B |
Deepdive
AI-generated from this story and its cited sources. Not investment advice.



