Micron's $50 Billion Quarter Meets an AI Stack That Moves Its Cache

Micron reports Wednesday on a $50 billion revenue guide at an 86% gross margin, and even the Overweight camp expects smaller estimate hikes. New serving measurements show frontier models relocating cache from GPU HBM to host DRAM instead of shrinking it, a shift this week's earnings previews do not touch.

In this storyMUMSSKHY
Vincent JiangVincent Jiang · 3 min read
Share
Portrait of Micron Technology chief executive Sanjay Mehrotra
1 / 6Slide 1 of 6
Micron Technology President and Chief Executive Officer Sanjay Mehrotra

Micron reports fiscal Q4 after the close on Wednesday, 30 September 2026, guided to $50 billion of revenue plus or minus $1 billion at an 86% gross margin 123. Last quarter printed $41.46 billion at an 84.6% gross margin 27. The stock has more than tripled this year, up over 230% 1, a run priced for another round of big estimate hikes.

Consensus sits above the guide at $51.24 billion, up 353%, the fastest on record 4. Morgan Stanley stays Overweight at a $1,200 target and still expects estimates to rise, just "less so than prior quarters" 15. Fiscal 2027 profit estimates jumped from $95.80 to $154.70 in three months, then rose just 1.2% in the month through early August 6. For holders, the size of the next hike is the trade.

From $4 billion to a guided $50 billion in three years

  • Revenue
  • Estimate
$0B$20B$40B$60BQ4 FY23Q2 FY24Q4 FY24Q2 FY25Q4 FY25Q2 FY26Q4 FY26 guide$50B$49–51B$41.46B
Data
Revenue
Q4 FY23$4.01B
Q1 FY24$4.73B
Q2 FY24$5.82B
Q3 FY24$6.81B
Q4 FY24$7.75B
Q1 FY25$8.71B
Q2 FY25$8.05B
Q3 FY25$9.3B
Q4 FY25$11.32B
Q1 FY26$13.64B
Q2 FY26$23.86B
Q3 FY26$41.46B
Q4 FY26 guide (estimate)$50B ($49–51B)
Micron quarterly revenue by fiscal quarter, USD billions, from SEC filings via Sharadar; the final point is company guidance for the quarter reported 30 September 2026.7,1,6

Frontier models push cache into host memory

None of that math prices what the newest frontier serving stack does to memory. Z.ai's GLM-5.x models, 744 billion parameters with 40 billion active per token, use sparse attention to read less of the cache at each step 8. The catch: picking which tokens to read requires the full context to sit in HBM first, so the capacity bottleneck survives 8.

SGLang's HiSparse layer offloads cache entries to host DRAM, and Z.ai says GLM-5.3 adds local storage as a further tier 8. Measured on a B200 serving GLM-5.2, doubling concurrency from 8 to 16 requests cut GPU-memory token reuse from 90.3% to 54.8% while host-memory reuse rose from 6.0% to 40.3%, with the overall hit rate above 95% 8. Agents, which recycle long histories and write little, are the workload doing this 8.

Sparse attention does not shrink the memory bill. It forwards the bill to a cheaper tier.

Doubling concurrency moved B200 cache reuse from GPU memory to host DRAM

  • 8 concurrent requests
  • 16 concurrent requests
0%50%100%GPU memory (HBM)Host DRAMOverall hit rate96.3%
Data
8 concurrent requests16 concurrent requests
GPU memory (HBM)90.3%54.8%
Host DRAM6%40.3%
Overall hit rate96.3%95.1%
Share of prompt tokens reused from each memory tier while serving GLM-5.2 on a B200 system at 8 versus 16 concurrent requests, with overall hit rate above 95%. Source: SemiAnalysis / InferenceX measurements, 28 September 2026.8

Conventional DRAM delivers the volume while HBM keeps the margin

Micron sells both ends of that ladder, and its growth already lives at the bottom. Its June-quarter DRAM revenue rose 65.5% from the prior quarter to $36.0 billion, the fastest of the big three, while its share of the HBM market slid from 21% to 18% 9.

The premium stays with HBM. SK Hynix holds about half that market and runs a 76% operating margin, and each HBM4 wafer eats roughly three DDR5 wafers of capacity, part of why conventional DRAM prices jumped 58% to 63% in a quarter 9.

SK Hynix DDR5 server memory modules on display
Server DDR5 memory modules: conventional DRAM pricing jumped over 50% in a single quarter as capacity was diverted toward HBM. · 4300streetcar

Micron's own investor site pitches agentic AI under a white paper titled "Agentic AI in Data Centers: The Critical Role of High-Bandwidth Memory" 10; the serving measurements put agentic traffic one tier down.

Long-term purchase pacts cannot resolve where future margins settle

The bull case has receipts: 16 strategic customer agreements worth about $100 billion through 2030, with roughly $22 billion of customer cash and other commitments 23. It also has arithmetic: $100 billion is about two quarters of revenue at the guided pace 2. Synovus sees supply tight through fiscal 2027 1, and SK Hynix is weighing a Solidigm IPO that values its NAND unit at up to $150 billion 11.

Even the Overweight camp concedes there is "no real way to prove that near term" 1. The beat will take care of itself; at an 86% margin, scarcity is not the question. Which memory the agentic era pays a premium for is. The test on Wednesday's call: whether Micron's HBM share keeps sliding while conventional DRAM does the growing, and whether management's DRAM pricing talk credits scarcity or substitution 9.

How this brief was made

01Gathered & sourced407 channels · 1,800 articles▾

Agents swept 407 channels and ingested 1,800 articles, then de-duplicated and ranked them for signal.

02Verified & cross-validated11 claims · 36 data feeds▾
03Reviewed & edited1 human editor▾

One editor read the draft against the evidence, tuned the framing, and signed off before it shipped.

Become a contributor

Reporting on the business of AI and want it read? We take pitches from outside contributors who bring primary sources and a number worth arguing about.

Share

Deepdive

AI-generated from this story and its cited sources. Not investment advice.

Reader comments

0 comments

    Sign up

    Get your curated digest

    After email confirmation, you will receive a daily digest of the most relevant news that matter to your portfolio