DocuSign demoted the frontier model to judgment calls and cut its AI bill 90%

DocuSign says small task-specific models now do its routine contract reading for 90% less, eight times faster. The same week, OpenAI halved prices, and its cheapest new tier targets the exact clerical work DocuSign just moved off the frontier.

Vincent JiangVincent Jiang · 3 min read
Share
Sam Altman speaking on stage at TechCrunch Disrupt San Francisco 2019
1 / 7Slide 1 of 7
OpenAI chief executive Sam Altman, whose GPT-6 Sol and Luna tiers halved token prices on 22 September.

DocuSign reads more than 1 million contracts a day, and that volume ran on frontier models until the bill became its biggest infrastructure line 1. The fix was not a better rate card. Cheap task-specific models now handle routine extraction at 90% lower cost and 8 times the throughput, and the frontier model is saved for judgment calls, seeing only the passages that matter 1.

Ninety percent off, on the company's word

The saving is company-claimed, reported 21 September 2026 and unaudited 1. The scale behind it is public: second-quarter revenue of $875.75 million and net income of $77.72 million, filed in early September 2. The Iris assistant and agents, unveiled in May, run on what the small models extract first, so every agent inherits the lower bill 1. July extended the lineup with workflow agents and AI-assisted web forms 8.

The labs cut prices the same week

On 22 September, OpenAI released GPT-6 Sol and GPT-6 Luna at half the old rates: Sol at $2 per million input tokens and $10 output, Luna at $0.10 and $0.50 3. Sam Altman called them "half the price per token, and even less per task" 3.

The prices are permanent, not promotional, and they landed mid-war 4. xAI had shipped Grok 4.7 at $2 and $6 per million tokens the day before, and Anthropic answered with Opus 5.5 at $4 and $20 ninety minutes ahead of OpenAI 4. The LLM Token Expenditure Index had already hit 97 cents in early September, its lowest reading ever, more than half below its summer high 5.

Luna undercuts every rival tier in the September price war

  • Input
  • Output
$0 /M tokens$5 /M tokens$10 /M tokens$15 /M tokens$20 /M tokensGrok 4.7 (xAI)Opus 5.5 (Anthropic)GPT-6 Sol (OpenAI)GPT-6 Luna (OpenAI)$0.1 /M tokens
Data
InputOutput
Grok 4.7 (xAI)$2 /M tokens$6 /M tokens
Opus 5.5 (Anthropic)$4 /M tokens$20 /M tokens
GPT-6 Sol (OpenAI)$2 /M tokens$10 /M tokens
GPT-6 Luna (OpenAI)$0.1 /M tokens$0.5 /M tokens
List price per million tokens at launch, 21 to 22 September 2026, input and output rates side by side. Sources: IJR News; Yahoo Finance.3,4

All of it lands while OpenAI holds early investor talks on a round near a $1.2 trillion valuation, up from $852 billion in March 6. That price is a wager that enterprise token bills keep scaling with usage.

OpenAI's new GPT-6 tiers arrived at half price or less

  • GPT-5.6 rate
  • GPT-6 rate
$0 /M tokens$5 /M tokens$10 /M tokens$15 /M tokens$20 /M tokensSol input$2.0 /M tokens$4.0 /M tokens-50%Sol output$10.0 /M tokens$20.0 /M tokens-50%Luna input$0.1 /M tokens$0.2 /M tokens-50%Luna output$0.5 /M tokens$1.2 /M tokens-58%
Data
GPT-5.6 rateGPT-6 rateChange
Sol input$4.0 /M tokens$2.0 /M tokens-50.0%
Sol output$20.0 /M tokens$10.0 /M tokens-50.0%
Luna input$0.2 /M tokens$0.1 /M tokens-50.0%
Luna output$1.2 /M tokens$0.5 /M tokens-58.3%
API list price per million tokens: prior GPT-5.6 rates versus GPT-6 rates announced 22 September 2026, which OpenAI says are permanent. Sources: IJR News; Yahoo Finance.3,4

OpenAI counters by occupying the cheap tier

OpenAI's counter is to occupy the cheap tier itself. Luna is built for high-volume clerical work, document summarization and information extraction, exactly the jobs DocuSign just routed off the frontier 3. OpenAI attributes the cuts to caching and inference savings, passed directly through 3.

Charles-Henry Monchau of Syz Group reads the same record as margin math: "Token deflation compresses the revenue line while compute commitments stay fixed" 5. The moat, he writes, must shift "toward distribution, memory and context" 5. Silicon Data's Steve Hou goes further, saying supply may already "provide sufficient capabilities for most tasks" 5.

DocuSign holds that distribution card in its own market. On 30 September it opens its MCP server, and any agent inside Claude, ChatGPT, Gemini, Copilot or Slack can call Iris intelligence directly 7. A lab can halve its price. It cannot halve its way into data it never sees.

Needing fewer tokens is the labs' real problem

No public number ties DocuSign's savings to any one lab's revenue, and no independent benchmark says the 90% came free of quality loss 1. Anthropic is expected to begin marketing its IPO in mid-October at the earliest 6, as token deflation threatens lab profits on the way to the public market 5. The labs cut the price of tokens. Their customers have learned to need fewer.

How this brief was made

Become a contributor

Reporting on the business of AI and want it read? We take pitches from outside contributors who bring primary sources and a number worth arguing about.

Share

Deepdive

AI-generated from this story and its cited sources. Not investment advice.

Reader comments

0 comments

    Sign up

    Get your curated digest

    After email confirmation, you will receive a daily digest of the most relevant news that matter to your portfolio