OpenAI Is Moving Into Legal AI. Harvey Has to Keep Moving Too.
Astra for Law raises the baseline. Harvey's same-day research points to the harder contest: whose system learns to do work lawyers will keep paying for.
Richard TangSeptember 17, 2026 · 4 min read
OpenAI introduced Astra for Law on September 17, combining GPT-6 Astra with legal search and specialized instructions. Harvey and Legora are named as future API users; selected firms get initial access through ChatGPT and Codex. The supplier is also offering a direct route to the customer. 1
I would take that threat seriously. I would not conclude that Harvey's business has just disappeared.
That conclusion assumes Harvey stands still while OpenAI advances. The useful question is what Harvey can learn and build as capabilities that once required specialist engineering become available from its supplier. A better foundation can strengthen a product and make parts of it harder to charge for at the same time.
There is a concrete reason to ask that question today. Alongside OpenAI's launch, Harvey published research on using lawyers' judgments to evaluate and improve legal AI. Its research exposes a difficult constraint: getting an answer judged good enough for a particular task. 2
The baseline has moved. The spending is still unknown.
OpenAI's index spans more than 230 million URLs. The announcement describes it as complementary to licensed content from providers including Thomson Reuters. It supplies no price, paid-seat total, or revenue for Astra for Law. It also does not establish that using the index is free. 1
Those omissions matter. A large corpus tells us about the material a system can search. It does not tell us what a firm can cancel, what it will pay for the replacement, or whether the replacement completes the same work.
Our earlier Harvey piece examined the company's financing and valuation. 3 Harvey's own September 9 announcement reported a $550 million raise at a $15.5 billion valuation, and said 80% of Am Law 100 firms used its products. Those are company-reported figures; usage does not tell us the size or profitability of those relationships. 4
The new launch creates a reason to revisit that business. It does not give us a new revenue multiple or evidence that customers are leaving.
The difficult part is deciding what better means
Harvey describes a study covering 24 legal-agent tasks, with three lawyers rating each task. Reviewers were unanimous in only 25% of matchups. The company says much of the disagreement concerned how to weigh the same strengths and weaknesses. 2
This is where the problem gets interesting. Imagine two draft answers: one is easier to read, while another handles an important exception more carefully. Which should the system learn to produce? Counting positive ratings without understanding the reason could teach it the wrong lesson.
Harvey is experimenting with generative reward models: AI evaluators that compare outputs against criteria shaped by legal experts. It reports that these evaluators sometimes accepted serious errors in exchange for other strengths, or selected a winner when lawyers considered both answers inadequate. These are findings from Harvey's own research, not independent validation. 2
A system can produce more answers without producing a reliable signal about which answers deserve to be repeated. That is the bottleneck I would investigate before celebrating a data flywheel.

The work around the model matters: what the evaluator can inspect, how a lawyer's correction is captured, what counts as a serious failure, and whether the next version improves on unfamiliar work. More usage would only be an advantage if the company could turn the relevant experience into dependable improvement, within the permissions governing that information.
Having customers is not, by itself, evidence that this loop exists. Showing the loop improve customer outcomes would be much more persuasive than describing the product as specialized.
OpenAI can keep learning too
In OpenAI's own test on 200 private validation questions, Astra for Law passed the overall correctness check on 54.0%, versus 38.7% for Astra using web search. OpenAI also describes firm-specific applications built with Sullivan & Cromwell, Ropes & Gray, and Cooley. 1
The benchmark is a narrow comparison, not a measure of completed client matters. But the firm projects make the competitive question harder. It would be a mistake to assume the foundation provider must remain distant from the work while specialists alone understand customers.
Harvey cannot rely on today's division of labor remaining intact. Nor should we assume OpenAI's launch is the final product against which Harvey will compete. Both sides can improve.
If I were building the specialist, I would use the stronger foundation where it helped and concentrate effort on the difficult work customers still could not complete satisfactorily. I would want evidence of fewer consequential corrections, less review effort, and repeat use on meaningful tasks. Those are proposed tests, not results either announcement establishes.
The next commercial evidence should be equally specific: a renewal, a procurement decision, disclosed pricing, or a customer explaining which work moved and why. Continued payment would tell us customers still found value. We would then need to understand what they were buying, rather than deciding on their behalf that a general-purpose alternative ought to be enough.
My judgment is that Astra for Law makes standing still more dangerous. It does not make building a specialist pointless. Harvey's opportunity is to turn a stronger foundation into a better working product faster than the foundation provider can close that remaining gap. The company will have to keep earning that gap. Its valuation cannot do the work for it.
How this brief was made
01Gathered & sourced373 channels · 1,030 articles▾
Agents swept 373 channels and ingested 1,030 articles, then de-duplicated and ranked them for signal.
02Verified & cross-validated4 claims · 28 data feeds▾
Every one of 4 load-bearing claims was checked against primary sources, with 28 live data feeds reconciling the figures and charts.
- 1OpenAI, Introducing Astra for Law, September 17, 2026
- 2Julio Pereyra, Harvey, Augmenting Human Preference in Complex Domains, September 17, 2026
- 3The Inference, 5x in 19 Months: Inside Harvey's $15.5 Billion Legal AI Machine, September 10, 2026
- 4Harvey Team, Harvey Raises $550M at a $15.5B Valuation to Help Legal Teams Own Their Intelligence, September 9, 2026
03Reviewed & edited1 human editor▾
One editor read the draft against the evidence, tuned the framing, and signed off before it shipped.
Become a contributor
Reporting on the business of AI and want it read? We take pitches from outside contributors who bring primary sources and a number worth arguing about.
Deepdive
AI-generated from this story and its cited sources. Not investment advice.


