How Harvey AI is Changing the Legal Industry with Winston Weinberg
Summary
- Harvey was born from a deceptively strong GPT-3 result: 86 of 100 landlord-tenant answers passed a blind “send without edits” test by three attorneys. The team had brute-forced context and early chain-of-thought prompts, then cold-emailed OpenAI’s general counsel, whose reply was, “I had no idea the models were this good at legal.” Weinberg’s real bet was the slope—better models plus better context, evaluation, and application engineering—not that GPT-3 could already one-shot complex law.
- Commercially, Harvey says it has more than 250 clients and $50 million in ARR after raising more than $500 million, but its distribution insight is more important than the headline scale. Instead of starting with startups or mid-market users, it pursued A&O Sherman, PwC, and other demanding institutions, then partnered with Lexis, which was involved in the round, to combine its industry goodwill, trust, and products with Harvey’s AI. “The best way to do that is to actually go after the hardest people first”: elite design partners help define workflows, establish legitimacy, and open the rest of a conservative market.
- The product strategy is to expand into narrow, high-quality workflows and then collapse them into a simple interface. Harvey is building 30–50 reusable “AI patterns”—case-law research is one example—across broad productivity tools and start-to-finish specialists, then using orchestration so an uploaded share purchase agreement can trigger seven relevant actions instead of exposing “10,000 workflows.” The aspiration is an end-to-end S-4 filing; another example assesses antitrust requirements across 72 countries.
- Reliability is not one bar: general tools win by making imperfect work cheap to verify, while specialists need much higher minimum quality but are easier to evaluate recursively. For the first, Harvey shows its work in a junior-to-partner review pyramid with explanations and inline citations—“show your work.” For the second, every scoped step can be tested; this is why Weinberg calls generic benchmarks “completely useless for us” and hires experienced lawyers to design tasks and judge outputs.
- Weinberg expects AI to displace legal tasks, compress apprenticeship, and split law-firm economics—not simply erase lawyers or the billable hour. Repetitive work that delays strategic exposure for five to ten years can become fixed-fee, lawyer-in-the-loop output, while scarce senior advice will stay hourly and may become more expensive—he floated, with uncertainty, perhaps 10× a junior’s rate rather than 3×. Firms can also encode their practice into Harvey, sell it as software, and profitably offer work that currently serves as a discounted or loss-leading path to major transactions.
- Reasoning gains push Harvey’s workflow frontier outward, while falling inference prices let it increase quality across every user base rather than optimize for cost. Reasoning models can extend an antitrust analysis from deciding where to file toward preparing the filings; lower prices let Harvey spend more compute on quality. Because legal and tax diligence both “apply all of these rules to these documents,” the same patterns can move into tax, audit, deals, and other professional services.
- The broader application-layer thesis is to target work where the “price per token” is high, then learn the real workflow from practitioners rather than ideating inside tech. Weinberg urges founders to observe industries outside Silicon Valley and predicts specialized work-completion breakthroughs in coding and medicine that feel like a new ChatGPT moment to experts. His adoption read is equally important: professionals reject abstract “Skynet” replacement, but “when folks can see and they actually use these tools, they want this.”
Deep dive
1. An 86% blind test made GPT-3 look like a company
Weinberg had no startup plan when his eventual co-founder, Gabe, showed him GPT-3. He was “incredibly surprised that no one was talking about GPT-3,” while his own legal work supplied an immediate testing ground for whether the model could handle expensive professional tasks.
On r/legaladvice—where Weinberg joked that almost every response is “So who do I sue?”—they collected roughly 100 landlord-tenant questions and developed step-by-step prompts before chain-of-thought was widely discussed. Three landlord-tenant attorneys, told nothing about AI, judged whether each answer was ethical and good enough to send unchanged; 86 passed.
A cold email carrying those results prompted OpenAI’s general counsel to reply, “I had no idea the models were this good at legal,” followed by a C-suite meeting weeks later. The lasting conviction came from persistence: friends stopped after one imperfect attempt, while Weinberg would “hammer on this until it works,” including spending 24 hours testing everything GPT-4 could do that GPT-3 could not.
2. Harvey expands workflows, then collapses the interface
Harvey’s deliberately broad mission is to become an AI platform for legal and professional services. Weinberg’s premise is categorical: if someone cannot imagine applying AI to an industry and transforming the whole industry, “I don’t think you’re thinking ambitiously enough.”
The constraint is that models cannot one-shot the most complex tasks. Harvey therefore builds a platform that is “constantly expanding and constantly collapsing”: specialized features and agentic workflows proliferate underneath, then combine into a simple experience rather than a “tentacle monster of a platform.”
Internally, the reusable layer consists of roughly 30–50 “AI patterns.” A strong case-law research system, for example, can become one component of a summary-judgment motion and a billion different types of litigation use cases; one team develops the pattern while others implement it across products.
Asked which end-to-end task excited him most, Weinberg chose filing an S-4 because it combines external data, internal data, and “a million steps.” His broader model of professional work joins those inputs with general process knowledge and client-specific practice—such as how one private-equity firm handles side-letter compliance or what it considers market for a clause.
3. Trust comes from auditable output and elite design partners
For broad productivity tools, Weinberg accepts a lower minimum viable quality because law firms already operate through hierarchical review. Harvey should show its work in the way a senior associate reviews a junior associate’s work: explain why it did something, surface the information it used, and link citations to the exact source sentence—“show your work.” Guo’s compression: make checking cheap, because partially correct work can still help.
A specialized workflow carries a higher quality threshold because it aims to produce the final output from start to finish. Yet its narrow scope makes that standard more attainable: each step can be evaluated separately, then the complete process tested recursively.
Harvey regards most generic benchmarks as “completely useless for us.” Domain experts teach the model what steps matter, specify the output practitioners need, and perform the final evaluation; those evaluators cannot be too junior because, as Weinberg put it, “if they were too junior and they were able to eval it, they would be senior.”
Guo challenged Harvey’s reversal of software orthodoxy: why begin with conservative institutions such as A&O Sherman and PwC rather than easier mid-market customers? To illustrate A&O Sherman’s legacy, Weinberg noted that it is almost 100 years old and became famous advising on what he thought was King Edward VII’s abdication. His answer was that industry transformation requires partnership and institutional trust. Harvey considered self-serve and product-led growth, but concluded it would eventually need credibility from the hardest buyers—and partnered with Lexis, which was involved in the round, to combine its goodwill, trust, and products with Harvey’s AI.
4. Orchestration turns workflow sprawl back into email
Guo’s pushback: enterprise software historically accumulates complexity until training costs make simpler markets unreachable. She sees AI potentially changing that equation because models make interface construction and continual simplification cheaper; Weinberg agreed that orchestration is the enabling capability.
Harvey can separately build systems to extract representations and warranties from a share purchase agreement, summarize them, and identify closing conditions. When an SPA is uploaded, orchestration can recognize it and ask whether the user wants one of seven relevant actions, hiding the fragmented machinery beneath.
Weinberg’s desired endpoint is strikingly plain: “The UI for professional services is just email.” Users should not have to search through 10,000 workflows—the condition he sees in existing business software, where feature sprawl makes systems painful to learn.
Guo supplied the commercial markers: more than 250 clients and $50 million in ARR. Weinberg said automation fear was greatest among people who had not used the product; nevertheless, early in the prior year Harvey still had power users alongside customers who were blocked or confused. Follow-up questions, practice-area profiles, and bets on improving models helped usage make a “complete 180” within six or seven months.
5. AI compresses apprenticeship before it removes lawyers
Junior lawyers often reach elite firms after studying material disconnected from actual practice, then spend years reviewing discovery or data-room documents. Strategic work may not arrive until ten years into a career—“if you’re lucky, five.” Weinberg expects AI to compress that timeline so juniors can engage clients and higher-level judgment earlier.
His distinction is “not job displacement, it is task displacement.” Legal work is too messy for simply training on all legal data and declaring the problem solved; lawyers will remain collaborators even as specific repetitive tasks disappear.
For someone hoping to become a star lawyer in five or ten years, Weinberg’s advice barely changed from before Harvey: maximize hands-on experience handling client requests and determining what is actually best for the client. A smaller matter can teach that better than a prestigious $100 billion merger where the junior may not get real responsibility.
The likely economic structure is mixed. Automatable, lawyer-reviewed tasks move toward fixed fees, while high-level advice remains hourly and may become more expensive: perhaps a pharmaceutical-merger specialist deserves 10× a junior’s rate, not 3×, though Weinberg said he did not know. Firms can also encode their expertise into Harvey and sell software to clients, turning discounted or loss-leading work used to win an LBO or M&A mandate into a higher-margin product.
6. Better reasoning extends the frontier into adjacent professions
Reasoning models unlock the next subproblem in workflows Harvey has already decomposed. An antitrust system might first combine acquirer and target financials to determine filing obligations across countries; improved reasoning can push it toward completing the filings themselves. Harvey keeps building at the edge, and every model improvement moves that edge outward.
Falling inference cost is another direct tailwind because Harvey is “not optimizing for cost at all times right now” but for quality. Legal then serves as the “tip of the spear” for tax, audit, and deals: legal and tax diligence share the high-level operation of applying complex rules across document sets, allowing systems to be adapted rather than rebuilt from zero.
7. Harvey hires for agency while its founder learns to let go
Engineers may not arrive understanding take-private deals or tax diligence, but Weinberg values respect for domain complexity and agency over a perfect résumé. Everything changes every six months, so the winning hire is smart, hungry, decisive, and willing to ship, admit failure, iterate, and keep tracking every major model provider.
For potential leaders, the strongest signals are care and ownership: obsession or refusal to make excuses, willingness to name what went wrong, and asking what to improve rather than seeking praise. Weinberg sees self-reflection plus a desire to improve as especially predictive.
Weinberg’s clearest founder mistake was waiting too long to “scale yourself.” He still believes founders should perform each role before hiring—many of his mis-hires came from misunderstanding the job—but a company cannot scale if its founder touches every customer and product decision. Harvey began the prior year with roughly 40 people, making context-sharing and delegated judgment unavoidable.
The cultural expectation comes from Kobe Bryant’s “The job’s not finished.” Weinberg views AI as an unusually compressed moment in which intense effort can have enduring impact; employees must recognize or trust that premise, sustain the standard without reminders, and adopt Guo’s parallel formulation: “I can’t sit this one out.”
On technical imposter syndrome, Weinberg’s answer is immersion rather than credentials. Asking admired people isolated questions is overvalued because they have only a tiny fraction of the founder’s context; spending substantial time around exceptional practitioners is undervalued because their intuition gradually becomes yours. Product bets that work—such as Harvey’s usage turnaround—then supply pragmatic evidence without eliminating uncertainty.
8. The next ChatGPT moment will arrive inside a profession
Weinberg’s application test is “how expensive is the token?” Every fragment of a 50-page merger agreement can be costly to produce, making legal work fertile ground for application value. Founders should not reject an industry because technology has barely touched it or because they lack expertise; repeated conversations can reveal how the work actually functions.
Guo contrasted field observation with founders “ideating in their underwear in their house.” Weinberg urged people to spend time outside Silicon Valley and beyond their immediate networks, even while maintaining that a technology company should be based in San Francisco. Important industries contain workflows that most technology workers simply have never seen.
His near-term prediction is sophisticated work completion in coding, medicine, research, and other specialties—a new ChatGPT-like shock experienced first by vertical experts. Conservative professionals may recoil when AI is framed as “Skynet replacing your job,” but once capability is shaped into a useful task, many actively want it: technology can remove the accumulated “junk on top” while preserving the profession’s underlying spirit.