Pioneers Insight Method Research Author
Inside Legora's Tech Stack: Why Token Maxing is Failing Enterprise Startups | Legora CTO
Back to Episodes

Inside Legora's Tech Stack: Why Token Maxing is Failing Enterprise Startups | Legora CTO

Summary

  • The 100-year bottleneck in software just moved. Jacob’s core frame: product work → writing code → review, and “the rate limiter was how quickly can you write code. That is now super cheap.” At Legora, Claude and Cursor are the top two sources of generated code — roughly 2% apart, “miles above the next engineer,” and “way above 50%” — so the constraints are now review and product synthesis, and orgs should restructure around those.
  • Review tooling is his explicit startup ask — current tools force humans to stare at lines of code when what matters is impact on architecture, stability, and security boundaries: “if you’re gonna do a startup, please do something that solves the review thing.” Same logic kills the editor: “the current shape of an IDE will die” — the next one is maybe graphical, planning at the systems layer while agents execute.
  • Token maxing is the wrong enterprise metric: leaderboards plus token usage in performance reviews make “people just burn tokens just to look good — a really stupid way to do anything.” Reward output via hack days and demos instead. Yet his own AI budget is driven primarily by opportunity cost: “I don’t want to say infinite, but… the cost of not doing it is extremely high. This almost outweighs any sort of token cost.”
  • Cursor’s post-acquisition neutrality question — its “reason to exist” was third-party token optimization (routing to cheap open-source models, setting limits), which Harry argues is harder now that it’s tied to X. Harry says model-independent Cognition and Factory will do very well; Jacob was “a bit surprised and a bit sad”: “a shame that the industry is vertically integrated in that way.”
  • Shallow SaaS is in the blast radius. Legora vibe-codes internal HR, ATS and payroll-adjacent tools because off-the-shelf software “always basically never really works”; a public-company CEO’s chief of staff took three weeks and replaced Kooper (likely Coupa). Rule: shallow surface + heavy customization → build; deep complexity → don’t build it yourself. He predicts a new enterprise role — internal AI systems teams — will flower out of IT.
  • Moats survive the vibe-coding clones: “it’s very quick to get to the 90% where it looks the same… It’s the other 90% that are difficult” — edge cases, audit logging, RBAC at scale. Model dependence is overrated too: “let’s say you took away a model from Legora, people would still pick Legora.” Legora runs ~10 models, choosing on performance and latency over cost — “you can probably wait an hour more for the output if it’s better.”
  • The scaling scoreboard: $100M ARR in 18 months (per Harry’s intro), year-end revenue “above 250” (Harry’s bet: 272), 80 engineers today going to a predicted 270 by end-2027 — after Jacob’s famous 300-Spartans slide capped the team at 20. Given six months’ exclusive access, he’d take superior engineers over a superior model “for sure”: models improve anyway; great engineers “build a system that exponentially improves.”

Deep dive

1. Code is cheap now — the bottleneck moved to review and product work

  • Jacob’s framing of software’s three phases: product work (translating “user pain, user dreams, nightmares into something tangible”), writing the code, then review and merge — and phase two “was the primary bottleneck for the past 100 years.” AI compressed it: “the rate limiter was how quickly can you write code. That is now super cheap.” Everything downstream of that claim — org design, hiring, tooling — follows from asking “what’s the bottleneck to our velocity” and attacking it.
  • The internal evidence is stark: Claude and Cursor sit atop Legora’s generated-code stats, about 2% apart and “miles above the next engineer” — “way above 50% of all code.” Both tools coexist (“the cursor harness is quite good”; some engineers avoid Claude Code because its harness “is annoying sometimes”).
  • Processes compress too: incidents now get “an incident agent” unleashed on logs and telemetry — “it’s surprisingly good… the postmortem basically almost writes itself.” Engineers still wake up at night, but well-equipped.

2. Review is broken — the startup plea, and the death of the IDE

  • Legora runs AI review bots — security reviewers, specialized reviewers — that iterate with the AI coder: “agents fighting each other until they arrive at something.” But it’s not enough: “I keep telling people at all events… if you’re gonna do a startup, please do something that solves the review thing.”
  • What review should actually check: “no one wants to look at all the lines of code” — the questions are impact on systems architecture, stability, security boundaries. If there’s no strategic trade-off, “maybe you don’t have to review it at all. Just unleash the agent.”
  • The engineer’s job shifts one abstraction up — what does the system look like, what bets are we making, what’s reusable — plus a new discipline he calls meta-engineering: agent-enablement teams (soon an explicit role at Legora) that set guardrails and data loops so you can “unleash agents and say, hey, increase conversion rate on my e-commerce store.”
  • The corollaries: “the current shape of an IDE will die… it’s not reading lines of code. Maybe it’s graphical” — review and plan at the architecture level while agents run off. And his self-declared crazy take on law mirrors it: “eventually lawyers will not be nitty-gritty about the language of the contracts. They will work a level above — what’s our negotiation stance, what risks are we okay taking.” Hedged as spoken: “I’m not sure if this is true, but this is my hunch.”

3. Security is the bill coming due for AI-generated code

  • Despite the speed religion, Legora still human-reviews every single PR “just because we have to be sure.” Jacob calls it inefficient and wants risk scores instead, but concedes Harry’s worry is right: “threat actors are extremely efficient now… we need just as good defense and I’m not sure if we’re there yet.”
  • It’s not hypothetical — a vendor security incident the day before the recording forced a credential rotation (no client impact, he stresses). “I think we’re going to see more of them.”

4. Moats when anyone can vibe-code your clone

  • People are vibe-coding Legora, Salesforce and DocuSign clones today, and it doesn’t change his product strategy: “it’s very quick to get to the 90% where it looks the same… It’s the other 90% that are difficult” — edge cases, unhappy paths, audit logging, RBAC, “the weird scenarios you end up at at a certain scale.”
  • On taste as defense against sameness: “if you don’t have taste, then you let AI slop converge to grayness… if you’re just letting AI rip, you’re going to look the same as everyone else.” Design survives one level above individual features — design language, hierarchy, “the opinionated stance of who we are” — stored in Figma, which Harry needles is now just “a storage feature” (Jacob concedes: “it could be something else”).
  • The customer-speed problem: three different speeds — AI, product, humans — and Legora’s business is “translating the immense speed of AI development into a user base that’s been historically underserved.” That’s why forward-deployed engineers are still required in enterprise legal — “that’s the price you pay for being on the forefront” — though “in 5 years maybe we don’t at all.”
  • The real threat isn’t Harvey: “the thing that’s going to kill us is if we lose the ability to constantly react and readjust and reinvent ourselves.” And for founders facing incumbents: “just work harder than the 800lb gorilla… no one in the 800lb gorilla is extremely excited to be there.”

5. Vibe-code the back office — bad news for shallow SaaS

  • Legora’s internal-enablement team reimagines the company from first principles: going 200 → 1,000 employees, “can we just vibe code our HR system? Our talent acquisition system? Our payroll system?” Off-the-shelf tools “always basically never really work” and building is “so cheap” now. The example as told: Ryan vibe-coded a Canada→Sweden relocation app for an incoming team in a day; a public-company CEO’s chief of staff took three weeks and replaced Kooper (likely Coupa) — “and it works and it’s brilliant.”
  • His build-vs-buy rule has two axes — surface area and depth. Shallow app plus heavy customization → build it yourself (“that’s probably actually the right thing to do”); deep app hiding real complexity → don’t build it yourself, “there’s just too much stuff for you to build.”
  • The role that follows, his pick for a job that’s common in five years: internal IT has “a flowering moment” into an internal-AI-systems team building custom tools — “if enterprises don’t create that role I will get really annoyed.”

6. PMs survive — precisely because product work is the new bottleneck

  • Against the product-and-engineering-are-converging consensus: at Legora “the bottleneck is no longer coding, which means the bottleneck is the product work.” A PM spending 50% of their time coding means “we’re missing out on so much product work” — it’s pure opportunity cost. The carve-out: devtools and consumer, where engineers are their own clients — “you haven’t needed a PM before AI either there.”
  • The nuance: PMs should vibe-code prototypes — high-fidelity artifacts kill handover cost, and they can iterate with users before engineering ever gets involved. Co-location in Stockholm runs on the same logic: PM, designer and engineer together means “you can almost not even have the handover.” Great engineers who want remote “probably also want very isolated problems… not for us right now.”
  • His advice to a new CS grad: “the most important thing is you need to learn how to learn… if you can just learn faster than everyone else, then over time you win.”

7. The model layer: ~10 models, performance and latency over cost, open source’s moment

  • Legora runs maybe 10 models across tasks, “swinging between OpenAI and Anthropic” as “the best model changes almost weekly.” Selection prioritizes performance and latency, “not so much cost” — “if you’re a lawyer… you can probably wait an hour more for the output if it’s better.” And model dependence is “much less than most people think”: the value is legal primitives, enterprise features, optimal routing — “take away a model from Legora, people would still pick Legora.”
  • Open source is “having a great moment”: a local model helps him code offline on flights, transcription runs on his iPhone, and he thinks Whisper Flow — his pick for most underrated AI company — should “just go local.” His worry: European and American open-source models have been lacking — “game theoretically we won’t end up in a great place if there’s a duopoly or monopoly on the models.” Europe’s place in the model race: “it should, but it doesn’t yet.”
  • On the efficiency frontier: a sub-quadratic, huge-context model dropped two days before recording — “super exciting… I don’t know if I trust the benchmark yet” — and “maybe the current LLM architecture is not the one to take us all the way.”

8. Token maxing is the wrong metric — and Cursor’s post-acquisition neutrality question

  • His advice to public-company boards asking about token maxing: leaderboards plus token usage in performance reviews mean “people just burn tokens just to look good. That’s a really stupid way to do anything.” Do hack days and demos instead — “reward them for being effective and efficient and having more output, not for necessarily using AI.”
  • His own budget answer, on what percent of developer salary to spend on tooling: “I don’t want to say infinite, but for me it’s a question of opportunity cost… the cost of not doing it is extremely high. This almost outweighs any sort of token cost” — with the caveat that a company with lower opportunity cost should answer differently.
  • Jacob said Cursor’s “reason to exist” was neutral third-party token optimization — routing to cheap open-source models, setting limits across vendors. Harry argued that model-independent Cognition and Factory “do very well.” After Jacob asked whether the concern was Cursor being tied to X, Harry said “100%.” Jacob was “a bit surprised and a bit sad… I just think it’s a shame that the industry is vertically integrated in that way.”

9. Hiring: the 300-Spartans mistake, 80 → 270 engineers, and the ego filter

  • The honest answer to what Harvey did better: “they’ve been more aggressive with hiring.” Jacob showed the whole company a 300-Spartans-versus-Persians slide ~18 months ago predicting Legora would “cap out at 20 engineers” — “I got that really wrong.” Today: ~80, “way too small still,” and the cost is unbuilt features, not slowness. His on-record bet: 270 engineers by end-2027 (“I’m going to say a number that’s too low”), with the discipline that “I’d rather miss my number and have A players than hit it with B players.”
  • Biggest mind-change in 12 months: hiring — “as long as adding someone is net positive, we should add someone.” Acqui-hires are the fast path: one great founder attracts five A-players, “you get five in one week”; the acquired code usually gets rebuilt into Legora. Integration is “surprisingly easy if you hire people with low ego” — and ego reveals itself in salary-and-title negotiation: “great people want more money. They don’t mind so much about the title.”
  • Where his hires went wrong — introspectively: with very senior people, “I’m not confident enough in saying this person who’s seen more than me is wrong,” then two-to-six weeks in it unravels. His feedback at two weeks: “you’re not going to stay if you don’t change this” — no one has recovered from it yet. Hardest role today: technical engineering managers, because “anyone who’s seen scale typically is no longer technical”; Legora’s EMs spend most of their time coding but are accountable for team health in six-person teams that “decide their own road map.”
  • Europe versus US: US candidates are “less risk averse” but “more transactional”; Europeans take longer to convince “but once they’re convinced, they really stay.” Equity still needs explaining in Sweden and Europe — “they’re like, wait a minute, you’re giving me half a million… not literally, but like in the future” — an ecosystem cost the next Stockholm startup hopefully won’t pay.

10. “Slapped in the face every 3 months” — build for 100x

  • His two named regrets: underestimating growth, and staffing developer experience too late. The three-person DevEx team (still “too few”) should have existed “when Opus 4.5 came out” — he described engineer productivity as “10x, let’s say,” so if you can make everyone 20% more efficient, it’s even more gains. What they built: a background coding agent giving each engineer ~10 concurrent agents, custom review agents, CI flows that wait until everything’s green before raising to a human, and README-driven onboarding where new engineers “just ask their Claude Code.”
  • The scaling rule changed: “I used to say 10x, but that was not enough. So now everything needs to scale to 100x the usage.” The example as told: tabular view bulk extraction, where 100,000 cells versus 10,000 spikes the system instantly — solved with fair queuing: the 100k-cell user “can go grab a coffee,” the 10-cell job stays fast.
  • Given six months’ exclusive access, he’d take superior engineers over a superior model “for sure” — models get better on their own, but great engineers “build a system that exponentially improves, and that’s worth a lot more.”
  • The revenue exchange: Jacob says “above 250”; Harry bets 272 and says he thinks it will be above that too; Jacob then retreats — “I don’t know numbers… David’s going to call me really mad soon.” Harry’s intro frames the base: $100M ARR in 18 months, “the fastest growing enterprise company in history.”