George Sivulka, Co-Founder & CEO @Hebbia: The Future of Foundation Models | E1250
George Sivulka, Co-Founder & CEO @Hebbia: The Future of Foundation Models | E1250
Summary
- Sivulka’s core market call: all AI companies are undervalued — and so is the S&P 500. His logic: if the computer created ~$100 trillion of stock-market value over 60 years, AI compute will create another $100 trillion over the next 60, with “more than 50% of GDP contributed by agentic applications in the next few decades” — while Harry thinks it happens faster than that. It’s additional value, not cannibalized value, and it lifts non-AI incumbents too: “computers made legacy businesses better if you use them correctly.”
- The spiciest relative-value take: at OpenAI ~160, Anthropic ~40, xAI ~50, “xAI is the most undervalued company” — and it “might overtake OpenAI and Anthropic in value over the next 12 to 24 months.” His reasoning is Elon’s geopolitical positioning, operational talent, and a leaner business with less “administrative bloat” — plus a belief that governments are among the largest AI users.
- The model layer will become commoditized; value will accrue at hardware and the application/agent layer. Nvidia’s real moat is people (CUDA-trained ML PhDs), which holds for training but less so for inference — so the macro shift from training to inference “will destabilize slightly the dominance of Nvidia chips.” His public-markets pick for the AI wave: “I would probably buy Nvidia… or AMD rather,” because AMD benefits from inference scaling “in an outsized way.”
- From the man who first productionized RAG in 2020: “we actually don’t think RAG works at all.” ~90% of real enterprise queries aren’t findable in documents — they’re about documents (“is this company a good investment”) — and “90% of enterprise AI right now is almost like this vapor… fugazi fugazi,” including a likely Klarna staff-cut story (“I think it’s BS… an amazing marketing story”).
- His replacement thesis, and what the $130M from Index and Thiel funds: scaling laws at inference — run hundreds or thousands of sub-model calls over every document rather than waiting for bigger models. The proof-by-anecdote: BloombergGPT, trained on “the best financial services training set of all time,” was destroyed “at every single finance task” by GPT-4 a few weeks later, though he said he didn’t know the exact timeline — verticalized fine-tuning “will ever catch up” to scaling, never.
- Against the slow-enterprise-adoption consensus (his neighbor Daniel Dines’), he cites Excel hitting 90% penetration in finance in 18-24 months (1985-86, off the HP12C): finance is “the slowest moving, most lethargic leviathan… unless you’re providing outsized alpha, in which case finance moves faster than any other industry.”
- Quickfire tells worth keeping: he would not sell Hebbia for $2B today, answered “no” flat when asked if he trusts Sam Altman, disagrees “completely” with SaaS’s claim that business apps collapse into agents, and thinks chat is the wrong interface — “it’s like asking if the TI-84 was the right interface for computers.”
Deep dive
1. Three founder archetypes — and the chip that never leaves
- Sivulka’s opening taxonomy: “you can bucket great founders into three backgrounds — the most common is that you had kind of a messed-up childhood, the second most common would be you’re gay, and the third most common would be you were adopted.” His evidence: Elon (messed-up childhood), Bezos and Jobs (adopted), Thiel and Altman (publicly gay). The mechanism: early misalignment breeds the drive to prove yourself.
- His own version: Staten Island-born to a mother he described as probably a “Mafia child” and a Slovak father who escaped the Iron Curtain, both intending to be professional athletes — and their only son was “chasing butterflies on the soccer pitch,” a math kid at a school where his parents “barely even knew what Stanford was.” Asked if he felt like a disappointment: “the answer is yes… it’s kind of like the ugly duckling.” Harry matched him with his own — the closing beat of the episode: “the chip remains though, it’s not going anywhere.”
2. The NASA snow day — persistence as doctrine
- The formative story, as told: rejected five times for a NASA internship offered only to undergrads, the 15-year-old showed up uninvited at NASA Goddard in Manhattan on a snow day, got kicked to the curb, sat crying in the snow — and his mother, a medical salesperson, told him “you sit your ass down and you call every single number that you can get into the building.” Someone picked up; he pitched himself for two hours, botched the interview, memorized every poster title on the professor’s wall overnight, came back cold, worked for free, and published internationally recognized research — which got him into Stanford.
- The generalized doctrine, categorical as stated: “you can bring a lemonade stand to $100 million in ARR… you could literally brute-force anything in the world” — the only variable persistence changes is the rate at which you get there. On getting into Stanford: “on to the next one… the next day I was like, okay, how do I become the youngest PhD student in my school’s history?”
3. GPT-3 stole his thesis — so he built the product instead
- As one of the youngest PhD students in Stanford’s history working on meta-learning, GPT-3 landed in June 2020 — the paper title was effectively his research agenda (“large language models are multitask or meta learners”). His reaction: “they just stole the most important thing I could work on right out from under my hands — well, if I can’t build the most important technology, how can I build the most important product?”
- The product insight that still frames Hebbia: “I don’t even think ChatGPT is a really good product — it’s like a calculator… it’s not a product like Excel that lets you just build whatever you’d like with it.” The wedge came from watching friends go into banking and PE and return “the least happy versions of themselves” — “there was more pain in financial services around processing unstructured data than anything I’d ever seen.”
4. The closet, the Thiel breakfast, and the funding cascade
- The founding conditions: on leave from a ~$42K Stanford Graduate Fellowship, he rented the master-bedroom closet of an East Palo Alto house for $500-600, rotating a dorm mattress and a Home Depot folding table on the floor, working 16-18 hour days “like a monk,” waking mid-night to check training runs to conserve GPU credits. A former boss nearly cried on a Zoom pitch: “just come work here — what are you doing to yourself?” His own verdict, hedged honestly: “I probably went too hard… detrimental to my health, but that was a crucible.”
- The Thiel story as told: introduced by a friend who’d interned at Founders Fund, he lost the lunch-vs-breakfast negotiation, drank “18 cups of coffee” and drove a $4,000 Craigslist 2006 Audi from 3am to 8am to Thiel’s house. Thiel arrived nearly an hour late; a 30-minute breakfast ran four to five hours across the business’s flaws, math, and “deep esoteric philosophy,” ending with “I’m not investing… but I’d love to put in a check.” Driving away playing Kanye: “I felt like I was inducted into the Illuminati.”
- Why Thiel is Thiel, in Sivulka’s framing: “ontologically smart” (world-model and pattern-matching) plus “phenomenologically smart” (how humans and processes behave) — always asking ex ante “could I have predicted this ahead of time?” The cascade: ~$1M pre-seed (Thiel, Floodgate), then Mike Volpi at Index — who heard about Hebbia from his Stanford-student daughter — adding ~$2-2.5M, a $30M Series A from Index at ~$1M revenue, and eventually the $130M round. The seed partner call happened with clothes hanging behind him: “Mike’s like, look, he’s living in a closet — and everyone’s like, ah, great founder.”
5. RAG’s creator: RAG doesn’t work
- The plot twist he delivers deliberately: Hebbia was “the first to productionize retrieval-augmented generation” in 2020 — and now “we actually don’t think RAG works at all.” The data behind the reversal: looking at deployed queries at major finance firms, “almost 90% of the questions people were asking these systems weren’t answerable by search through the documents” — they aren’t in the data, they’re about the data. His example: “is this company a good investment” over marketing materials that are “often a load of crap” — the job is distilling truth, not finding quotes.
- The industry indictment that follows: “90% of enterprise AI right now is almost like this vapor… look at this amazing demo — and the minute they actually try to use it in a real-world example, it completely fails. A lot of these usage statistics are all kind of — one of my favorite phrases is fugazi fugazi.” Hebbia’s counter-tagline: “stop experimenting with AI, start driving value.”
- On the RPA boundary — he’s “not a big believer in RPA,” calling it AI “in the old, ten-years-ago sense.” Hebbia’s queries run over 800-page credit agreements and 230-page CIMs: “tell me where there’s an event of default that we can trigger.” He calls Daniel Dines’ framing “a nice phrasing” and says Hebbia captures high-level, ambiguous decision-making, tracing it “all the way down back to individual citations.”
6. Scaling at inference — the new scaling law, and why fine-tuning always loses
- The idea he says he changed his mind toward ~18 months ago: if you can’t train bigger models fast enough, “take whatever is state-of-the-art and run it more times — hundreds or even thousands of sub-models over every single document to answer the same question.” OpenAI’s o1 recursively runs a model more before answering; Hebbia Matrix scales inference at the orchestration layer. His metaphor: not a bigger engine but “a Tesla — made of a bunch of smaller electromechanical motors that make a lot of torque.”
- The load-bearing anecdote: Bloomberg, holding “the best financial services training set of all time,” trained BloombergGPT, a GPT-3.5-class model — and GPT-4—released, he thought, a few weeks later, though he said he didn’t know the exact timeline—“just destroyed BloombergGPT at every single finance task.” The conclusion he draws is categorical: refined verticalized models “always would lose to scaling laws,” and “nothing that other players can do to fine-tune models will ever catch up” to inference scaling.
- Two honest hedges worth preserving: on data exhaustion, when Harry pushed back with video and synthetic data, he conceded “that’s a gut feel — I’m not particularly in data collection myself.” And on capital efficiency: “the cost of intelligence will go to zero” — inference cost per fixed parameter count is down “seven orders of magnitude in four years,” so “yes, we run more LLM calls than anyone might say would ever be necessary, but we have the best accuracy in the business… every single quarter our margin goes — we’re not spending money fast enough.”
- Hebbia is deliberately model-agnostic: Anthropic works better for dense legal and colloquial documents, o1 or GPT-4o for others, with tasks decomposed across OpenAI, Anthropic, and Gemini on accuracy-vs-speed tradeoffs.
7. The $100 trillion thesis — everything is undervalued, including the S&P
- Harry’s blunt challenge — “I can’t quite get my head around that” — drew the full claim: if the computer’s introduction created ~$100 trillion of stock-market value over 60-80 years; AI compute does the same for the next 60, with “more than 50% of GDP contributed by agentic applications in the next few decades” — with Harry adding, “I actually think it’ll happen faster.” It’s additional value, not substitution: “maybe I’m too techno-optimist, but all these companies are massively undervalued, including the non-AI companies” that ride the wave.
- Harry’s sharper pushback — prior transitions took 10+ years while AI adoption is “instant,” so doesn’t that change who wins? Sivulka’s answer via fire-to-torch: “encapsulating and building a useful product on top of a technology change is the thing that takes more time… if Excel was that product for compute, Hebbia has built that product for AI.” Today’s chatbots give “surface-level value — it’ll help your kid cheat on their homework, but whether something is a good investment is a much richer problem.”
8. Enterprise adoption: finance moves fastest when the alpha is real
- Against Daniel Dines’ we-underestimate-enterprise-inertia view (his neighbor and “gym buddy… the one who taught me guns”), the historical counter: Excel went to 90% market penetration in finance in 18-24 months (1985-86, everyone off the HP12C), and credit-card-data-driven investing took two years. “Finance is the slowest-moving, most lethargic leviathan… unless you’re providing outsized alpha or real value, in which case finance moves faster than any other industry. So I’m actually making a bit of a bet.”
- Harry’s worry about the lost apprenticeship — juniors never going “through the shit” of analyzing companies, leaving future decision-makers ungraduated. Sivulka’s rebuttal (“maybe it’s naive”): the partner leans on 20 remembered deals from 40 years; the junior with Matrix reads “every deal in our company’s history” and can say quantifiably “this company is 90th percentile across all this investing criteria — we should pay a 90% premium.” His claim, stated as genuine belief: AI “makes humans better,” will “increase the AUM of the firms that use it” and drive more employment — as Excel changed jobs rather than deleting them.
- The test he applies to loud AI stories: “I think it’s BS… an amazing marketing story. When you’re screaming it from the rooftops, that almost always means internally you’re freaking out about something — the behavior itself negates the content.”
9. Models will commoditize, clouds stay sticky
- The layer call (“not a hot take anymore — I’ve been saying it for a few years”): the model layer will become commoditized; value will accrue to hardware, infrastructure, and the application/agent layer. Harry’s cloud analogy — commoditized yet a great business — gets a structural rebuttal: cloud is an “OPEC oligopoly” of few entrenched players where switching costs a startup $10-20M, “whereas here it’s a very simple API key… there will be an entire industry of being able to switch models from OpenAI to Anthropic when OpenAI goes down.” He agrees clouds happily run models as loss leaders for their moats (Anthropic-Amazon, OpenAI-Microsoft).
- On Nvidia: “the best moats aren’t technological moats, they’re not data moats — they’re people moats.” CUDA is how every ML PhD learned to train, which protects training — but “for inference, what you’re using matters less,” so the macro shift to inference “will destabilize slightly the dominance of Nvidia chips,” opening the door to AMD and custom ASICs. Net position: “still bullish on Nvidia, but even more bullish on other chipmakers” — likely large tech providers and AMD rather than a new Cerebras-style generation (“chips are hard”). His one-stock answer for the AI wave: “Nvidia — or AMD rather.”
10. xAI at 50 is the buy — plus Altman, Doge, and prayer as antivirus
- Given OpenAI at 160, Anthropic at 40, xAI at 50: “xAI is the most undervalued company… xAI might overtake OpenAI and Anthropic in value over the next 12 to 24 months, which is crazy — but I think they’re all undervalued.” The case: Elon’s geopolitical position (governments among the largest AI users, energy and nuclear as the bottleneck), operational talent, and less “administrative bloat” — “if this becomes commoditized, whoever can operationalize model creation and serving fastest might start to win.” The cited evidence is xAI’s ability to build the largest GPU cluster in a short time. On Doge, though: “it will be his greatest challenge… the largest organization in the world by spend, by headcount — it’s not going to be as simple as Twitter.”
- Against SaaS’s apps-collapse-into-agents line: “I think he’s completely wrong.” Hebbia’s design question instead — “what are the apps that AGI would want to use?” An AGI would rather use Matrix to diligence a company than read thousands of documents “by hand” in a long context window. And 10,000 agent employees create “a management problem” — the better agents get, the more human-legible they must be; chat “was always a useful feature… like asking if the TI-84 was the right interface for computers.” His self-image: “Hebbia is the Bell Labs of defining AI interfaces.”
- Commercial mechanics: 90% of the market is still in the experimental-budget phase; CTOs and IT “are actually the people that know the least about the business” — business users closest to the workflow should drive it. Hebbia prices per-seat, deliberately: “when you charge for consumption you’re disincentivizing the change — you’re penalizing every time you use an AI application. What the heck.”
- The quickfire is unusually revealing: wouldn’t sell for $2B today; asked “do you trust Sam Altman?” — “no,” full stop. He believes UFOs are real and “the US government has access to” fundamentally different propulsion technology. And the trait he hid: deep religiosity in an atheistic industry — he prays an hour every morning, “an antivirus for the human mind… a lot of the best ideas I’ve had at Hebbia have come from moments of silence” — plus 10-foot oil canvases as the other channel to the subconscious.