No Priors Ep. 124 | With SurgeAI Founder and CEO Edwin Chen
Summary
Surge says it surpassed $1 billion in revenue last year with a little over 100 employees, five years after launching in 2020, serving clients including Google, OpenAI, and Anthropic. Chen credits being profitable from the start: capital was not Surge’s bottleneck, so he avoided giving up control for fundraising as social proof and argues founders should “go out and build whatever you’re dreaming of” before raising.
Surge’s stated moat is not labor supply but a technology-mediated quality system for data whose ceiling rises with generative complexity. Hemingway and a second-grader can draw essentially the same bounding box, but poetry, proofs, code, and games admit outcomes with radically different quality. Without systems that measure those differences, vendors are merely “scaling up mediocrity.”
Surge’s work spans SFT and preference labels, verifiers, failure-mode analysis, and rich RL environments that simulate work over long horizons. Chen’s salesperson world spans Salesforce, Gmail, Slack, spreadsheets, documents, presentations, calendars, and even a car accident that changes meeting travel. He sees “no ceiling” on useful realism and doubts one terminal reward can capture a complicated trajectory.
Chen is categorical that human feedback “will never run out,” even if models become superhuman, because models still need external objectives and synthetic volume routinely fails quality tests. Customers may spend six months generating 10 million-20 million synthetic items only to find “99% of it just wasn’t useful”; he says 1,000 highly curated human examples can outperform 10 million synthetic ones.
LMArena can reward clickbait rather than capability, creating an incentive to make answers longer, more formatted, and more emoji-heavy. Five-to-10-second raters often do not check factuality or instruction-following, and Chen says researchers have accepted regressions in both to improve leaderboard rank. His alternative is expensive but direct: careful human evaluation with fact-checking, instruction verification, and taste.
Chen expects a plural frontier-model market rather than commodity convergence, and picks xAI as the underdog most likely to catch OpenAI, Anthropic, and DeepMind. Anthropic’s coding and enterprise focus, OpenAI’s consumer orientation, and Grok’s different behavioral boundaries create distinct products; eventual AGI “may encompass this all,” but companies can only sustain so many priorities at once.
The Meta–Scale deal increased Surge’s visibility among legacy Scale users, while public model evaluation is its next strategic expansion. Chen says those users had not known about the under-the-radar company, and Surge wants to expose benchmark-driven failure modes as frontier labs publish less. The larger thesis: independent quality measurement becomes more valuable as model providers optimize against increasingly gameable public signals.
Deep dive
1. Bootstrapping preserved control because capital was not Surge’s constraint
Chen traces Surge to the data bottleneck he repeatedly encountered doing ML at Google, Facebook, and Twitter: teams could barely source what they needed for a simple classifier, let alone the “next generation AI systems” they imagined. Founded in 2020, Surge reached its fifth anniversary with just over 100 employees and, Chen says, more than $1 billion in revenue during the prior year. Sarah says it serves clients including Google, OpenAI, and Anthropic.
Bootstrapping was initially practical: Surge was “profitable from the start” and did not need money, while fundraising meant surrendering control. Chen’s sharper complaint is cultural—too many founders optimize for a $10 million round and TechCrunch headline before identifying a problem. His first-principles default is to build first, then raise if an actual financial constraint appears.
Elad’s pushback—worth keeping: Silicon Valley may raise too readily, but founders elsewhere often underuse venture capital when it could unlock scale; Sarah adds that unknown founders may need external validation to recruit. Chen distinguishes founders who need subsistence capital from those with savings, then challenges assumed headcount: data scientists tune 2%-5%, while early startups need 10x-100x swings, and founders—not an early PM—should own product.
2. Generative data turns quality from a checkbox into the product
At the simplest level, “our product is our data”: for coding models, Surge supplies SFT solutions and unit tests, preference judgments between code or explanations, and verifiers such as whether a web app contains a login button or triggers the intended action. Evaluation accompanies training data through comparative model insights, loss patterns, and failure modes delivered back to labs.
Chen separates this from competitors he calls “body shops,” whose deliverable is “warm bodies” rather than measured output. His analogy illustrates the quality ceiling: Hemingway and a second-grader can draw essentially the same bounding box around a car, but Hemingway can write a much better poem—generative AI has “almost an unlimited ceiling” on quality.
Asked how Surge can evaluate enough evaluators, Chen compares its system to Google Search or YouTube ranking millions of pages and videos. Surge collects signals about annotators, their work, and their activity, then feeds them into internally built ML systems to measure quality.
The eight-line moon poem exposes why credentials and checklists fail: “is it a poem,” eight lines, and the word moon can all pass while the writing remains high-school-level. English-literature PhDs are not automatically poets, just as some MIT CS graduates are poor coders. Quality must embrace haiku, internal rhyme, emotion, and “a thousand ways” to prove the Pythagorean theorem—or the vendor is “scaling up mediocrity.”
3. Humans and models must co-produce data as RL worlds grow realistic
Rising model baselines do not remove humans; they change the interface through “scalable oversight,” which Chen defines as humans and models producing data better than either can alone. An SFT story that once began from a blank page can now start with a model’s generic skeleton, leaving the person to perform substantial creative editing instead of low-value “cruft.”
RL environments make that collaboration materially harder. Chen’s salesperson simulation spans Salesforce, Gmail leads, Slack conversations, Excel tracking, Google Docs, PowerPoint, and a calendar, then adds evolving time and external events—a car accident should alter travel to a customer meeting. Thousands of messages and hundreds of emails must remain realistic, interesting, and mutually consistent.
On realism, Chen’s answer is categorical: “there’s no ceiling.” More diversity, richness, and longer time horizons give models more to learn from; synthetically generating an environment is insufficient if the resulting world lacks coherent events, creativity, or the tools needed to represent an entire job.
His five-to-10-year demand forecast is “all of the above,” not an RL-only substitution for other data. Environments can create extremely long, rich trajectories, but a single terminal reward may not capture all the work involved in achieving a complicated goal. Chen expects expert reasoning, environments, verifiers, and multiple rewards to remain complementary.
4. Synthetic scale and public leaderboards can optimize the wrong objective
Chen is categorical that human feedback “will never run out.” Customers sometimes arrive after six months and 10 million-20 million synthetic items, only to conclude that “99% of it just wasn’t useful” and search for the small usable slice. His deliberately extreme comparison: 1,000 highly curated human examples can be more valuable than 10 million synthetic points.
The need is not merely for better examples but for an external objective. Chen cites a top frontier model that, in roughly 10% of his uses, inserts random Hindi or Russian characters into otherwise unrelated answers about subjects such as Donald Trump or Barack Obama. The model is not self-consistent enough to flag the error, so a person must still say, “this is wrong.”
His indictment of LMArena is that raters spend five-to-10 seconds choosing whichever response “looks better,” rewarding formatting, bolding, emojis, and length without checking factuality or instruction-following. Training to that signal becomes the model equivalent of clickbait; Chen says the easiest way to improve rank is simply to make responses longer.
The alternative is slow, proper human evaluation: fact-check the answer, verify every instruction, and use people with taste to judge writing. Sarah’s cross-domain analogy sharpens the risk: protein evolution can select bizarre, unanticipated activities against a narrow catalytic signal, just as model training drives toward a local maximum shaped by whatever feedback designers chose.
5. A plural model market raises the value of independent evaluation
On the Meta–Scale deal, Chen says the effect was beneficial for Surge, which he describes as already the largest player: legacy teams that had not heard of the under-the-radar company became aware of it. His broader concern is category damage—low-quality vendors burn labs on human data, pushing them toward slower methods with poor objectives and ultimately slowing model progress.
Asked which underdog could catch OpenAI, Anthropic, and DeepMind, Chen chooses xAI because it is “very hungry and mission-oriented.” He expects more frontier providers rather than commodity convergence: Anthropic has excelled at coding and enterprise, OpenAI carries a consumer orientation through ChatGPT, and Grok has different boundaries around what it will say and build.
The mechanism is organizational focus and model personality: companies can prioritize only so many principles at once, producing different skills and behavior. Chen already switches among models depending on the task and expects that habit to deepen across personal and professional life, even while conceding that eventual AGI “may encompass this all.”
Surge’s next strategic expansion is public research and external model evaluation, partly because frontier labs publish less. Chen says researchers have explicitly told him their VPs accept worse factuality and instruction-following if LMArena rank rises; six months of longer, flashier answers can therefore equal “zero progress.” Meanwhile, IFEval rewards contrived tasks such as capitalizing five letters in every mention of Abraham Lincoln rather than demonstrating real-world usefulness.