Pioneers Insight Method Research Author
Surge CEO & Co-Founder, Edwin Chen: Scaling to $1BN+ in Revenue with NO Funding
Back to Episodes

Surge CEO & Co-Founder, Edwin Chen: Scaling to $1BN+ in Revenue with NO Funding

Summary

  • Surge AI hit $1B revenue from a 2020 start with zero funding — founder Edwin built the V1 himself “in a couple weeks” right after GPT-3 launched, posted it on his blog, and was profitable from month one without a sales team. He refuses to sell for $30B or even $100B: “I definitely wouldn’t sell for 30 billion or even 100 billion… getting acquired would be this admission of failure.”
  • Edwin’s core moat claim: rivals are “body shops or body shops masquerading as technology companies” — they recruit warm bodies by resume-filtering for PhDs and have no way to measure or improve data quality. Quality control is genuinely adversarial: half of the people who graduate with a CS degree “can’t even code,” and the ones who can “are going to try to cheat you” — selling accounts abroad, using LLMs to generate the data.
  • His bottleneck ranking for AI progress: data quality first, compute second, algorithms third. Throwing compute at bad data yields “progress that actually isn’t there” — labs repeatedly discover after 6-12 months that training and eval data were junk and some models got worse. LM Arena is the poster child: voters reward emojis, bolding, and length, and the #1 model on the leaderboard insists Pope Francis is still alive.
  • Synthetic data is overrated: models trained heavily on it are good at “synthetic problems, not real ones,” and customers report “a thousand or a couple thousand pieces of really high quality human data… worth more than 10 million pieces of synthetic data.” A lot of Surge’s work is cleaning up synthetic-data damage.
  • The Scale acquisition brought Surge a wave of interest: it was an “open secret” among top researchers that Surge was “the biggest and the best in the space,” and teams still on Scale for legacy reasons produced “a massive wave of interest” — echoing Handshake’s Garrett telling Harry about a “tidal wave” of migrating Scale customers.
  • The efficiency thesis behind the whole story: at Google/Facebook/Twitter, 90% of people work on useless problems built to impress a VP, so a company 1/10th the size moves 10x faster. He believes 100x engineers exist (multiply 2-3x on speed, ideas, work rate, fewer meetings), that AI “disproportionally favors people who are already the 10x engineers,” and that a $1B single-person company “exists one day.”
  • Predictions through the quick-fire: 2028 “if you’re talking about automating the job of the average engineer,” 2038 “if you’re talking about curing cancer”; the biggest model providers may not be founded yet because “we’re only 2% or 5% of the way” to AGI; multiple frontier AGIs will coexist with different personalities; and today’s benchmark hacking is a live paperclip-maximizer problem that gets dangerous as models grow more powerful.

Deep dive

1. $0 raised, $1B in revenue — the MVP-first playbook

  • The mechanics were almost anticlimactic: having worked the data problem for years as an ML engineer, Edwin had “a very clear vision,” so instead of hiring 10 engineers or raising “$10 million, $20 million, or $30 million,” he built the V1 himself in a couple weeks in 2020 right after GPT-3 launched, posted it on his blog, and “there actually was this giant demand for the data already.”
  • Why not raise once demand exploded? “There was nothing that raising would help us with. We were very lucky to be profitable from month one.” He explicitly refused a sales team: he wanted customers who bought “precisely because they understood the value of high quality data,” since early customers shape the product — not people reached by a team emailing 10,000 prospects.
  • With today’s tooling (Lovable and likely Replit), he sees no excuse to raise pre-MVP for “90 to 95% of startups” — hardware and a few capital-heavy categories excepted. His day-one advice to himself: “focus always on the 10x improvements that you can make as opposed to worrying about 10% realities.”

2. Ninety percent of big tech is working on useless problems

  • The formative observation from Google, Facebook, and Twitter: magically remove the 90% not working on interesting problems and you get a company that’s 1/10th the size but moving 10x faster with a 10x better product — less interviewing, fewer meetings, no “updates for the sake of updates,” higher talent density, faster idea percolation.
  • His diagnosis of why: priorities at big companies are “divorced from the end customer” — you build to impress your VP for promotion. His reductio: improve an internal tool → people get 5% more productive → because they spend 10-20% of their time interviewing → because the org is “growing for the sake of growing.” Many managers’ real goal “is to tell their friends they’re a VP of a thousand-person org.”
  • The hiring filter follows: strong candidates interview him about the product (“why don’t you improve these things… what if you guys did this instead”), while the tell for empire-builders is “if I join, will I be able to hire 20 more people to support me?”
  • He runs zero one-on-ones and shows people his near-blank Calendly: “It’s almost like a negative sign if you’re having a one-on-one weekly meeting because it means that you just don’t know what’s going on with these people.” On whether revenue-per-head is now the valley’s metric, an honest hedge: “I can believe in it. I can hope for it. I don’t know if it’s true right now.”

3. 100x engineers are real — and AI favors them

  • His arithmetic for the mythical figure: some people code 2-3x faster, have 2-3x better ideas, work 2-3x harder, sit in 2-3x fewer meetings — “2 to 3x is often actually an underestimate… multiply all those things out and yeah, you get to 100.”
  • On what AI does to the distribution: it mostly removes drudgery, and since “good people have so many ideas that they just don’t have time to implement,” it “disproportionally favors people who are already the 10x engineers.”
  • The $1B single-person company: “I absolutely believe that that company exists one day” — solo startups already do $10M in revenue, and AI efficiency “multiplying 100x” gets you there.

4. Competitors are “body shops masquerading as technology companies”

  • Pressed by Harry on whether rivals were mismanaged or Surge is exceptional — “both.” Many peers “don’t have any technology”: no way of measuring data quality, no way of improving it, sometimes no worker platform at all. They resume-filter — “anybody with a PhD, they’ll just instantly hire them” — and pass the body, not the data, to frontier labs, so they can’t A/B test quality algorithms or tooling changes.
  • Quality control is adversarial in a way people underestimate: “I went to MIT, but half of the people who graduate with a CS degree can’t even code.” And the ones who can “are actually just going to try to cheat you — sell their accounts to somebody in a third world country, use LLMs to generate the data for you.”
  • The consequence for the throw-humans-at-it approach: “the teams I know who try this actually end up moving 10 times slower than anybody else without realizing it.”

5. The founding wound: Twitter’s data pipeline was two people from Craigslist

  • Building a sentiment classifier at Twitter needed just 10,000 labeled tweets, but the human-data system was “literally just two people we hired off of Craigslist working 9 to 5” — a month’s wait to start, another month labeling in a spreadsheet, and the output was junk: they didn’t understand slang (“she’s such a bad [__] … they were labeling this negative” when it’s actually really positive). He spent a week labeling tweets himself because it was faster.
  • The deeper unmet need: training the chronological-era recommendation algorithms on clicks and retweets created “this incredibly negative feedback loop… lots of girls in bikinis, lots of listicles about 10 horrifying skin diseases.” He wanted raters labeling against product principles — and if Twitter couldn’t even get sentiment right, they definitely couldn’t get richer data at the quality scale they needed. Surge launched in 2020 because post-GPT-3 “there was just so much more that you could see the industry moving towards.”

6. Raising is a status game; the risk is the point

  • His sharpest cultural critique: “people are just raising for the sake of raising… their goal is to tell all their friends that they raised $10 million” and get a headline about raising $10 million — pivoting weekly, tweeting hot takes, attending VC dinners until something lands a thousand retweets. “You’re not taking any risks. You’re just somebody looking to make a quick buck.”
  • He does believe founders should pursue ideas unique to them: a commodity idea can get “a decent medium-sized company,” but a generational one “really should be about an idea that is almost unique to you.”
  • Quality is the enforcement mechanism internally — every joiner is told “quality is the most important thing… if you have to make a deadline slip… if we have to say no to a project,” so be it. And on lowering the hiring bar under pressure: that urgent hire is “probably building a feature that nobody cares about” anyway.

7. Not for sale at $100B — the business already is the prize

  • Harry’s $30B-to-$50B escalation got a flat ceiling: “I definitely wouldn’t sell for 30 billion or even 100 billion. I already have everything I want. We’re profitable. I have complete control of our destiny… Getting acquired would be really limiting. It would be this admission of failure.”
  • What he’s in it for: AGI — “when you’re a kid you literally dream of building AI that can do all these amazing things” — and the moments when labs launch a model and “one of their first things… they’ll reach out to me and be like, hey, we couldn’t have done this without you.” That extends to 2-3 a.m. customer calls: “our models are freaking out, I need a bunch of data by 6 a.m.” — “nothing makes me happier than knowing we can deliver 10,000 data points in the next few hours.” His caveat: don’t confuse hours with value — “the best ideas come to me when I’m just walking around.”
  • His confessed blind spot is disarming: “I could not tell you what EBIT is… the difference between that and revenue and profit and net margin — I actually just don’t know any of these terms.” His dream north-star metric: are models progressing in fundamental ways, and how much of that is attributable to Surge — the closest proxy today being the variety and complexity of projects on the platform.

8. ChatGPT was the inflection; Scale’s acquisition brought interest

  • Growth was “very very strong” from month one, but ChatGPT was the inflection point “because people just saw how incredibly valuable human data and RLHF was.”
  • On the Scale acquisition: “it was an open secret where a lot of top researchers already knew we were the biggest and the best in the space,” so teams using Scale “for legacy reasons” became a source of “a massive wave of interest.” Harry’s corroboration: Handshake’s Garrett told him he was “staying up all night” under a “tidal wave of Scale customers.”
  • The onboarding pattern from converts: elsewhere they’d “spend months trying to improve the data quality for really basic stuff and it will look like it’s better for a month but then it’ll just quickly regress” — versus Surge’s principle of producing “data that you simply couldn’t get anywhere else.”

9. Data quality beats compute — and LM Arena is clickbait

  • Asked to rank the bottlenecks to progress: “data quality first, followed by compute, followed by the algorithms.” He rejects the compute-scales-everything premise outright: without the right data and objectives “you’re just going to fall into this trap of seeing progress that actually isn’t there” — labs tell him repeatedly that after 6-12 months they realized “their training data was [junk], their evaluation data was [junk]” and some models were worse than at the start.
  • The LM Arena anatomy: participants don’t fact-check — “one response will just be a complete hallucination, but because it has an emoji and a couple words bolded, people will vote for it.” The easiest way to climb is longer responses; the #1 model on the leaderboard, asked when the Pope died, insists Pope Francis is still alive and dismisses the search results as “rumors and misinformation.” Labs climbing it are “training their models to produce better clickbait.”
  • On Grok’s benchmark wins, he defers to the source: on the Grok 4 livestream “Elon himself” said the models are really good at “homework problems” — “the equivalent of making them really good on SAT problems, but not at problems people are actually facing.” Yet he finds xAI itself genuinely impressive: “it’ll be 11:00 p.m. and I’ll DM them… they’re in the office and there’s a ton of people behind them.”
  • His quick-fire safety take ties back to this: people who call AI safety overblown “ignore the paperclip maximizer problem” — accidental optimization toward wrong objectives is already happening via benchmark hacking, and “in the future when the models are more powerful… literally building the code for some trillion dollar company, the consequences can be much worse.”

10. Synthetic data is overrated, PhDs aren’t enough, and the frontier isn’t settled

  • On the perennial threat question: synthetic data has made models “good at synthetic problems, not real ones” — companies spent a year on it and are now “throwing a lot of it out,” telling Surge that “a thousand or a couple thousand pieces of really high quality human data… have actually been worth more than 10 million pieces of synthetic data” because models “collapse on this very narrow scope of similarity.” His fresh example: a 2025 frontier model “randomly outputting Russian characters and Hindi characters” mid-response — “a mistake that would be obvious to any second grader,” which is why “you always need this external value system as a safeguard.”
  • On PhD-level data: Surge has “Harvard professors and Stanford PhD students and Princeton computer science theorists” — “way more PhDs than Google or Meta or Microsoft combined doing work for us in a single day.” But “80% of the computer science PhDs I know write shitty code… and Ernest Hemingway didn’t have a PhD.” You need street smarts plus technology to find the top 1-2% — his analogy: “Vimeo has a lot of so-called high-quality videos, but they don’t have any algorithms, so YouTube’s videos are way higher quality in the end.”
  • Two frontier calls worth logging: he now expects multiple frontier AGI companies — “Claude is really really good at coding… at enterprise and instruction following, whereas ChatGPT is more optimized for consumer… Grok is willing to be a little bit transgressive” — like poets and mathematicians, no single greatest. And the biggest model providers may not be founded yet: “we’re only 2% or 5% of the way towards AGI… it’s like asking 10 years ago whether Google would be the final search engine.”
  • The timeline stakes: 2028 if you’re talking about “automating the job of the average engineer,” 2038 if you’re talking about “curing cancer” (real-world experiments take time); he pushes back on Robinhood/Benioff’s 50%-of-code claims — “if you’re really working on meaningful problems, I don’t think models today can write 50% of the code”; he “absolutely” believes in 10% GDP gains within a decade; and Google’s dilemma is that Sundar “has to be willing to take a short-term hit to all of their advertising revenue to build something better — and that’s just really hard.”