Pioneers Insight Method Research Author
ChatGPT – The Super Assistant Era | BG2 Guest Interview
Back to Episodes

ChatGPT – The Super Assistant Era | BG2 Guest Interview

Summary

  • ChatGPT is at 900M weekly actives (“about 10% of the world coming to us now, 90% left to go”), and Nick Turley allocates “all my points” to long-term retention as the north star: “things like revenue, they follow from that.” The proof point: moving GPT-4-class intelligence from behind the paywall to free (4o) was “totally revenue positive and retention positive.”
  • Historical growth splits one-third / one-third / one-third: classic friction removal (killing the login wall — “Sam will say I told you so”), core product built jointly with research (search, personalization, post-trained into the model), and pure model gains — both step changes (3.5→4, 4 paywalled→4o free) and unsplashy iteration like 5.3 and 5.4.
  • The next billion requires going beyond chat: today’s product is “a raw appliance… too much like a computer terminal,” and the unlock is to “productize reasoning in a way that works on people’s behalf without them even knowing” — long-horizon tasks users never encounter as a concept. Actions plus proactivity compound into the super assistant; the codebase is literally named “SA server.”
  • Agent timing is the tell: ChatGPT agent was “slightly too early” — without escape velocity “users don’t learn to trust it, they don’t even try” — but general-purpose agents are near the partial-credit threshold where hill-climbing begins: “on task I think we’re close.” Codex already has escape velocity (“so many engineers who don’t open their IDE like ever”), and quantitative knowledge work follows because it’s “testable… very RL friendly.”
  • Pricing will evolve: power users are getting “almost too much value,” and “having unlimited plan is like having unlimited electricity plan… there’s a reason you can’t buy that.” Ads are framed as an access play for markets without credit cards, principles (answer independence, privacy) published before the pilots scaled — and the top support inbound isn’t how to disable ads but “how do I run an ad?”
  • GPUs are the binding constraint with no end in sight: “demand keeps going up even as prices go down,” tokens-per-user charts are “mindboggling,” and planning “working backwards from GPUs is usually pretty good idea” — humans are hireable and agent-leveraged, “but GPUs are zero sum.”
  • Code Red is over, exited “which we knew we would” with 5.3 (everyday users) and 5.4 (“workhorse” for knowledge work); the premortem for OpenAI missing its mission “is probably focus,” and the biggest differentiation is the team — “anything we build will get copied.” His long idea: hands-on AI professional services inside real companies, because “we’ve saturated all the emails.”

Deep dive

ChatGPT started as a demo — retention leads

  • The origin, in Turley’s words: ChatGPT “was intended to be a demo and we were going to wind it down after a month.” Subscriptions shipped “simply because it could shape the demand… a way of gracefully turning users away” at capacity — the business model was stumbled into by solving for the user, then kept because “we had consistently more tech that we couldn’t scale.”
  • Asked to allocate 100 points across his dashboard metrics: “I care a lot about long-term retention and I would put all my points there… the sign of durable value is whether or not people are coming back in three months… things like revenue, they follow from that.”
  • The principled-decision exhibit: GPT-4 sat behind the paywall until an inference breakthrough let OpenAI give it “suddenly available to everyone” — and that free giveaway “ended up being totally revenue positive and retention positive.”

2. Why the retention curves smile — and where growth actually came from

  • On the host’s “smiling” retention chart: no single lever — for many users “it’s a multi-month process for them to understand how can this thing help me.” Search and personalization were important levers: ChatGPT “used to be a pretty worky product” with usage dropping weekends and summers; now it’s mobile-first and personal.
  • His growth attribution is “roughly one-third, one-third, one-third”: classic friction removal — removing the authentication wall was one of the highest pure-impact moves (“Sam will say I told you so… that was his feedback from day one”) — core product investments where research and product “came together” to post-train search and personalization into the model, and model improvements.
  • Model gains include both step changes (GPT-3.5→GPT-4; paywalled GPT-4→4o free) and “iteration that isn’t splashy, that doesn’t warrant a named release” — “I’m really excited about the updates we just made with 5.3, 5.4,” which methodically address user feedback and show up in retention.

The next billion: beyond the computer terminal

  • Today’s product is “a raw appliance… a power tool — it doesn’t tell you what it’s for,” discovered via prompts on Twitter and Instagram. “Everyone in the world has intelligence-constrained problems… but you need to frame that to people. We’re a little bit too much like a computer terminal and it needs to feel more like an operating system of software.”
  • The bet he’s most animated by: reasoning today “is relevant to a very small group of people,” but “if you can figure out how to productize reasoning in a way that works on people’s behalf without them even knowing — the model doing long-horizon tasks on your behalf — it doesn’t mean you encounter the concept, it just means it’s benefiting you.”
  • Also on the goal sheet: going deeper with the existing billion — “actually helping them achieve their goals, not just answering questions” — because “we’re going to go beyond pure chatbots pretty fast.”

4. Actions + proactivity compound into the super assistant — and timing was the miss

  • On why agents haven’t landed yet: ChatGPT agent “was just slightly too early. The models weren’t quite good enough to hit real escape velocity — and the problem is users don’t learn to trust it, they don’t even try.” Only niche things worked (migrating a file server to the cloud). The flywheel starts when “it works well enough that you get at least partial credit — because you’re getting partial credit you get really good tasks back and then the magic begins.” His hedge, kept as hedged: “on task I think we’re close,” but “even people inside OpenAI would have had a hard time predicting exactly when this gets good.”
  • Pulse was the first inversion — “you’re not prompting the model, the model’s prompting you” — but it’s “limited in the value it can provide because it’s not connected to your life and it can’t take action.” With both pieces: “hey, you just landed… I’m going to call a cab for you,” or at work, “I proactively ran this analysis because I saw your metrics dropped.” Even fitness becomes agentic — prompting the host’s quip, “You’re going to give Ozanic a run for the money” (likely Ozempic).
  • Domain-specific agents already work: Codex has escape velocity — “we’ve got so many engineers who don’t open their IDE like ever.” Next up, he wouldn’t be surprised to see “other forms of quantitative knowledge work… it’s testable, you know if it worked or not — very RL friendly.” Brad described Deep Research as “a consumer product” and “our first agentic thing,” adding that “what consumers want is I can just ask it anything” without retraining.

5. Chat is the intent layer, not the deliverable

  • The codebase name gives away the vision: “SA server” — super assistant server — “proof that this was always the vision.” Chat “is a great way of expressing your intent… but it’s not a great output. What you want back is an artifact — here’s your plan for your trip, here is the analysis… I just made you five bucks.”
  • The behavioral shift he flags as underestimated: ChatGPT is “increasingly a true thought partner… a sparring partner” — from relationship advice to work analysis, “a second brain of sorts” — trending toward “a teammate in the workplace and a super assistant at home.”
  • The host’s own high-stakes use case — a new baby crying at 3am — draws the episode’s warmest exchange: “incremental hours of sleep” as north star. Turley: “spiritually that is pretty close to what we hope we can do — help you reach whatever you consider self-actualization.”

6. Power users do the product discovery — and pricing will price like electricity

  • Build for the extremes: the busy non-user “forces you to really nail the interface,” while power users “teach us what’s possible… it’s actually impossible for us to do all the product discovery on our own” given how empirical the tech is. Brad’s macOS example frames the design goal, “where the complexity is progressively disclosed” — magical for novices, terminal and knobs for developers.
  • On power users getting “almost too much value”—the host described some as getting thousands or tens of thousands of value from a $200 subscription: “there’s no world in which pricing doesn’t significantly evolve… it’s possible that in the current era having unlimited plan is like having unlimited electricity plan — there’s a reason you can’t buy that.”
  • On ads, given Sam’s historical reluctance: the north star is access — subscriptions fail where “people don’t have credit cards” — and the principles were set before the pilots: “it’s very important that the answer of ChatGPT be independent,” plus privacy. The tell from support data: “the most common inquiry about ads is not how do I disable ads — it’s how do I run an ad?” Same ecosystem appetite shows up in shopping (organic, works, but needs visual discovery) and partnerships, which he’ll only do if the experience is “truly awesome” and accretive.

7. GPUs are zero-sum — and there’s no line of sight to that ending

  • On allocation between ChatGPT, Codex, and research: “I’ll let you know when I figure it out — just kidding.” The pain is real user demand you can’t serve — “if you only ever worked in software, that’s an entirely unusual dynamic.”
  • He rejects the “naive business-school” incremental-revenue-per-GPU approach because breakthrough capabilities are zero-to-one: “we couldn’t have told you there’s going to be consumer demand for a research product — but if you don’t productize it, you will never know.” Research funding sits with Mark (likely Mark Chen).
  • The planning heuristic: humans can be hired and agent-leveraged, “but GPUs are zero sum… starting working backwards from GPUs is usually pretty good idea.” And the constraint isn’t easing: “demand keeps going up even as prices go down,” tokens-per-user charts are “mindboggling,” and internal employee usage “is a pretty good indicator for what’s about to happen.”

8. Code Red is over — focus endures

  • The context, as the host framed it: Google had a great model, “likely Marc Benioff switching very vocally to Gemini,” OpenAI delaying ads, health agents, and shopping. Turley’s version: a tool “to create focus… we need to show up for our users” on basics — reliability, performance, “the way that talking to the model feels,” personalization. “We just exited the code red — which we knew we would — with the launch of 5.3, a great model for the everyday user, and 5.4, a workhorse if you’re trying to do real knowledge work.” Not the new normal, but a tool he’ll reuse.
  • The competitive frame worth keeping: “if you were to premortem why a company like OpenAI does not achieve its mission, it’s probably focus.” And the key differentiation: “the biggest differentiation of ChatGPT is the team behind it, because we’re not static — anything we build will get copied.”
  • On hiring Peter of OpenClaw (likely Peter Steinberger): the project was “so inspiring” — AI that is “fully embodied, exists across different UIs, has state, has an interaction pattern that feels a little bit more like talking to a human… very curt” texting. “There’s a lot more to come.”

9. Rapid fire: long hands-on AI services, curiosity, and writing

  • His long: companies “going into companies… doing effectively professional services with AI — because we’ve saturated all the emails and you need to get proximate to the problems.” The logic: labs made so much progress on math and coding because those are the domains lab people are proximate to; “if you get proximate, you can build something transformative… the obvious problems have been solved by the models.”
  • Credit to a rival: “NotebookLM is awesome and differentiated.” The underrated AI capability it exemplifies: “transform things into a different medium” — text to visual, soon visual to video — echoed in ChatGPT’s new dynamic math blocks.
  • For students: “the most important perma skill in this era is curiosity — if the machine can answer all your questions, you better have good questions.” The job that gains value: entrepreneur is the easy answer; the non-obvious one is writing — “it forces you to be very clear of what you have to say,” and even as prompt engineering dies, “expressing what you want to a machine requires you to be a very precise writer.” Plus “a permanent need for high-quality, trusted, authoritative content.”

10. Feeling the AGI — repeatedly, and it doesn’t wear off

  • His most instructive moment was a negative one: freshly trained GPT-4 “didn’t impress me at all… because we hadn’t figured out how to post-train it” — then it became a step change. “Profoundly humbling — it might not look like we are close to really powerful, useful AI, but we probably are.” The GPT-4 shockers: poetry, compiling code, and simulating “an entire computer terminal.”
  • The best anecdote: demoing reasoning to the whole company, the streamed chain of thought swore — “Oh, damn it, maybe I have to adjust because I realized I had made a mistake in the puzzle” — “entirely emergent from the RL process… completely blew my mind.” Most recently: Codex users “walking around with their computer open because they don’t want the task to end.”
  • His closing epistemics on timing, verbatim and worth pricing in: “it’s quite possible to predict where things will end up… but it’s really hard for me to make statements on anything between sort of eventually and in three months.”