How to Understand the Next Wave of AI Before Everyone Else | Tibo Interview
How to Understand the Next Wave of AI Before Everyone Else | Tibo Interview
Summary
- Codex has hit 20 million users on a growth curve Berman describes as suddenly vertical, after a period when Anthropic was “sucking all the oxygen out of the room.” Tibo declines to engage the rivalry directly — “I don’t tend to look at the competition that much” — and instead credits the growth to merging Codex into ChatGPT, putting the coding agent in front of a massive existing user base.
- The forecast in Tibo’s tweet that “Codex will seem primitive in two to three months” points to the laptop itself as the coming constraint. Laptops were designed around human limits (typing speed, attention, open windows); “the model doesn’t have the same constraints” and may eventually handle “100 applications opened at the same time perfectly fine,” pointing toward access to resources beyond a laptop.
- Berman frames ultra-fast mode as roughly 10–14× normal token speeds; it could shift solo-developer work from 10–15 parallel agents toward staying in flow. Berman suggests that might mean three or four agents at a time, while Tibo emphasizes designing around attention. Generation-heavy work can feel roughly 10× faster, but tool-call-heavy trajectories see only “a 3× or 4× speedup” as the bottleneck moves to network and agent overhead. Tibo expects these speeds to be “if not the default, very close to the default” in “maybe a year or two,” with a costlier tier always one level up.
- The 80% Luna price cut cited by Berman reflected planned compute capacity plus efficiency gains; Tibo does not quantify the split. Frontier models helped re-engineer the serving stack, allowing a “very significant increase in throughput” in “the same compute envelope”—an early form of recursive self-improvement. Non-ultra-fast speeds are up roughly 60% versus three months ago, and Soul is “significantly more token-efficient than Terra.” Tibo’s commitment: “not just pocket that gain.”
- On the pause Berman attributes to Sam Altman of “the absolute frontier of RL,” Tibo describes it as necessary to “harden all parts of the system” before restarting training “with full command.” The safety team reached “a fairly clear set of principles” through what he calls a debate and discovery process, amid a “huge surge” in alignment investment.
- The usage-reset button — now a physical object — is a deliberately informal goodwill practice. Berman links the ability to offer resets to capacity planning; Tibo says, “It’s not done in partnership with marketing or finance… I can press the button whenever I want.” The contrast he draws: “you can pay lip service and say that you care, or you can be like we actually care.”
- The ChatGPT-Codex merge ends in one adaptive interface for everyone: “you and your mom will use the same thing.” It will be each person’s “personal AGI.” Job labels like software engineer or designer are “just human concepts we have invented to deal with abstractions” — the interface should tailor itself, not force users to self-sort.
- The DeepMind backstory is the cultural thesis: LaMDA chat existed roughly a year before ChatGPT, but “DeepMind was not set up to ship product.” OpenAI’s bottom-up culture has “very little stop energy” and a willingness to disrupt itself “even though it might mean reallocating resources from the main gig.” Berman frames that as the contrast with Google’s inability to make the same move; Tibo concedes Google had a plan but says it was not the right place.
Deep dive
1. DeepMind built LaMDA chat a year before ChatGPT — and couldn’t ship it
- Tibo’s account of the pre-ChatGPT era: DeepMind’s language-model group had internal “LaMDA chat” with ambitions to release it publicly roughly a year before ChatGPT, but per his own tweet, Google was too nervous and DeepMind was blocked from shipping products that could disrupt Google. “DeepMind was not set up to ship products”; at OpenAI, “research and products collaborate extremely closely,” with “a big bias toward shipping.” Asked if it felt special: “the first time you realized you could get coherent text… initially it was more funny than helpful.”
- His founder advice distilled from the contrast: conviction, fast iteration from real users, and “being willing to disrupt yourself… even though it might mean reallocating resources from the main gig” — “it’s very hard but it’s super important.” He concedes fairness to Google: “They had a plan… it was all part of a big plan,” but says it was not the right place for him.
- The counterweight to bottom-up shipping: “you don’t want a hodgepodge of features” — simplicity and pride in quality, with the ChatGPT iOS app “one of the best apps out there” as the standard.
2. “Codex will seem primitive in two to three months”
- The harness critique: sophisticated users have normalized clunkiness — skill files are “hard to maintain over time,” memory is imperfect, and sub-agent networks can break “the illusion” at various points. The target is “something that deeply understands you,” a proactive “perfect little partner” that helps with day-to-day work without breaking that illusion.
- The laptop-as-constraint argument: laptops were designed around human throughput — typing speed, thinking speed, and how many apps you need open. “The model doesn’t have the same constraints”; it may eventually handle “100 applications opened at the same time perfectly fine,” so future models “will need access to more than the resources of your laptop.”
- Tibo splits the field into two problems: the personal AGI (“deeply rooted in the understanding of you as a human”) versus full-on automation — auto-patching regressions from production logs, or in cybersecurity closing a scanner-found vulnerability window “to almost zero” with humans only approving high-risk actions.
- Voice already changed his own behavior: the new voice model supports tool use, so “in the morning, I just sit there with my phone and dictate a couple of things to do for ChatGPT, and then it just goes and does it. It has access to all my tools.”
3. Ultra-fast collapses the multi-agent workflow back into flow
- Berman’s setup: current speeds push him to 10–15 parallel agents with 30–45-minute turnarounds, creating “pretty significant cognitive overhead.” He suggests ultra-fast might lead him to use three or four agents at a time instead. Tibo’s answer focuses on attention: with ultra-fast plus voice, “this thing can operate at the same speed, if not faster, than you,” so “what I was doing before, multitasking 10 agents—I don’t really want to go back to that.” The design principle is being “friendly to your attention” while letting the technology adapt to you.
- The honest caveat on the speedup: it works “amazingly well when there aren’t that many tool calls involved” — prototyping a website or video game can feel roughly 10× faster — but tool-call-heavy trajectories yield “only a 3× or 4× speedup” as overhead shifts to the network or another part of the agent stack.
- Internal allocation is telling: ultra-fast goes to incident commanders during outages (“every second matters”) and teams that “believe they are working on something very critical” — but “the vast majority is reserved for customers.” OpenAI employees are restricted because they could otherwise “gobble up all of our production GPUs.”
- On pricing: efficiency compounds through inference hardware, token efficiency, and stack engineering, so “maybe in a year or two these speeds will become, if not the default, very close to the default,” though “you will always have one tier up” with different, costlier hardware trade-offs.
4. The merge thesis: one interface, “your personal AGI”
- Why merge Codex and ChatGPT despite user pushback? “Our future models want us to be merged… it’s the same technology under the hood, the same harness” — highly multimodal and voice-first. Labels like software engineer or designer “are just human concepts that we have invented to deal with abstractions”; the end state is explicit: “you and your mom will use the same thing. It will be your personal AGI,” connected to different tools and tailored to each life.
- The “illusion” he keeps invoking is ambient, human-rooted computing — write on the office whiteboard and it should be able to understand that too. Since the new ChatGPT voice shipped, voice interaction is “growing very fast,” which he generalizes: “every time you lean into something more natural, humans just choose the path of least resistance.”
- On Anthropic — Berman notes they were “sucking all the oxygen out of the room” before something changed — Tibo declines to focus on the competition, saying he looks at “what we can do uniquely well.” The growth driver he emphasizes instead is making the technology usable by product managers, designers, sales, marketing, and communications, then distributing it “very quickly through ChatGPT where we already have a ton of users.”
5. The reset button: informal goodwill, underwritten by compute
- Origin story: compensating users when iteration broke things — “here’s some extra usage because we happened to break it for 30 minutes.” Now institutionalized but deliberately informal: “there isn’t really a whole lot of scrutiny behind it… not done in partnership with marketing or finance. I can press the button whenever I want.” Yes, there is now an actual physical button.
- Berman’s analogy — Amazon’s return policy as trust-building — lands; Tibo’s version: “you can pay lip service and say that you care, or you can be like we actually care.” Resets also mark good moments: “go explore this new thing… you haven’t used ultra yet, here’s some extra usage.”
- Berman says resets require capacity planning; Tibo says “we planned compute way ahead.” OpenAI was questioned two years ago for over-investing in compute, a bet Berman calls “one of those crazy good bets,” which Tibo affirms.
6. Soul optimizing Luna, the RL pause, and the efficiency flywheel
- Berman cites an 80% Luna price drop and a Terra price drop, then asks how much came from algorithmic gains versus strategic capacity planning. Tibo says compute was planned far ahead, while frontier models helped OpenAI “serve, restructure, or re-engineer” its stack for significant efficiency and performance gains. The result was a “very significant increase in throughput” in “the same compute envelope”—“almost not even a trade-off.” Outside ultra-fast, speeds are “60% faster than what it used to be three months ago,” and the stated commitment is to share gains rather than “just pocket that interesting gain.”
- Tibo broadens recursive self-improvement beyond models developing models: models can build “the infrastructure that is on the critical path of using those models.” Berman points to the inference stack, optimization, CUDA kernels, products, and cloud agents; Tibo says this is recursive self-improvement in some sense, but “much more infrastructure,” with the ability to point that work back at itself. “If we were not doing that, I think that would be pretty silly.”
- On the pause Berman attributes to Sam Altman of “the absolute frontier of RL” — raised alongside a passing reference to “the Hugging Face incident,” which is not explained — Tibo says the pause was “necessary to allow the teams and individuals to really understand and harden all parts of the system” before restarting training “with full command.” Rather than naming a single unpause metric, he says the safety team reached “a fairly clear set of principles” through “a debate and a discovery process,” amid a “huge surge in investment” in alignment.
- His closing pitch to the AI-nervous: Luna “would have sat at the frontier” six months ago and is now “crazy cheap.” Berman mentions a Replit-related free mode; Tibo echoes that it is being given away in a free mode. So “whatever is at the frontier now will become way, way cheaper to run in six months.” The personal case: health and finance tools he uses himself, leaving him “more informed when I go see my doctor.”