Open AI Insider on GPT-5, AGI & the Great AI Race w/ Kevin Weil & Dave Blundin
Summary
GPT-5 was presented as OpenAI’s effort to combine frontier capabilities into one broadly useful product, while the interview left AGI unresolved. Kevin Weil called it OpenAI’s smartest model and emphasized coding, health, complex instruction-following, tool use, and agentic work; pricing came in at “less than half” the prior level. Yet OpenAI still cannot reliably predict what a model will unlock: capabilities appear “through the mist,” sometimes only after release.
Compute, not customer demand, is OpenAI’s binding constraint. Weil said the company uses essentially the same model settings as customers, remains “completely maxed out at all times,” and finds an immediate use for every new GPU—lower latency, faster tokens, wider product access, or more research experiments. Stargate’s more than $500 billion of planned infrastructure is therefore a capacity expansion into what he called “basically infinite demand for GPUs within these walls.”
OpenAI’s distribution strategy reverses conventional software monetization: expensive capabilities begin in paid tiers, then migrate toward free. Deep Research moved from Pro to Plus and eventually limited free access; India received a heavily discounted plan with roughly 10x the usage of a free account. Diamandis said models are now about 100x cheaper than GPT-4 was at launch even as intelligence increased, while Weil said paid tiers will retain the most computationally intensive work.
The startup test is whether better foundation models strengthen the product or erase it. Weil advised founders to build where current models show “little glimmers of hope,” so the next release makes the application “sing,” rather than patching a limitation that OpenAI may soon remove. His premise is sweeping: practically every scaled product, service, and device predates AI and “they’re all going to be reinvented.”
OpenAI’s envisioned AGI product is ambient, proactive, and capable of generating disposable software—not merely a smarter chat window. Weil expects interfaces to be created in real time, routine work to be completed before the user asks, and an assistant that can “see what you see” and keep “chugging away behind the scenes.” The Jony Ive question remained unanswered when the interview was cut.
Reasoning adds a second scaling axis beyond pretraining: how long a model can work while staying on track. ChatGPT reasoning may run for roughly 60 seconds and Deep Research for 20–30 minutes, but Weil sees no reason models could not work for days, months, or years; “the longer the models think, the smarter they get.” Well-specified problems such as chip layout can then turn compute into iterative gains because every candidate design can be scored.
AGI may arrive as a rising capability gradient rather than a clean threshold event. Weil argued that today’s models already outperform him in some domains and remain clearly inferior in others, while “the level of water is rising”; the hosts compared that transition with society barely noticing that AI had passed the Turing test. The measurement challenge is shifting from saturated, easily graded benchmarks toward ambiguous but economically valuable work.
The human hedge in an increasingly automated economy is purpose, personal connection, and deployment into consequential institutions. Weil rejected futures where people merely “eat grapes and write poetry and receive our UBI,” arguing that saved time gets redirected toward larger goals and that in-person connection will matter more. His parallel role as an Army lieutenant colonel reflects the episode’s broader deployment thesis: superior models provide little advantage “if they’re sitting on the shelf” while rivals integrate weaker systems everywhere.
Deep dive
1. GPT-5 trades spectacle for breadth, price, and usability
Weil described GPT-5 as probably “the most anticipated AI launch of all time,” but framed the substance as breadth: OpenAI’s smartest model, a strong coding system, and a model capable of complex instruction chains, numerous tool calls, service integration, and agentic execution “without losing the plot.”
Health received unusual product attention because users already bring ChatGPT everything from a child’s new symptom to cancer diagnoses and associated data. Weil’s boundary was explicit: “It doesn’t replace a doctor,” but a capable system available for conversation 24/7 can still help people think through information.
Diamandis contrasted expectations of an AGI reveal with the delivered product: a more useful unified system at less than half the prior price point. The hosts likened its one-product simplicity to the original Google search box—an important interface advantage for technology whose possible uses are otherwise difficult to explain.
OpenAI itself does not know the complete capability envelope in advance. Researchers may see a model “coming through the mist,” but Weil said emergent properties can surprise the team during development or even become visible only after hundreds of millions of users begin testing it in unanticipated ways.
2. Seven hundred million users turned launch day into the real evaluation
Internal and external testing cannot reproduce what happens when a model reaches 700 million people. OpenAI therefore combines feedback from Reddit, Twitter, LinkedIn, customer support, friends, and usage data; that data showed one of the company’s biggest waves of Plus upgrades around GPT-5.
Commercial strength did not negate product misses. Weil conceded that GPT-5’s initial personality felt “a little bit wooden,” and the team rapidly shipped a warmer version—an example of deployment operating as continued product development rather than a terminal release.
The hosts argued that voice has crossed a qualitative threshold: instead of abandoning a conversation after a minute or two, a user can stay engaged for hours on a drive. Weil’s children treat a 20-minute voice conversation with ChatGPT as entirely ordinary, illustrating how quickly interface expectations can reset.
The broader uncertainty is not limited to personality. Weil rejected the idea that anyone at OpenAI knows exactly what will exist three months or a year ahead: research ideas can begin working, but what those capabilities enable becomes clear only as the model and product come together.
3. Iterative deployment makes free access the destination
Weil rejected the premise that OpenAI meters out mature capabilities on a secret schedule. Its stated mission leads toward “iterative development, iterative deployment”: release as soon as a system is ready, subject to safety checks, then let users build, expose shortcomings, and feed those lessons into the next model.
Internal systems are ahead of public products, but Weil characterized them as research and development—not finished intelligence being strategically withheld. OpenAI would rather release early than wait to bestow something “fully formed upon the world,” because external experimentation is itself part of gaining mastery over a capability.
The monetization path runs opposite to conventional software. Features begin in Pro or the $20-per-month Plus tier because they are expensive; OpenAI then works to move them outward. Deep Research began in Pro, reached Plus, and eventually gave free users a monthly allowance.
Free distribution and premium economics can coexist, in Weil’s view. OpenAI will maximize the useful intelligence available without charge, while increasingly valuable, compute-heavy tasks remain paid—especially as ChatGPT moves from a reactive system awaiting prompts toward one that works proactively while the user sleeps.
4. India is the clearest test of intelligence becoming mass infrastructure
India matters to OpenAI not merely as a large market but as a youthful population with extensive latent building capacity. Weil described strong engagement among developers and API customers, while Diamandis emphasized the country’s 1.4 billion people and potential applications across education, healthcare, and governance.
Weil’s software thesis begins with roughly 30 million people worldwide who can currently code. AI coding models could expand effective software creation to 300 million and eventually three billion people; each order-of-magnitude increase in access to such a general-purpose capability, he argued, changes the world.
On the morning of the interview, OpenAI launched an India-only paid tier priced far below its standard plan and offering about 10x the access of a free account. The discount expresses the access mission, but Weil kept the constraint visible: GPUs cost money and remain limited.
For countries asking about “sovereign AI,” Weil returned to falling access costs: ChatGPT can be used free on a phone without signing in, while Diamandis said current intelligence is roughly 100x cheaper than GPT-4 was at launch. Weil endorsed the broader cost curve rather than detailing a separate sovereign architecture.
5. Competition accelerates OpenAI, but focus is its claimed moat
Weil readily acknowledged Google as a fast-moving competitor building good models, alongside Anthropic and “a bunch of other players.” Competition benefits consumers and businesses and motivates OpenAI to move faster; as he summarized a friend’s maxim, “capitalism is undefeated.”
OpenAI watches competing products because “there are smart people within these walls and there’s a lot more smart people outside of these walls.” Rival features sit alongside customer requests and observed use as inputs, but Weil said the decisive variable remains internal execution against a clear AGI mission.
Diamandis raised Google’s roughly $100 billion in free cash flow, long infrastructure history, and massive data centers as a structural vulnerability for OpenAI. Weil’s rebuttal was concentration: Google supports many products, while OpenAI has “one product” and “one mission”; building AGI is existential rather than one initiative among many.
6. Founders should build for the next model, not defend against it
Weil’s opportunity map starts with a reset: virtually every product, service, and device operating at scale was designed before AI. Adding AI is analogous to adding electricity to mechanical systems, and history suggests many incumbents will fail to perform the reinvention themselves.
OpenAI will participate in that rebuilding, but cannot cover medicine, material science, technology, and every other industry. That leaves an enormous startup surface even if the model provider expands horizontally.
The best position is at the model frontier, where an application only barely works and founders can see “little glimmers of hope.” If a new model arrives two months later and makes that fragile capability “sing,” the application was aligned with progress rather than exposed to it.
The dangerous position is a wrapper whose value consists of repairing today’s model deficiency. Weil’s blunt test: if the next OpenAI launch could “obviate the need for the thing you’re building,” do not build it. Diamandis connected this to Perplexity founder Aravind Srinivas’s phrase “AI complete”: ride the capability wave instead of being swamped by it.
7. The AGI interface becomes generated, ambient, and potentially neural
Chat remains powerful because it approximates the generality of human communication—writing, speaking, and exchanging visual context—but Weil does not see it as the only AGI interface. A sufficiently capable model should generate the most economical UI for each task in real time.
The discussion extended that vision to ephemeral software: purpose-built tools can be created for an immediate need and discarded because recreating them takes an instant. Real-time image and video generation could also visualize designs, construct scenes, or produce entertainment around the user.
Weil’s preferred metaphor is the “jewel in your ear” from Ender’s Game: an intelligence that can see what the user sees, hear what the user hears, and access broad information. ChatGPT should not wait inside an app; with the context the user gives it, it should keep “constantly chugging away behind the scenes,” acting and saving time.
BCI produced a real disagreement. Weil initially argued that AI already moves too quickly for additional input bandwidth to matter, while Diamandis argued that visual and neural access could still help. Weil later said he would personally adopt a BCI, and both qualified that interest by the question of safety. The unresolved question is whether neural access expands cognition or merely accelerates an interface humans already struggle to absorb.
8. Automation raises the value of shared culture, physical presence, and purpose
AI-generated media likely becomes hyperpersonalized, with films made for a single viewer. Weil preserved the tradeoff: individualized content is compelling, but society loses the shared experience of everyone seeing the same news broadcast or theatrical release; at the opposite extreme, explicitly human-made work may become more valuable.
The hosts proposed that the virtual world could change at “warp speed” while physical infrastructure changes more slowly. Weil’s invariant was human nature: face-to-face work, handwritten notes, and personal connection still matter, and may matter more when digital production becomes abundant.
Purpose is the dividing line between Weil’s “Star Trek universe” and “Mad Max universe.” He rejected the image of people sitting back to “eat grapes and write poetry and receive our UBI”; previous labor-saving tools did not eliminate ambition, and AI should redirect saved time toward more interesting work.
Every one of eight billion people still receives seven days a week, 24 hours a day, and 365 days a year. Diamandis called AI the greatest time multiplier; Weil’s corollary was that the human drive to pursue something larger and leave the world better “is not going to change”—AI can only supercharge it.
9. OpenAI’s operating advantage is a tight research-to-product loop
Weil estimated the research organization only loosely as “a few hundred.” Its work spans highly academic experiments that may fail or remain unseen for years, post-training tied closely to customer needs, and every point between those poles.
OpenAI’s distinctive mechanism, in his telling, is the loop among research, engineering, and product: create a capability, turn it into a product, collect feedback, and return that evidence to model development. Deep Research and agent functionality emerged from this pattern.
Responsibility is divided rather than concentrated in the CPO role: Weil oversees products such as ChatGPT and the API; Mark leads research; Greg leads data-center infrastructure and scaling. He also emphasized co-location and whiteboard work in what remains “a very research-oriented culture.”
A physicist by background, Weil had nearly joined a fusion company after Sam introduced him to companies in the field. He later contrasted his previous employers with OpenAI: Facebook and Instagram moved quickly in his experience, but “nothing in my experience compares to OpenAI.”
10. Saturated benchmarks push AGI measurement toward messy economic work
Traditional AI benchmarks are now “almost all saturated,” forcing researchers toward harder tests such as FrontierMath and ARC-AGI. Saturation demonstrates model progress, but also removes the simple yardsticks that made improvement legible.
OpenAI increasingly evaluates tasks connected to economic value: interpreting health cases, constructing a company financial model, or performing work associated with doctors, lawyers, and bankers. These resemble one definition of AGI—performing economically valuable tasks—but lack a single canonical answer.
The grading problem grows as capability broadens. Mathematics is comparatively easy to score; financial models can be built several valid ways, and creative writing has no uniquely correct output. The models are entering precisely those useful domains where evaluation becomes “softer” and self-improvement harder to supervise.
The hosts observed that AI passed the Turing test with little ceremony after it had stood for 50 or 60 years as a landmark. Weil expects AGI to feel similarly gradual: models are already far smarter than him in some areas and clearly weaker in others, but “the level of water is rising.” Diamandis cited an IQ measurement near 148 for GPT-5 Pro; he also said GPT-5 had won an IMO gold medal, while Weil noted second place in one programming competition.
11. Messy model names are a side effect of shipping capabilities early
Weil accepted “well-deserved flak” for a catalog that simultaneously contained o4-mini and 4o-mini. The confusion reflects iterative deployment: OpenAI prefers releasing a specialized breakthrough quickly rather than waiting until every capability fits an elegant universal product.
GPT-4o introduced broader interaction, including speech; the separate o1 line introduced reasoning—trying hypotheses, rejecting failures, and continuing rather than answering immediately. Early reasoning models excelled at hard scientific problems but were not necessarily the right choice for relationship advice or routine factual questions.
Specialization allowed faster observation of where reasoning worked, what users attempted, and where it failed. GPT-5 became the point at which OpenAI recombined those strands into one experience, but Weil expects future experimental offshoots followed by reintegration once the company understands them.
Reasoning also creates a second axis of intelligence scaling. Pretraining remains one axis; test-time compute asks how long a model may think while remaining coherent. With o1, o3, and ChatGPT, models may think for about 60 seconds, while Deep Research can spend 20–30 minutes gathering, identifying gaps, and returning for more evidence.
12. Closed-loop chip design can convert more compute into more compute
Weil sees no stated reason models could not work for two days, two months, or two years. Andrew Wiles did not solve Fermat’s Last Theorem in five minutes—he worked for seven years—and OpenAI continues to see evidence that longer reasoning enables harder solutions.
Chip design is especially attractive because it is constrained and gradeable. A system can propose a layout, simulate it, score the chip’s speed, and iterate; that creates the evaluation loop that open-ended creative domains lack.
The hosts argued that AI can also translate algorithmic intent into hardware, narrowing the organizational distance between model and chip teams. Weil said OpenAI is designing its own chips, using AI to improve design and layout, and working with manufacturing partners.
His most aggressive conditional claim concerned well-specified problems: apply arbitrary compute to repeated scored iterations and, “so far,” expect arbitrary improvement. He still considered the field fertile for technically strong startups, with material science another underappreciated domain for the same mechanism.
13. Education should use AI to raise the assignment, then aim at frontiers
Diamandis rejected school bans on AI because students will use it everywhere after graduation. His proposed redesign was to stop assigning eighth-grade work that AI makes meaningless and instead ask an eighth grader to tackle graduate-level challenges, such as designing a starship with AI support.
Weil agreed: assume ChatGPT exists, teach students to use it, take them deeper than a conventional classroom could, and raise expectations for the final work. He cited Ethan Mollick’s Co-Intelligence and Mollick’s practice of redesigning classes around universal AI use rather than pretending it is absent.
Even if progress froze at GPT-5, Weil believes existing capabilities would transform society over the next decade; progress, however, is unlikely to freeze. Diamandis extended that premise toward grand challenges in space, energy, and science, arguing that widespread tools let people pursue missions once reserved for kings, queens, and robber barons.
Diamandis offered his own conditional frontier forecasts: recurring Starship trips to the Moon by the end of 2026, “definitely” 2027; robot boots on Mars around 2030; Helion targeting 2028 and CFS around 2030. Weil’s broader point was that competition among dense clusters of startups and AI-empowered founders raises the probability that such ambitious programs succeed.
14. In defense and startups alike, deployment beats intelligence on the shelf
Weil said he was inducted into the Army as a lieutenant colonel on June 13 alongside Shyam Sankar of Palantir, Meta CTO Boz, and former OpenAI research head Bob McGrew. The program is designed to bring technology and the Department of Defense closer by combining industry expertise with institutional knowledge.
Diamandis argued that wearing the uniform makes the group part of the institution rather than outside consultants. Weil agreed that being inside should make them more effective; their work is more ad hoc than a conventional reserve schedule, uses need-to-know compartmentation, and divides focus areas to manage conflicts. Weil expects to concentrate partly on AI for monitoring and improving physical performance.
The strategic concern is integration speed: “It does us no good to have the best models in the world if they’re sitting on the shelf” while the PLA uses inferior models but integrates them everywhere. The comment echoed the episode’s broader iterative-deployment logic—usable systems and feedback loops matter more than latent capability.
Weil’s final advice to entrepreneurs was to “lean into the AI in every way possible” and assume the steep curve continues. Imagine what should exist but cannot yet be built, then work toward the capabilities likely to arrive in six months, one year, or two; in his categorical formulation, founders who build for that future will benefit because “the models are going to get there.”