OpenAI vs. Grok: The Race to Build the Everything App w/ Emad Mostaque, Dave Blundin & AWG | EP #199
Summary
OpenAI’s reach is becoming a two-sided control point: 800 million weekly ChatGPT users on one end and massive compute demand on the other. Developers doubled from 2 million in 2023 to 4 million, while API throughput jumped from 300 million to 6 billion tokens per minute. Alexander Wissner-Gross annualized that to 3 quadrillion tokens and projected 30 quadrillion next year—approaching humanity’s estimated 50 quadrillion spoken annually—while GPUs, and separately energy, remain binding constraints.
The everything-app contest is fundamentally a battle for finite attention, with an app-store phase that may be transitional. OpenAI’s Apps SDK puts Booking.com, Figma, Coursera, and Zillow inside ChatGPT, while Meta, Google, and X pursue the same conversational real estate for their agents. Mostaque’s progression is investable shorthand: consumption became cheap, creation is becoming cheap, and “the valuable thing is curation and attention.”
Agentic software development is crossing from code assistance into recursive production. OpenAI said Agent Builder was completed in under six weeks with Codex writing 80% of its pull requests; Mostaque noted the Codex CLI receives two updates a week and interpreted Dario Amodei’s claim that 90% of code would be written by AI as meaning it can be written by AI. Visual workflow boxes are viewed as transitional because “code is just a human translation layer”; the end state is a voice-and-image Jarvis that can explain its own continuous changes.
Sora 2 turns generative video into a product-design API, with pricing already exposing the labor-substitution curve. Mattel’s demo converted a sketch into a photorealistic toy video with apparent physics, while Alexander Wissner-Gross called it “mechanical design getting solved.” At $0.10 per generated second, or $360 per hour, his assumed 10x annual cost deflation would make API-based design dramatically cheaper—provided compute supply catches demand.
OpenAI’s $20 subscription faces compression even as its installed base expands. Mostaque said breakthroughs from DeepSeek, Grok 4, and others have cut token costs 20–30x, reducing a basic chat experience from roughly $200 a year to “a couple of bucks a year.” That forces OpenAI toward advertising, commerce, likeness-driven media, and economically valuable agent workflows while it conducts a global user “land grab.”
The AI capex trade is spreading from GPUs into the entire industrial stack. OpenAI’s 6 GW AMD agreement follows 10 GW with Nvidia; at Mostaque’s estimate of $50 billion per gigawatt, that is roughly $800 billion of buildout. BlackRock’s reported $40 billion pursuit of 78 data centers totaling 5 GW, plus Corning optics, liquid cooling, valves, power, and fab inputs, shows the breadth—but Blundin says the calculable ceiling remains TSMC, Intel, and Samsung manufacturing capacity.
Digital computer use and embodied autonomy are converging into one labor platform. Anthropic was projected to reach superhuman OSWorld performance within months; FSD 14.1 adds 10x more AI parameters and neural-network routing, while Gemini Robotics-ER 1.5 and Optimus point toward common vision-language-action stacks. Mostaque put the tipping point in “the next like six months,” and Wissner-Gross supplied the recursive endgame: robots building robots, then data centers, which produce digital superintelligence that improves the materials and energy efficiency of the whole system.
Deep dive
1. Token production is approaching a human-scale crossover
Sam Altman’s Dev Day comparison set the scale: from 2 million weekly developers, 100 million weekly ChatGPT users, and 300 million API tokens per minute in 2023 to 4 million developers, more than 800 million users, and over 6 billion tokens per minute. His pitch: “It has never been faster to go from idea to product.”
Dave Blundin called the 300-million-to-6-billion increase the standout number and argued demand will accelerate as individual developers consume 10,000-plus tokens while coding. His verdict on the announcements was deliberately expansive: “Some of the most staggering things in human history got announced yesterday and they still undersold it.”
Alexander Wissner-Gross translated 6 billion tokens per minute into roughly 3 quadrillion annually, against an estimated 50 quadrillion tokens spoken by all humans. He expects OpenAI to reach 30 quadrillion next year and potentially surpass annual human speech the year after: “The number of AI tokens coming into the world is about to overtake humans.”
Wissner-Gross called the current OpenAI-generated-token share roughly 6%, noting about 4 billion smartphone users still lack access to this “superintelligence.” A 5x–7x expansion could underpin “transformative economic changes at a planetary scale,” while Mostaque identified the immediate token bottleneck plainly: AI can produce economically valuable tokens without a human limit, except for the GPUs.
2. Conversational attention is becoming the new operating system
ChatGPT’s Apps SDK lets users address services such as Booking.com, Figma, Coursera, and Zillow inside one conversation. Diamandis framed it as the ultimate interface: ask for a trip, diagram, lesson, or property search without navigating the underlying applications.
Mostaque’s framing: “Attention is all they need.” OpenAI, Meta through WhatsApp and Instagram, Google, and eventually Musk through X all want to own the conversation from which MCP-enabled agents execute tasks. “Everyone’s trying to be the everything app.”
The economic sequence matters: consumption was expensive and became cheap; creation was expensive and is becoming cheap. Mostaque therefore sees curation and attention—the pixels seen and sounds heard—as the remaining scarce assets through which these platforms will monetize.
Wissner-Gross accepted the app-store analogy as a natural market movement but called the current phase transitional because “every pixel is going to be generated.” He compared today’s composable ChatGPT canvas to Apple’s 1987 Knowledge Navigator concept: “We caught up with the future approximately 40 years later.”
3. Cheap creation points toward autonomous corporations
OpenAI’s staged example began with a dog-walking idea, generated imagery and a name, then asked Canva to turn it into a fundraising deck. Wissner-Gross extended the workflow to its logical destination: “the multi-trillion-dollar endgame” is an autonomous corporation.
His personal specimen was less theatrical: stopped at a Cambridge traffic light, he used AI to create a business plan and begin recruiting a Princeton team. Diamandis described the larger trajectory as “mind to materialization”—state an intention, then let software assemble the business and eventually route its revenue.
Diamandis cited a tweet claiming that the platform launch had eliminated “a million different startups.” Blundin rejected it outright, citing Replit founder Amjad Masad’s willingness to discard an early foundation model while retaining the team: “You’re not reinventing your business constantly, you’re dead on arrival.”
4. Visual agent builders are a bridge to voice-native software
OpenAI demonstrated Agent Builder by giving the presenter eight minutes to build and ship an agent through visually connected workflow nodes. Wissner-Gross saw value as an enterprise safety net, but compared the format to early filmed vaudeville: a new medium temporarily reproducing the conventions of the old one.
Mostaque was blunter: “Where we’re going, we don’t need nodes and spaghetti.” Within one or two years, he expects users to converse with a Jarvis-like agent that generates whatever Kanban board, Gantt chart, workflow, or application is useful at that moment. A recent Claude release already hinted at applications programmed on the fly.
Blundin’s pushback—worth keeping—is that telling Claude or Cursor to upgrade an account should not produce instructions for navigating settings when the agent has MCP access and can simply act. “The whole interface to AI is going to be voice,” supplemented by images, not boxes connected with lines.
An OpenAI employee said Agent Builder itself was produced in under six weeks, with Codex writing 80% of the pull requests. Mostaque added that the Codex CLI receives two updates weekly; Wissner-Gross carried recursive improvement one step further to “negative speed where the software is just written preemptively.”
5. Codex is beginning to connect language with the physical environment
In the demo, Codex inferred an Xbox controller’s joystick mapping, inspected an audience through a camera, and moved stage lights in response to speech. Diamandis saw a specific product opportunity: an AI that makes audiovisual systems foolproof by connecting screens, calls, lighting, slides, and live internet material conversationally.
Blundin argued such a product was not hypothetical: a team could release it “in like six months or less,” and every stage presentation would become its own demonstration. Voice and gesture recognition could replace backstage signaling while making the interaction visible to the audience.
Codex’s expanding name also illustrated a branding problem. Wissner-Gross described it as an umbrella for a coding model, web-based software agent, and command-line tool; Diamandis said exponentially expanding AI capabilities were pushing products toward thematic “grab bag” brands.
Diamandis’s desired endpoint was one AI that assumes anything is doable and hides every backend call. Mostaque’s assessment was emphatic: the physical Iron Man version might be one to three years away, but “within the virtual world, that’s today”—the components exist even if nobody has productized the whole.
6. Sora 2 makes product design a deflationary API call
OpenAI previewed Sora 2 in the API through a Mattel workflow: a hand sketch became a photorealistic video of a Hot Wheels- or Matchbox-style toy descending ramps. Diamandis initially missed the point because the synthetic object and its physics looked real: “That toy doesn’t even exist.”
Mostaque said storyboard-like inputs can specify changes scene by scene, yet the model does not explicitly decompose the task; it “literally interpolates the concept to the video.” With matching audio, 3D extensions, and adaptation, these systems are “genuinely world models,” though he stressed that users are only scratching the surface.
Diamandis extended the interface from sketch to verbal specification: describe a handled container for hot liquid, alter its dimensions, compare cost and insulation, then request manufacturing and online distribution. His best consumer example was a child describing a unique toy with a parent before an N-of-1 manufacturer prints it.
Wissner-Gross’s economic call was categorical: “This is mechanical design getting solved.” Sora 2’s base pricing was $0.10 per second, or $360 per generated hour; assuming 10x year-over-year cost deflation, he expects design outsourcing to become cheaper than human labor. Today’s five- or six-minute generation wait still reveals the compute bottleneck.
7. OpenAI must monetize workflows as basic chat commoditizes
Mostaque argued that breakthroughs from DeepSeek, Grok 4, and others cut per-million-token costs 20–30x, leaving “the basic chat experience” insufficient. He estimated that an experience costing roughly $200 annually just over a year ago now costs a couple of dollars, while valuable workflows expand from 2,000 tokens through 20,000 and 200,000 toward 2 million.
OpenAI therefore needs advertising, commerce, Sora likenesses, and agentic verticals capable of supporting much higher token consumption. Google and Meta already possess advertising cash flow; competitors can increasingly offer for free what began the year as OpenAI’s $20-per-month flagship product.
Diamandis described expansion into India, the UK, Greece, and elsewhere as a global land grab conducted alongside Chinese open-source models. Mostaque’s strategic reading was that Sam Altman is securing the two “foundational points of control”: an installed base of users and enormous compute, trusting everything between them to fill in.
8. Rival models are attacking both computer control and attention
Anthropic’s progress was presented through OSWorld, a Salesforce-originated benchmark covering hundreds of everyday tasks across Ubuntu, Windows, and macOS using screenshots, keyboard, and mouse. Wissner-Gross projected that its straight-line progress could cross human performance by year-end or within the next few months.
When Diamandis said Perplexity Comet had given him a vague, “complete garbage” explanation, Wissner-Gross clarified the benchmark as practical cross-application computer control. Mostaque called its roughly 360 tasks evidence that generalist models, reinforced through richer environments, are becoming capable of most standard human digital work: “This is the takeoff point.”
Diamandis answered that takeoff with “what could possibly go wrong”; Wissner-Gross deliberately reversed the framing to “what could possibly go right.” Blundin added that the same digital control is already being applied to science factories running hundreds or thousands of experimental devices continuously.
Grok Imagine’s move from version 0.1 to 0.9 emphasized 15-second clips, “speed and fun,” and gaming. Wissner-Gross proposed video as a first-class reasoning modality—models imagining clips within their chain of thought—while Mostaque contrasted OpenAI’s $4.3 billion first-half revenue with gaming’s $200 billion last year. Whether the games will be good remained explicitly uncertain.
9. Compute demand is reprogramming the industrial base
OpenAI’s agreement to deploy 6 GW of AMD GPUs moved AMD shares by roughly 30%, according to Blundin, and would give OpenAI 10% ownership upon milestones “for basically no price.” He viewed AMD’s access to TSMC manufacturing capacity as the deeper strategic prize.
Mostaque combined that 6 GW with 10 GW they are doing with Nvidia and estimated $50 billion of buildout per gigawatt—about $800 billion. To absorb it, he expects OpenAI and peers to sell fully autonomous workers priced from $10,000 through $100,000, effectively pursuing the whole software-labor market.
BlackRock was reported to be considering buying up to 78 data centers totaling 5 GW for $40 billion, including brownfield sites such as a former Ohio coal plant. The downstream beneficiaries discussed included Corning optics, silicon, glass, photonics, liquid cooling, valves, and other infrastructure. Alex Wissner-Gross said one operator had bought a million valves to isolate leaks around high-value 1U equipment.
Blundin’s limiting model was “infinite demand” bounded by chip fabs whose capacity can be seen four years ahead at TSMC, Intel, and Samsung. Wissner-Gross’s condition was different: capex can continue while AI drives service costs toward zero and produces valuable discoveries. He described colocated natural-gas, SMR, and future fusion plants as a likely off-grid path; Diamandis separately raised gigawatt data centers in space powered by solar.
10. Embodied AI closes the loop between labor and compute
Tesla’s FSD 14.1 arrived with 10x more AI parameters, neural-network navigation and routing, obstacle-aware detours, and selectable arrival behavior; Musk’s phrase was “V14 feels alive.” Wissner-Gross expects convergence with the robotaxi stack and, over two to three years, an end-to-end vision-language-action model shared with Optimus.
Gemini Robotics-ER 1.5 was described as outperforming GPT-5 on embodied reasoning and pointing accuracy. Mostaque emphasized the efficiency curve: tasks that required high-performance, roughly 1,000-watt chips a couple of years ago may move to edge compute, enabling specialized models and hardware that can learn broad physical tasks.
DeepMind’s ASIMOV safety benchmark delighted Wissner-Gross because it compares Isaac Asimov’s three laws of robotics against alternative constitutions for embodied models—and, according to his reading, finds better ones. The conversation’s concrete deployment was a recycling company called Andy Systems, using vision to recover valuable metals for the next generation of chips.
Diamandis reported that Musk had said the showcased Optimus kung-fu session was autonomous rather than teleoperated; no countervailing evidence was offered in the discussion. Mostaque predicted a learned skill such as kung fu might occupy only a few megabytes; Wissner-Gross said roughly two-thirds of service labor connects to physical work and called the industry “painfully close” to unlocking it.
The closing loop tied labor back to infrastructure: reaching 250 GW by the early 2030s may require robots to construct the data centers. Wissner-Gross’s “innermost loop” was recursive self-improvement—robots building robots first, those robots building data centers, and the resulting digital superintelligence improving the materials and energy efficiency of both robots and data centers—while Mostaque placed the broad tipping point “in the next like six months.”