Pioneers Insight Method Research Author
AI Experts React: Elon’s Grok 4 Is Now #1 in AI —This Changes Everything w/ Emad, Salim & Dave #182
Back to Episodes

AI Experts React: Elon’s Grok 4 Is Now #1 in AI —This Changes Everything w/ Emad, Salim & Dave #182

Summary

  • Grok 4 has pushed academic benchmarks close enough to saturation that the competitive question is shifting from raw intelligence to usable agency. It scored 100% on AIME 2025, while Grok 4 Heavy reached 44.4% on Humanity’s Last Exam versus 26.9% for Gemini 2.5 and 21% for o3. Emad Mostaque’s crucial distinction: the model “is reasoning, but it’s not planning” — leaving planning, memory and agentic integration as the next bottlenecks.

  • xAI’s lead is as much an infrastructure and execution story as a model story. Founded only 28 months earlier, it scaled to a cited 340,000 GPUs after Elon Musk delivered on a seemingly implausible promise to operate 100,000 H100s; the panel put the installed hardware near $10 billion and said xAI was targeting 1 million GPUs. Dave Blundin recalled that experts said coherence at that scale was impossible, then reacted, “Oh, god dang, he did it.”

  • The training-cost mix has flipped, creating a new flywheel around synthetic, structured reasoning data. Mostaque said post-training once consumed roughly 1% of compute, rose to 10% with DeepSeek and is now approximately equal to pre-training, partly because frontier models can generate data for their successors. That does not guarantee a “more sane” Grok — mode collapse remains possible — while Peter Diamandis said model competition is becoming “an engineering and quality challenge” rather than pure brute force.

  • Inference is rapidly demonetizing even as premium intelligence may command higher prices in high-value workflows. Grok 4 was quoted at $3 per million input tokens and $15 per million output tokens; Mostaque estimated about $20 for “a million very good words” and projected equivalent intelligence could get 5–10 times cheaper annually, potentially reaching $1 per million words. Diamandis argued developers will pay materially more for marginal gains that compress engineering time, while Mostaque suspects the $300 SuperGrok Heavy plan is a loss leader for enterprise conversion.

  • Enterprise value will initially come from augmentation, error reduction and data assimilation, not instant wholesale replacement. Emad said the Arc Institute was testing Grok 4 across millions of experiment logs and CRISPR workflows; Peter cited approximate medical-study results in which AI alone beat both physicians and physician-plus-AI combinations. Mostaque nevertheless stressed that replacement remains “way off” on liability, while Salim Ismail emphasized processing scans and sensor data no human could integrate.

  • Coding, games and video expose the same opportunity: generation is arriving before planning, feedback and distribution are solved. xAI showed a first-person game produced in four hours and said a specialized coding model was weeks away; the release speaker forecast the first really good AI game and first watchable AI movie next year, with a half-hour of watchable AI television potentially this year. Mostaque’s warning for incumbents: lower production costs benefit companies, but “for the individuals working in the industry, this is terrible.”

  • If frontier models converge on one capability plateau, the durable bottlenecks become interfaces, agents, chip access and distribution. Mostaque expects Grok 5 to coordinate anywhere from 60 to 6,000 agents, use professional tools and behave like a remote worker that “just gets the job done and it doesn’t sleep.” The panel discussed Google’s roughly 3 million chips, million-chip ambitions at xAI and Meta, and supply constraints across the sector; efficient edge models such as Liquid AI could provide the everyday intelligence tier.

Deep dive

1. Grok 4 can reason at postgraduate level, but it still cannot plan

  • The release video framed Grok 4 as “better than PhD level in every subject, no exceptions” on academic questions, while cautioning that it may lack common sense and has not yet invented technologies or discovered new physics. The release speaker thought invention might arrive later this year and said he would be “shocked” if it had not happened next year.

  • Blundin called this a “golden moment” resembling Iron Man’s JARVIS: the assistant can build the suit, but the human must decide “how you’re going to save the world.” It can solve extraordinarily hard problems, yet does not independently determine what should be built or why.

  • Mostaque sharpened the limitation: “I think it is reasoning, but it’s not planning as yet.” He associated the next model scale with roughly a ronnaFLOP, or 10²⁷ FLOPs, and said improvements should continue through a combination of compute and data.

  • Diamandis asked whether postgraduate capability across every subject already qualifies as AGI, saying, “We passed through the Turing test without noticing. Are we going to pass through AGI without noticing, too?” Mostaque replied that society adapts quickly to powerful tools, but full agentic capability still needs additional building blocks.

2. Benchmark saturation moves the frontier from answering to discovering

  • AIME 2025 was the cleanest signal: Grok 4 scored 100%. Diamandis remarked, “You’re literally running out of benchmarks,” while also noting that compute, data and algorithms are still improving and that quality is increasingly the differentiator.

  • Humanity’s Last Exam contains 2,700 questions spanning domains so broadly that the best humans were estimated to score roughly 5%, perhaps 10% maximum, within the domains they understand. Mostaque gave the example, “Compute the reduced 12th-dimensional spin bordism of the classifying space of the Lie group G₂,” illustrating why no individual can match the model’s breadth. Blundin supplied a separate example from five-dimensional gravitational theory.

  • On the published comparison, o3 scored 21%, Grok 4 25.4%, Gemini 2.5 26.9% and Grok 4 Heavy 44.4%. Asked when the test reaches 100%, Mostaque answered “two years max,” probably next year. Diamandis found the measurement problem unsettling: AI may soon ask and answer questions humans cannot understand, leaving people unable to judge how quickly it is advancing.

  • The panel’s speculative prize was new science. Mostaque cited the possibility of models progressing from mathematics to physics, chemistry and biology, and said it would not surprise him if new physics appeared this year or by the end of next year. Blundin relayed Alex Wissner-Gross’s speculation that solving questions around quantum teleportation could reveal other intelligences in the universe; this was speculation, not a demonstrated Grok 4 capability.

3. xAI turned extreme scale into an engineering advantage

  • Diamandis cited xAI’s March 2023 founding and a claim that it had become the number-one model in 28 months. He recalled Musk promising 100,000 H100s by the end of that summer while listeners said “no freaking way” — then delivering them.

  • The cited cluster now contained 340,000 GPUs costing around $30,000 or more apiece, prompting a rough $10 billion calculation. Diamandis also said xAI was planning 1 million GPUs by year-end. The panel linked that appetite to roughly $1 billion per day entering AI and Jensen Huang’s projection of $1 trillion annually by 2030.

  • Blundin described the technical upset: experts said a cluster that large could not maintain “power laws and coherence.” Musk returned to first principles, created new chip connections and made it work, producing the industry reaction: “Oh, god dang, he did it.”

  • Mostaque recalled that in 2022 Amazon built his team a 4,000-A100 system, then the tenth-fastest public supercomputer, and that hundreds of chips melted during scaling. Hardware and model scaling have since become engineering problems that can be attacked directly.

4. Post-training now matters as much as pre-training

  • Mostaque described old pre-training as throwing an internet snapshot into a “giant supercomputer mixer,” yielding something like “a disheveled graduate student without his coffee.” Reinforcement-learning cleanup then consumed only about 1% of compute; DeepSeek raised that to 10%, and the Grok 4 presentation said post-training spending had reached parity with pre-training.

  • The mechanism is a data flywheel: frontier models generate structured reasoning traces and training material for their successors, reducing dependence on indiscriminate internet scrapes. Peter’s broader conclusion was that better data, algorithms and engineering quality now differentiate leading models more than simply “chucking everything into a pot.”

  • Asked whether this produces a saner Grok, Mostaque’s honest answer was “Fingers crossed.” More post-training does not eliminate mode collapse or latent-space failures; it gives the model a more structured curriculum than raw Reddit- and internet-scale data.

  • Published pricing was $3 per million input tokens and $15 per million output tokens, with a stated 56,000-token context; the API discussion separately cited a 256K context. Blundin cautioned that many advertised context windows are not fully usable, though a genuine large window could let a model process roughly 100 books of information concurrently in one pass.

5. Cheap intelligence and premium subscriptions can coexist

  • Mostaque compared Grok 4’s cost with Claude 4 Sonnet and o3 while calling it better than both. At roughly 0.7 words per token, he estimated “a million very good words that are smart” costs about $20.

  • With Vera Rubin hardware alone, he expected inference to become three to four times cheaper next year; adding algorithmic gains, equivalent intelligence might decline 5–10 times in cost annually. His endpoint: “It’ll be a buck for a million amazing words.”

  • Diamandis argued that SuperGrok Heavy’s $300 monthly price could still be compelling for code, mechanical design and other consequential work. If a marginal capability gain saves expensive engineering time or turns a wrong answer into a right one, buyers could tolerate another tenfold price increase despite commoditizing rivals.

  • Mostaque guessed the premium tier loses money, paralleling what OpenAI had said about its Pro level. He saw it as an enterprise loss leader: solve the “UI problem,” bring team data into the model through what Andrej Karpathy calls context engineering, then upsell an organization for which $300 per high-level knowledge worker is negligible.

6. Enterprise adoption begins with research and medical augmentation

  • Mostaque said the Arc Institute, a portfolio neighbor and leading biomedical research center, was already testing Grok 4 on research workflows. It could sift through millions of experiment logs and select a promising hypothesis “within a split second,” including for CRISPR research. He added that the xAI enterprise effort had started only two months earlier and that Grok would be available through hyperscalers.

  • Mostaque expects medicine to follow “augmentation first”: reduce errors, improve outcomes and only eventually replace professionals. Diamandis argued that regulation would be the main block, while Mostaque emphasized that the liability profile for full replacement remains “way off.”

  • The panel cited approximate results from a Google medical-AI study. Diamandis described a physician alone at roughly 80%, a physician-plus-AI “centaur” near 87% and AI alone in the low 90s. Mostaque recalled the physician-alone figure as 70%, referring to Daniel Kraft’s estimate that doctors give a wrong diagnosis about 30% of the time. The figures were presented as approximate, not as a single settled result.

  • Mostaque stressed that human bias can also affect the output. Ismail’s broader point was not merely beating one doctor: available scans and sensors already generate more information than a person can absorb, and AI can incorporate data that “never could have gotten into the diagnosis before.”

7. Generative entertainment expands supply, but attention and distribution stay scarce

  • The release video demonstrated a first-person-shooter game reportedly built in four hours. The release speaker stressed that core game logic is not the only hard part: sourcing textures, files and other assets has historically constrained visually convincing production.

  • The release speaker expected AI to generate art, map it onto 3D models and create executables through engines such as Unreal or Unity, probably this year or certainly next. The timing calls were a good AI video game next year, a half-hour of watchable television this year and a watchable AI movie next year.

  • Diamandis anticipated radical fragmentation: games could iterate every four hours, while friends might watch personalized versions of a film with different endings. He expected interactive media, including characters and voices that respond directly to viewers, to grow faster than passive watching.

  • Mostaque pushed back that human attention will not grow. He cited games at roughly $450 billion versus movies at $70 billion, and noted that video games had grown from about $170 billion to $500 billion while average Metacritic scores rose from 69% to 74%; the movie industry had grown much less and had an average IMDb score around 6.3.

  • Mostaque argued that marquee shared stories will survive because “distribution, distribution, distribution” still determines reach. Lower costs are good for companies and can help individual creators tell richer stories, but “for the individuals working in the industry, this is terrible.”

8. Coding’s next abstraction is context, not more hand-written code

  • The release speaker said xAI had recently trained a specialized coding model designed to be both fast and smart, with release expected “in a few weeks.” Mostaque said the existing Grok models already wrote clean code and expected the dedicated model to improve further.

  • Diamandis revisited Mostaque’s earlier prediction of “no more coders in five years,” which had generated hate mail in India. He then asked whether elite coders might simply produce 100 times more code. Mostaque answered that the next role would be “really good context engineers” directing systems.

  • Mostaque’s logic is that code is an intermediate language created because compilers could not understand the complexity of human intent. He cited Cursor reaching $500 million in revenue in a year and Anthropic roughly $4 billion, probably with two-thirds of that tied to code.

  • The remaining weakness is orchestration. Diamandis said models can already produce modules and dashboards, sometimes anticipating requirements he had not considered, but creative projects still need planning, coordination, multi-agent systems and tighter UI feedback loops.

9. Grok 5’s contest will be agents, world models and scarce compute

  • The release speaker said xAI would begin training a video model on more than 100,000 GB200s within three or four weeks. Mostaque contrasted that with the first state-of-the-art video model his team trained on 700 H100s and with contemporary leading video efforts using roughly 2,000–4,000 chips.

  • Mostaque called video models “world models”: learning visual change also teaches representations of physics, enabling 3D assets, simulated worlds and self-driving training data. He expected xAI’s system to begin as a separate model, though language, image and video systems might eventually converge into one model.

  • His Grok 5 sketch was a multi-agent system with “60 or 600 or 6,000” workers, a world model, system connectivity and command of Maya, physics simulators and Lean. At sufficient scale it becomes an “incredibly versatile worker”; interaction with Grok 5 or Grok 6 may simply look like a Zoom call.

  • Mostaque expects Gemini 3, GPT-5 and Grok 4-class systems to occupy roughly the same intelligence plateau. The differentiator becomes an interface resembling a remote colleague: message it, assign work, receive check-ins when uncertain and let it operate beyond the cited seven-hour task horizon. His preferred AGI definition is “actually useful intelligence” that “just gets the job done and it doesn’t sleep.”

  • Capital is not the primary constraint. The panel cited Google at roughly 3 million chips, million-chip ambitions at xAI and Meta, OpenAI capital and Stargate-scale infrastructure, and Amazon Trainium support for Anthropic. The next jump is 10 million chips in a world the panel estimated contains only 20 million, making access, packaging and supply decisive.

  • Diamandis noted that Apple had not entered the discussion. Blundin said Apple controls about a third of TSMC manufacturing capacity for its M2 and M3 lines and could become a major data-center player. Mostaque argued that once a model is good enough, it starts to resemble a utility, making efficient edge models increasingly important.

  • Nvidia remained the default — “you don’t get fired getting Nvidia” — but buyers will take capable chips wherever available because virtual workers are far cheaper than human teams. Diamandis said Liquid AI’s edge models run on M3-class chips and car hardware and are claimed to be about 100 times more efficient than brute-force transformers, potentially supplying everyday intelligence while scarce frontier systems handle “genius” work.