Grok 4 Wows, The Bitter Lesson, Third Party, AI Browsers, SCOTUS backs POTUS on RIFs
Grok 4 Wows, The Bitter Lesson, Third Party, AI Browsers, SCOTUS backs POTUS on RIFs
Summary
- Grok 4 tops the benchmarks — Artificial Analysis now ranks it above OpenAI’s o3-pro and Gemini 2.5 Pro, less than two and a half years after starting in March 2023. Chamath reads it as vindication of Rich Sutton’s Bitter Lesson: general compute scaled on Colossus (100k → 250k → 1M GPUs) beats human-knowledge approaches — and puts a question mark over Llama’s $15B for 49% of Scale AI, “exactly a bet on human knowledge.”
- Human-labeled data has a short half-life, Keith Rabois warns: “there may be a year, two years, three years max, when anybody uses human labeled data for maybe anything” — a direct shot at investors chasing revenue traction at Scale, Mercor, and Surge. His caveat on the Bitter Lesson itself: it only holds where data is abundant — “physical world AI is lacking in data, and so you just try to approximate humans.”
- Travis Kalanick has been doing “vibe physics” at 4 a.m. with GPT and Grok — “I’ve gotten pretty damn close to some interesting breakthroughs” — but is honest that Grok 3 and existing ChatGPT can’t originate ideas: getting one past conventional wisdom is “like pulling a donkey.” The prize is a scientific-method machine: “game the F over. You just light up more GPUs, and you just got, like, 1,000 more PhD students working for you.”
- Building a browser is “an absolutely stupid capital allocation decision” in 2025, per Chamath — “a glorified markup reader” — and Perplexity’s real prize is replacing Bloomberg’s “atrocious” $25k/year terminal, “a $100 billion enterprise” there for the taking. Keith is blunter: ChatGPT is becoming the verb at a billion users, “there’s nothing left of Perplexity if they can’t pull this off” — and “Google Search is toast” too.
- Elon’s third party splits the panel: Keith calls him “probably a replacement level politician” — Michael Jordan playing baseball — notes no true third-party Senate win since 1970, and concedes only “a few House races.” Travis is all in (“Elon is almost always right… I’m on the Elon train”), and Chamath says the mechanics changed: 2023 FEC guidance lets super PACs run full ground games, and 3-5 independent candidates equals real leverage with the filibuster “on borrowed time.”
- SCOTUS backed Trump’s RIF plans 8-1. Chamath: the “CEO of the United States” must be able to fire people, and had DOGE launched after this ruling it would have gone through “like a hot knife through butter” — though Keith counters “that ruling doesn’t happen without DOGE,” and warns only planning was approved; implementation “may not be 8-1.”
- Travis’s robot kitchens drop labor from 30-35% of revenue to 7-10% — a 60-sq-ft machine doing 300 bowls an hour, aimed at an “internet food court” where anything can be made in an 8,000-sq-ft facility. On the Pony.ai chatter: “there’s no real deal right now, but there is definitely some inbound.”
Deep dive
1. Travis’s robot kitchens: labor from 30% of revenue to 7-10%
- On the Pony.ai speculation — Jason’s frame: a Chinese autonomy player with cars on the road, a Middle East footprint, and an Uber deal — Travis stays coy but confirms the shape: with Waymo scaling city by city and Tesla doing it “the hard way, classic Elon style,” people who want alternatives “have reached out to me… there’s no real deal right now, but there is definitely some inbound.” His angle: solve autonomy once and it moves both people and food — “autonomous burritos.”
- The Lab37 machine is the concrete business: 60 square feet, 300 bowls an hour, 18 dispensers and 10 sauces, running a Chipotle/Sweetgreen-style make line. Staff prep the food, load it, and leave — “this restaurant runs itself for many hours without anybody there.” Delivery-kitchen labor runs 30-35% of revenue (higher in brick-and-mortar); running the machine, “it’s between 7 and 10% of revenue.”
- His lesson from failed food robots — including, Keith needles, what Friedberg tried (“he died as a vegan martyr”): automation must be end-to-end. “You have a million-dollar pizza machine, and then on the left you have a guy feeding ingredients in and on the right a guy taking the pizza out… instead of one guy making pizzas, I have a million-dollar machine and two guys making pizza.”
- The endgame is the “internet food court” — Amazon’s everything store for food in an 8,000-sq-ft facility where combinatorial dispensers make anything. Sizing: ~85% of US meals are eaten at home and Uber Eats plus DoorDash are only “1.8% or 2% of all meals,” so this isn’t killing restaurants — it’s replacing home cooking as a service. “You don’t have to be wealthy to be healthy.” Keith’s home-run version: a robot private chef in every house.
2. Grok 4 takes the benchmark crown — and vindicates the Bitter Lesson
- The release: Grok 4 shipped Wednesday night — $30/month base, $300/month “heavy” with a multi-agent feature where agents attack the same problem in parallel and reach a study-group consensus. Per Artificial Analysis benchmarks it has passed OpenAI’s o3-pro and Gemini 2.5 Pro as the most intelligent model — book-smart, Jason cautions, not street-smart, and it “got a little frisky” on X and “needed to be red teamed a little bit more decisively.”
- Chamath’s core argument, via Rich Sutton’s 2019 essay: whenever a general learning approach that scales with computation competes with a human-knowledge approach — chess, Go, speech recognition, computer vision — “the general computation problem always won.” The bitterness is psychological: “there was a psychological need for humans to believe we were part of the answer… You just have to let go, give up control.”
- What makes xAI the proof case: starting March 2023 — under two and a half years — Elon bet on the 100,000-GPU Colossus cluster, then 250,000, then a million, and the results show general compute “can actually get to the answer and better answers faster.” The precedent: Tesla FSD with cameras only versus Waymo’s lidar — collect billions of driving miles and apply general compute rather than the “laborious and very expensive approach.”
- The implication he wants investors to sit with: Llama “just spent 15 billion to buy 49% of Scale AI. That’s exactly a bet on human knowledge.” And the compounding math: if interjecting humans creates “300 to 1,000 basis points of lag per successive iteration, then over two or three iterations you’ve totally lost.” So what do Gemini, OpenAI, and Anthropic do when they wake up to these results?
3. The pushback: human knowledge isn’t obsolete where data is scarce
- Travis’s correction — worth keeping: Tesla’s autonomy approach IS human knowledge — “the whole idea is to approximate human driving.” Elon’s insight was almost humanist: “I’ve got two eyes. Why can’t my car?… I don’t have any lidar spinning around on my head.” The compute half of Chamath’s story arrives with Hardware 5, “coming out on Tesla probably next year.”
- Keith’s sharper version: “It’s not quite that binary.” LLMs — the most important unlock in AI — are trained entirely on human writing: “it’s not like there was some tablets floating in space that weren’t drafted by humans that we’ve trained on.” Non-LLM models may prove Chamath right, but “almost no one’s really using non-LLM based models at scale.” And humans turn out to be near-ideal drivers except when distracted, so training against them was the right call.
- His VC takeaway is the tradeable one: everyone debating investments in Scale, Mercor, or Surge is “just looking at revenue traction” and missing that there’s a very short half-life on human-labeled data — “a year, two years, three years max, when anybody uses human labeled data for maybe anything.” Self-driving labeling is already machine-done at some AV software players.
- Keith’s boundary condition on the whole essay: the Bitter Lesson “is only true when you have enough data.” Where the data doesn’t exist — and “physical world AI is lacking in data” — you may need to “hack your way there through human interactions” for years or decades.
4. Synthetic data and “vibe physics”: AI as a scientific-method machine
- A voice in the discussion states the next phase: “The cumulative sum of human knowledge has been exhausted in AI training. That happened basically last year” — so models will write essays, grade themselves, and self-learn on synthetic data. Chamath: subsequent Grok versions won’t train on any traditional dataset in the wild — agents creating synthetic data from scratch — at which point “everything gets flipped upside down.”
- Travis’s confession: up at 4 or 5 a.m., going down Quora physics threads with GPT or Grok — “the equivalent of vibe coding, except it’s vibe physics… I’ve gotten pretty damn close to some interesting breakthroughs.” But his honest limit on Grok 3 and existing ChatGPT: “it cannot come up with the new idea. These things are so wedded to what is known… It’s like pulling a donkey” — and even when it concedes you’ve got something, “you have to double and triple check.”
- Where this converges for Travis: “If you have a foundational model that is the best in the world at the scientific method, game the F over. You just light up more GPUs, and you just got, like, 1,000 more PhD students working for you” — though hypotheses still need physical-world testing, so imagine labs wired to these systems. Jason: “What could go wrong?”
- Keith adds why speed compounds — every second of lag breaks the recursive dive — and the evidence it’s already real: “Models trained solely on science tend to expose connections that no human has ever had before.” The example that floors the table: Navier–Stokes, one of the seven most difficult or important problems in math — “We use it to design airplanes, to design everything. It hasn’t been proved.” Point a computer at it and “who knows what’s possible? Teleportation.”
5. How does a better model judo-flip OpenAI’s juggernaut?
- Keith’s challenge to the table: OpenAI is “marching steadfastly towards a billion MAU… It’s a juggernaut. So how do you use the better product to judo flip the less better product?”
- Travis’s answer is culture: the Elon way — “missionary engineers that work twice as hard,” a “culture that is ultra fierce truth-seeking,” no politics. “You start winning on truth.” But his concession is notable: “the product of OpenAI, the product department, those guys are crushing… They’re just leading in a lot of different ways.”
- Jason’s tactical add: people forget Elon’s factory edge — Jensen Huang marveled at how Colossus was stood up at all, and “the factory is the product” at Tesla. Plus the grudge: Sam Altman “hoodwinked” Elon on OpenAI’s missionary open-source origins — “he did him dirty.”
- Keith’s frame is vertical integration: Apple’s AI performance is “absolutely miserable on the most important technology created through the last 70 years,” yet it’s still worth trillions because it’s vertically integrated. OpenAI can’t compete at the factory level, so “they’ve gotta ship the device… But then if they do that, they’re like Apple plus AI.”
6. Agentic browsers: new category or glorified markup reader?
- Jason demos Perplexity’s Comet (launched on the $200/month tier): an agent pops open a browser window, searches flights, can fill Amazon carts, and can connect to Gmail and OpenTable — the pitch being you’re already authenticated and don’t look like a bot, “it’s your browser doing the work.”
- Keith’s verdict: “a great Hail Mary attempt by Perplexity” — necessary because ChatGPT “is becoming the verb” and “there’s nothing left of Perplexity if they can’t pull this off.” The same logic indicts Google: “Google Search is toast,” and with Chrome plus Gemini they should be building exactly this — “the assets that are best at Google right now have nothing to do with search.”
- Travis reports the mood among consumer software CEOs — “big boys… with real stuff” — asking how they survive when agents take over: “every consumer software CEO that has an app in the App Store is tripping. Sometimes I’m doing almost like therapy sessions with them.” Chamath’s jab: “So you’re lying to them. You’re doing hospice care.”
- Jason calls building a browser “an absolutely stupid capital allocation decision” in 2025. Keith calls the browser itself doomed middleware — “a browser is like the dumbest thing to build in 2025… a glorified markup reader” handling HTML and rendering “under the water,” when the endpoint is something you speak to or think at. “Seeing a bunch of visual diarrhea is not elegant. It’s lazy.”
7. Perplexity’s real prize is Bloomberg — and Apple can’t buy its way out
- Jason’s alternative strategy: “building a browser is an absolutely stupid capital allocation decision… Perplexity’s path to a legacy business is to replace Bloomberg.” Having paid $25,000 a year for it: “the terminal is atrocious… anybody that could build a better product would take over a $100 billion enterprise because it’s there for the taking” — screens max out at five companies, and the real lock-in is messaging (“my team has traded huge positions via text message on Bloomberg”).
- Keith endorses it as “actually a pretty coherent” strategy — pick a vertical, own it, exploit unique data sources that may never license to OpenAI. On the Apple-acquires-Perplexity speculation: “Apple has missed every possible window on AI, and continues to miss it” — CEO challenges, cultural challenges, infrastructure challenges — and “what are you gonna spend, like, a billion dollars for product taste?”
- Keith’s corollary about Mark: “Grok 4 shows that Mark really does need to spend money to build a whole new team, ‘cause everything they’ve done in AI has also missed the boat.” Travis, meanwhile, wants in on the Bloomberg clone — offering Chamath naming rights: “the Poly Hypatia.”
8. Elon as “replacement level politician” — the third-party debate
- Keith’s metaphor — the episode’s sharpest: it’s Michael Jordan playing baseball. “Elon is probably a replacement level politician. He’s Michael Jordan for entrepreneurial stuff, but the third-party stuff is not going to work.” Jason’s approval chart is “a flaw of averages”: Trump has “the highest approval rate of any Republican ever measured” at 95% (Reagan peaked at 93) — Democrats just hate him, and “being polarizing is an ingredient to being successful.”
- The structural case against: MAGA already was “a third-party takeover of the Republican Party,” and you can’t run that play twice in a compressed window; “smart parties absorb the best ideas of third parties,” starving them of oxygen; no true third-party Senate winner since 1970 (Buckley’s brother); and people vote for people, not ideas — Elon “literally can’t constitutionally” be the party’s face.
- Travis, unbothered: “I have this axiom that I’m making up right now… Elon is almost always right.” Nobody has ever had this much capital to be a “party boss” outside the system, and “the threat of that happening can make good things happen separately, even if it doesn’t go all the way. I’m on the Elon train.”
- Jason’s version of the path: a couple of House seats at ~$2M a race, maybe a Senate seat at ~$25M, $250M deployed every two years — a Joe Manchin-style swing caucus with a Grover Norquist-style pledge. Polymarket has 55% odds Elon registers the American Party by year-end, and of his top 10-20 friends, “50% will join Elon’s party day one.” Keith’s concession lands as the headline: “I think you can win a few House races. I don’t think you can win a Senate race.” Jason: “That is the biggest mistake you’ve ever made. He’s now gonna win two.”
9. The mechanics: FEC ground games, 2019 spending levels, a doomed filibuster
- Chamath’s under-appreciated enabler: 2023 FEC guidance transformed what super PACs can do — no longer just ads, but ground operations: door knocking, phone banking, get-out-the-vote. “A super PAC became more like a full campaign machine, and Trump showed the blueprint” across the swing states. To the extent Elon uses those rules, the three-to-five-seat leverage play is “the only path.”
- Keith’s fiscal correction of Elon: “he’s actually wrong about the reason why we have a deficit… It’s not because we’re undertaxed. It’s we’re massively overspending. If you just held federal spending to 2019 levels, with our current tax revenues we would be in a surplus” (Travis: “500 billion”) — which is why Elon should have backed the big beautiful bill and helped get Republicans to 60 votes. Or skip it: “the filibuster is an artifact of history… at some point, some majority leader’s just gonna say, ‘We’re done with the filibuster,’” and steamroll cuts at 50-51. Chamath agrees it’s “on borrowed time.”
- On candidate selection, a live disagreement: Chamath wants people who “transcend politics and policy… straight up bosses with enormous name recognition” — Schwarzenegger in the Gray Davis recall, rank X and Instagram followers. Keith winces: “let’s not get more celebrities as politicians… let’s get people who’ve led large, complex things” — though if Elon can tune his epic talent algorithm to politics, “that might work.” Jason’s filter: they must survive the podcast test — Kamala “couldn’t hang for two hours in an intellectual discussion. If you can’t hang, you’re out.”
10. SCOTUS backs RIFs 8-1 — but only the planning, for now
- The setup: Trump’s February EO implementing the DOGE Workforce Optimization Initiative told agencies to prepare RIFs; the AFGE (820,000 members) sued, arguing large-scale workforce changes require Congress; a San Francisco federal judge (a Clinton appointee) blocked it; eight of nine justices overturned the block.
- Chamath calls it “incredibly important, incredibly right”: put ~3 million employees and contractors into 2,000+ federal agencies with outdated technology and you get slow processes and runaway rulemaking — regulations exploding since 1993 until “we all succumb to an infinite number of rules that we all end up violating and not even know it.” “If the CEO of the United States isn’t allowed to fire people, all of that stuff just compounds.”
- His counterfactual — “I wish Elon had come in and created DOGE now… with that Supreme Court ruling in hand, these guys probably would’ve been like a hot knife through butter.” Keith’s rejoinder, worth keeping: “Except that ruling doesn’t happen without DOGE. DOGE caused that ruling to occur.”
- Keith’s legal nuance: the Constitution vests all executive power in the president, “period,” but this EO was only approved “to allow for the planning” — when a specific plan litigates, “it may not be 8-1.” The looming fight is whether a president can refuse to spend money Congress explicitly appropriated. And some targets need legislation: the Department of Education was created by statute in 1979 — “every single educational stat has got worse since the department was created… but there is a law on the books, so you may have to repeal that.”