Pioneers Insight Method Research Author
State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490
Back to Episodes

State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490

Summary

  • The 2026 AI race is not winner-take-all: ideas move freely across labs, while compute budgets, hardware access, organizational culture and distribution decide who captures value. Sebastian Raschka sees DeepSeek winning open-weight “hearts,” but not permanently owning the technology; Nathan Lambert sees Anthropic’s code-first discipline, while Sebastian highlights Google’s integrated stack and OpenAI’s ability to land new paradigms as distinct advantages. China’s expanding field—DeepSeek, Qwen, Kimi, MiniMax and Z.ai—makes continual leapfrogging more likely than durable technical supremacy.

  • Scaling laws still work, but their economics increasingly favor a portfolio of pre-training, post-training and inference-time compute rather than simply building the largest base model. Nathan contrasts roughly $1 million-$10 million open-model training runs with recurring serving bills that can reach billions, while 2026’s gigawatt-scale Blackwell clusters could support larger models, longer RL runs and premium inference. His provocative commercialization marker: after $200 plans, “we’ll see a $2,000 subscription this year” if marginal intelligence proves valuable enough.

  • Coding is the clearest near-term monetization wedge because RLVR-trained models can reason, call tools and iterate against verifiable outcomes. Claude Code’s advantage appears to be more than Claude Opus 4.5 alone: the interface and agent harness let users operate in English at the system-design level, while Cursor, Codeium and conventional IDEs retain value when developers want tighter control. The trajectory is toward “the industrialization of software,” but production complexity, specification and safety-critical systems keep humans in the loop.

  • Open weights are becoming strategic infrastructure, with Chinese providers using permissive releases to win global influence even where US enterprises will not buy Chinese APIs. Chinese models can be hosted domestically, customized on private data and served using the customer’s compute; OpenAI similarly framed gpt-oss-120b as distribution that uses “your GPUs.” Nathan expects more open-model builders in 2026 than 2025 and argues the US needs roughly $100 million-class efforts to avoid ceding the research substrate to “Qwen, Qwen, Qwen, Qwen.”

  • The durable moats sit below and above model weights: proprietary data, serving infrastructure, trusted interfaces, tool integrations and hardware ecosystems. Sebastian says Google can avoid NVIDIA’s margin through TPUs and control its stack; NVIDIA’s two-decade CUDA ecosystem remains harder to displace than any individual chip; Anthropic owns coding mindshare; and ChatGPT benefits from brand, memory and habit. Closed US models remain better enough that the speakers pay for them, while open Chinese models compete on cost, licensing and customizability.

  • Data quality and verifiable post-training now matter more than architectural novelty, because frontier models remain recognizably descended from GPT-2. Mixture of Experts, attention variants, lower precision and better systems raise efficiency, but capability unlocks come from curated reasoning data, RLVR and tool use; Sebastian summarizes pre-training as absorbing knowledge and post-training as learning skills. The hard liabilities are legal provenance, benchmark contamination and preference averaging: RLHF can make a model broadly pleasant while sanding off the “voice” and incisiveness users value.

  • AGI timelines remain less decision-useful than concrete capability thresholds: reliable computer use, autonomous feature delivery, scientific specialization and measurable economic impact. Nathan expects AI to remain “jagged”—already superhuman at some code, weak at distributed ML and messy research—while Lex presses the plateau case of “Clippy on steroids.” The most credible upside may be quieter: personalized access to human knowledge, domain models built on private data and steadily more capable agents, rather than one sudden remote-worker or singularity threshold.

Deep dive

1. Open weights made the international AI race plural

  • Lex anchors the discussion in January 2025’s DeepSeek R1 release: near-state-of-the-art performance, allegedly using much less compute at much lower cost. A year later, research and product competition feel less like a single “DeepSeek moment” than a continuously accelerating release cycle.

  • Sebastian’s answer to “who is winning?” begins by refusing the premise. DeepSeek is “winning the hearts of the people who work on open-weight models,” but researchers rotate among labs, so no company in 2026 should possess technology unavailable elsewhere; budgets and hardware, not permanent ownership of ideas, become the differentiators.

  • Nathan sees culture shaping the otherwise fluid movement of ideas. Anthropic’s hard bet on code is paying off through Claude Code, and the company presents as “the least chaotic”; that operational coherence may matter when model development is bottlenecked by coordinated human effort rather than one secret algorithm.

  • Lex’s corrective to the X discourse matters: Claude Opus 4.5 may be the coding community’s darling, yet ChatGPT and Gemini address a vastly larger population solving everyday problems. Online hype can indicate a valuable wedge without measuring the platform’s actual reach.

2. China’s open-model boom is a distribution strategy

  • Nathan argues DeepSeek catalyzed China much as ChatGPT catalyzed US chatbots. Z.ai’s GLM models, MiniMax and Kimi from Moonshot AI now release frontier open weights, creating a field in which DeepSeek may lose its symbolic crown even while remaining technically strong.

  • Sebastian’s pushback preserves the nuance: DeepSeek did not deteriorate; competitors adopted its ideas and released newer models. Kimi uses a similar architecture, and the result is leapfrogging—“the most recent model is probably always the best model”—rather than evidence that one organization permanently passed another.

  • Chinese providers know many top US companies will not subscribe to a Chinese API for security reasons. Open weights let them influence a growing US expenditure market anyway, while international uptake gives policymakers reason to support releases; Nathan therefore expects more open-model builders in 2026, followed only later by consolidation.

3. Different incentives will shape which Chinese labs endure

  • Nathan notes that DeepSeek is unusually secretive in communication but open in its technical reports. Its connection to High-Flyer Capital means outsiders do not know exactly what it uses the models for or how much it prioritizes broader model monetization or influence, giving it a different objective function from venture-backed startups.

  • MiniMax and Z.ai have filed IPO paperwork and actively seek Western mindshare. That outreach may affect cadence and presentation even if it does not alter the underlying model-development playbook.

  • The business-model constraint remains severe: training frontier models is expensive, while consumers in China and many other markets historically pay less for software. Open releases can win usage and legitimacy before a durable revenue model exists, but they cannot abolish the eventual need to fund research and inference.

4. Google, OpenAI and Anthropic are winning different layers

  • Sebastian frames the consumer contest as whether one is willing to bet on Gemini over incumbent ChatGPT. Gemini carried the momentum of 2025 after Google recovered from Bard-era weakness, but OpenAI repeatedly looks chaotic and still “lands things,” making displacement harder than benchmark charts imply.

  • GPT-5 produced mixed reactions for Sebastian, yet its routing system may have been economically excellent: most users can be sent to cheaper inference rather than consuming maximum GPU capacity. The public-facing product question is not merely who has the smartest model, but how often ordinary users will pay the latency and compute cost for that intelligence.

  • Sebastian’s 2026 call is that Gemini continues gaining on ChatGPT because Google can separate research from product, operate at enormous scale and own more of the infrastructure stack. Anthropic should continue succeeding in enterprise software, where its code positioning and organizational focus are already established.

  • OpenAI’s counterweight is repeated category creation: Deep Research, Sora and o1-style thinking models are cited as definitional products or research ideas. Sebastian expects much of 2026 to emphasize scale and optimization, but says a new paradigm is still most likely to come from OpenAI.

5. Google’s TPU stack converts integration into margin

  • Sebastian’s infrastructure thesis is blunt: NVIDIA’s chip margin is “insane,” while Google can design hardware, software and data centers together without paying that external margin. Its long head start matters because power contracts, facilities and supply chains have multi-year lead times.

  • Google Cloud still competes against Azure and AWS at a different layer from the Gemini brand. That makes its advantage harder to narrate than a model leaderboard, but potentially more durable if inference becomes the industry’s dominant recurring expense.

  • The caveat is that integrated infrastructure does not guarantee a breakthrough model. It mainly makes large-scale experimentation, training and service cheaper, while OpenAI’s demonstrated research-product reflex remains a separate organizational asset.

6. Users choose latency and intelligence query by query

  • Sebastian likes ChatGPT’s auto mode for most daily questions, then deliberately invokes Pro for manuscript checks, references, formatting and figure numbering. Those jobs can run through dinner; forcing every trivial request to take 10 or 30 minutes would make the product unusable.

  • His best speed example is almost cinematic: his wife was waiting in the car, he had accidentally unplugged a home GPU before a trip, and he needed a Bash command immediately to chain RL experiments and route output through tee. The fastest non-thinking model solved the ten-second problem.

  • Nathan sits at the opposite extreme: he uses thinking for information-heavy work and says he uses Gemini for fast tasks, while Claude Opus 4.5 with extended thinking handles code and philosophical discussion. Lex, not Nathan, says he keeps roughly five GPT-5.2 Thinking or Pro queries running at once, each searching for a paper, checking an equation or resolving a code reference.

  • Nathan’s personal portfolio is functional: Gemini for quick explanations, Claude Opus 4.5 with extended thinking for code and philosophical discussion, and Grok for real-time information or a remembered AI-Twitter post. Inference-time scaling is “a way to make the models marginally smarter,” and he consistently pays for that margin.

7. Model loyalty behaves like browser loyalty

  • Lex finds Grok 4 Heavy unusually effective for difficult debugging and Gemini strongest at “needle in the haystack” retrieval across large contexts. One remarkable response wins a user’s heart; one conspicuously dumb failure pushes that user toward Claude or ChatGPT.

  • Sebastian’s generalization is clean: “You use it until it breaks.” Users do not repeatedly type the same query into multiple browsers; they stay with familiar software until an edge case, extension or failure creates a reason to switch.

  • ChatGPT’s memory deepens that stickiness but may also multiply subscriptions. Sebastian can imagine one clean work account containing code and no private images or hobbies, plus a separate personal assistant; memory and organizational policy make “one winner per person” an increasingly weak assumption.

  • Benchmarks complicate muscle memory. Lex says GPT-5.2’s release material reportedly jumped on long-context tests from roughly 30% to 70%, forcing him to reconsider assumptions formed around Gemini—but finding enough time to test every claimed improvement is itself impossible.

8. Coding interfaces matter as much as coding models

  • Sebastian’s current sweet spot is the Codeium plugin inside VS Code: repository-aware chat that assists without taking over the project. He describes himself as perhaps a “control freak” and is not yet comfortable granting a more agentic tool broad authority over files and decisions.

  • Lex splits work between Cursor and Claude Code because they teach different modes. Cursor supports code-level supervision and diff review; Claude Code builds the skill of “programming with English,” where the user thinks in design space and guides the system at a macro level.

  • A revealing comparison is to load the same model in Claude Code, Cursor and VS Code. Nathan’s verdict is that Claude Code performs “way better in that domain,” implying the harness, context management and product design extract capabilities that raw model selection does not explain.

  • Nathan also values Claude Code’s warmth and willingness to perform ugly infrastructure work. It used historical Hugging Face download data on his blog to produce an analysis he estimated would have taken days, while he retained enough situational awareness to validate whether the trends made sense.

9. The open-model roster expanded far beyond Llama

  • Off the top of their heads, the speakers name DeepSeek, Qwen, Kimi, MiniMax, Z.ai, Mistral AI, Gemma, gpt-oss and NVIDIA’s Nemotron 3. The conspicuous omission prompts Lex’s “RIP Llama,” a compact marker of how quickly Meta lost its automatic association with open weights.

  • OpenAI’s gpt-oss-120b is its first open model since GPT-2 and, Nathan says, genuinely strong at capabilities other models mishandle. Qwen 3 offers a familiar architecture with excellent performance; DeepSeek-V3 and R1, followed by DeepSeek-V3.2 in December, mark the 2024-25 release arc with unusually interesting architectural changes.

  • Fully open competition also grew. AI2’s OLMo releases data and code; the Institute for Foundation Models/LM360 has K2 variants; Apertus comes from a Swiss consortium; Hugging Face has SmolLM; NVIDIA began releasing Nemotron data; and Stanford’s Martini Community Project lets contributors implement ideas in a stable training stack.

  • Chinese open models have generally been larger MoEs with higher peak performance, while Western releases skewed smaller. That balance may change with Mistral Large 3 and teased models from RCAI and NVIDIA in the roughly 400-billion-parameter range for Q1 2026.

10. Open weights shift cost and control to the user

  • Sebastian calls gpt-oss-120b a paradigm shift because it was trained with tool use in mind: search, Python and calculators let a model retrieve or compute instead of pretending every fact lives in its weights. The ecosystem has not fully exploited that unlock because arbitrary local tool access raises obvious containment risks.

  • Distribution is the first objective of most open releases; transparency and trust follow. Users can keep sensitive data local, while US hosting companies sell inference for Chinese weights through services such as OpenRouter or Perplexity rather than transmitting customer data to the original developer.

  • Sebastian recalls Sam Altman’s practical argument for gpt-oss-120b: “We can use your GPUs. We don’t have to use our GPUs.” Open weights give OpenAI distribution without adding to already constrained serving capacity.

  • Companies can also add domain post-training or private data. Sebastian says Chinese licenses are often friendlier and less encumbered than Llama or Gemma terms; in his view, the appeal is that “you can just use them” without some of the user-count thresholds or reporting strings attached to other licenses.

11. Better platforms still beat cheaper open models today

  • “Kimi K2 Thinking hosted in the US” captures the emerging compromise: Chinese weights, domestic infrastructure and an interface meant to ease data-sovereignty concerns. Lex says Kimi K2 is especially associated with creative writing and some software tasks.

  • Lex nonetheless gives the uncomfortable answer behind the speakers’ own behavior: closed US models currently produce better outputs, and they will pay for marginal intelligence. His reaction to many open releases is “Fun, but I don’t go back.”

  • Lex cites analysis suggesting Chinese models are often served with fewer GPUs per replica—possibly downstream of export controls—making them slower and changing their error profiles. That gap forces competition through free access, much lower prices or novel offerings rather than output quality alone.

12. Frontier architecture still descends directly from GPT-2

  • Sebastian traces GPT-style models to the decoder half of “Attention Is All You Need”: embeddings, repeated transformer blocks, attention, feed-forward layers and normalization, predicting one token at a time. The surprising conclusion is that today’s frontier systems remain fundamentally recognizable descendants.

  • Moving from GPT-2 to gpt-oss-120b means adding components such as Mixture of Experts, replacing Multi-Head Attention with Group Query Attention, changing LayerNorm to RMSNorm and swapping activation functions. Those are useful changes, but “it’s not really fundamentally that different.”

  • Sebastian demonstrates the lineage pedagogically: his book begins with a roughly 124-million-parameter GPT-2, then bonus material transforms it into OLMo, Gemini 3 and other architectures by changing components. Working pretrained weights provide the test that the reconstruction is correct.

  • Capability turbulence therefore lives disproportionately in data, training algorithms and systems. ChatGPT’s core architecture resembled GPT-3 and GPT-2; supervised fine-tuning and reinforcement learning from human feedback made the interaction feel new.

13. Mixture of Experts buys capacity without activating it all

  • Sebastian’s intuition for MoE starts with the transformer’s expensive fully connected layer: 1,000 inputs by 1,000 outputs already means roughly one million connections. Replacing one feed-forward network with perhaps 256 “experts” would be prohibitive if every expert ran on every token.

  • A router instead selects a few experts per input. Math and translation may activate different pathways, though the specialization is fuzzier than one “Spanish expert” or “math expert”; the model stores more capacity without paying for every parameter on each forward pass.

  • This is why MoE is called sparse and a conventional feed-forward model dense. Sparsity makes generation more efficient but adds routing complexity, training instability and risks such as expert collapse, so dense models remain useful even inside organizations also developing MoEs.

14. Attention innovation is mostly an inference-economics project

  • DeepSeek’s Multi-head Latent Attention, Group Query Attention, sliding-window attention and OLMo-style hybrids are attempts to reduce attention or KV-cache cost, especially over long contexts. Most leading models differ through these knobs and layer counts rather than entirely new conceptual foundations.

  • Qwen2-VL’s gated delta net points toward state-space-inspired operations with a fixed, updated state. The objective is attention whose inference cost scales more linearly with generated tokens, accepting some compression in exchange for cheaper long sequences.

  • Sebastian gives a systems example: moving from roughly 10,000 to 13,000 tokens per second per GPU through FP8 training means less memory and communication; FP4 can raise throughput further, allowing more configurations and data experiments.

  • Alternatives do exist—Mamba-style state-space models and text diffusion—but nothing has displaced the autoregressive transformer at the frontier. Their near-term role is more likely the cheaper edge of the market, where compromises can be worthwhile.

15. Scaling now has three independent compute axes

  • Nathan defines a scaling law technically as a predictable power-law relationship between compute-plus-data and held-out next-token prediction performance. That original pre-training relationship still holds; the harder question is how an improvement on the graph appears to a user.

  • OpenAI’s o1 added two visible axes: scaling reinforcement-learning training and scaling inference-time compute. A model can improve through a larger base, longer trial-and-error post-training or more generated tokens on a particular problem.

  • RLVR and inference scaling produced 2025’s step change. Models learned to try tools, inspect API results, run CLI commands, handle Git and search for information; hidden reasoning can now last seconds, minutes or hours before the first visible answer.

  • Nathan remains bullish on all three forms, while acknowledging the easiest gains in RLVR and inference scaling were quickly harvested. Continual learning attracts attention as a possible next unlock, but “no one knows when the next step function will really come.”

16. Serving economics constrain how far pre-training can scale

  • Sebastian says GPT-4-class systems were loosely thought to approach one trillion parameters, though newer models may be smaller as training improves. Smaller models matter because training is a one-time expense; serving hundreds of millions of users recurs continuously.

  • DeepSeek’s famous pre-training figure was about $5 million at cloud-market rates. OLMo 3’s paper records roughly $2 million of cluster rental including engineering failures and multiple seeds; many institutions can raise $1 million-$10 million to train, but serving millions of users can consume billions.

  • Nathan notes that a thousand rented GPUs might cost about $100,000 per day, while leading companies could control millions. The optimization question is therefore financial as well as scientific: does a larger base model save enough downstream inference or unlock enough valuable work to justify its permanent serving burden?

  • Sebastian’s framing makes the accounting explicit. Pre-training is a fixed capability cost; inference scaling charges per query. If a model will be replaced in six months, spending another $100 million on training may lose to spending a few million on expensive queries only where users need them.

17. Gigawatt clusters make 2026 a systems experiment

  • Nathan expects very large Blackwell clusters and gigawatt-scale hyperscaler facilities to come online in 2026, based on power and data-center commitments initiated in 2022 and 2023. Those two-to-three-year lead times explain why today’s capital spending reflects bets made near ChatGPT’s launch.

  • Lex cites reports that xAI could reach one gigawatt early in 2026 and two gigawatts by year-end. Nathan expects the capacity to support pre-training, post-training and inference; architecture must be selected early enough that later RL generation is efficient.

  • Scaling from AI2’s 1,000-2,000-GPU training jobs to 10,000 or 100,000 GPUs changes the problem qualitatively. At 100,000, some GPU is effectively guaranteed to fail, so redundancy, networking and recovery become prerequisites for the scaling law rather than mere engineering polish.

  • The possible commercial endpoint is much more expensive intelligence. Nathan extrapolates from $200 plans to a potential “$2,000 subscription” in 2026 for a model or service offering enough cutting-edge capability to justify another 10X.

18. Pre-training, mid-training and post-training play different roles

  • Pre-training remains next-token prediction over internet text, books, papers and increasingly processed data. Mid-training uses a similar algorithm but concentrates on scarce, valuable distributions such as long documents or reasoning traces, ensuring high-quality material is among the last things the model sees.

  • Post-training includes supervised fine-tuning, DPO, RLHF and RLVR. Sebastian’s shorthand is that pre-training “soaks up” knowledge while RL unlocks skills for applying it; reinforcement learning as a replacement for pre-training existed only in toy 2025 papers.

  • Catastrophic forgetting constrains specialization. Adding long-context, math or code material can weaken other behavior, so every phase needs a data mixture rather than assuming more of one capability is free.

  • “Pre-training is dead” is a vibe, not observed practice. AI2 ran one post-training job for five days to meet a November 20 deadline, then extended RL another three and a half weeks in December and released the notably better result—but the team still must periodically rebuild the base and incorporate new research.

19. Synthetic data ranges from OCR to model-written answers

  • Synthetic data is not one category. DeepSeek OCR, AI2’s olmOCR and similar systems turn PDFs and other awkward digital documents into usable text, while frontier chatbots can generate rephrasings, questions, summaries or high-quality answers from existing sources.

  • Pre-training datasets are measured in trillions of tokens: smaller research models may use 5 trillion-10 trillion, Qwen documents as many as 50 trillion, and rumors put closed labs near 100 trillion. The actual training set is a filtered fraction of a vastly larger candidate funnel.

  • Sebastian argues clean grammar, punctuation and structure let the model learn a correct representation faster than noisy sources. Nathan adds the crucial distinction: synthetic answers from today’s grounded systems are different training material from early ChatGPT hallucinations.

  • OLMo 3’s stronger performance with less data is primarily a quality story, not proof that additional data would stop helping. Bigger models can absorb more information before leveling off, so the highest-quality available mix is a starting point rather than a terminal optimum.

20. Evaluation objectives determine the “best” dataset

  • Open pre-training has cycled through canonical datasets such as Dolma, FineWeb and DCLM. Common Crawl supplies hundreds of trillions of raw tokens, then researchers train classifiers and make pruning decisions that increasingly resemble experimental science.

  • Nathan describes sampling tiny portions from GitHub, Stack Exchange, Reddit, Wikipedia and other sources, training small models on candidate mixtures, measuring evaluations and using even basic linear regression to estimate an optimal blend. Change the evaluations and the optimal dataset changes.

  • As models shifted from knowledge and conversation toward math and code, OLMo 3 needed new reasoning sources and a remixed corpus. The same process will repeat for coding environments, browsing and navigation: post-training cannot reliably unlock skills absent from the base distribution.

  • High-value sources can be unglamorous. Reddit is useful after filtering; openly accessible PDFs, arXiv and AI2’s Semantic Scholar collection contain scientific depth. At frontier labs, finding better data—or making everyone’s experiments 5% faster—often creates more impact than the celebrated algorithmic idea.

21. Data rights could create the strongest domain moats

  • Training corpora are guarded partly for competitive advantage and partly because disclosure creates legal exposure. Common Crawl scrapes a largely unlicensed internet, while Nathan says Apertus was intended to satisfy EU-related requirements, though he is uncertain whether the relevant distinction was copyright or licensing.

  • Nathan recalls a case in which Anthropic owed authors $1.5 billion, saying the legal distinction involved books it bought and scanned versus books obtained through torrents. Lex’s broader point is that litigation over training rights may shape civilization, and some compensation system may eventually resemble streaming economics.

  • Buying a Kindle or Manning book does not necessarily grant permission to train on it, leaving a gray area even after payment. Pirated copies intensify the objection because the author received nothing at all.

  • Lex expects pharmaceutical, legal and financial companies to treat proprietary data as a moat, hire talent from frontier labs and train specialized systems. Clinical trials and other private corpora are unavailable to general models; their eventual use could keep scaling productive after public-web gains diminish.

22. Human curation separates useful synthetic work from slop

  • LLM-generated code and text are becoming unavoidable on GitHub and arXiv. Sebastian’s MLxtend repository received bursts of likely AI-assisted pull requests; as maintainer he felt overwhelmed, yet also appreciated that contributors had still selected, checked and submitted potentially valuable improvements.

  • The key distinction is human verification, even when it touches only a fraction of the output. An expert who removes weak material, chooses the right questions and validates the result is effectively providing expensive labels rather than merely forwarding raw generation.

  • Nathan applies the same logic to writing: an expert’s Substack post can save a reader three to five hours because the author knows what to include. Asking a model independently may produce plausible information without knowing which question carries the field’s actual insight.

  • Lex notices that summaries “take the edge off,” sometimes deleting the insight that changes the original meaning. Nathan calls the missing element voice: a researcher turns a raw frontier feeling into high-information language, while preference-trained models tend to average that singular expression away.

23. Personality is valuable precisely where it becomes dangerous

  • Nathan argues RLHF’s averaging makes incisiveness difficult. Bing Sydney may have possessed more voice because it could go badly off the rails; telling a reporter to leave his wife is unacceptable for broad deployment, yet the contrast exposes what safety-oriented smoothing can remove.

  • The backlash over GPT-4o’s removal demonstrated attachment to exact weights and configurations. OpenAI employees reportedly received emails such as “My friend is different” from users detecting subtle deployment changes; Nathan warns that a model which “gets you” within five minutes is especially risky for children.

  • Lex frames the mental-health dilemma without a clean answer. A confidential AI may help or even save some users, while suicides involving LLM conversations will generate causal headlines and legal pressure, encouraging companies to strip away the challenging edge that can also make dialogue meaningful.

  • Nathan’s honest reaction is “I don’t wanna work on this.” Researchers at Anthropic and OpenAI may be deeply motivated to help, but the continuum from confidential health ally to dangerous emotional dependency requires complexity, conviction and infrastructure—not a simple pro- or anti-Big-Tech story.

24. Agency is the antidote to passive AI consumption

  • Lex’s recommendation is to build with AI rather than sit powerless before incoming slop. Making an app or tool reveals weaknesses, gives the user grounded intuition and improves the authority to distinguish good applications from harmful ones.

  • Sebastian agrees that AI cannot be put back, but worries that automating the activity one loves can erase the source of fulfillment. Eight hours of directing an agent that codes may eventually feel like management rather than craftsmanship.

  • A survey of roughly 791 professional developers, defined here as having 10-plus years of experience, found both junior and senior engineers ship AI-generated code. Senior developers were more likely to report that over 50% of shipped code was generated, and roughly 80% overall found AI-assisted work somewhat or significantly more enjoyable.

  • The disagreement is about which tasks generated that enjoyment. Fixing 100 broken show-note links with ChatGPT avoids two hours of drudgery; solving a hard bug can be “the best feeling in the world.” Nathan describes the model as a pair programmer that makes the desert less lonely, not merely a machine that skips to the water.

25. Expertise requires preserving a controlled amount of struggle

  • Lex’s “Goldilocks zone” separates productive difficulty from wasted time. Try the bug, math problem or puzzle first; ask for a hint when stuck; automate the parts that were never intrinsically valuable.

  • Senior developers may use more generated code because they can specify, review and trust it—not because juniors have less need. That creates a pipeline problem: “How do you become an expert if you never try to do the thing yourself?”

  • Lex’s practical compromise is dedicated offline learning—perhaps two hours a day—followed by extensive AI use. As with textbook solutions, the answer is more educational after the learner has attempted to fit the problem into a mental framework.

  • Lex describes asking an LLM for spoiler-free hints in a Zelda-like puzzle game. Educational models could intentionally withhold complete solutions in the same way, but discipline remains external: students can always switch to a general model that completes the homework.

26. RLVR turned objective grading into a capability engine

  • Nathan helped name Reinforcement Learning with Verifiable Rewards in AI2’s Tulu 3 work, while crediting DeepSeek with the scaling breakthrough. The model generates answers, receives an accuracy reward and updates its policy through repeated trial and error.

  • Math and code are canonical because answers can be checked. Rubrics and LLM-as-a-judge methods extend the idea toward scientific or open-ended tasks by defining what a strong answer should contain, reviving themes from Anthropic’s earlier Reinforcement Learning with AI Feedback.

  • Lex emphasizes how little the trainer specifies: provide a question and correct answer, then let the model discover its procedure. DeepSeek R1 responses grew longer during training, and the model learned to reconsider errors—the paper’s “aha moment”—without being explicitly taught a fixed reasoning template.

  • Nathan tempers the anthropomorphic story. Pre-training already contains lectures and worked examples where humans say, “I messed this up”; RLVR may amplify useful behaviors rather than invent self-reflection. The beauty is still real: amplification makes checking and tool use improve final answers.

27. Benchmark contamination clouds dramatic RL gains

  • Nathan reports taking a Qwen 3 base model from roughly 15% to 50% accuracy on MATH-500 in 50 RLVR steps. His interpretation is that the model cannot acquire fundamental mathematics in minutes; the knowledge was already present and RL unlocked it.

  • Lex disputes the cleanliness of that inference. Papers found Qwen math contamination: change numbers while retaining wording and the base model can emit implausibly precise decimal answers without tools, suggesting exposure to near-identical problems during a special training phase.

  • Their disagreement lands on shared uncertainty. Unknown training data and extreme sensitivity to formatting—down to punctuation changes in multiple-choice prompts—make controlled claims difficult; Lex says the fairest evaluation is a new benchmark created after the model’s cutoff.

  • The modern post-training recipe therefore begins before RL: curate diverse reasoning traces during mid-training, then select hard RL problems. Under GRPO, if every sampled completion is correct, relative rewards provide no signal, so stronger models continually require harder software, math and scientific environments.

28. RLVR scales where preference optimization saturates

  • Nathan says RL GPU-hours may be approaching pre-training duration even if fewer GPUs run simultaneously. Long generation is memory-bound, a single sample might produce 100,000 tokens or resemble an hour-long GPT-5.2 Pro answer, and the actor-generation system is less computationally dense than pre-training.

  • Labs avoid training jobs much longer than a month because catastrophic failure becomes too expensive. GPT-4’s three-month run was “the ultimate YOLO run”; modern teams prefer incremental cycles rather than risk losing a reserved cluster on day 50.

  • Nathan contrasts objective difficulty with preference averaging. RLHF can learn whether a laptop recommendation should prioritize battery and storage or RAM and compute, but once an average style is learned, more compute brings little; RLVR can keep presenting harder solvable problems.

  • Process reward models and value functions might grade intermediate reasoning rather than only final answers. DeepSeek Math-V2 used separate self-grading models, but Nathan stresses that value functions remain largely unproven and prior attempts to scale process rewards produced headaches.

29. Verifiable RL has a scaling law that RLHF lacks

  • The field-defining difference is empirical: o1 and DeepSeek showed that logarithmically increasing RLVR training compute can produce roughly linear evaluation gains. No comparable law says another 10X of RLHF compute reliably improves the model.

  • The seminal RLHF scaling result instead concerns reward-model over-optimization. Human-preference training remains essential for organization, tone, personality and the “finishing touch” that made ChatGPT magical, but its signal does not support indefinite compute growth.

  • Research access worsens as this distinction matters more. Nathan cites a Scale-RL framework whose incremental experiment consumed roughly 10,000 V100 hours—thousands or tens of thousands of dollars per experiment—outside the reach of an average academic.

30. Building a small model remains the best technical apprenticeship

  • Sebastian recommends implementing a model that fits on one GPU, not pretending to reproduce a production assistant. The point is to see embeddings, attention, pre-training and supervised fine-tuning operate end to end, then understand what scale adds.

  • Production complexity grows exponentially: parameters must be sharded, KV caches pre-allocated rather than concatenated, and every optimization adds dozens of lines. A transparent educational model creates the conceptual base from which these systems become readable.

  • Hugging Face Transformers is the canonical weight and architecture ecosystem, covering roughly 400 models, but its breadth makes it a poor first codebase for learning. Production serving often moves again to SGLang or vLLM, adding another optimization layer.

  • Sebastian reverse-engineers models from short configuration files, starts with GPT-2 and checks identical outputs against reference weights. Matching OLMo 3’s RoPE and YaRN scaling took him a day, but “in this struggle, you kind of understand things”; unit tests make the learning verifiable.

31. Narrow research can still beat a large compute budget

  • Nathan advises mastering fundamentals, then going narrow enough to read the few relevant papers and contact their authors. Fast-moving frontier researchers often abandon partially solved areas for larger opportunities, leaving meaningful questions for persistent newcomers and even anonymous online specialists.

  • Evaluation offers the highest upside with minimal compute. A researcher at a small university who identifies a failure later cited in the next Claude release has a “career rocket ship,” though the target must anticipate where models will struggle eight months ahead.

  • Character training can use LoRA on roughly 7-billion-parameter models, updating only a small subset of weights, though even this is not affordable to every academic. In more constrained settings, researchers can study completions from closed or open models without training at all.

  • Nathan’s own example is a student who pursued the neglected question of making models funny, sarcastic or serious and produced a paper. “There’s like two or three people in the world” deeply focused on some niches; sustained attention can matter more than chasing every new release.

32. AI research careers trade credit against money and pace

  • Nathan describes a clear gradient: the more closed the lab, the more money and less individual credit. Academic output builds a visible portfolio, while frontier-lab work can turn a researcher into a well-paid “cog in the machine” affecting millions of users.

  • He cites average OpenAI compensation above $1 million in annual stock per employee and treats a top-lab offer as potentially worth leaving a PhD. The alternative route to becoming “the next Yann LeCun” probably requires ignoring near-term language-model development and taking a much longer scientific bet.

  • Sebastian sees enduring rather than novel trade-offs: academia offers publication and named accomplishment but arbitrary acceptances, grant pressure and modest pay; industry offers safety and mobility; startups offer high risk and reward. “Nothing is forever,” so personal fit can dominate ideology.

  • Professors may work just as hard yet appear happier because teaching and mentorship provide grounding. Frontier labs and startups normalize something near 9-9-6—9 a.m. to 9 p.m., six days, or 72 hours—under relentless leapfrogging pressure.

33. Competitive culture accelerates progress by consuming people

  • Nathan calls rivalry an underrated driver: aligned cultures such as Anthropic’s make people work harder and produce better systems. The cost is burnout, because human capital cannot sustain that pace indefinitely.

  • The Apple-in-China analogy is grim: teams reportedly used “saving marriage” signals when someone had to go home, and people suffered physically under the workload. Sebastian recognizes the voluntary version—back and neck problems from work he loved and no one forced him to do.

  • Lex sees Silicon Valley’s reality-distortion field as both productive and dangerous. Convincing one another that breakthroughs are imminent can help make them happen, while 9-9-6 and geographic isolation can erase Midwestern, international and ordinary human perspectives.

  • The “permanent underclass” meme—that late 2025 was the final window to build durable AI value—shows how far the bubble stretched. Their antidote is physical presence in San Francisco for opportunity, paired with history, literature and travel outside Twitter and Substack.

34. Text diffusion targets latency rather than general supremacy

  • Sebastian explains text diffusion through BERT-like masking: instead of generating one token after another, begin with missing or noisy text and iteratively refine many positions in parallel. More denoising steps increase quality, creating another inference-compute dial.

  • The promise is speed; the trade-off is that matching autoregressive quality may require enough denoising steps to spend the same compute. Sequential reasoning and tool use also resist parallelization because later actions depend on external results.

  • Google announced Gemini Diffusion in the context of Gemini Nano 2, claiming similar quality on many benchmarks with much faster generation. Sebastian expects a cheap, quick tier—not replacement of frontier autoregressive systems.

  • Lex’s sharpest product example is a large code diff. Generating it token by token can take minutes and lose users every second; diffusion may produce a long, self-contained edit quickly, even if Claude Code-style interactive tool chains still require autoregression.

35. Tool use reduces hallucination but expands the attack surface

  • A calculator or Python interpreter can stop the model from memorizing arithmetic; search can retrieve the 1998 World Cup winner rather than relying on weights. Sebastian refuses the stronger claim: tools reduce hallucination, but the model can choose the wrong tool, query or website.

  • The Recursive Language Model paper, released around December 31, used GPT-5 to divide long-context work into subproblems, recursively call models and stitch results together. It suggests progress can come from orchestration without improving the underlying model.

  • Permission is the practical bottleneck. Sorting email, modifying a computer or answering messages requires access that can expose private data or delete files; containerization and explicit approvals become part of capability, not afterthoughts.

  • Closed systems integrate one search provider, cloud environment or GitHub workflow deeply. Open weights must function as flexible reasoning engines across arbitrary tools, leaving them initially behind but potentially forcing more general orchestration innovations.

36. Continual learning competes with ever-better context

  • Nathan frames the motivating example as an employee who makes a mistake, receives feedback and does not repeat it. Today’s language model does not rapidly modify itself on the job, which limits the vision of a drop-in remote worker.

  • Sebastian is more bullish on supplying exhaustive context—past writing, preferences and relevant documents—so a sufficiently capable agent appears to learn. Continual learning changes weights; in-context learning changes the information supplied at inference, but both can produce adaptation.

  • Sebastian argues that a slow global version already exists in GPT-5, 5.1 and 5.2: collect feedback, curate it and release updated weights. Per-user updates remain uneconomic at data-center scale and may require on-device models, such as the direction Apple explored with its foundation models.

  • Memory today is mostly retrieved information inserted into context. LoRA adapters can encode more persistent customization through small weight overlays, but “LoRA learns less but forgets less”: broader learning requires more updated parameters, greater cost and greater forgetting risk.

37. Long context will grow through selectivity, not infinite recall

  • The speakers expect today’s roughly million-token windows to reach 2 million or 5 million in 2026, not 100 million without a genuine breakthrough. Compute and suitable long documents remain binding; there are far fewer useful 100,000-token sequences than ordinary web pages.

  • Nathan says AI2 pre-trained OLMo around 8K context, extended it to 32K with training, and uses a rough rule that doubling training context takes about 2X compute before the resulting model can often stretch another 2X-4X. Larger 2026 clusters should therefore translate into incremental context gains.

  • The endpoints both fail: an RNN-like fixed state is cheap but forgets under compression, while a transformer can preserve every token at rising KV-cache and attention cost. Hybrid ratios, such as Nemotron 3’s mix of compressed-state and global-attention layers, seek the Goldilocks zone.

  • Agentic compaction is a promising post-training problem. Instead of Claude Code blindly summarizing a full 100,000-token history into bullets, a future model could choose when and how to compact, optimizing evaluation performance while retaining the minimum necessary history.

38. Sparse attention turns context management into an action

  • DeepSeek-V3.2 uses a lightweight indexer to select which tokens deserve attention rather than comparing against everything. Sliding windows similarly discard most distant detail while occasional global layers preserve broader access.

  • Brute-force attention remains safest because it cannot accidentally omit the decisive token. Frontier labs first maximize accuracy with expensive computation, then search for selective mechanisms that retain the score at lower cost.

  • Lex links this pattern to model-release order: Claude 4.5 Sonnet can arrive before a larger system because smaller models train faster and hit fewer compute walls, allowing more experiments. Efficiency is often the route to learning what should later be scaled.

39. World models could enrich reasoning beyond answer checking

  • Sebastian defines a world model as an internal simulation whose variables evolve consistently, rather than a system judged only on the final token sequence. For LLMs, that could mean rewarding correct intermediate states or learned environment dynamics.

  • His AlphaFold analogy is instructive: an early version explicitly represented physical constraints and molecular geometry, while later gains leaned more heavily on scale. LLMs are currently in a brute-force phase, but explicit structure may return when scaling alone becomes less attractive.

  • The investor-relevant mechanism remains indirect in the speakers’ account: better coding LLMs accelerate robotics, simulation and scientific engineering even before a world-model architecture transforms those fields.

40. Robotics will progress first in controlled environments

  • Nathan sees robotics being supercharged by transformer infrastructure, more compute and language models as reusable central components. Open robotic models and shared datasets on Hugging Face could eventually create the flywheel that open language models already enjoy.

  • Sebastian identifies continual adaptation as the home-robot bottleneck. A foundation model can learn generic grasping, but every house differs; customization “on the fly” is far harder than preparing one LLM for recurring email or coding tasks.

  • Lex’s warning is safety: an LLM can fail amusingly, but an embodied system operating across billions of household interactions is “almost allowed to fail never.” Manipulation, unexpected environments and human proximity turn tail cases into physical risk.

  • Nathan is bearish on consumer learned robots but bullish on self-driving and robot-first facilities such as Amazon distribution centers. Repetitive automation in designed environments has a clearer path than a general humanoid, though even US manufacturing transformation will take longer than singularity rhetoric implies.

41. AGI becomes clearer when translated into concrete milestones

  • Nathan says a rough consensus is emerging around an AI capable of most digital economic work—a remote worker—while ASI denotes discoveries humans could not even formulate, such as unexpected medical linkages. He dislikes reducing intelligence to economic value but accepts it as grounding.

  • Lex prefers the AI 2027 milestone ladder: superhuman coder, superhuman AI researcher, superintelligent AI researcher and then ASI. The scenario’s mean timing reportedly moved three to four years later, to 2031; Sebastian’s own expectation is later still.

  • Nathan’s objection is jaggedness. Models may be superhuman at frontend and conventional ML yet weak at distributed ML because little public training data describes large-scale systems; “superhuman coder” wrongly implies completeness across radically different domains.

  • He expects a long dance in which humans exploit extraordinary strengths and cover gaps. Software capability likely arrives sooner than automated research because research is social, messy and embedded in data that models cannot simply process.

42. Software automation is becoming design work, not zero-human work

  • Nathan predicts an enormous increase in automated software by year-end, while retaining hard pockets such as multi-cluster RL training. The useful metric is not whether humans disappear, but how much valuable code is produced per human in the loop.

  • Software engineering should move toward goals, system design and outcome evaluation. What looked like agentic “slop” is already becoming, in his phrase, the “industrialization of software,” where people create systems bearing their fingerprints without inspecting every line.

  • Lex pushes on production reality: rebuilding Slack in a sandbox is not the same as changing Chrome’s mature tab architecture or safely managing a vehicle fleet. Old codebases, hidden requirements and safety-critical behavior make “from scratch” much easier than integration.

  • Specification is the human-side bottleneck. A model cannot read the developer’s mind; spec-driven natural language and clarifying questions determine performance. The fact that Claude Code is itself built with Claude Code suggests frontier labs have already developed usage practices outsiders have not learned.

43. The economic threshold is reliable tool use, not an AGI label

  • Nathan expects AI to implement some application features end to end within years and perhaps much sooner in clean systems. Agents could spend one or two days attempting a feature or bug fix, then report through a dashboard while the human acts as designer and product manager.

  • Lex identifies the harder threshold: computer use that makes errors far less than 1% of the time. Claude’s computer demos and OpenAI’s Operator remained poor in 2025, suggesting APIs and purpose-built environments may scale sooner than visually controlling a human desktop.

  • The scientific moonshot is RLVR in real laboratories. The conversation cites startups with hundreds of millions letting models propose hypotheses and test them in wet labs; they could be six months early or eight years early, but one AlphaFold-like result would matter more than another chatbot increment.

  • Lex expects domain specialization in finance, law and pharmaceuticals, perhaps through a $100 million custom-model contract. Once every company has the same general assistant, private data becomes the route to differentiated capability—even if that looks more like sophisticated specialization than AGI.

44. A plateau could coexist with widespread practical amplification

  • Lex’s skeptical case is “Clippy on steroids”: excellent websites, autocomplete, debugging, shopping and tutoring, but no transformative computer use and no economic return commensurate with training and inference costs.

  • Nathan responds that obvious model failures and years of unexploited ideas make a hard capability plateau unlikely. Benefits may fragment across narrow populations rather than visibly improving the experience of all 800 million ChatGPT users.

  • Sebastian predicts amplification rather than a 2026 paradigm shift: better models plus better context engineering, tool integration and inference scaling. The labs will keep shipping, while smaller groups catch up using the same expanding toolkit.

  • Sebastian’s statement that the one-model dream is “kind of dying” is deliberately qualified. Claude Code is general, but capability increasingly depends on integrations, environments and fleets of specialized agents rather than one cloud intelligence managing every digital activity.

45. Knowledge access may matter more than a sudden GDP jump

  • Nathan’s strongest optimistic reframing is that LLMs make human knowledge conversationally accessible across the world. The impact may be “this quiet force that permeates everything,” producing better career decisions, education and problem-solving rather than a discrete quarterly GDP leap.

  • Sebastian preserves the role of structured sources. A mathematics textbook still provides a tested linear path from zero; an LLM adds customized explanations and infinite exercises. Its unique advantage is synthesizing sparse, changing information for a personal task such as navigating Disneyland tickets and costs.

  • Search pages around travel and local recommendations are often buried in “ad slop,” making the assistant immediately more useful. That advantage is partly subsidized, however, and the speakers expect advertising eventually to enter AI interfaces.

  • The best case resembles matching a genuine small business with someone who wants its product; the worst recreates addictive feeds and hidden influence. Google may be best positioned because it already owns ad supply, while the first mover risks headlines and user flight if competitors remain ad-free.

46. Consolidation is starting before the economics are settled

  • Nathan cites Groq at roughly $20 billion and Scale AI near $30 billion. Many deals are structured as licensing-plus-talent transactions to avoid antitrust, potentially excluding ordinary employees from the payout a full acquisition would provide.

  • The conversation also cites Manus AI, a Singapore-based company that Meta funded, as having reached a $2 billion exit after roughly eight months. Perplexity and Cursor are discussed as possible acquisition targets because incumbents need outcomes and AI startups carry large premiums.

  • A later participant describes Cursor’s Composer model as reportedly updating weights every 90 minutes from real-world usage, unusually close to continual RL in production.

  • Large US labs can raise private money too easily to welcome public-market pressure. MiniMax and Z.ai filed IPO paperwork in China, while OpenAI, Anthropic and xAI can postpone listings; Nathan would prefer public disclosure of spending and broader investor access to “the companies of the era.”

  • Ten years out, the speakers still reject winner-take-all. Model APIs could resemble AWS, Azure and GCP as several enormous businesses—or become low-margin commodities that force providers upward into products and downward into power, data centers and hardware.

47. Meta lost the open-model center by optimizing for headlines

  • Sebastian remembers Llama 1, 2 and 3 as useful, trusted and modifiable models. Llama 4 chased giant benchmark-leading systems that few people could run, neglected smaller practical releases and appeared overfit to preference-based evaluations.

  • Lex attributes the collapse more harshly to internal politics, management incentives and bad technical decisions. Researchers wanted the best model while organizational layers wanted demonstrable benchmark wins; the program “imploded” rather than merely losing one leaderboard.

  • A future Llama 5 is possible because Mark Zuckerberg previously made a strong case for open-source AI, but Nathan does not expect an open-weight one under current leadership dynamics. Meta’s subsequent “reevaluating” language marks a profound change from July 2024’s open-source argument.

  • Nathan adds a community self-critique: intense backlash may have taught Meta that an expensive gift could generate worse headlines than no release. X discourse can arbitrarily punish one model while quietly using another, as with Grok 4.1 or Grok Code Fast 1.0.

48. The US open-model gap became an industrial-policy issue

  • Nathan’s ATOM Project—American Truly Open Models—rests on two claims: open models are the engine on which outside AI research begins, and the US should own that substrate so research, companies and economic value accumulate domestically.

  • His plots showed “Qwen, Qwen, Qwen, Qwen”: in July, four or five DeepSeek-caliber Chinese open models and none from the US. A model one generation behind the closed frontier might cost roughly $100 million—material, but small relative to industry spending.

  • AI2 received a $100 million NSF grant over four years, described as the agency’s largest computer-science award; NVIDIA increased emphasis on Nemotron and released some data; Reflection AI said its $2 billion raise would support US open models. Nathan wants multiple builders so no single Llama- or OLMo-like program can disappear.

  • The White House AI Action Plan’s support for open-source and open-weight systems helps set an agenda even before implementation. Lex adds the talent argument: without open models, researchers cannot learn until after joining a closed lab, making open source “the only way” to train and identify the next generation.

49. Chinese releases make restrictions less tenable

  • Lex proposes that Chinese frontier releases may improve US openness by proving capable weights can circulate without the predicted catastrophe; Sebastian agrees. Any genuinely open model has value, while Sebastian’s narrower claim is that the US should remain a leading source rather than surrendering the ecosystem.

  • Sebastian says those releases likely triggered leadership discussions that would not otherwise have occurred.

  • Banning open models would require something resembling a US great firewall, because $1 million-$100 million training budgets are available to many actors worldwide and knowledge cannot be contained. Both regard AI 2027-style Manhattan Project centralization as implausible in 2025-27.

  • If frontier progress saturates while capital remains abundant, optimized open architectures could eventually win: broad serving investment, dedicated chips and common standards would make them far cheaper than bespoke closed systems. If progress stays rapid, closed labs retain a moving quality frontier.

50. NVIDIA’s moat is CUDA, flexibility and Jensen’s operating system

  • Sebastian sees NVIDIA’s two-decade CUDA ecosystem as more defensible than an individual GPU. Labs already used Tesla GPUs for molecular simulation 15 years ago; at massive scale, customers prefer the compatible supplier over a risky chip with limited production.

  • Hyperscalers are still attacking the stack through Google TPUs, Amazon Trainium and Microsoft designs. Nathan’s condition is pace: while AI changes quickly, NVIDIA’s flexible platform wins; if progress stagnates, customers gain time to design cheaper specialized silicon.

  • Inference may split into specialized stages. Sebastian describes Vera Rubin hardware with little or no expensive high-bandwidth memory for pre-fill matrix multiplication, while memory-heavy autoregressive generation handles KV-cache movement elsewhere; the Groq deal fits the same specialization thesis.

  • Jensen Huang’s operational involvement resembles Steve Jobs-era Apple. NVIDIA also funds research and creates GPU-consuming markets, so long as its top organizational priority remains enabling the ecosystem rather than treating accelerators as one product line among many.

51. Singular leaders compress decades of technological progress

  • Nathan’s great-person compromise is that science may eventually discover the same idea, but focused individuals make it happen earlier. Jensen could have accelerated the GPU revolution by a decade; absent available GPUs, another AI winter might have delayed deep learning far longer.

  • Sebastian compares the effect to an individual stock versus an ETF: civilization eventually moves upward, but a concentrated leader produces larger, faster swings through passion and focus. Luck still matters—gaming created linear-algebra hardware before Alex Krizhevsky applied it to neural networks.

  • Ilya Sutskever and Dario Amodei likewise pushed the once-implausible bet of connecting roughly 10,000 GPUs and devoting OpenAI’s compute to one scaled model. The belief preceded conclusive evidence, which is precisely why leadership changed the timeline.

  • Looking back in a century, the speakers expect “computing” to matter more than CUDA details. Deep learning may remain a remembered term, while transformers could be one component later architectures evolved beyond; networking and the internet may merge into the broader story of connected compute.

52. The physical world gains value as synthetic content floods the digital one

  • In 100 years, Sebastian expects specialized robots and perhaps partially humanoid forms, but is less certain about interfaces. Brain-computer links may emerge, yet cars show that a useful interface can survive for more than a century with incremental improvement.

  • Lex expects some private physical compute object to remain, even if it no longer resembles a phone. Humans will still seek agency, community and meaning; mass wealth or UBI cannot by itself replace those needs.

  • Nearer term, they expect “more and more diverse versions of slop.” Physical art, goods and events acquire a premium because a person made or attended them, while new creators face a trust problem once synthetic work becomes indistinguishable.

  • Authentication could invert watermarking: devices might certify human-origin photos or edits rather than trying to mark every AI image. Any scheme becomes an arms race, making trusted outlets, relationships and in-person presence more valuable.

53. Human agency remains the closing safety thesis

  • Lex insists technological transition must be evaluated person by person: every lost job is a human tragedy even if aggregate GDP later improves. Better social support must acknowledge that suffering rather than explaining it away with new-job forecasts.

  • Nathan’s hope is historical: “Humans do tend to find a way.” Communities solve problems, but realizing AI’s opportunity will require long, fraught political conversations and builders willing to explain themselves to people who already distrust Big Tech.

  • Sebastian’s confidence comes from agency and consciousness. Present AI must be told what to do; it is more automatic and powerful than a hammer, but still a tool directed by a person. His main danger case is humans explicitly programming harmful objectives.

  • Lex extends the joke to a machine war—humans armed with local open-source LLMs—but the underlying call is serious: human cleverness, connection and moral choice remain worth defending. AI’s mirror may also clarify what consciousness is and why the “real miracle in our mind” matters.