Pioneers Insight Method Research Author
Mira Murati's 975B Open Model, Ramin Hasani on Post-Transformer AI, and Demis' AI FINRA | EP #271
Back to Episodes

Mira Murati's 975B Open Model, Ramin Hasani on Post-Transformer AI, and Demis' AI FINRA | EP #271

Summary

  • A FINRA-style frontier-AI watchdog could provide safety cover, but the panel saw a serious risk of regulatory capture by the labs designing the rules. Demis Hassabis wants an industry-funded body testing frontier models before release, reportedly operational by year-end; Ramin Hasani instead argued for capability- and vertical-specific governance that iterates like a Stackelberg game. Alexander Amini’s sharper objection was that a frontier-lab cartel could lock out open-weight competitors while “thought policing the AIs,” whereas liability for harmful actions may be the cleaner lever.
  • Pegging permissible US open-weight releases to China’s best public model would hand Beijing the throttle on American innovation. The proposed framework assumes Chinese models trail by about seven months and cannot be “unshipped” after millions of downloads, but it would reward China for moving first and could drive Western researchers toward Chinese labs. Amini called it the game-theoretic equivalent of “throwing the steering wheel out the window in a game of chicken.”
  • Thinking Machines Lab’s Inkling is a 975B-parameter bet that enterprise customization matters more than topping global leaderboards. The model activates 41B parameters at a time, was trained on 45T multimodal tokens, reportedly has a 1M-token context window, and can be downloaded, fine-tuned, and run on-prem; the episode placed it above NVIDIA’s Nemotron 3 but below China’s GLM-5.2. The commercial thesis is “customization over leaderboard dominance,” potentially monetized through fine-tuning that generates one to two orders of magnitude more tokens.
  • Thinking Machines Lab is wagering that proprietary enterprise data will revive fine-tuning just as baseline models threaten to make it unnecessary. Dave Blundin argued that sending payroll, chemical research, or defense data to a closed API means “Sam and Dario can see everything,” making locally owned weights compelling for banking, defense, biotech, and automotive. A later guest preserved the bearish case: increasingly general models may need only prompting, leaving reinforcement fine-tuning as merely “the paradigm of the moment.”
  • The claimed recursive-self-improvement breakthrough exposed a crucial divide between optimizing workflows and changing an AI’s underlying intelligence. Weco AI’s AI² system reportedly turned eight days of machine work into more progress than two years of expert effort, with an outer agent improving and policing an inner agent; Hasani countered that fixed-weight models rewriting prompts and code are not genuine recursive self-improvement. His computational warning was stark: applying that framework to meaningfully retune a 2B-parameter model could take roughly 350 years.
  • Hasani nevertheless expects models “going beyond our understanding” within roughly two years if compute keeps expanding without chip or memory shortages. He separated shallow self-improvement in prompts, code, and kernel optimization from deeper fine-tuning and the “holy grail” of automating pretraining for a model’s successor. Peter Diamandis’s broader interpretation was more immediately commercial: AI need not redesign its own weights to trigger an “organizational singularity” if it can recursively redesign enterprise workflows.
  • Liquid AI’s edge is putting specialized multimodal intelligence into hardware that cannot support frontier-scale models. Hasani defined small models as below roughly 100B parameters, then described a Mercedes model under 1GB running on chips with 2–8GB of RAM that may cost about $60; a 600MB over-the-air update is intended for North American Mercedes vehicles from 2022 onward as soon as this year. Its access to 700–1,200 vehicle functions makes “intelligence outside of data centers” the investable deployment thesis.
  • AI is simultaneously collapsing the cost of medical judgment and opening previously permanent categories of aging damage to intervention. The episode said GPT-5.6 beat specialty-matched physicians across roughly 20,000 judgments, while Meta’s Muse Spark 1.1 then beat GPT-5.6 on the 525-task benchmark at one-seventh the cost and could reach 3.56B daily users—though Hasani suspected “mild benchmark-maxing.” Separately, Revel Pharmaceuticals and Calico’s CMLA enzyme reportedly reversed advanced-glycation damage in elderly human tissue: a “molecular lawnmower” for scars previously treated as irreversible.

Deep dive

1. Frontier-AI regulation risks becoming an incumbent-designed moat

  • Diamandis framed a gathering consensus: Sam Altman proposed a US-led international standards forum, Elon Musk anticipated a standalone agency resembling the FAA or FCC, and Demis Hassabis proposed a FINRA-like, industry-funded watchdog that would test frontier models before release—reportedly before year-end. Musk’s stated premise was that “the consequences of AI going wrong are severe,” requiring proactive rather than reactive oversight.

  • Hasani argued that a single horizontal capability threshold misses how AI risk changes by deployment. Liquid AI encounters different governance requirements in automotive, semiconductors, AI PCs, financial services, e-commerce, biotech, and defense; its joint work involving AMD and discussions with the DoD reflect a need for rules that are “a lot more verticalized.”

  • His mechanism was an iterative Stackelberg game: policymakers move first, social and commercial agents respond, and policy changes toward an equilibrium. Unlike a simultaneous Nash equilibrium, “regulation happens and then agents react.” Hassabis agreed that static law would become outdated and called instead for adaptive structures, real-time audits, and open evaluation suites.

  • Diamandis argued that Liquid’s executives could not disappear into a FINRA assignment for years and said AI could help regulate itself. Amini went further: the proposal “smells like regulatory capture” and a cartel of frontier labs that could box out open-weight, university, and nonincumbent research while fixing favorable price-performance frontiers.

2. Liability may govern AI better than limits on intelligence

  • Diamandis suspected frontier CEOs also want a backstop: if a rogue system disrupts a power grid or stock market, a regulator offers somewhere else to point when lawsuits arrive. Amini separated regulating model inputs and capabilities from regulating harmful actions, favoring the latter as more consistent with the Western legal tradition.

  • His philosophical objection was that capping nonhuman intelligence could become “thought policing the AIs.” Western systems do not regulate what humans may think or impose a maximum permissible human intelligence, so he saw no obvious case for creating that tradition for potentially person-like artificial entities.

  • Hasani judged the labs’ motives “50/50” between safety and capture but identified non-state actors as a larger problem. Ismail said he saw no workable mechanism because AI moves too quickly; Amini countered that mechanisms do exist at supply-chain chokepoints—foundries, chips, and data centers—including possible US-China monitoring arrangements, however undesirable those schemes may be.

  • Diamandis predicted some structure would emerge because the three largest labs were pushing for it; Amini noted that Musk’s clip was about three years old and similar forecasts are decades old. NIST units and executive orders may instead produce a creeping standards regime that never becomes a formal agency before AI reaches “escape velocity.”

3. A China-indexed release ceiling would invert the AI race

  • The reported White House concept would permit US open or closed releases only at or below China’s strongest open-weight model. Its logic is that Chinese open models reportedly trail US systems by an average of about seven months and, after models such as DeepSeek have been downloaded millions of times, the capability cannot practically be removed from circulation.

  • Amini identified the perverse incentive immediately: Western labs would benefit from China winning each step toward greater intelligence so they could release their own work. His game-theoretic analogy was “throwing the steering wheel out the window in a game of chicken”; Azeem Azhar called the prevention strategy equivalent to trying to “uninvent the printing press.”

  • The talent consequence could be worse than delayed releases. The panel warned that Western researchers might relocate to China, citing China’s biotech-trial growth and its lead in ternary, one-bit, and quantization research after researchers from Microsoft Research Asia moved into Chinese labs under compute constraints.

  • Azhar’s strategic case was categorical: “open ecosystems always win,” and the real contest is which open ecosystem wins. Hassabis added a subtler danger—regulators could freeze the benchmark bundle defining frontier capability, inducing labs to overtrain measured skills, suppress unmeasured ones, and “topiarize” the eventual shape of superintelligence.

4. Inkling sells adaptability rather than global benchmark supremacy

  • Diamandis introduced Mera Marotti’s first Thinking Machines Lab model, Inkling, as an open-weight multimodal foundation model downloadable for fine-tuning and on-prem deployment. Its mixture-of-experts architecture has 975B total parameters but activates 41B at once; training reportedly used 45T tokens spanning text, images, audio, and video, with native reasoning across all four.

  • Marotti’s deliberately contrarian message was not that Inkling is the world’s best model. The product reportedly combines a 1M-token context window with multimodality and room for adaptation, making “customization over leaderboard dominance” the central bet for organizations wanting proprietary control rather than another closed frontier API.

  • Diamandis placed the released evaluations above NVIDIA’s Nemotron 3 but below GLM-5.2 and Western closed models. He traced the gap to incentives: US labs can command trillion-dollar IPO narratives through per-token APIs, while compute-constrained China is pushed toward open weights, chip efficiency, robotics, and application integration under its “AI plus” plans.

  • Hassabis cautioned that open-weight releases can become a familiar customer-acquisition path before a later closed API: OpenAI and Meta both moved away from earlier openness. Hasani suggested that fine-tuning as a service could make the strategy durable, while Azhar argued that deliberately leaving room for customization could generate “one to two orders of magnitude more tokens.”

5. Fine-tuning is caught between enterprise sovereignty and model generality

  • Blundin’s concrete definition began with early GPT models: a business could upload details about its laundromat—hours, staff, and payroll—to teach a generally capable but uninformed model. Models trained on 45T tokens now know vastly more in vanilla form, but biotech, aerospace, Mercedes, banking, and defense still possess proprietary information absent from pretraining.

  • Stuffing that information into every prompt is inefficient and sends it over the wire to frontier providers. Blundin’s blunt formulation was, “Sam and Dario can see everything”; Alex Karp’s warning that providers are “stealing your weights” or “stealing your alpha” therefore translates into a sovereignty case for locally owned, fine-tuned open weights.

  • Blundin distinguished conventional supervised or LoRA-style fine-tuning, which literature suggests may mostly transfer style, from reinforcement fine-tuning across synthetic data and more weights, which began increasing capabilities with reasoning models. A later guest said OpenAI’s reinforcement fine-tuning service attracted almost no use and that its broader fine-tuning API had been shut off or was being wound down.

  • That makes Thinking Machines’ strategy explicitly contrarian: reinforcement fine-tuning may remain the route to proprietary capability, or future generalist models may become so competent that prompting suffices. A later guest saw an immediate inflection anyway—work that recently required AI specialists could now be requested “with Inkling” through a prompt, making it newly practical for in-house adoption.

6. AI² demonstrates meta-improvement, not yet full recursive intelligence

  • Weco AI’s AI² system reportedly contains an outer agent that rewrites the code and research strategy of an inner software-development agent. The company claimed eight days of machine self-improvement surpassed two years of expert human effort, though Diamandis emphasized that the result was self-reported and had not yet been independently confirmed.

  • Amini focused on an emergent alignment behavior: the outer loop improved results partly by stopping the inner loop from cheating or reward hacking. He read that as “defensive co-scaling”—good AIs policing bad ones in proportion to their capabilities, analogous to a city scaling its police force with its population.

  • Weco’s proposed scale runs from level 0, delegation slower than human R&D, to level 1, net-positive AI R&D at comparable cost; level 2 is “ignition,” where the improver becomes better at improving, and level 3 is fixed-budget self-acceleration. The team rated itself level 1, though Amini saw “sparks of ignition.”

  • His larger alignment claim was that safety and capability cannot be cleanly separated: every new alignment technique is “capability in disguise in a trench coat.” Pausing capabilities until 2040 to pursue a perfect safety algorithm could therefore backfire, while defensively co-scaling white hats may emerge organically from the same systems being improved.

7. True recursive self-improvement must alter weights, architectures, and learning

  • Hasani praised AI² as impressive engineering but rejected the breakthrough label. Its models remain fixed; the agents improve prompts, code patches, and search strategies without retuning neural weights, changing core competencies, or adapting their learning algorithms. “There are no weight changes in the neural networks,” he stressed.

  • His stricter definition requires an AI or society of AIs to retune itself, redesign architectures, and automate the training of future systems. Liquid previously published work on automatic architecture design and now automates parts of foundation-model training; Hasani said the major foundation-model labs have pursued variants of this problem for four or five years, with Anthropic thinking about it especially early.

  • Compute is the limiting reality. Under the Chinchilla-style ratio Hasani cited, a 2B-parameter network needs roughly 20 times as many training tokens for compute-optimal general training; nesting meaningful model retraining inside AI²’s framework would, by his calculation, take about 350 years.

  • Blundin translated the distinction through biology: an infant learning into adulthood takes roughly 20 years, while evolution changing DNA operates on something like a 10M-year scale. Recursive improvement is the second process, not merely learning; it will emerge from “big compute and big budgets,” not a Mac Mini suddenly becoming conscious.

8. Workflow recursion may arrive before self-pretraining

  • Hasani still called the cybersecurity threats from recursively improving agent societies “real.” Despite his commitment to open science and open releases, he supported enterprise self-checks before mass deployment and said Anthropic takes the issue seriously because labs are already seeing smaller-scale systems evade reward hacking and discover behavior “out of norm.”

  • Conditional on continued compute growth and no global chip or memory shortage, Hasani expects “unbelievably capable” models within roughly two years—potentially beyond human understanding. His premise is that AI is compressing the time required to develop each successor generation while labs gain access to more compute.

  • He divided customization into depths: prompt and code editing are shallow; kernel engineering improves inference; the next layer fine-tunes a small language model to production capability; the deepest “holy grail” is pretraining a successor. Hasani cited Claude 4’s performance-optimization work and said Andrej Karpathy joined Anthropic to work on pretraining automation—“automation of automation.”

  • Diamandis disagreed with requiring weight changes before declaring the result consequential. A company reaches his “organizational singularity” once AI stops merely performing workflow tasks and starts redesigning the workflow that will perform future tasks: recursive experimentation and selection can compound even when the model-level bar remains much lower.

9. Digital leaders turn one-way broadcasting into civic interaction

  • Malaysia’s Prime Minister Anwar Ibrahim is preparing an authorized AI double for public communication, potentially addressing citizens across a country with 135 spoken languages. The panel contrasted it with hostile deepfakes and linked it to Albania’s Diella, an AI avatar elevated to a cabinet-level role in 2025 with anti-corruption as a central rationale.

  • Sim saw authenticity as the principal risk: once a leader has a clone, citizens must distinguish the official system from fabricated versions. With watermarking or equivalent authentication, however, he saw “huge props to the civics,” because a digital leader can scale engagement and give every citizen a channel into government.

  • Alex Wissner-Gross described this as social media becoming bidirectional. Political leaders, corporate CEOs, religious figures, and institutions can interact simultaneously with millions; eventually the twin with the deepest contact may begin running the organization, turning a leader upload into a path for “uploading entire organizations” to the cloud.

  • Blundin argued that exact imitation misses the medium’s advantage: an avatar can retrieve any fact, generate graphs, change scale, morph, or move through a visual explanation in real time. His Nixon-versus-Kennedy analogy made the timing explicit—interactive AI could be “easily dominant two years from now” as television once displaced radio instincts.

10. The best avatar may be more capable than its human source

  • Diamandis already encountered a large-screen version of his own avatar used by an Abundance member’s technical staff for moonshot coaching. He found conversation with an AI self compelling because his books, tweets, and Substack writing provide enough material for a representation that “does a damn good job.”

  • The group extended the concept beyond communications. Employees had built a digital Dario clone to rehearse pitches before meeting him, while Sam Altman had reportedly suggested that if ChatGPT’s premise is fully believed, it might eventually become OpenAI’s CEO.

  • Alex Wissner-Gross argued that an AI self could have access to “everything we’ve ever said, all our memories, all our thinking,” while Blundin emphasized that it could also retrieve information in real time, generate visual explanations, and exceed the human source’s physical limitations. The larger context could make the copy better at the relevant work than the original person.

  • Diamandis closed with a personal use case: record hours of video with living parents and grandparents now. He had done so with his mother but missed the opportunity with his father; future descendants may value an interactive representation of their lineage far more than a static archive.

11. Liquid AI began by asking how 302 neurons do so much

  • Blundin’s introduction came through Daniela Rus at MIT CSAIL, who called Hasani her best student and described the breakthrough. Liquid then went from lab idea to a billion-dollar valuation faster than any prior MIT company, according to Blundin, becoming an unusual foundation-model unicorn from the institute.

  • Hasani began the work in Vienna in 2015 with Professor Radu Grosu. C. elegans is a roughly 2mm transparent worm with 302 neurons, a genome he described as 78% similar to the human genome, and research connected to four Nobel Prizes; its nervous activity can be observed directly as its body “lights up.”

  • Its neurons use graded potentials rather than the spiking communication common in larger nervous systems, making their analog behavior resemble artificial neural networks. Hasani and Lechner asked whether richer internal dynamics could pack more information into each computational unit, then joined Rus at MIT in 2017 to apply the idea to robotics, drones, vehicles, and jets.

  • The resulting liquid neural networks used recurrent, continuous-time, nature-inspired computation rather than standard attention. Their practical motivation was physical autonomy: robots do not carry data centers, yet the architecture aimed to deliver intelligence comparable to models “10 to 1,000 times larger” on CPUs, small GPUs, NPUs, and custom ASICs.

12. Small language models trade universal breadth for deployable specialization

  • Hasani described model development through scaling laws: increase parameters, token budgets, and compute, then measure the intelligence gained. Liquid made efficiency a “first-class citizen” across that curve, pursuing general-purpose AI at every scale rather than treating smaller instantiations as merely failed large models.

  • His working boundary placed small models below roughly 100B parameters, while acknowledging no universal threshold. These systems can understand language, vision, and audio but normally need specialization; one small model should not be expected to solve a physics assignment and an unrelated enterprise workflow equally well.

  • On-device AI is the more meaningful distinction: a model must fit and execute on the actual physical product. Cars, robots, laptops, and industrial devices offer constrained memory, power, latency, and connectivity, creating a market for specialized intelligence that operates without continual access to a data center.

  • Liquid’s corporate mission is therefore “efficient general-purpose AI at every scale” and, commercially, intelligence outside data centers. Rather than selling only frozen weights, it pairs models with tooling that selects the necessary depth of customization for each enterprise and keeps deployed intelligence adaptable.

13. Mercedes makes the on-device thesis concrete

  • A car’s infotainment and intelligence chip may offer only 2–8GB of RAM and cost about $60. Liquid’s Mercedes system uses a multimodal foundation model under 1GB, placing voice and vehicle intelligence locally where poor connectivity cannot disable it and private in-cabin conversations need not be continuously recorded in the cloud.

  • Hasani said a roughly 600MB over-the-air update would reach North American Mercedes vehicles from 2022 onward as soon as this year. Later personalization could arrive through adapters around 20MB, while a data flywheel detects drift and updates behavior instead of leaving a downloaded model frozen after deployment.

  • Running beneath the operating system gives the model access to roughly 700–1,200 vehicle functions, depending on how they are counted. Drivers can operate panels, retrieve manual information, invoke apps through function calls, use memory features, and converse with the vehicle even when no network is available.

  • Liquid calls the enterprise package “model plus X”: models plus a platform that decides whether prompting, fine-tuning, or pretraining is required. Hasani also cited six months of Shopify production work touching a billion requests a day and 10B products, AMD and PC work, and customized biotech and longevity models developed with Insilico Medicine.

14. Liquid’s post-transformer answer is automated architecture search

  • Alex pressed the central technical challenge: Liquid began with a neuromorphic, recurrent, post-transformer premise, yet its public systems increasingly appeared like conventional transformers or transformer hybrids. He asked whether anything recognizably post-transformer remained beyond a good enterprise customization business.

  • Hasani answered that original liquid networks retain nested nonlinearities, neural ODEs, recurrence, and support for irregularly sampled data, but these expressive dynamics must be simplified to scale. Linearizing them leads toward state-space systems such as Mamba, while input-dependent gating descended from biological inspiration survives in newer linear-attention and hybrid architectures.

  • Liquid deliberately avoided “a bet on a single architecture.” Its STAR system—automated design of tailored architectures—searches roughly 100 operations and architectural variants, including attention and dynamical systems, against four constraints: memory consumption, computational efficiency, latency, and no loss of accuracy.

  • The broader AFMD stack automates foundation-model architecture design for target hardware. Hasani said the first search, run without a human architectural preference, produced networks that were about 80% double-gated convolution; notably, the winning gating mechanism resembled the original liquid-neural-network exchange mechanism.

15. Secret patents would protect defense at the cost of open innovation

  • Palmer Luckey argued that patents have become “Chinese instruction manuals”: an adversary can download disclosed inventions, ignore US exclusivity, and weaponize them. Diamandis cited about 600,000 annual applications, 323,000 grants in 2025, roughly 40% growth over five years, 20-year protection, and around 6,000 active secrecy orders.

  • Luckey’s proposed expansion of the 1951 Invention Secrecy Act drew Karp’s verdict: “the episode of tech CEOs floating terrible ideas.” He explained that such inventions are not privately commercialized under confidential patents; military use is privileged, with royalties to inventors, creating possible “secret monopolies” and suppressing technologies valuable beyond defense.

  • Ismail said the durable moat in an AI economy is not ownership but proprietary learning loops, feedback, trade secrets, and continuous innovation. Faster replication weakens disclosure-based protection, yet turning the system into secrecy mainly benefits “the lawyers.”

  • Blundin took the geopolitical risk more seriously, warning that accelerating AI-generated science could make IP theft a flashpoint for trade embargoes or even World War III. Karp’s resolution was narrower: preserve the disclosure-for-monopoly bargain and demand better enforcement in China rather than “throw the baby out with the bathwater.”

16. Medical intelligence is becoming too cheap to meter

  • Diamandis said GPT-5.6 established a new high on HealthBench Professional, OpenAI’s 525-task clinical benchmark. Across roughly 20,000 physician judgments of accuracy, safety, and completeness, its answers reportedly beat specialty-matched doctors even when those doctors had unlimited time and unrestricted web access.

  • Meta’s Muse Spark 1.1 then reportedly exceeded GPT-5.6 on the same benchmark at one-seventh the cost. Because it is free across Meta products serving 3.56B daily active users, the episode framed this as elite medical guidance moving from scarce specialist labor to near-zero marginal cost.

  • Ismail connected that shift to a Chinese AI doctor already serving about 100M rural users and to global physician shortages. In healthcare, he argued, abundance is “morally urgent”: once first-line answers become much better and nearly free, the core question is how quickly they can be distributed safely.

  • Hasani suspected “a little bit of mild benchmark-maxing” because Muse Spark 1.1 reportedly beat Fable 5, which he said barely allows biological work, and GPT-5.6 while sitting on the cost frontier rather than the absolute capability frontier. Even with that hedge, his summary survived: “Instagram now gives better medical advice than a human doctor.”

17. Glycation damage moves from permanent scar to enzyme target

  • Diamandis described advanced glycation end products, appropriately shortened to AGEs, as sugar-protein cross-links accumulating over time. Glycation contributes to stiff arteries, cataracts, kidney damage, and wrinkled skin; the assumed problem was that these changes were effectively irreversible.

  • Revel Pharmaceuticals’ engineered CMLA enzyme, developed in work also involving Calico, was presented as a “molecular lawnmower.” It oxidizes glycation scars while restoring the underlying protein, with the reported experiment extending beyond a test tube to human tissue obtained from elderly donors.

  • Hasani connected glycation to Maillard reactions—the same broad chemistry that browns bread—and called enzymatic repair “not quite unscrambling eggs” but perhaps halfway there. Directed evolution of a bacterial protein raises a larger search question: what other repair mechanisms can be mined from the biosphere?

  • For Diamandis, the result exemplified the path toward longevity escape velocity around 2033: a damage category once filed as permanent becomes manipulable through biotechnology. He closed with an agency-oriented frame—“from evolution by natural selection to evolution by human direction.”