Pioneers Insight Method Research Author
Back to Pioneers
Ramin Hasani
Founders 3 Curated Dialogues

Ramin Hasani

Liquid AI · CEO & Co-Founder

Frontier Insights

Core Frontier Thesis: Frontier AI’s future decouples from centralized clouds; native edge intelligence driven by non-Transformer architectures (LFM2) delivers brain-scale efficiency, zero hosting costs, and real-time privacy across automotive and industrial endpoints.

Strategic Moves: Liquid AI bypassed brute-force scaling, leveraging automated hardware-in-the-loop operator searches (AFMD) to deploy sub-1GB models directly onto client hardware, securing enterprise validation with Mercedes-Benz and Shopify.

Risks & Warnings: Severe hardware-model overcoupling could bottleneck cross-platform agility; bridging the capability gap to frontier cloud models remains perilous; and industry-wide production adoption hovers at a fragile ~11%, demanding radical workflow restructuring.

Key Views & Dialogues

Mira Murati’s 975B Open Model, Ramin Hasani on Post-Transformer AI, and Demis’ AI FINRA | EP #271

  • 🗓️ Date2026-07-17 | 🎙️ Show:Moonshots

Frontier AI governance is becoming a contested market structure: Demis Hassabis’s industry-funded, FINRA-style watchdog could improve pre-release testing, while critics warn of regulatory capture and exclusion of open-weight competitors. Thinking Machines Lab’s 975B-parameter Inkling bets on on-prem customization, activating 41B parameters with a 1M-token context window. Liquid AI’s sub-1GB Mercedes deployment and AI²’s self-reported, independently unconfirmed claims leave deployment economics, governance, and genuine recursive capability as key watchpoints.

View Dialogue Notes & Key Takeaways
  • A FINRA-style frontier-AI watchdog could provide safety cover, but the panel saw a serious risk of regulatory capture by the labs designing the rules. Demis Hassabis wants an industry-funded body testing frontier models before release, reportedly operational by year-end; Ramin Hasani instead argued for capability- and vertical-specific governance that iterates like a Stackelberg game. Alexander Amini’s sharper objection was that a frontier-lab cartel could lock out open-weight competitors while “thought policing the AIs,” whereas liability for harmful actions may be the cleaner lever.

  • Pegging permissible US open-weight releases to China’s best public model would hand Beijing the throttle on American innovation. The proposed framework assumes Chinese models trail by about seven months and cannot be “unshipped” after millions of downloads, but it would reward China for moving first and could drive Western researchers toward Chinese labs. Amini called it the game-theoretic equivalent of “throwing the steering wheel out the window in a game of chicken.”

  • Thinking Machines Lab’s Inkling is a 975B-parameter bet that enterprise customization matters more than topping global leaderboards. The model activates 41B parameters at a time, was trained on 45T multimodal tokens, reportedly has a 1M-token context window, and can be downloaded, fine-tuned, and run on-prem; the episode placed it above NVIDIA’s Nemotron 3 but below China’s GLM-5.2. The commercial thesis is “customization over leaderboard dominance,” potentially monetized through fine-tuning that generates one to two orders of magnitude more tokens.

  • Thinking Machines Lab is wagering that proprietary enterprise data will revive fine-tuning just as baseline models threaten to make it unnecessary. Dave Blundin argued that sending payroll, chemical research, or defense data to a closed API means “Sam and Dario can see everything,” making locally owned weights compelling for banking, defense, biotech, and automotive. A later guest preserved the bearish case: increasingly general models may need only prompting, leaving reinforcement fine-tuning as merely “the paradigm of the moment.”

  • The claimed recursive-self-improvement breakthrough exposed a crucial divide between optimizing workflows and changing an AI’s underlying intelligence. Weco AI’s AI² system reportedly turned eight days of machine work into more progress than two years of expert effort, with an outer agent improving and policing an inner agent; Hasani countered that fixed-weight models rewriting prompts and code are not genuine recursive self-improvement. His computational warning was stark: applying that framework to meaningfully retune a 2B-parameter model could take roughly 350 years.

  • Hasani nevertheless expects models “going beyond our understanding” within roughly two years if compute keeps expanding without chip or memory shortages. He separated shallow self-improvement in prompts, code, and kernel optimization from deeper fine-tuning and the “holy grail” of automating pretraining for a model’s successor. Peter Diamandis’s broader interpretation was more immediately commercial: AI need not redesign its own weights to trigger an “organizational singularity” if it can recursively redesign enterprise workflows.

  • Liquid AI’s edge is putting specialized multimodal intelligence into hardware that cannot support frontier-scale models. Hasani defined small models as below roughly 100B parameters, then described a Mercedes model under 1GB running on chips with 2–8GB of RAM that may cost about $60; a 600MB over-the-air update is intended for North American Mercedes vehicles from 2022 onward as soon as this year. Its access to 700–1,200 vehicle functions makes “intelligence outside of data centers” the investable deployment thesis.

  • AI is simultaneously collapsing the cost of medical judgment and opening previously permanent categories of aging damage to intervention. The episode said GPT-5.6 beat specialty-matched physicians across roughly 20,000 judgments, while Meta’s Muse Spark 1.1 then beat GPT-5.6 on the 525-task benchmark at one-seventh the cost and could reach 3.56B daily users—though Hasani suspected “mild benchmark-maxing.” Separately, Revel Pharmaceuticals and Calico’s CMLA enzyme reportedly reversed advanced-glycation damage in elderly human tissue: a “molecular lawnmower” for scars previously treated as irreversible.

  • 🔗 Original source & video: Mira Murati’s 975B Open Model, Ramin Hasani on Post-Transformer AI, and Demis’ AI FINRA | EP #271

Listen to full conversation →


Intelligence on the Edge: Liquid AI’s Ramin Hasani on the Search for Device-Native Foundation Models

  • 🗓️ Date2026-07-04 | 🎙️ Show:The Cognitive Revolution

Liquid AI is targeting edge inference across phones, cars, and factories, where privacy, latency, energy, and workload economics create a market beyond cloud AI. Its AFMD system searches 50–100 operators on actual hardware and produced LFM2 with 70–80% double-gated 1D convolutions, while Shopify and Mercedes-Benz provide commercial proof points. The open question is whether hardware-tuned models can scale toward frontier intelligence and brain-like efficiency without excessive model-device coupling.

View Dialogue Notes & Key Takeaways
  • Liquid AI’s commercial thesis is that edge hardware is an underused inference market, not merely a cheaper copy of cloud AI. Hasani sized smartphones alone at roughly $500 billion annually; Labenz framed phones plus laptops as about $1 trillion of compute shipped each year, while the introduction used roughly $800 billion. Energy limits, privacy, latency, and workload economics all support Liquid’s aim to “build an intelligence layer on top of the diverse formats of hardware” already in pockets, cars, factories, and other devices.

  • Liquid’s central technical finding is that the best architecture depends on scale, specialization, and the hardware constraint. Attention remains the richest, least-structured mechanism and may justify its (n^2) cost at frontier scale, while smaller or narrower models benefit from recurrence, convolutions, gating, or domain-specific dynamics. “The larger the network becomes, the more unstructured you can make it”; conversely, constrained systems can trade generality for substantially better speed and memory efficiency.

  • The scientific lineage began with biologically inspired differential equations that produced striking control systems from only tens of neurons. Liquid networks parallel-parked a small car with 12 neurons, drove with 19, and flew a drone with 30; later work extended the approach to jets and other predictive systems. Their advantage was input-dependent dynamics and out-of-distribution adaptability—not magic or continual learning: “There’s no free lunch,” and the trained parameters remain fixed.

  • Liquid’s automated foundation-model design system turns architecture selection into a hardware-grounded search problem. AFMD evaluates roughly 50–100 operators and hybrid combinations using an evolution strategy, actual target processors, memory and latency constraints, and around 100 downstream benchmarks rather than perplexity alone. After searches spanning roughly 10 million to 72 billion parameters, the lesson was “You have to give it to the algorithms,” including when the algorithms discard Liquid’s founders’ own preferred mechanisms.

  • LFM2’s winning CPU design is surprisingly simple: 70–80% double-gated 1D convolutional layers, plus a smaller allocation to attention. The gate makes computation input-dependent; the convolution supplies a cheap, unstructured operator, replacing much of attention’s memory and quadratic cost without the elaborate hand-tuned machinery found in many alternative architectures. Hasani’s blunt account of the search result: “All of this has to go away” when extra gates and human-selected features fail the full efficiency objective.

  • Commercial proof points suggest Liquid has moved beyond an architecture experiment. Hasani reported more than 1 million weekly Hugging Face downloads, the No. 5 position among U.S. organizations behind Google, Meta, Microsoft, and NVIDIA, more than 50 model instantiations used in enterprises, and only about 1,000 in-house GPUs. Shopify uses Liquid models in production across commerce workloads, Mercedes-Benz signed a contract for a roughly 600-megabyte in-car audio and visual system, and Labenz found the 1-billion-parameter Apollo model fast enough on an iPhone for private document search and classification.

  • Silicon vendors may need to own a tunable intelligence layer, not stop at chips and kernels. Liquid is working with companies including AMD and Qualcomm around processor road maps, while Hasani points to NVIDIA’s Nemotron effort as evidence that models optimized for hardware can improve the enterprise proposition. Labenz challenged whether this creates excessive model-hardware coupling; Hasani’s answer was essentially, “Why do you want to change the model?”—provided the default is fast, tunable, and does not block the wider open-source ecosystem.

  • The immediate local-agent opportunity is orchestration, while true intelligence-per-watt miniaturization needs new learning paradigms. A proposed local LFM2 24B A2B system can route requests, filter PII, invoke specialized models, and escalate difficult work to the cloud, but Hasani says no current local model matches frontier quality without fine-tuning; forthcoming tooling could create production models for tens to low thousands of dollars. Longer term, current architectures will not approach the brain’s roughly 20-watt efficiency because “intelligence for me is an emergent property,” requiring objectives that induce multiple ways of learning—not next-token prediction alone.

  • 🔗 Original source & video: Intelligence on the Edge: Liquid AI’s Ramin Hasani on the Search for Device-Native Foundation Models

Listen to full conversation →


AI Leaders Reveal the Next Wave of AI Breakthroughs (At FII Miami 2025) | EP #150

  • 🗓️ Date2025-02-20 | 🎙️ Show:Moonshots

Stability AI is converting Stable Diffusion’s 270 million downloads into 50–60 specialized production models for film, TV, gaming and advertising, with convincing on-demand video estimated in six to 12 months. Liquid AI’s private, on-device systems promise “$0” hosting costs across phones, cars, satellites and jets. Tenstorrent targets systems 5–10 times cheaper through open infrastructure, while enterprise adoption remains the risk: only 11% of use cases reached production, and generative AI reached “maybe 7% at the best.”

View Dialogue Notes & Key Takeaways
  • Stability AI is shifting Stable Diffusion’s distribution lead—launched in August 2022, with 270 million downloads versus 9 million for the next model—into specialized production tools for film, TV, gaming and advertising. Prem Akkaraju said roughly two dozen of a planned 50–60 “ultra-narrow AI” models address workflows such as rig removal, paint and rotoscoping, and camera match-plate construction. He estimated convincing on-demand video within six to 12 months and argued that Hollywood is “confusing headwind with tailwind”; he pointed to tools appearing in Avatar 3, 4 and 5.

  • Liquid AI’s bet is that private, on-device intelligence can expand the market beyond GPU-equipped clouds. Ramin Hasani described a non-transformer architecture derived from liquid neural networks that can deliver ChatGPT-like experiences locally on phones and laptops, while also powering cars, satellites and jets. Once deployed on a device, he said, hosting costs “$0”; future robots could have local AI brains rather than rely on the cloud. Peter said Liquid AI had gone from zero to a $2 billion valuation in about two years, and noted a quarter-billion-dollar round led in part by G42.

  • SandboxAQ is targeting quantitative industries where language models cannot perform the underlying physics. Jack Hidary said the company has raised $850 million and argued that drug and materials discovery needs models “trained on molecules and atoms,” with quantum equations translated into GPU-compatible matrix algebra. Quantum computers may join GPUs in a hybrid GPU/QPU cloud in five to seven years, but useful large quantitative models, or LQMs, run on GPUs today; Aramco is its newest announced customer.

  • Tenstorrent is attacking AI’s compute cost and lock-in with native tensor processors and an open software stack. Peter noted its recent $700 million Series D. Jim Keller’s target is systems 5–10 times cheaper than current systems, spanning small television-chip configurations through large-model training machines. His thesis is that AI need not be “unbelievably expensive, unbelievably big, [or] unbelievably proprietary,” and that open infrastructure will broaden adoption and innovation.

  • Enterprise AI remains more pilot theater than production, according to figures Sukharevsky cited. He said only 11% of use cases reached production over the past five years, with generative AI at “maybe 7% at the best,” because companies insert technology into broken processes instead of redesigning them. QuantumBlack, he said, has 5,000 people in 50 countries, five R&D centers and roughly 43 products deployed globally. Senior sponsorship, data, architecture and organizational politics are central constraints.

  • The panel offered two execution prescriptions: focused teams that start with concrete gains, and institution-wide commitment. Hidary cited an 11-person team solving GPS-denied navigation with AI and quantum sensors and urged responsible, faster adoption to tackle diseases and battery storage. Keller asked his software organization to double productivity and first produce code with fewer bugs. Sukharevsky agreed with Peter’s warning that companies failing to use AI may be out of business by decade’s end, but said transformation requires top-level commitment: “you cannot go small—you need to go big to succeed.”

  • 🔗 Original source & video: AI Leaders Reveal the Next Wave of AI Breakthroughs (At FII Miami 2025) | EP #150

Listen to full conversation →