Pioneers Insight Method Research Author
Back to Pioneers
Verifiable Rewards
Innovators 1 Curated Dialogues

Verifiable Rewards

Key Views & Dialogues

🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences

  • 🗓️ Date2026-07-16 | 🎙️ Show:Latent Space

Lila is betting that controlled experiments can become AI’s next internet-scale training corpus, with nature providing verifiable rewards and an information-gain-driven lab turning experiments into a compounding model moat. Lila reports a six-month, two- or three-person in vivo CAR-T program versus roughly six years and $100 million, but clinical translation, scale-up, regulation, and 5–6% model FLOPs utilization remain key watchpoints.

View Dialogue Notes & Key Takeaways
  • Lila’s core bet is that controlled experiments can become AI’s next internet-scale training corpus, with nature itself supplying verifiable rewards. The internet was “the fossil fuel we fracked,” while scientific reinforcement learning lets models propose experiments, observe reality, and create better training data. The resulting flywheel—not any single drug or material—is the company’s intended moat.

  • The operating system is optimized for information gain and iteration speed, not maximal robotic throughput. Instruments form a graph connected by a PCI-bus-like transport layer, while every action is an API call whose executor might be “a robot arm” or “a human arm.” Lila calls itself “token generation maximalists and flexibility maximalists”: the next experiment must teach the model something valuable, not merely add another low-information sample.

  • Early results suggest genuine capability lift, but the line between foolish and novel remains deliberately porous. In expression and gene-editing tasks, Lila reports the model getting “like 80%” zero-shot versus humans at 0%; proposed non-platinum-group electrocatalysts progressed from boring to apparently stupid, then became its best performers. That upside comes with rigorous reruns, environmental telemetry, tool restrictions, and acceptance that informative false positives can waste time.

  • The clearest commercial proof point is an in vivo CAR-T program compressed into six months by two or three people. The comparison point involved roughly six years and $100 million of prior R&D, while Lila reports “monster” UTRs at about 10× Moderna and Pfizer references and superior non-human-primate B-cell depletion and durability relative to the Capstan data. Lila will not run the clinical trial; it converts such proof points into fee-plus-upside “zero FTE startup” partnerships.

  • Generalization across scientific domains is the economic thesis, not a branding flourish. Lila has generated 10 trillion model-produced, experimentally verified reasoning tokens spanning life sciences, chemistry, and materials, and says its general model often beats domain-specific alternatives. Chemistry learned in small-molecule drug discovery has transferred into applications such as metal-organic frameworks. Rafa says language need not be the necessary representation for every scientific modality, while Andy emphasizes token-based reasoning with tool use.

  • The physical scaling target is a lights-out laboratory whose economics resemble cloud infrastructure. Today’s system includes custom drivers, deliberately voided warranties, and even a vision-language model operating Windows 95; the destination is 24/7 uptime, dense vertical stacking, autonomous transport, and maximized tokens per unit volume. A 100,000-square-foot Massachusetts facility is an intermediate step toward a lab that “should feel like a data center.”

  • The largest risks sit after discovery and at the seams between simulation, hardware, regulation, and economics. A host cites roughly 5–8% of clinical programs advancing from IND to approval, materials require scale-up and qualification, simulated materials data often fail to predict reality, and Andy says reinforcement-learning workloads achieve only about 5–6% model FLOPs utilization. Lila’s narrower promise is to “make the die as loaded as possible,” not abolish downstream risk—and the team repeatedly concedes that its ambitious hardware and onboarding assumptions might fail.

  • 🔗 Original source & video: 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences

Listen to full conversation →