Pioneers Insight Method Research Author
Periodic Labs: Training AI Scientists, with Liam Fedus & Ekin Dogus Cubuk (from a16z)
Back to Episodes

Periodic Labs: Training AI Scientists, with Liam Fedus & Ekin Dogus Cubuk (from a16z)

Summary

  • Periodic Labs’ $300 million seed, led by Andreessen Horowitz, backs a capital-intensive thesis: an AI scientist must learn by acting on physical reality, not merely by absorbing the internet. Liam Fedus calls experiment a “physically grounded reward function,” with simulations and language models as tools but nature as the final error-corrector: “Nature is our RL environment.” The investable proposition is a lab-built loop of hypotheses, automated experiments, positive and negative outcomes, and improved models.

  • Scaling laws may continue to hold while still failing to deliver useful scientific discovery on the required timescale. Fedus’s challenge is “what is this y-axis?”: scaling against internet or coding distributions does not manufacture missing physics knowledge, and “that model is not going to then cure cancer.” Ekin Dogus Cubuk adds that out-of-domain performance can improve as a power law yet have such a shallow slope that reaching the target might take centuries.

  • The existing scientific corpus is not just too small; it is noisy, selectively published, and missing the iterations that teach scientific judgment. Reported physical properties can span orders of magnitude, formation-enthalpy errors can defeat prediction, and superconductivity datasets have a high noise floor. Because negative results are rarely published, a model trained on literature can at best reproduce a distorted distribution rather than learn why an experiment failed.

  • High-temperature superconductivity gives Periodic a measurable north star and forces it to build the entire autonomous-science stack. The stated ambient-pressure benchmark is about 135 Kelvin; exceeding it would provide an unambiguous score, while a hypothetical 200 Kelvin superconductor would, Cubuk argues, update humanity’s view of the universe even before commercialization. Getting there requires autonomous synthesis, characterization, simulation, and experiment selection—capabilities that can be tested for transfer into magnetism and other physical domains.

  • The near-term business is an intelligence layer for advanced manufacturing, not a wait for a miraculous superconductor. Fedus targets copilots for researchers and engineers in semiconductors, space, defense, and other industries with “massive R&D budgets,” reducing iteration time across literature review, simulation, design, and experiment. Deployment follows a land-and-expand motion: solve one critical, well-scoped problem with clear evaluations rather than promise to transform an entire fabrication line on day one.

  • Periodic’s model strategy goes beyond retrieval by encoding private scientific and industrial knowledge through mid-training and high-compute reinforcement learning. Mid-training means continuing pre-training on knowledge absent from the base model—from crystal structures and simulation outputs to descriptions of how materials were made—then connecting those distributions so one dataset improves performance on another. That deeper encoding creates an enterprise challenge too: knowledge may need to be bucketed into separate systems when some data is accessible only to senior leadership.

  • The organizational moat is a roughly 30-person, cross-disciplinary team built around curiosity, translation, and urgency rather than credentials. Weekly teaching sessions let ML researchers explain RL and data cleaning while physicists and chemists teach quantum mechanics and scientific history; the cultural rule is “no stupid questions.” Advanced degrees are explicitly unnecessary because even the best specialist knows far less than the combined physics, chemistry, synthesis, and characterization the mission demands—and the founders want progress “ASAP,” not in 10 years.

Deep dive

1. Nature replaces the digital grader

  • Fedus recalls meeting Cubuk eight years earlier at Google Brain while helping him flip a tire too heavy for one person. Their later conversations repeatedly returned to quantum mechanics and superconductivity, while Cubuk watched LLMs become useful for recovering forgotten chemistry, writing simulation code, and eventually serving as “a first class citizen” in physics research.

  • Their technology trees converged around reasoning, high-compute reinforcement learning, and scaling behavior in simulation and experiment. Chatbots were “a great milestone along the way,” Fedus says, but the intended destination was technology that accelerates science and physical R&D; physics offered verifiable outcomes, relatively fast iteration, and simulators for broad classes of systems.

  • Fedus’s reconstruction of early ChatGPT clarifies the missing ingredient. Supervised examples turned an autocomplete model into an assistant, then RLHF optimized a reward model built from human preferences—but that reward mostly encoded “be a friendly assistant” and could not tell whether mathematics was correct or code valid.

  • Modern digital graders improved correctness, yet science ultimately answers to experiment. Periodic therefore pairs familiar tools such as Python and a browser with simulations and quantum-mechanical tools, but treats the laboratory result as the final reward and error-corrector when a simulator is deficient: “Nature is our RL environment.”

2. Scientific discovery requires action, failures, and target-domain data

  • Cubuk’s simplest diagnosis is that science is iterative: even brilliant humans fail repeatedly before discovering something, and “you put a human in a room without any chance to iterate on something, they won’t discover anything important.” Models likewise need to calculate, simulate, experiment, observe an unwanted result, and revise their understanding.

  • Literature alone provides a compromised learning signal. One Periodic engineer found reported values for a physical property spanning many orders of magnitude; a model trained on that record is not magical and can merely reproduce its distribution, rather than collapse the “reducible uncertainty” that only a new experiment resolves.

  • Valid negative results are especially valuable and especially scarce. Most published results are positive, while failures are omitted or dismissed as sloppy science; Cubuk adds that a negative result can depend on experimental context and become positive when another researcher changes the procedure.

  • The founders reduce their laboratory thesis to three deficits: noisy data, missing negative results, and no ability to act. A high-throughput laboratory addresses all three by producing high-quality data, including negative results, and letting an agent iterate on what to do next.

3. Scaling works only on the distribution being scaled

  • Anjney Midha’s pushback—worth keeping—is that continued pre-training and post-training might eventually crack physics through brute scaling and emergent capabilities. Fedus accepts the empirical premise that scaling laws still hold, but asks, “What is this y-axis?” Performance scales predictably on a chosen test distribution, not automatically on every capability humanity wants.

  • His coding analogy draws the boundary: online code is abundant, unit tests provide rewards, and a system can improve at producing software that passes tests. It may accelerate its own software development and help cancer researchers analyze data, but “that model is not going to then cure cancer”; the necessary knowledge and environmental iteration are absent.

  • Cubuk sharpens the argument mathematically: in-domain and out-of-domain performance can both improve monotonically as power laws, yet the out-of-domain slope may become so shallow, depending on distance from the training data, that one might otherwise “spend centuries” reaching the desired capability. Periodic therefore intends to move the target closer to the training distribution by changing the training set toward what it wants to do.

  • Some required datasets simply do not exist at usable quality. Cubuk cites formation-enthalpy labels whose errors make the next material insufficiently predictable and superconductivity datasets whose noise floor often defeats training. Their stance is not anti-scaling: “Do a baseline for the thing you care about.”

4. Superconductivity forces a complete autonomous laboratory

  • Periodic begins at the quantum-mechanical energy scale where chemistry, biology, and materials operate. One planned lab will automate powder synthesis: mix existing-material powders, heat them to a chosen temperature, and form a new material—a process Cubuk says can be handled by machinery roughly comparable to an airport coffee robot.

  • The output space remains scientifically rich despite that mechanical simplicity: powder synthesis can produce candidate superconductors, magnets, and other technologically important materials. Models will couple those experiments to theoretical calculations and simulations, learning a “foundation model” intuition for quantum mechanics rather than merely generating scientific prose.

  • Progress has a hard benchmark: the best stated ambient-pressure superconducting temperature is about 135 Kelvin, so surpassing it is directly measurable. Applied work can be scored just as concretely through ductility, toughness, strength, and whether a requested material property can actually be discovered and produced; the signal is “hard to hack.”

  • A 200 Kelvin superconductor would matter even before a product because observing quantum effects at that temperature would revise how people see the universe. Superconductivity is also technically attractive as a phase transition usually governed more by fundamental crystal properties than by defects or microstructure that current simulations cannot capture.

5. The scientific north star contains commercially useful subgoals

  • Midha presses the startup question: if AI programming became the commercial wedge on the path toward a general AI researcher, what pays for Periodic’s journey? Fedus’s answer is “copilots for engineers, researchers in advanced industries,” especially space, defense, semiconductors, and manufacturing workflows that must repeatedly alter materials and physical designs.

  • High-temperature superconductivity is less a single bet than a forcing function. Reaching it requires autonomous synthesis, autonomous characterization, correct simulation use, literature comprehension, theoretical reasoning, and experiment execution—each a useful capability for customers whose R&D staff follow the same read-simulate-build-learn loop.

  • Fedus is explicit that “technology and capital are intertwined.” A wildly successful commercial entity can maximally accelerate science, so Periodic aims to become an intelligence layer that shortens industrial iteration cycles and improves the solutions reached by researchers and engineers.

  • The deployment thesis is deliberately narrow at first. Companies are seeking an AI strategy and sometimes losing senior expertise, but Periodic will not propose transforming a fabrication line immediately; it will co-define one urgent problem, specify clear evaluations, match it to current technical strengths, find internal promoters, and “land and expand.”

6. Mid-training turns proprietary knowledge into model capability

  • One prospective customer spends substantial time training employees to operate essential simulations. Its needs include automating those simulations, assisting the design process, matching file formats, feeding outputs into existing pipelines, and treating scattered datasets together—mundane integration details that determine whether advanced models alter real workflows.

  • Fedus distinguishes retrieval from understanding. Retrieval can respect each employee’s access privileges and fetch internal documents, but ChatGPT showed that pre-training knowledge “into the weights” creates something richer than search; the customer still lacks a way to distill its accumulated expertise into one model or coordinated set of models.

  • Mid-training is the bridge: take a pre-trained model and continue pre-training on missing knowledge before conventional supervised or reinforcement-learning post-training. For Periodic, inputs range from low-level crystal structures to higher-level accounts of how material XYZ was made, alongside simulation and experimental data the lab needs to produce.

  • Simply mixing datasets A, B, and C does not guarantee generalization; Periodic wants each addition to improve performance on the others. It also benefits from better base LLMs, open simulation tools, and specialist neural networks used as agent tools, avoiding the need to reproduce every capability internally.

7. Cross-disciplinary culture and academia complete the system

  • Building the company required an “N-of-1 team” spanning LLMs, experiments, and simulation, with fractal specialties beneath each: solid-state chemistry and physics, automation and facilities, theoretical and computational simulation, mid-training, RL, and infrastructure. The roughly 30-person group includes bridge researchers living between the pure-specialist corners of that simplex.

  • Weekly lessons and a “no stupid questions” culture make translation operational. Computer scientists repeatedly map scientific claims into APIs—“what’s the input, what’s the output, what’s the target?”—while scientists teach the physical concepts and history needed to choose better training tasks and reward reasoning strategies.

  • Advanced degrees are not required. Cubuk inverts the NBA comparison that an elite player is closer to LeBron James than an amateur is to him: even Periodic’s best physicist has vastly more physics, chemistry, materials science, synthesis, and characterization left to learn than they already know, so collaboration and intense curiosity outrank credential boundaries.

  • Academia supplies tools, tasks, and ways of thinking industry may miss. Periodic is forming an advisory board including Zhi-Xun Shen, Steven Kivelson, Mercouri Kanatzidis, and Chris Wolverton, plus a grant program for LLM agents, synthesis, materials discovery, and physics modeling; one physicist’s correction captures the value: “It should be thinking in terms of symmetries.”