Pioneers Insight Method Research Author
Material Abundance: Radical AI’s Closed-Loop Lab Automates Scientific Discovery
Back to Episodes

Material Abundance: Radical AI’s Closed-Loop Lab Automates Scientific Discovery

Summary

  • Radical AI’s core thesis is that materials innovation is trapped in a costly commercialization gap: a new system can require “north of $100 million” and “10-plus years,” sometimes 20–25, to reach market. Academia pursues fundamental understanding while corporate R&D targets 1%, 2%, or 5% gains, leaving a “wide-open white space” for breakthroughs tied directly to commercial needs.

  • The proposed moat is a closed-loop data flywheel joining AI recommendations to a robotic laboratory targeting 100 physical experiments per day. Krause ran roughly 50 experiments in a full year at the Army Research Lab; even government programs producing 400–500 annually would represent about a week for Radical AI. Synthesis results, failures, microscopic images, and property tests feed the next experiment through active learning.

  • The business model extends through scale-up and manufacturing because discovery and commercialization cannot remain separate. A quarter-sized laboratory “button” may need to become a 400-pound part without losing its properties. Radical AI therefore intends to sell materials at scale, protecting composition where useful but treating low-cost processing know-how as the deeper trade secret.

  • Language models sit at the center of Radical AI’s bet on reproducing scientific intuition, supported by GNN-based MLIPs, generative models, computer vision, simulation, and laboratory-analysis models. The key task is not merely predicting atomistic forces; it is reasoning across papers, patents, experiments, images, successes, and failures to decide what experiment should happen next. Colindres calls the scientific lab the best possible “ground truth” evaluation.

  • High-entropy alloys are the initial commercial wedge because they combine properties relevant to hypersonics, fusion, and other extreme environments. A hypersonic material must survive speeds above Mach 5, heat, pressure, oxidation, cost constraints, and supply-chain limits simultaneously. Radical AI’s recent Air Force Direct-to-Phase-II project applies high-throughput experimentation to alloys for hypersonic systems.

  • Experimental data—not another marginal model architecture—is presented as the scarce strategic asset. Krause estimates that “90%” of his Army Research Lab work did not work and was never captured; even existing notebooks are unlabeled, incomplete, and impossible to interpret without knowing the scientist’s intent. Radical AI argues that controlling the lab turns those otherwise lost traces into structured, contextual data.

  • Management has attached aggressive clocks to the scientific upside while keeping the claims conditional. Colindres expects a materials equivalent of AlphaGo’s Move 37 within 12–24 months “if we follow and hit our objectives and our roadmap”; he expects a natively multimodal internal model within six months and a likely publication or release within 6–12 months. The eventual ambition is transfer across material classes without requiring millions of examples in each one.

  • The principal execution risks remain physical automation, genuine model understanding, and factory-scale reproducibility. Some legacy instruments still require scientists, some decades-old tools cannot expose their pumps, chambers, or heat sources to software, and Colindres says they cannot yet demonstrate that models possess principled understanding rather than memorized heuristics. “You don’t have a new material until you can make it in a lab,” Krause insists—and not a commercial one until it scales.

Deep dive

1. Materials discovery is stranded between academic science and incremental R&D

  • Krause’s baseline is unusually stark: developing a material system can cost “north of 100 million,” take “10-plus years,” and sometimes stretch to “20, 25” years before commercialization. Materials underpin automotive, aerospace, manufacturing, defense, climate, energy, semiconductors, and electronics, so the delay propagates across the industrial economy.

  • The fragmentation mechanism matters. Academic researchers optimize for fundamental understanding, usually without commercial application as the objective; corporate laboratories optimize products already in market, pursuing “one, two, 5% performance gains” that improve margins and can be reported to Wall Street.

  • Between them sits what Krause calls a “wide open white space”: fundamental discovery selected from the outset for commercial value. Radical AI is positioning itself there, rather than as either a research software vendor or another internal optimization laboratory.

  • Labenz’s pushback was whether rising costs simply reflect depleted low-hanging fruit, as in drug discovery. The founders’ answer: both fields face impossible search spaces and fragmented development, but materials substitute scale-up and processing for clinical trials—the difficult conversion of a small reaction into tons of repeatable product.

2. Scale-up is the materials industry’s clinical-trial bottleneck

  • Colindres locates the hardest problem “quite literally in scale up.” A research sample may be a quarter-sized “button”; the commercial target may be a 400-pound piece. Changing the process in many ways, including environmentally and chemically, can erase the properties that made the original discovery valuable.

  • Radical AI wants a continuous data thread from the button through progressively larger forms. When performance degrades, the system should identify what was introduced “environmentally, chemically, or so on and so forth,” then alter the next process rather than treating manufacturing as a disconnected downstream handoff.

  • Krause says this is probably the bulk of the company’s work: candidates die between computation, first synthesis, replication, and industrial manufacturing. Even silicon’s use in transistors took substantial time to mature and continues evolving with semiconductor node sizes; more speculative materials may never survive the first reproduction attempt.

  • Labenz asked whether a sufficiently urgent “warp speed” program could compress the process to one year. Krause joked that Colindres already makes the team treat sub-year development as mandatory, but their substantive answer was that autonomous experimentation must accelerate both discovery and the learning required during processing.

3. LK-99 showed that synthesis details determine whether a discovery exists

  • Labenz challenged Krause’s initial wording on LK-99: his understanding was that the apparent superconductivity was a mistake, not a real material awaiting replication. Krause agreed that the latest explanation he had read involved an accidentally created impurity, possibly copper sulfide, while retaining the hedge: “Whether LK-99 is real or not, I don’t know.”

  • The broader lesson survived that correction. Reported temperature ranges, container details, and other synthesis conditions were insufficiently specified, yet each can change the result. Ingredients alone do not define a material; the trajectory through time, heat, and equipment is part of what was made.

  • Krause’s kitchen analogy carries the point: give two people identical bread ingredients, but let one score the loaf and allow the starter to rise while the other rushes it for 30 minutes at 1,000 degrees. The nominal recipe is similar; one produces bread and the other “a rock.”

  • His categorical standard is useful: “You don’t have a new material until you can make it in a lab, and you don’t have a new commercial material until you can take what you made in a lab and scale that up.” A claim that cannot be reproduced does not yet qualify as a discovery.

4. Customer pull can turn decades-old shelf science into a platform material

  • Colindres’s favorite example is Corning’s Chemcor, developed in the 1950s or early 1960s and left on a shelf for roughly half a century. About 2007, Steve Jobs sought a glass screen with better feel and appearance than the plastic then used in phones; Chemcor was adapted into Gorilla Glass.

  • In his telling, Gorilla Glass subsequently entered Apple, Samsung, and LG products and likely became Corning’s largest revenue driver. The point is not that the underlying science suddenly appeared—it is that a customer finally articulated the properties and product context that allowed the existing material to become valuable.

  • Radical AI therefore does not want to “make materials in a vacuum.” Customer requirements enter before model optimization: desired elements, oxidation resistance, strength, ductility, operating environment, cost, and supply chain are treated as part of the scientific target rather than commercialization details added afterward.

5. Enabling materials matter more than marginally better paint

  • The company’s north star is an “enabling material” that creates or unlocks an industry, not one that merely improves an existing product. Krause contrasts this with corporate R&D’s small percentage gains; Colindres describes a world of floating trains, interplanetary travel, and electric vehicles capable of driving New York to Boston without recharging.

  • A room-temperature superconductor is their canonical example because its second- and third-order consequences cannot be specified in advance. Just as touchscreen glass did not directly predict Uber or DoorDash, Colindres says, “We won’t actually know what the world will look like” after such a superconductor—but it would permit unprecedented technologies.

  • This ambition also imposes focus. Krause acknowledges that an early-stage company cannot “boil the ocean” across every material and industry; it needs a beachhead, execution, and then expansion. Room-temperature superconductors remain on the long-run roadmap rather than the company’s immediate operating program.

6. High-entropy alloys turn materials design into multi-objective optimization

  • Radical AI’s current beachhead is high-entropy alloys, which Krause says can combine mechanical and thermal performance rather than forcing a simple trade-off. The attraction is high strength retained at high temperature, with potential use across otherwise distinct extreme-environment applications.

  • For hypersonics, the material must operate above Mach 5 while tolerating high temperature, pressure, atmospheric oxidation, and corrosive conditions. In a fusion reactor, a related alloy might need to withstand persistent radiation and erosion that current materials such as tungsten cannot tolerate for long periods.

  • Customers almost never request one maximized property. The real specification is mechanical performance at a stated temperature and pressure, with required oxidation resistance, acceptable cost, and a viable supply chain for the constituent elements. “You’re not just optimizing on thermal expansion,” Krause stresses.

  • The range runs from silver nanoparticles inside Lululemon workout shirts to radiation-blocking shields for electronics on Mars. Their target class is not a “5% improvement game,” Krause says, but potentially “100x better” or “50x better” performance that makes a new application possible.

7. The flywheel connects property requests to 100 experiments per day

  • At the highest level, the system contains an AI engine and a robotic self-driving laboratory. The customer supplies desired properties and constraints; the engine performs “property-driven optimization,” proposes compositions and synthesis procedures, and sends a design of experiments through Radical AI’s operating system.

  • The laboratory then synthesizes the material and characterizes what it made. Techniques include X-ray diffraction and scanning-electron microscopy, followed by property tests for hardness, strength, stress-strain behavior, creep, and crack propagation—the measurements that connect composition and microstructure to application performance.

  • Machine-learning submodels analyze those streams in real time. The resulting insight returns to the AI engine, which modifies experiment two using experiment one’s outcome. Colindres emphasizes that this is “not some insane new change to the way we do the scientific method,” but a much faster way to extract and reuse its information.

  • Throughput changes the feasible learning curve. Radical AI targets 100 experiments per day; Krause estimates he completed about 50 in a full year at the Army Research Lab. Even directed government programs reaching 400–500 experiments annually amount to approximately one week of Radical AI’s intended operation.

8. Language models coordinate a heterogeneous scientific model stack

  • Colindres lists GNNs for atomistic modeling, material-science machine-learning interatomic potentials or MLIPs, generative models, the GPT family of language models, and computer vision in the lab. Different tasks require different inputs, outputs, architectures, and levels of physical structure.

  • The computational layer is the portion he is most confident will work: MLIPs are established and already useful. Radical AI’s differentiating bet is the language layer—reasoning over scientific knowledge, integrating modalities and tools, generating hypotheses, designing experiments, then learning from their successes and failures.

  • The architecture is therefore closer to an orchestrating scientist than one universal predictor. An LLM might call a GNN to resolve atomistic uncertainty, consult literature or patents, and inspect SEM images through computer vision before recommending a synthesis. Other models “bolster” the language system as it develops scientific intuition.

9. Human scientific intuition is largely accumulated failure that nobody records

  • Colindres points to Radical AI’s third co-founder, Gerbrand Ceder, as an exceptional materials scientist who “knows things that the rest of us don’t know.” Yet his own explanation is mundane: roughly 40 years of experiments, student work, and colleagues’ papers produced a vast internal catalog of what succeeds and fails.

  • Krause gives a narrower version from his transition into neuromorphic computing at the Army Research Lab. After a year working on ion-gated semiconductor channels, he says he was better than probably 99% of other scientists at thinking about how to design new semiconductor materials for that application.

  • The missing asset is the unsuccessful work. Krause estimates “90% of the stuff I did didn’t work, quote unquote,” and none of those e-beam, PVD, or ALD experiments became a reusable dataset. Only the successful final material left the scientist’s head and entered the formal record.

  • Colindres calls the latent corpus an “exascale set of data,” but historically the access method has been a 40- or 50-year academic career focused on perhaps one or two domains. Radical AI wants to compress the acquisition of that intuition to less than a decade and transfer it across more problems.

10. The discovery system must be rewarded for surprise, not just fit

  • Krause’s monolayer-film story shows why visual and accidental evidence matters. His group noticed that more transparent 2D semiconductor films conducted better but did not understand why until an outside expert recognized them as monolayer TMD films—an explanation they had never thought to seek in that experimental setting.

  • Colindres sees a fundamental machine-learning tension: models trained on distributions naturally dismiss outliers, while major scientific discoveries can hide in “some weird corner of the chemical universe that everyone thought was kind of ugly.” Systems must learn normal structure while remaining “primed toward novelty and surprise.”

  • Radical AI uses Bayesian approaches, Monte Carlo methods, diversification, noise sampling, active learning, reinforcement learning, and initially humans in the loop to support what Colindres calls “hunch hunting.” The goal is to pursue a strange result at 10x, 20x, or 100x a scientist’s scale without turning the process into undirected randomness.

11. The physical laboratory serves as both ground truth and steering mechanism

  • Labenz connected novelty to the broader tension among base models, RLHF, and reward-only reinforcement learning: imitation can sand away creative edges, while unconstrained reward optimization can produce unsafe or pathological strategies. Colindres’s answer centered less on ideology than on evaluation quality.

  • “There honestly is no better eval than the lab,” he argues. A physical result establishes whether the model achieved the requested property and whether it sacrificed another requirement. That ground truth supports active learning and customer-specific steering in a domain where safety also matters.

  • Updating must be more dynamic than periodically retraining weights. Colindres describes a surrogate or agentic layer that uses current models, reacts immediately to new experimental evidence, and selects tools while the underlying models improve. Mechanistic interpretability is meant to reveal why recommendations move toward or away from the intended objective.

12. Materials data is scarce, unlabeled, and inseparable from intent

  • Krause contrasts materials with language and software: “We don’t have Newton’s notebook,” and there is no GitHub-like open-source community publishing millions of structured experimental traces. Papers expose successful conclusions, not the scratched-out attempts that taught the scientist how to reach them.

  • Labenz suggested paying graduate students for their notebooks, but Krause’s response indicates that even a complete archive would require interviewing its authors because observations are unlabeled, references are missing, and identical measurements mean different things across specialties.

  • His XRD example is precise: a scientist studying crystalline materials looks for diffraction peaks, while an amorphous-materials researcher may use the absence of peaks as the confirmation. The same output appears empty to one researcher and successful to another unless the target and interpretation are attached.

  • Even a timestamp can be misleading. Krause sometimes recorded only the halfway point when he switched off heat and allowed uncontrolled cooling; “one minute and six seconds” is useless without that procedural context. Simulation data is easier to organize, but Colindres argues it addresses only “like 10% of the problem.”

13. The search space is astronomical, but customer properties provide an index

  • Krause invokes roughly (10^{80}) observable atoms in the universe as an intuition pump for the combinatorial space, describing the desired material as “a single grain of sand” on a beach the size of Earth. He contrasts that with a stated protein search space around (10^{130}), even larger but similarly impossible to enumerate.

  • Alloy development alone may draw from roughly 60 elements. Krause says that simply taking five elements produces “5.5 million and change” potential combinations in the equal-balance framing; allowing each element’s fraction to vary from 0% to 100%, adding decimal-level composition changes, or altering the processing path makes the space much larger.

  • Radical AI explicitly rejects brute-force combinatorial science as sufficient. Property requirements eliminate incompatible regions first; specialized models then narrow candidates before active learning reaches the expensive laboratory loop. The aim is to “index the search space,” not test every nominal composition.

14. Legacy scientific instruments—not generic robotics—limit autonomy today

  • Krause says the difficult part is usually not the rover or sample-handling track; it is the scientific instrument. Vendors have spent 25, 50, 75, or 100 years designing equipment around human operators, with physical switches, knobs, and interfaces that expose outputs but not full software control.

  • A nominal interface may return measurements while withholding control of the turbo pump, heat source, vacuum chamber, pressure, loading, or contamination management. Radical AI retrofits those functions, commissions custom equipment with vendors, or expects particularly inflexible tools to be displaced over the coming decade.

  • The company avoids redesigning the underlying physics of X-ray diffraction or other characterization methods. It automates loading, chamber state, laser timing, and sample movement because “Radical AI is not a materials-tools company”; the intended products are discovered and manufactured materials.

  • Some scientists still orchestrate experiments and bridge instruments today. Krause’s qualification is temporal: every tool begins with a scientist demonstrating its operation and should progress toward autonomous use; if retrofitting lacks sufficient value, Radical AI will work with someone else to custom-build a replacement.

15. Active learning repeats the explore-versus-exploit decision at every scale

  • Colindres explains that there is no single inference-time scaling rule because the question changes by layer. Atomistic exploration asks which structures and force predictions deserve deeper calculation; laboratory exploitation asks which composition can actually be synthesized and which pressure, temperature, or process adjustment should follow.

  • At the atomistic level, uncertainty in predicted forces may trigger a DFT calculation that updates the model. In the lab, an experiment may reveal that today’s pressure or temperature differs slightly from the assumed condition, prompting an immediate adjustment to the next run.

  • Simulation should screen hypotheses before scarce physical capacity is consumed, much as a scientist thinks, writes, and models before entering a laboratory. But Colindres repeatedly preserves the caveat: science is not perfectly logical or smooth, so every layer must “abide by the general rules” and remain able to break them.

16. Radical AI has timelines for Move 37, but not proof of understanding

  • Asked whether models have progressed from memorized heuristics to a principled representation of materials, Colindres gives an honest non-answer: “I wouldn’t say that we are in a place where we can demonstrably point to a very clear and definitive sense of understanding.” Hallucination and memorization are visible; genuine grokking is not yet proven.

  • Work with Goodfire and CEO Eric Ho is intended to make learned representations inspectable and steerable. Krause is excited that interpretability might expose useful heuristics a scientist never knew to request, including relationships that transfer between domains without carrying the scientist’s prior bias.

  • No earth-shattering materials equivalent of AlphaGo’s Move 37 has appeared yet. Colindres predicts Radical AI could produce one within 12–24 months, explicitly conditioned on hitting its objectives and roadmap; the “holy grail” is later transfer from alloy experiments into areas such as superconductors without millions of class-specific examples.

  • Colindres also expects a natively multimodal system with encoders for the relevant scientific data types to be used internally within six months, followed by a likely publication or release within 6–12 months. He stops short of claiming one model will permanently replace every specialist tool.

17. Vertical integration makes processing know-how the durable business

  • The founders began with what the materials problem requires, not the easiest venture-backed product. Their conclusion was blunt: if the company only sells software, “you’re never going to make real money,” while licensing every discovery demands repeated new discoveries and remains exposed to expiration and reverse engineering.

  • The model instead resembles large materials companies such as 3M, Applied Materials, BASF, and Materion: sell material at scale. Radical AI differentiates its target selection by pursuing materials intended to enable fusion, new transportation, hypersonics, or interplanetary systems rather than another paint variation.

  • Compositions can be patented, but Krause places the deeper IP in processing: the trade secrets required to manufacture a material reliably and at the lowest cost. Knowing the exact composition does not teach a competitor how to preserve its microstructure and properties through industrial scale-up.

  • Model policy follows that commercial structure. Some machine-learning systems may be published, while others will remain internal; the enduring proprietary assets are the experimental dataset, customer-linked discovery loop, processing recipes, and manufacturing capability.

18. The Air Force project is the first execution test for the platform

  • Radical AI’s recently announced Air Force Direct-to-Phase-II project targets high-entropy alloys for hypersonic applications. The work combines high-throughput experimentation with methods for testing and optimizing candidate materials before the Air Force considers them for hypersonic systems.

  • Krause frames the strategic setting directly: China and Russia already field hypersonic systems, while China will build a manufacturing hub around a new material specifically to learn how to scale it. His contention is that the U.S. has struggled to sustain a comparable emphasis on the materials foundation of important programs.

  • The company is advocating public-private coordination with the Office of Science and Technology Policy, National Science Foundation, Department of Defense, Department of Energy, and the Hill. Whether the application is hypersonics, fusion, or model-training data, Krause argues that government demand and private execution must work together.

  • The operating culture is designed around repeated experimental failure. “We fail every single day and we will continue to fail purposefully every single day for the next decade,” Krause says. His recruiting filter is equally explicit: candidates seeking a job should not apply; candidates prepared to dedicate their work to the mission should.