Pioneers Insight Method Research Author
šŸ”¬ The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI
Back to Episodes

šŸ”¬ The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI

Ā· Source link Ā· AI Summary archive

Summary

  • Radical AI’s core wager is that materials AI gains its edge by closing the experimental loop, because composition generation is only the beginning and ā€œthe ground truth is the material itself.ā€ A useful material must be synthesized, characterized, processed, manufactured, and qualified against cost, supply-chain, and application constraints. Krause’s differentiation from Lila, Citrine, Periodic, and others is therefore experimental data and lab infrastructure—not merely a better model.

  • The early throughput numbers suggest a real step-change, though Radical remains far from end-to-end manufacturing. Krause gave two different windows for producing roughly 1,200 alloys: five or six months in one account and three months later. About 300 were absent from the literature and ā€œprobably 10ā€ had especially exciting performance. Current throughput is 8-20 alloys per day at roughly $60-$300 each, with a stated target of 100 per day by June or July—versus the MACH program’s cited benchmark of 500 alloys in 12 months.

  • Commercialization risk moves downstream, where qualification and manufacturing intuition could absorb much of the discovery advantage. Aerospace qualification typically takes about 10 years, and Radical currently operates at grams to 200-500 g—not the 300-pound or 10-ton scales raised in the conversation. Krause sees a plausible 3-5-year path into defense or space applications, but not manned-flight turbines; semiconductor integration remains ā€œstill pretty long.ā€

  • The addressable opportunity is not just discovering stronger alloys, but designing materials concurrently with the products that need them. High-entropy alloys containing five to seven roughly equal elements could target temperatures north of 2,000°C, even 3,000°C, as well as pressure, oxidation, corrosion, or neutron bombardment. Krause’s borrowed framing is ā€œconcurrent engineeringā€: instead of designing a rocket, turbine, or chip around decades-old materials, engineers iterate the material and product together.

  • Radical’s self-driving lab is an orchestration problem spanning software, robotics, perception, scientific judgment, and physical tooling. An automated lab is like hands-free highway driving; a self-driving lab is a Waymo that chooses the route and runs an entire research campaign. Custom grippers must pry 3,000-4,000°C alloy ā€œbuttonsā€ from trays, models must judge whether melting is complete, and an operating system must decide whether a failed sample should proceed or be killed.

  • Krause argues that materials science is experiment-constrained rather than compute-constrained, making laboratory throughput and experimental data the economic moat. The relevant search space may contain roughly 10^40 possible alloys, yet useful discovery signals are already appearing from hundreds of experiments because high-quality experimental data are scarce. His blunt formulation: ā€œWe think in science models aren’t the moat, experiments areā€ā€”hence Radical can open-source models while treating its experiments and data as the edge.

  • The broader strategic thesis is that self-driving labs could multiply scarce scientific labor and help the US compete with China without copying China’s system in which one entity can control public and private activity. Radical says one metallurgy PhD can oversee 10 campaigns at once, reversing the traditional ratio of roughly 10 researchers to one campaign. Krause’s proposed counterweight combines national-lab data, HPC and instrumentation with private software and autonomy—but the hosts noted that China can deploy the same productivity tools, so execution and scale-up infrastructure remain decisive.

Deep dive

1. Materials AI must close the loop from hypothesis to physical truth

  • Challenged to differentiate Radical AI from Lila, Citrine, Periodic, and an increasingly crowded field, Krause returned to one conviction: ā€œthe ground truth is the material itself.ā€ Models matter, but a proposed composition is only the beginning.

  • Radical’s intended closed loop has an AI scientist propose candidates, the lab synthesize and characterize them, and experimental results feed the next campaign. Automation is extensive, but humans remain involved, especially in synthesis and scientific annotation. The objective is not a plausible digital structure; it is a material that can eventually enter an industrial application.

  • Krause’s causal argument is that performance often emerges after composition selection. Microstructure, heat treatment, post-processing, additive manufacturing versus casting, and scale determine whether the nominally same alloy is strong, ductile, oxidation-resistant, manufacturable, or useless.

  • The industry’s cited 15-30-year timelines reflect fragmentation: academia discovers, government-backed programs conduct light testing, and large companies optimize existing systems by 5% or 10%. Data rarely travel across those handoffs, severing discovery from manufacturing.

2. Radical has automated discovery-scale work, not the material’s full lifespan

  • Radical currently covers hypothesis generation, synthesis, characterization, and early property testing. Its characterization suite includes SEM, EDS, XRD, XRF, and TGA—different instruments for identifying structures, phases, chemistry, and thermal behavior.

  • Property testing includes oxidation performance, tensile stress-strain curves, and microindentation. Vickers hardness is measured directly, while the lab’s ductility signal is only a proxy; Krause explicitly cautioned that it is ā€œnot an exact measurement of ductility.ā€

  • The reported output is roughly 1,200 alloys. Krause described the window as five or six months in one account and three months later in the conversation. Around 300 were novel relative to the literature, and ā€œprobably 10ā€ produced performance exciting enough for deeper industry discussions and patent work.

  • The scale boundary is material: Radical works with grams, commonly 200 g or 500 g, not the 300-pound or 10-ton scales raised in the conversation. Wind-tunnel, torch, and other expertise-heavy aerospace tests remain with third parties, while full manufacturing data have not yet entered the loop.

3. High-entropy alloys make the case for concurrent engineering

  • Radical is not merely permuting a mature recipe book, Krause argued. Its high-entropy alloys combine five to seven elements at approximately equal atomic shares, creating candidates for extreme temperatures—often north of 2,000°C, even 3,000°C—high pressure, oxidation, corrosion, and neutron exposure.

  • The opportunity exists because aerospace and other industries still rely heavily on alloys developed in the 1950s through 1970s, sometimes augmented by later coatings. Long development cycles make incumbents rationally favor incremental optimization over unfamiliar material families.

  • Krause borrowed SpaceX materials executive Charles’s phrase ā€œconcurrent engineeringā€: design the material while designing the rocket booster, turbine, missile, solar cell, or other product. Performance specifications become inputs to material discovery instead of constraints inherited from whatever qualified alloy already exists.

4. Qualification, supply chains, and unit economics decide what survives

  • The hosts compared downstream materials qualification with drug development. For aerospace and defense alloys, qualification under the FAA or military specifications can require multiple ingots, standardized tests, and roughly 10 years before use in safety-critical systems.

  • The hosts asked whether qualification could receive an ā€œOperation Warp Speedā€ treatment by parallelizing tests. Krause pointed to DARPA work using additive manufacturing and layer-by-layer analysis, but did not claim the sequence had been solved; the goal is to achieve the same result through a newer mechanism.

  • His safety distinction is deliberately stark: a bending iPhone can be rejected or, disastrously, recalled; a turbine material on a 787 must meet a much higher bar. ā€œNo one would want that bar to be removedā€ā€”the outdated mechanism, not the safety requirement, is the target.

  • Supply-chain shocks can rewrite the objective midstream. Krause cited hafnium rising 10-15x because China controls a majority of the supply chain; C103 contains about 10% hafnium by weight, and Radical has worked on removing that element while preserving performance. Space tolerates higher cost for performance, whereas consumer electronics and medical devices are far more price-sensitive.

5. A self-driving lab chooses the route, not merely the speed

  • Krause’s clean distinction: an automated lab executes human-defined experiments at high throughput, like hands-free driving that still requires the driver to make a turn. A self-driving lab runs research campaigns, like a Waymo whose passenger specifies only the destination.

  • Radical divides that system into difficult sample manipulation and tooling, a lab operating system, and connected automation. The software tracks samples, controls instruments, consumes sensor data, performs quality checks, and can terminate a bad experiment before wasting time on XRD, SEM, and later tests.

  • Physical manipulation is deceptively difficult. Alloy ā€œbuttonsā€ blasted at 3,000-4,000°C stick to their trays; a person intuitively uses a chisel, while Radical needed custom robotic actuators that could remove them without damaging the sample or altering its microstructure.

  • Humans remain important teachers. Metallurgists annotate SEM imagesā€”ā€œI see dendritic formation on this image in these locationsā€ā€”so the AI scientist can absorb the judgment a PhD applies almost unconsciously when reading microstructures.

6. Radical narrowed its platform ambition to earn vertical depth

  • Krause’s change of mind is central: the founding plan was seven labs across seven material systems. Customer questions about specialized tests, manufacturing, and scaling to 300 pounds exposed the shallowness of that plan, so Radical prioritized vertical depth in alloys before polymers or ceramics—possibly without ever expanding.

  • Semiconductors remain an adjacent program because new back-end-of-line interconnect materials might reduce integration losses and energy costs. Krause said some systems recommended today could deliver roughly 2x to 5x improvements, potentially exceeding 10x later, but withheld the material details and admitted, ā€œI don’t know to what level.ā€

  • Integration into an iPhone or an NVIDIA GPU remains ā€œstill pretty long,ā€ compounded by scarce chip capacity and the need to build testing infrastructure from scratch. Krause was more confident about a 3-5-year alloy path into defense or space systems, explicitly excluding manned flight as the likely first use.

7. Active learning proceeds campaign by campaign, with selective human control

  • Radical’s AI scientist designs a campaign, selects a batch of candidates at its chosen confidence level, and sends them through synthesis, characterization, and early property testing. Machine-learning models analyze some outputs; scientists annotate others; all results return to the database for the next campaign.

  • Updates occur by campaign, not after every specimen. Krause said the lab could probably run approximately seven to 10 campaigns across different systems, with results revising hypotheses daily or every other day: ā€œWe actually want to take a few shots and get enough data back to change our hypothesis.ā€

  • Characterization is fully automated, as are oxidation and microindentation. Tensile testing is nearly automated, while synthesis still uses metallurgy PhDs because casting requires judgments such as whether a corner has fully melted; Krause expected the custom synthesis tool to be automated by the summer.

  • Candidate generation is already assigned to the AI scientist. Human researchers occasionally submit competing compositions—effectively red-teaming it—and may be rejected as insufficiently strong, but they also learn from unexpected elements the model introduces.

8. Throughput lets the AI explore where human intuition says not to look

  • Radical’s literature map shows published alloy families overlaid with regions its AI scientist explored. Human scientists explained their omissions candidly: they expected certain elements to evaporate, fail to cast, form poor grains, or damage mechanical properties—yet some combinations synthesized successfully.

  • The hosts’ pushback is worth preserving: perhaps the machine is simply receiving more trials and a higher ā€œtemperature,ā€ while an unconstrained human team might explore similarly. Krause conceded that literature often pulls the system toward known successes and that throughput is ā€œan important number.ā€

  • His answer is behavioral as much as algorithmic. A PhD researcher might perform around 50 experiments annually, spend roughly two weeks at a time fabricating and testing each one, and therefore treat every choice as precious; an AI scientist making eight, 20, or eventually 100 samples daily can afford speculative ā€œshots on goal.ā€

  • Current experiments cost roughly $60-$300 depending on elements such as platinum, palladium, aluminum, or titanium. Throughput ranges from eight refractory alloys to 20 easier systems per day, with a rough June-July target of 100 daily regardless of system.

  • The AI scientist also operates in parallel rather than in the serial way of a human researcher: Krause said it could compare 100,000 publications with 100,000 SEM images in real time, whereas a human cannot retain and directly compare that volume.

9. The bottleneck is experiments, not compute or search-space size

  • Krause contrasted Radical with the DARPA-GE Aerospace MACH program, which he described as producing 500 alloys in about 12 months after AI and simulation screening. Radical’s target is 500 in five business days, an order-of-magnitude change even before manufacturing-scale work.

  • The search space still overwhelms brute force: Krause estimated roughly 10^40 possible alloys and said humans would take seven million years to synthesize them all. AI remains valuable for screening, but its feedback quality depends on experiments that the industry historically has not captured or shared.

  • The hosts compared tens or hundreds of alloy samples with biological assays that can scale to millions or billions, questioning whether sparse local patterns would beat expert design. Krause’s empirical rebuttal was that Radical sees meaningful results from 100, 200, or 300 experiments, with 300 new alloys among the 1,200 run to date and roughly 50-150 data points per alloy.

  • ā€œWe’re not compute constrained in the materials industry. We’re experiment constrained.ā€ Radical’s ambition is consequently a ā€œprotein data bank for materials,ā€ but one containing processing, microstructure, properties, cost, and application context—not merely crystallographic structures.

10. There is no single AlphaFold moment for an industrial material

  • Krause agreed that ā€œthere is no AlphaFold for materialsā€ at the full-system level. Narrow AlphaFold-like moments are possible: segmentation models can read SEM images, identify dendrites, cracks, or defects, and relate crack propagation to mechanical behavior.

  • Krause’s broader contrast with biology is that SELFIES and SMILES can represent molecular elements and bonds in a string, while an alloy’s supply chain, cost, microstructure, processing, and additive manufacturing versus casting cannot be captured that way. No single model can one-shot a material that ends up in an iPhone or on Starship.

  • That capability does not answer whether a material can be atomized into powder, additively manufactured, cast, scaled, or integrated. Worse, scale-up may reveal variables nobody knew to test, making the desired dataset incomplete by definition.

  • Krause’s standard for discovery is intentionally severe: hypothesis, synthesis, and characterization are milestones, not the finish. ā€œWe count a new discovery when you pick up your phone and there’s a new material sitting inside it.ā€

  • A 35-year 3M advisor crystallized the manufacturing problem: the vital data may reside in an operator who knows exactly when and how far to turn a knob. Krause’s honest non-answerā€”ā€œwe have not solved that problem yetā€ā€”leads to partnerships with established manufacturers while Radical learns to instrument and automate those tacit processes.

11. Hardware friction is becoming an infrastructure moat

  • One early war story involved instruments whose software did not expose an interface, forcing a two-week software sprint to work out programmatic control. Radical was paying for the software and had to be strategic about obtaining access; Krause said the team found ways to control what it needed.

  • Building autonomous inorganic science splintered the team into experimental and computational materials science, mechanical engineering, mechatronics, full-stack software, applied ML, robotics, path planning, perception, and computer vision. ā€œIt’s not about a robot in front of a toolā€ā€”every downstream integration problem appears once that arm is installed.

  • Krause sees three reasons the timing now works: machine-learned interatomic potentials can speed parts of the computational funnel; robots, grippers, and actuation are cheaper and better; and instrument vendors increasingly maintain software teams that support automation interfaces.

  • The fundamental limit remains long physical feedback loops. A facility containing 1,000 XRDs or SEMs could compress them through parallelism; the more transformational fix would be vendors rebuilding instruments ā€œfor agents and robots,ā€ so researchers operate the scientific system rather than train individually on every machine.

12. National scale and open models reinforce an experiment-first moat

  • Krause described China’s manufacturing innovation hubs as an advantage the US ā€œshould not mimicā€ institutionally but must answer operationally. In his description, whether public or private, one entity can control the relevant pieces and support a new material through scale-up, directly attacking the 25-year gap Radical wants to shorten.

  • Radical’s productivity example is one metallurgy PhD running 10 campaigns, versus roughly 10 scientists historically focused on one research problem. The hosts noted that China can do the same; Krause’s answer was sustained investment plus a distinct public-private model.

  • National labs contribute HPC, researchers, instruments, and, Krause said, perhaps the world’s deepest store of experimental data. He cited self-driving or semi-autonomous work at Berkeley, Argonne, Ames, Livermore, and Oak Ridge, alongside the Genesis Mission and hundreds of millions of dollars in investment.

  • Radical’s AI scientist is itself multi-agent: an orchestrator proposes and tests hypotheses; a literature agent extracts relevant figures; paid industry-standard datasets and previous experiments ground campaigns; and MATRIX, a VLM fine-tuned on ā€œQuinn,ā€ reads laboratory images. Krause said the public dataset showed, he believed, 5%-16% gains in general scientific reasoning, with math the stated exception.

  • His workforce call is specialization, not retraining everyone into the same hybrid role: ā€œDon’t try to become a materials scientist. Be an MLE that works in materials science.ā€ Unfamiliar ML perspectives can challenge decades of inherited laboratory procedure while domain scientists supply the intuition models lack.

  • Open source follows the same logic. Radical released MATRIX and its benchmark, and spun TorchSim into a nonprofit to harness community feedback, while its own experimental data remain proprietary. Better external foundation models are welcome because ā€œwe don’t sell modelsā€; Radical expects experiments, automation, and accumulated physical evidence to remain the edge.