Pioneers Insight Method Research Author
AI Discovered Antibiotics: How Small Data & Small GNNs Led to Big Results, w/ MIT Prof. Jim Collins
Back to Episodes

AI Discovered Antibiotics: How Small Data & Small GNNs Led to Big Results, w/ MIT Prof. Jim Collins

Summary

  • Antibiotic resistance is already a mass-mortality problem, but the market penalizes anyone who solves it. More than 1 million people are estimated to die annually from treatment-resistant infections, while a UK commission projected up to 10 million deaths per year by 2050. Antibiotics cost roughly as much to develop as chronic-disease drugs, yet are taken briefly, sold for only a few dollars, and—if especially valuable—“put on a shelf and kept for when we really need it.”

  • A 2,500-compound training set was enough to turn antibiotic screening from a sub-1% search into a roughly coin-flip hit process. Collins’s team binarized E. coli growth inhibition at an 80% threshold, trained a graph neural network, and found a 51%-52% true-positive rate. Screening 6,100 additional molecules for antibacterial activity, human-cell safety, and novelty produced one qualifying compound: halicin, named because “HAL in the movie killed humans; halicin, our molecule, killed bacteria.”

  • The capital requirement is tiny relative to frontier-AI infrastructure spending, with clinical trials—not compute or discovery—the dominant expense. Collins estimates roughly $20 billion could develop 15-20 antibiotics and address resistance for decades; he contrasted that with a reported $20 billion xAI fundraising effort. ARPA-H has already committed $27 million to Fair Bio to move 15 candidates through preclinical development—“a little under $2 million per compound” to become IND-ready.

  • The pipeline can screen chemical spaces no physical laboratory could approach. The original system scored about 110 million compounds in three days. Work with Edamine was described first as a roughly 65-70 billion-compound REAL Space and later as 220 billion. Ensembles, novelty and toxicity filters, synthesizability checks, and other pipeline stages reduce that space to many hundreds worth considering and ultimately dozens worth making; Monte Carlo tree search helps analyze recurring substructures among the top-scoring compounds.

  • The outputs are not merely more antibiotics: some attack resistant pathogens through new mechanisms, spare beneficial bacteria, and delay resistance. Abaucin against Acinetobacter baumannii and a gonorrhea candidate proved narrow-spectrum even without explicit counter-training. Halicin produced no observed E. coli resistance over 30 days, versus “many hundred-fold” resistance to Cipro, probably because halicin hits multiple molecular targets.

  • The remaining model gap is downstream drug behavior, where sparse data and human judgment still matter. Current models are good at finding molecules that kill bacteria in a dish, but weaker on solubility, bioavailability, metabolism, distribution, PK/PD, and whether the molecule works in a mouse. In a 30,000-molecule synthesizability contest, a medicinal chemist beat the model by scanning for liabilities—an intuition Collins wants captured through “the medicinal chemist in the loop.”

  • This is a concrete small-model deployment opportunity, not an AGI-dependent promise, though it still has a dual-use boundary. The same phenotypic approach is being explored for antifungals, antivirals, antiparasitics, cancer, metabolic disease, and senolytics. Collins remains categorical that humans must stay involved: even E. coli has roughly 4,000 genes, about 1,500 still functionally unknown, while toxicity predictors can be used by bad actors to search for highly toxic compounds.

Deep dive

1. Antibiotics impose selective pressure rather than delivering a binary kill

  • Collins described antibiotics as small molecules that usually disrupt a bacterial protein involved in cell division, protein production, or DNA replication. That initial disruption triggers stress responses, energetic demand, toxic metabolic byproducts, and further damage to DNA, RNA, proteins, membranes, and lipids—a reinforcing cycle rather than Labenz’s imagined missile simply “smashing” the bacterium.

  • Labenz’s pushback—that efficacy appears scalar, not binary—surfaced two distinct drug classes. Bactericidal antibiotics are developed to kill at concentrations safe for humans, although organisms within one infection receive uneven and sometimes sublethal exposure; bacteriostatic drugs intentionally inhibit growth without killing. The discussion raised the possible role of the host’s immune system but did not establish it as the explanation.

  • The practical assay reflects that ambiguity: discovery screens overwhelmingly measure growth inhibition because killing assays are harder. A useful molecule must also discriminate between bacterial and human cells and, ideally, between the pathogen and healthy organisms in the gut, skin, and elsewhere.

  • Resistance arises as mutations and stress-driven changes alter the drug’s target or downstream response. Survivors gain a fitness advantage and propagate; overuse in humans and livestock has moved resistant “superbugs” beyond hospitals into “playing fields,” child-care centers, schools, shopping centers, and communities.

2. Resistance can be delayed, but Collins rejects claims that it can be abolished

  • Collins’s sharpest warning: an antibiotic researcher claiming a drug permits no resistance is “either lying to themselves or they’re lying to you.” Apply any antibiotic long enough and resistance will eventually develop; the realistic objective is to extend the runway.

  • AI offers two routes in that continuing “battle of our wits against the genes of these superbugs.” Researchers can continually discover molecules with mechanisms untouched by existing resistance, or deliberately design compounds whose resistance probability remains lower over a defined period.

  • Multi-target action is especially attractive. If efficacy independently depends on hitting several proteins, bacteria must accumulate useful mutations across multiple sites; a mutation against one target may not protect the organism from the remaining two, three, or four.

3. Antibiotics transformed medicine, then became commercially irrational

  • The modern antibiotic era is short: Alexander Fleming discovered penicillin in September 1928, and a group at Oxford developed and manufactured it as a drug only in the early 1940s. Antibiotics subsequently enabled routine surgery and made cuts, bruises, and blisters that once could be lethal broadly manageable.

  • Yet the discovery heyday occurred in the 1940s, ’50s, and ’60s—before the microbiology and biotech revolutions—followed by what Collins called a “discovery winter.” Development costs resemble those for cancer or blood-pressure drugs, but antibiotics sell for dollars, are used for days rather than years, and generate much less lifetime revenue.

  • The stewardship paradox makes the economics worse. Companies that carried a molecule through approval were sometimes told doctors would reserve it as a last line of defense; after their products were shelved, many businesses went bankrupt despite reaching the milestone the system asked them to reach.

  • Collins favors public-private structures resembling Operation Warp Speed and much greater philanthropic involvement. Resistant infections lack the months, ribbons, walks, and public consciousness attached to other diseases, even though he argued that every listener has likely lost someone to an antibiotic-resistant infection.

4. A bootlegged 2,500-compound experiment produced halicin

  • Collins’s laboratory had used machine learning for a little over 20 years to reverse-engineer bacterial networks, understand resistance, and find ways to boost existing antibiotics. The discovery pivot began after MIT launched a campus-wide AI initiative in March 2018 and Collins connected with Regina Barzilay and Tommy Jaakkola.

  • With no dedicated funding, the team “bootlegged the project” from what it could assemble: 1,700 FDA-approved drugs—including the known antibiotic universe—plus 800 natural compounds. Each was applied to E. coli, both a model organism and a pathogen associated with urinary-tract infections and food poisoning.

  • The researchers labeled compounds antibacterial if they achieved at least 80% growth inhibition, then trained a graph neural network to learn bond-by-bond and substructure-by-substructure associations. It subsequently screened the Broad Institute’s 6,100-compound drug-repurposing library for predicted efficacy, human-cell safety, and structural novelty.

  • Only halicin passed all three tests, and it proved “a remarkably potent new antibiotic.” The name carries the project’s inversion of science fiction: HAL in 2001: A Space Odyssey killed humans; the molecule inspired by it killed bacteria.

5. Binarization made unusually small data useful

  • AI colleagues initially dismissed the effort: “You have far too little data to do anything meaningful.” Labenz shared that intuition, contrasting 2,500 molecules with language models trained on roughly 1-15 trillion pretraining tokens, before additional post-training and reinforcement learning.

  • The counterintuitive move was to throw information away. A regression model predicting the full zero-to-one inhibition range lacked enough data to generalize to new structures, while binary labels let the network concentrate on structural features distinguishing clear antibacterial activity from no activity.

  • Collins did not claim data quantity stops mattering; a scalar predictor might require hundreds of thousands or millions of compounds. But the coarse dataset contained many clear negatives and a few hundred known-antibiotic positives, making its natural structure effectively binary already.

  • The resulting 51%-52% true-positive rate would be poor for classifying cats against dogs, Collins conceded, but exceptional for drug discovery, where random or large-scale empirical screens commonly yield well below 1%.

6. Small datasets are affordable; physical screening and synthesis set the ceiling

  • The Antibiotics-AI Project later assembled 37,000 more compounds for about $150,000. Vendors may charge $10-$20 per compound in volume and sometimes around $100, while screening a roughly 40,000-compound library costs about $20,000 in liquid-handling robot time.

  • That library has been tested against seven bacterial pathogens and three human cell lines—ten screens in total. Collins wished pharmaceutical companies would expose their million-molecule libraries for public-health problems after completing their proprietary mining.

  • Physical holdings remain modest beside computational space: academic laboratories typically possess tens of thousands of compounds, major research centers hundreds of thousands to around 1 million, and pharmaceutical companies low millions. The Broad has roughly 800,000-1 million compounds in a barcoded robotic system; Edamine may have 4-4.5 million stored and ready to send.

  • Data generation, synthesis, and animal experiments matter much more than inference expense. Collins treats compute as “a fixed cost in the background”; synthesizing a prediction commits money and time, while animal models force another costly decision about which compounds deserve advancement and analog development.

7. Chemprop ensembles can traverse billions of candidate molecules

  • The original architecture was Chemprop, a convolutional graph neural network developed by Barzilay and Jaakkola’s teams before this antibiotic project. It naturally consumes the bond diagrams familiar from chemistry classes and was available in late 2018, when transformers had only begun appearing.

  • Labenz said he understood the system to use 20 identically structured networks with different random initial conditions; Collins confirmed the broader “wisdom of crowds” strategy, while not independently specifying the count. Predictions are averaged so one initialization or overfit model does not control selection.

  • Small-molecule language models such as ChemBERTa and systems from NVIDIA and IBM performed reasonably but did not outperform the graph network. Hybrid modeling yielded only a modest bump; more recent work with the quantum-mechanics-pretrained graph model Uni-Mol reportedly performs significantly better than Chemprop, while 3D molecular representations remain an active direction.

  • The initial computational screen scored about 110 million structures in three days. Edamine’s synthesis recipes and building blocks were described first as expanding the credible REAL Space toward 65-70 billion molecules and later as reaching 220 billion—still microscopic beside the often-repeated estimate of (10^{60}) possible compounds.

8. Explainability and synthesis rules convert scores into actionable chemistry

  • Felix Wahl’s explainability work used Monte Carlo tree search to find recurring substructures among the highest-scoring candidates. Rather than merely returning opaque efficacy probabilities, it exposed “rationales” overrepresented across chemical families, suggesting new structural classes that might share a mechanism while retaining useful chemical diversity.

  • Collins hedged the grand (10^{60}) estimate: “I’ve never seen the calculation. I’ve heard the number.” He considered trillions plausible given building-block chemistry, but said AI may ultimately help define which theoretically imaginable molecules constitute a practical chemical space.

  • Candidate structures are evaluated with three increasingly demanding questions: can they be synthesized, can they be made in a reasonable number of steps, and can they be made affordably? Collins has consequently grown more comfortable screening Edamine-like libraries with credible synthesis routes than accepting exotic generative molecules whose manufacturability and efficacy are both uncertain.

9. Drug-like behavior—and medicinal-chemist intuition—remain the weak links

  • The pipeline generally applies efficacy models and filters sequentially, although multitask models have shown advantages where features overlap. Novelty and human-cell toxicity are tractable; the missing “DrugProp AI” layer needs data on solubility, bioavailability, metabolic liabilities, absorption, distribution, excretion, toxicity, and PK/PD.

  • Collins framed the gap plainly: models are good at finding molecules that kill a bug in a dish, but insufficiently good at predicting whether they kill it in a mouse. For systemically delivered antibiotics that do work effectively in mice, he cited an estimated 90% chance of translating well to humans—making that preclinical prediction especially valuable.

  • Human judgment still wins some direct contests. A medicinal-chemist postdoc reviewed 30,000 structures with a “right swipe, left swipe” process and beat a synthesizability model, largely by identifying liabilities: “bad, bad, bad,” then good only when no defect was visible.

  • Collins wants reinforcement learning with human feedback to encode that tacit expertise, despite medicinal chemists’ skepticism and occasional hostility toward AI. His analogy was Baidu’s pre-pandemic labeling effort: roughly 15,000 people, including 2,000 medical students labeling images, converted difficult-to-articulate human knowledge into machine-usable supervision.

10. The discoveries justify scaling now, even while biology and safety remain open

  • The candidates already address resistant strains, sometimes with unplanned specificity. Abaucin against Acinetobacter baumannii and a gonorrhea candidate spared many commensal bacteria despite no explicit counter-training; Collins thinks pathogen-specific lipoproteins and their transport may explain the narrow spectrum, while accepting Labenz’s novelty-filter hypothesis as possible.

  • In a 30-day E. coli comparison, Cipro produced significant resistance within days and “many hundred-fold” resistance by the end; halicin produced none that the experiment detected. Collins attributed that delay—not permanent immunity—to probable action against several membrane-level molecular targets.

  • Labenz’s conclusion was that society need not wait for better models: “clone your lab 10 times” and apply the validated workflow broadly. Collins agreed AI is already viable for early discovery, while a $27 million ARPA-H grant to Fair Bio is supporting 15 generative-AI-driven candidates through preclinical development.

  • The same approach is being used to identify senolytics against “zombie cells” and is being explored for antifungals, antivirals, antiparasitics, cancers, neurological conditions, and metabolic disease. But Collins rejected premature autonomous-scientist claims: E. coli has about 4,000 genes, roughly 1,500 of unknown function, and toxicity models published in 2024 could also be used by bad actors to seek highly toxic compounds, including ones for which there may be no countermeasures.