đŹ "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Summary
Genesis Molecular AI argues that the frontier of foundational AI research has shifted from familiar LLM architectures toward diffusion models for 3D molecular structure. GANs failed on proteins and proteinâligand systems, while diffusion supplied âthe right primitiveâ for iteratively generating physical structures. The hostsâ sharper recruiting pitch is that LLM labs still largely rearrange transformer layers published in 2017, whereas âsome of the most innovative diffusion researchâ now happens in structure prediction.
Genesis presents roughly 1 Ă accuracy as the useful threshold for proteinâligand prediction, because the conventional 2 Ă scale can conceal chemically fatal errors. At 2 Ă , an aromatic ring may flip while the output still looks plausible; hydrogen bonds occupy only a 2.7â3.3 Ă distance window. Evan Feinbergâs formulation is blunt: âDrug discovery really is a science of resolution,â and models at 1.8â1.9 RMSD risk producing agent-amplified âslop.â
Genesis has adapted parts of the LLM scaling stack to molecules: synthetic-data pre-training, iterative inference-time computation, and eventually reinforcement learning. The public structural corpus contains only roughly 200,000 entries, versus an estimated 10^60 drug-like small molecules, so physics simulations supply additional training data. During inference, the models âthink in terms of crystal structures,â repeatedly refining an internal representation while physics-based guidance steers diffusion.
The highest-leverage AI opportunity, in Genesisâs view, is the missing design layer between known disease biology and clinical testing. Knowing the responsible target is orthogonalâand sometimes inversely relatedâto being able to drug it; Feinberg calls the blanket 10% clinical-success statistic âreally a lowballâ for candidates with strong genetics, pharmacokinetics, safety, and translational models. The opportunity spans zero-to-one binders and better successors to existing drugs, as later-generation ALK inhibitors demonstrate.
A viable drug requires simultaneous optimization across more than 30 ADMET-related endpoints, not merely an accurate binding pose. Potency, selectivity, solubility, membrane permeability, cytochrome P450 inhibition, hERG liability, tissue exposure, and other properties can invalidate a molecule independently. Worse, objectives anti-correlate: making a compound greasier may improve binding while damaging solubility, and adding polarity may then prevent cellular entryââplaying whack-a-moleâ at molecular scale.
Genesisâs prospective-data advantage comes from coupling its models to real drug programs and laboratory feedback: Incyte provides disclosed partner programs, while Insitro supplies rapid compound production and measurements. Disclosed Incyte work ranges from advancing existing chemical matter toward a development candidate to finding the first known binders for a target with no patents, papers, or co-crystal structure. The Insitro collaboration creates repeated designâmakeâtestâanalyze cycles and training data spanning structure, potency, and ADMET, although synthesis and high-fidelity validation remain stubbornly difficult to automate.
Sapphire is Genesisâs attempt to turn specialist models into an always-on drug-discovery workforce without removing scientists from strategic control. An LLM orchestrates pose, potency, ADMET, and chemistry tools so medicinal chemists need not master every parameter; the envisioned result is âfleets of hundredsâ of virtual scientists operating 24/7. Feinberg rejects full human replacement: experts set direction and evaluate outcomes while agents absorb execution and repetitive tool use.
OpenBind supplied the external generalization test Genesis says private partner data had previously prevented it from showing. On the unseen EV-A71 3C protease target, whose flexible loop must move around the ligand, Pearl reportedly produced a much wider performance gap than public in-distribution benchmarks and was âbasically correct for every single pose.â The remaining constraint is compute: both guests named GPUs as their bottleneck, while Edunov argued that âthe amount of alpha left in pure LLM space is just getting a little questionableâ relative to life sciences.
Deep dive
1. Diffusion became the primitive molecular AI had been missing
Edunovâs route to Genesis ran from physics into software engineering, FAIR research, and leadership of Llama 2 and Llama 3 pre-training. Joining as CTO let him ârecover my roots in physicsâ while applying large-scale AI methods to molecular systems.
Feinberg came from the complementary direction: a physics-and-computer-science background, a medical family, and graph-machine-learning work in VJ Pondâs Stanford lab. While Edunov studied âa lot of big graphs,â Feinberg studied molecules as âmany small graphsâ of atoms, bonds, and spatial interactions.
Feinberg remembers declaring GANs the future of image generation around 2017â2018, then watching mode collapse make them ineffective for proteins and proteinâligand complexes. The field had to âwait for the right primitive,â and diffusion proved much more useful for generating three-dimensional structures.
The surprising result is that foundational diffusion research is no longer concentrated in consumer image generation. Feinbergâs claim: âSome of the most innovative diffusion research is happening in our field,â making 3D structure prediction a pillar nobody would have forecast a decade earlier.
2. Drug-discovery AI is compounding, not awaiting an iPhone moment
When Genesis began roughly seven years earlier, its founders worried they were late: incumbents had raised orders of magnitude more capital, and investors questioned whether another AI drug-discovery company was needed. Feinberg now calls that hindsight almost absurdââthought to be late, but turns out it was still early innings.â
His original thesis remains intact: roughly 20,000 protein-coding genes can contribute to disease, and no single advance can solve that universe. âThereâs been no single iPhone momentâ; even smartphones, ChatGPT, and autonomous vehicles became useful through repeated improvements rather than one clean zero-to-one event.
The expectation is continuing expansion of solvable targets, with âlarge leapsâ embedded in cumulative iteration. Feinberg says current systems are vastly more useful than those of a decade ago and predicts another exponential improvement during the next ten years.
The hostâs central challenge was generalization: molecular models historically became pattern matchers that told researchers what they already knew. Feinberg agreed this is the urgency wherever AI meets the physical worldâmodels naturally interpolate near training data, while drug discovery demands reliable extrapolation.
3. Molecular design is the missing middle between biology and trials
Feinbergâs working analogy casts the protein as a lock and the drug as a key intended to change its function. Binding is ânecessary but not sufficientâ: the molecule must avoid anti-targets, reach the correct tissue, remain safe, and satisfy roughly 30 additional developability properties.
A decade ago, researchers hypothesized that accurate 3D proteinâligand complexes would improve affinity and potency prediction. They could barely test it: computational pose predictions were poor, while experimental crystallography or cryo-EM could cost tens of thousands of dollars and consume months, years, or an entire thesis.
Genesis says the recent breakthrough is not merely generating attractive structures but showing that systematic pose-accuracy improvements carry into potency prediction. Better complexes could also expose druggable configurations in proteins previously considered undruggable.
Target identification, molecular design, regulatory preparation, patient segmentation, and clinical trials require distinct models despite sharing some machinery. Feinberg places the highest leverage in design: patients are often told clinicians know what caused their condition but still lack a selective therapy capable of acting on it.
4. Pearl attacks a 10^60 search space with scarce structural data
Pearl accepts a protein sequence and ligand representation, then predicts their joint 3D complexâa co-folding task. Genesis deliberately concentrates on small and medium-sized molecules, including orally available drugs, macrocycles, and peptides, rather than treating every biological interaction as one problem.
âSmallâ does not mean computationally easy. Edunov estimates about 10^60 drug-like small molecules, each admitting rotations and alternative conformations; after the hosts proposed finding a needle in a haystack, Feinberg inverted it to âfinding hay in a needle stack,â where most candidates either fail or are dangerous.
The nearest molecular equivalent to internet-scale pre-training is the RCSB Protein Data Bank, with only a couple hundred thousand structures, though each contains substantial latent information. New experimental structures arrive at a âglacial paceâ because they are expensive and difficult to produce.
Genesis supplements that corpus by simulating small molecules with physics. Proteinâprotein systems are much larger and costlier to model, but molecular dynamics and related calculations can generate lower-cost structural data for small-molecule pre-trainingâprovided their biases are handled carefully.
5. Pearl imports scaling laws without copying language models literally
Edunov maps the LLM recipe into three stages: pre-training scaling, post-training through fine-tuning or reinforcement learning, and inference-time scaling. Genesis creates synthetic pre-training data and then lets its models spend additional inference computation rather than immediately emitting one structure.
The analogy to reasoning tokens is functional, not linguistic. The model is âthinking in terms of crystal structuresââpartially materialized internal representations that it revisits as the diffusion head iteratively refines predicted coordinates.
Because diffusion already unfolds across multiple denoising steps, Genesis can inject physics-based guidance during generation. The hosts framed this as balancing learned and physical force fields; Edunov accepted the steering intuition but declined to claim that researchers know exactly what the network internally represents.
Feinberg treats physical priors as disciplined representation choices, analogous to encoding images as pixel grids or language as token sequences. Genesis uses physics in inputs, architecture, and output validation while trying not to force models to inherit every assumption held by âwe puny humans.â
6. One angstrom separates plausible pictures from useful chemistry
The discussion notes that at 2 Ă , a structure is like a generative image whose detail is wrongâbut unlike visible blur, the molecular error may look entirely credible. A whole aromatic or heterocyclic ring can flip while still passing a coarse structural metric.
Feinberg grounds the 1 Ă objective in hydrogen bonds: donor-to-acceptor heavy-atom distances typically span 2.7â3.3 Ă , only a 0.6 Ă window. Too short is a clash; too long rapidly weakens the interaction. âDrug discovery really is a science of resolution.â
The hosts offered a serious counterpoint: a single pose is an abstraction over a probability distribution, and affinity also includes enthalpic, entropic, and dynamical contributions. Feinberg accepted the abstraction but defended its inspectabilityâa scalar potency output âmight as well be completely hallucinatedâ if no structure lets scientists test whether it makes sense.
Feinberg distinguished a potent ligandâs tightly resolved binding core from solvent-exposed regions that may genuinely âflop around.â The objective is sub-angstrom correctness where protein interactions occur, while accepting dynamics elsewhere; that core then supports free-energy prediction and the practical question, âWhat molecule do I make next?â
7. The fieldâs two-angstrom benchmark created an eval crisis
Asked how Genesis crossed the threshold, Feinberg gave âan extremely boring answerâ: âdata, infrastructure and evals.â Because âyou can only improve what you measure,â optimizing sub-angstrom accuracy changes filtering, curriculum, architecture, losses, and many small engineering decisions that compound.
Real partner and internal programs made 2 Ă failures obvious in a way academic leaderboards did not. Feinberg traced RMSD below two to old docking studies built for publication and later inherited by AI, not to medicinal chemists establishing that it was sufficient for prospective design.
His SWE-bench analogy was intentionally provocative: a model can score well without becoming anyoneâs preferred coding tool. Molecular evaluation is likewise moving beyond RMSD toward physical validity, PoseBusters, and lDDT; Feinberg describes âan eval crisis in our field that is now in transition.â
8. Structure prediction is only one pillar of a 30-property problem
Feinberg pushed back on popular claims that AlphaFold-era structure prediction solved drug discovery. A static, relatively low-resolution structure omits dynamics, selectivity, exposure, safety, and the other endpoints required to turn a binder into a medicine.
ADMET is represented operationally by more than 30 assays. Feinberg cited solubility, oral bioavailability, cytochrome P450 variants, and hERG inhibition, where an excessive effect can create cardiotoxicity.
These endpoints vary in learnability. Some correspond directly to interaction with a particular protein and might be addressable through 3D modeling; others aggregate many pathways, while public datasets can be âcomically small.â
Genesisâs breadth predates Pearl: its Stanford lineage produced MoleculeNet and multitask graph networks for pharma-scale ADMET prediction. Feinbergâs emphasis is continuityâthe company has focused on all models required for drug discovery, including molecular generation, without expanding into unrelated target-identification or clinical-trial problems.
9. Known biology leaves both first-in-class and best-in-class opportunity
Shawn Wangâs business challenge was whether targets with known biology and tractable structure had already been picked over. Feinberg separated the variables: biological validation is orthogonal to ease of drugging and may appear anti-correlated because the most compelling disease targets can be exceptionally hard to bind selectively.
Feinberg also disputed the undifferentiated 10% clinical-success statistic. Candidates with close genetic linkage, understood biology, translating animal models, adequate predicted pharmacokinetics, and strong safety profiles have âfairly highâ approval rates from Phase 1 through Phase 3âthough he supplied no universal figure.
Opportunity therefore spans true zero-to-one programs with no known binder and one-to-ten programs improving imperfect chemical matter. His public analogy was ALK inhibitors: later generations produced qualitatively better survival curves, showing why âweâve drugged ALKâ did not mean development should stop.
Genesis focuses on small and medium-sized molecules; small molecules alone remain about 65% of FDA-approved drugs, Feinberg said. The intended market is consequently both new target access and replacement of suboptimal clinical or approved agents.
10. Partner data turns models into prospective drug programs
Genesis disclosed work with companies including Gilead and an expanded Incyte collaboration. One Incyte program began with a challenging target and existing chemical matter; Genesis fine-tuned foundation models on partner data to move the program closer to the binary milestone of selecting a development candidate.
The opposite bookend began with strong disease linkage but no known chemical matterâno patent, paper, or ligand co-crystal. The teams found initial hits, then advanced them into inhibitors active in biochemical assays and living-cell assays.
The commercial model pairs Genesisâs AI specialization with pharmaâs strengths in biology, development, and commercialization. Renaming Genesis Therapeutics as Genesis Molecular AI reflected that identity rather than abandoning medicines; internal programs still dogfood the models, generate candid medicinal-chemist feedback, and make the platform âbattle tested.â
Feinberg described the organization as a âdouble helixâ: deep AI research alongside experienced drug hunters, some of whom have helped produce multiple approved medicines. The goal is to place models directly with many drug developers while retaining enough in-house discovery to understand the work rather than behave like âkeyboard jockeys.â
11. Wet-lab feedback is valuable precisely because automation remains messy
Feinbergâs near-term reinforcement-learning path begins with physics-based rewards and could extend into laboratory-in-the-loop rollouts: predict molecules, synthesize them, measure downstream properties, and feed those outcomes back into training. He said Genesis has already seen early signs that RL works with its models.
Insitroâs rapid compound production and measurement enable continuous designâmakeâtestâanalyze cycles across structure, potency, and ADMET. Feinberg called the collaboration unusually important because it combines historical and prospective data for joint foundation-model training.
The physical work resists simplistic robotic-lab narratives. Synthesis requires compatible reagents, catalysts, solvents, temperatures, and protocols; compounds then need purification and confirmation through NMR, mass spectrometry, or related methods to prove the vial contains what researchers intended.
High-throughput screens and DNA-encoded libraries can test millions or billions of compounds, yet their correlation with de novo resynthesis and high-fidelity assays may have âa shockingly low R squared.â Automation also favors constrained chemistry, while drug discovery searches for novel Pareto outliers among anti-correlated objectivesâspeed can exact a harsh cost in molecular quality.
12. Agents can scale drug hunters only after the underlying models work
Feinberg compared molecular agents with coding agents: both amplify positive and negative value, so âagents are only as useful as the underlying models that theyâre orchestrating.â Pose systems around 1.8â1.9 RMSD would merely automate production of structures medicinal chemists reject as âslop.â
Sapphire, Genesisâs code name for an agentic platform, is envisioned as âfleets of hundredsâ of medicinal chemists and CADD scientists working 24/7. Its LLM orchestrator can call pose, potency, ADME, and chemistry tools while handling parameter choices that no human can master across an entire software stack.
Crystal structure can become another model modality: an agent might inspect an image, invoke specialized geometric tools, or consume a natively tokenized 3D representation. That lets it reason from structural predictions rather than treating each predictor as an opaque scalar oracle.
Feinberg expects the Cursor trajectoryâfirst autocomplete, then increasingly autonomous executionâbut âI donât believe in full automation or replacing humans.â Scientists remain strategic directors, while agents absorb routine execution; his version is that drug hunters become âgrand strategistsâ with hundreds of computational workers.
13. OpenBind exposed generalization; GPUs now constrain the upside
Public benchmarks often compress model performance into a narrow band because every team has optimized against them. OpenBindâs unseen EV-A71 3C protease target offered a cleaner test: Pearl had not trained on or been developed around it, and the targetâs flexible loop must move to accommodate the ligand.
Feinberg said Pearlâs numbers were âway higherâ than published open models and that it was âbasically correct for every single poseâ in modeling the loop movement. Genesis presented this as public confirmation of the larger gaps it says it observes on confidential partner targets.
Both guests named GPUs as the bottleneck. Edunov argued LLM companies are consuming capacity needed for medicine discovery; Feinbergâs ideal intervention would be an enormous H100 cluster, while noting NVIDIA has invested twice in Genesis and collaborated on kernels and Pearl.
Edunovâs allocation thesis is that medicines remain valuable through economic cycles, whereas âthe amount of alpha left in pure LLM space is just getting a little questionable.â The hosts contrasted mainstream LLM architectures, still closely related to the transformer published in 2017, with molecular diffusion as a more architecturally distinct field with direct clinical stakes.