Pioneers Insight Method Research Author
🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
Back to Episodes

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik

Summary

  • Noetik’s core thesis is that oncology’s 90%-95% clinical failure rate is primarily a patient-selection problem, not a molecule-making problem. Ron Alfa argues that trials often show no placebo effect in cancer, so responders indicate that relevant biology is active; conventional development cannot identify the right subgroup. Noetik wants to discover therapeutically meaningful patient subtypes directly from human tumors, then use them both for reverse-translated target discovery and trial design.

  • The company’s claimed moat is a purpose-built, intentionally controlled human-tumor dataset rather than public biological data assembled after the fact. It has generated spatially resolved transcriptomics for more than 100 million cells, paired with H&E histology, protein imaging, and genotype data—“at least an order of magnitude larger” than comparable datasets Noetik has seen. Training on only 40% or 10% materially worsened its models, particularly when generalizing into unseen cancer types.

  • Noetik trains on expensive multimodal assays but can run clinical inference from the standard H&E image already collected for almost every oncology patient. Travis McKie calls H&E the “lingua franca of pathology”: the model can separate responders from nonresponders, predict locally expressed genes, and expose biology beyond single-mutation or single-protein biomarkers. That creates a plausible path from archived trials to prospective diagnostics using the same inexpensive input across many drugs.

  • Its “virtual cell” is deliberately practical and top-down, not an attempt to simulate every biochemical reaction inside a cell. OctoVC asks what T cells, macrophages, tumors, or gene expression would do in a particular patient context; PerturbMap tests model predictions through roughly 100 barcoded genetic perturbations within mouse tumors. The ambition is to predict which drug fits which patient without first constructing a mechanistically perfect cell.

  • The modeling work is moving from masked reconstruction toward autoregressive spatial prediction, with tissue context emerging as the scaling variable. OctoVC uses extreme masking—roughly 99% in the hosts’ recollection—to prevent the model from merely completing local edges. TARiO’s next-token objective showed larger models outperforming mainly at longer context lengths, suggesting that more surrounding tissue, not parameter count alone, unlocks complex patient-level biology.

  • The $50 million GSK agreement is the clearest commercial validation, but its structure matters: that figure includes upfront payment and milestones, alongside a separate annual model-license fee. GSK receives OctoVC models already trained on lung and colon cancer and can fine-tune them on its own translational data. Akash Tiwari’s key framing is that this resembles a meaningful biopharma business-development deal, except “the substrate is actually a model,” not a molecule.

  • The investment risk is inseparable from the moat: Noetik spent roughly four years building infrastructure and closer to $10 million than $50 million before knowing the approach would work. It lacked enough data to train a model for at least 18 months, individual transcriptomics runs took two weeks for two slides, and the company has not yet disclosed the trial-reanalysis results it says are coming. Dan Barr says several hundred patients across each major and selected minor cancer indication might generalize across oncology, but explicitly hedges that all disease biology could require another order of magnitude of data.

Deep dive

1. Patient selection—not pharmacology—is the proposed failure point

  • Alfa opens with Noetik’s contrarian premise: “90%, 95% of cancer drugs fail in the clinic,” even though industry is better than ever at pharmacology, target selection, and making molecules. He locates the dominant failure downstream, in identifying which patients carry the biology a drug can exploit.

  • His evidence is the responder hidden inside the failed aggregate. Alfa says that cancer trials often have no placebo effect; when a patient responds, that indicates real activity. The trial may have failed because that biology was diluted across an indiscriminately enrolled population.

  • The platform therefore works in both directions. It can begin with human tumors and reverse-translate toward new targets, or analyze phase two and phase three biopsies to identify the biology predicting response and redesign the next study. Noetik says it is doing substantial work on the latter, although those results remain undisclosed.

  • Alfa pushes the subtype claim further: “Nobody actually knows what the subtypes are.” A category long treated as one lung-cancer subtype might contain three functionally distinct diseases, and those latent divisions may matter more therapeutically than the classifications pathologists have used for more than a century.

2. “Frankensteinian” preclinical models leave clinical teams guessing

  • Bear’s critique begins with immortalized cell lines that have persisted for 40 or 50 years, often carrying abnormal chromosome counts and expression programs unlike recognizable human cells. Researchers still label them colon or lung cancer, but he calls them “Frankensteinian cells” whose experimental response often fails to map back to patients.

  • Moving those cells into animals does not close the gap: oncology commonly implants them under the skin “in weird places,” tests hundreds through a contract research organization, and infers an indication from which nominal colon or ovarian lines respond. Even lines derived from colon cancer may lack mutations characteristic of human colon tumors.

  • The resulting clinical chain is punishing. With no preclinical guidance about patients, a team may enroll all eligible tumors into an open-label study of roughly 50 people while simultaneously learning dose, safety margin, and efficacy signal. If lung cancer alone hypothetically contains 10 relevant subtypes, few responders are statistically unsurprising—and the molecule may nevertheless be canceled.

  • The hosts’ restatement captures the proposed fix: stop assuming today’s indication labels are homogeneous, and find the subpopulation that responds. Bear adds that existing biomarkers—one mutation, one stained protein, or one gene signature—are “biased towards simplicity” and usually correlate only weakly with clinical success.

3. Biological training data must be designed before it can be scaled

  • Tony Bui rejects the idea that a useful biology corpus can simply be scraped together. The Protein Data Bank was intentionally accumulated over decades; Shawn Wang compares ImageNet’s carefully curated, labeled 1.2 million-image scale. The lesson is that “you really need to be intentional about the data that you generate” and anticipate the models it must support.

  • Bui treats scale as “necessary, if not sufficient.” Language displays extraordinary scaling partly because of its corpus size, but thousands of hours of video have not automatically produced equivalent behavior. Biology adds tens of thousands of genes and proteins, their spatial organization, healthy tissue, other diseases, and even other species.

  • Shawn Wang offers the counter-thesis: because biology arises from underlying physical processes, perhaps enough of the space becomes in-distribution earlier than language does. Bui’s answer stays hedged—his “hunch is that biology is pretty complex,” Noetik remains far from covering all of it, and “I don’t know.”

  • Another useful check comes from protein folding: the discussion notes that good models may use only a small fraction of PDB, and some researchers argued that its 1990s coverage was already sufficient for a sufficiently capable algorithm. Dan Barr’s narrower estimate is that several hundred patients across major and selected minor cancers might be enough to generalize broadly across oncology.

4. Noetik captures tissue, cells, molecules, and batches together

  • Lessons from Tony Bui’s six years at Recursion shaped the experimental design. Images are information-dense, allow many patients on one slide, and reduce marginal data-generation cost relative to sequencing, where each additional run can correspond to another patient set.

  • The clinical anchor is H&E: hematoxylin and eosin create the familiar purple-pink tissue contrast used by pathologists to classify nearly every resected tumor. It captures tissue architecture but cannot reliably identify all relevant cell types, so Noetik adds multiplex immunofluorescence for B cells and other components of the tumor microenvironment.

  • Spatial transcriptomics supplies the molecular layer, detecting roughly 1,000 to 19,000 genes at resolved locations in the same cells. A probe binds each RNA species and the instrument cycles through detection for weeks; the discussion likens the output to an image with “20,000 color channels” instead of RGB’s three. Genotyping adds the underlying DNA alterations.

  • Batch control is built into the physical samples. Noetik samples each tumor dozens of times, randomizes hundreds of patient specimens across arrays, and represents every patient on multiple slides and processing runs. That lets researchers ask whether a patient embedding reflects immunotherapy response or merely “staining batches.”

5. The virtual cell is a drug-making heuristic, not a complete cell replica

  • Noetik separates two definitions. A comprehensive virtual cell would simulate millions of intracellular chemical reactions after any external signal; the discussion considers that intellectually interesting but says today’s data modalities cannot solve it. Noetik instead wants “some heuristic that’s useful for making drugs.”

  • Most current virtual-cell work, as described in the discussion, predicts transcriptomic changes after a small-molecule or CRISPR perturbation in cultured cells. The objection is translational: if the endpoint is what happens in a patient, “modeling data that comes from a patient” is more likely to succeed than modeling an in-vitro abstraction.

  • Noetik’s model therefore learns genes, proteins, cells, tissue, and patient context through self-supervision. It intentionally has not leaned heavily on electronic health records: the team does not want the representation constrained by what a doctor, operating with the knowledge of that moment, happened to record.

6. One H&E image can support discovery, trial rescue, and diagnostics

  • Before a molecule reaches patients, Noetik can simulate its target across lung, colon, ovarian, or broader oncology cohorts. The result might be a negative indication choice: a target initially intended for lung cancer could appear biologically unimportant there but relevant in ovarian cancer.

  • OctoVC can ask what a T cell would express inside a particular tumor microenvironment, or what happens if a target gene or protein is removed. Useful outputs include increased immune function, reduced tumor growth, or another response believed to correlate with clinical success.

  • The cleanest retrospective use case starts with a treated cohort: if responders all occupy one self-supervised patient cluster and not the other nine, that yields a direct enrollment hypothesis. Although training uses the full multimodal stack, inference needs only H&E—including a digitized image from a trial conducted years earlier.

  • The model can also predict where genes are expressed from that H&E. Seeing a drug’s protein target enriched in responder tissue provides an interpretability check, while the surrounding multigene pattern explains why one-protein biomarkers miss response. Noetik is applying the same input across collaborations, including one announced with Agena, with an eventual diagnostic as the natural endpoint.

7. Scale, architecture, and interpretability compound the data moat

  • Academic paired datasets may contain only a hundred or a few hundred patients, with spatial transcriptomics often below even that. Noetik reports more than 100 million spatially resolved cells, all paired with H&E and protein imaging—“at least an order of magnitude larger” than alternatives it has examined.

  • Its internal ablations support the scale claim: reducing training data to 40% or 10% made models “a lot worse,” especially when generalizing into cancers excluded from training. A sudden early doubling of the dataset similarly produced an immediate performance jump.

  • The team does not present data volume as sufficient. Noetik builds custom multimodal architectures and self-supervised objectives aimed at patient differences, gene or protein counterfactuals, and readable biological outputs. They call these “world models” because the task is to predict what follows from an action, not merely classify an image.

8. PerturbMap reconnects human predictions to causal animal experiments

  • The discussion addresses the tension in Noetik’s human-first approach after the company acknowledges that it still uses mice and injected cells. Human data should train the primary models, but the FDA may still want evidence that a new mechanism works in an animal system, leaving developers to “back into this system” unless they deliberately build a bridge.

  • PerturbMap creates roughly 100 distinct CRISPR knockouts, each carrying a combinatorial protein barcode, and injects them together into mouse lungs. The result can be hundreds of tumors whose identity and spatial biology remain readable—not one cell line implanted beneath the skin as a stand-in for human diversity.

  • Noetik can map genes causally to immune-cold or immune-hot phenotypes, then layer pharmacology on top; the team describes panels such as 50 knockouts across 50 drugs. Human tumors without immune infiltration can be recreated genetically in the mouse, where they likewise lack immune cells and fail to respond to immunotherapy.

  • The more surprising step is “in silico humanizing the mouse”: a model trained on human H&E and spatial transcriptomics is run directly on mouse H&E and emits human-gene predictions. Known antigen-presentation knockouts correctly appear cold, and several genes from one signaling pathway produce similar inferred phenotypes. The team acknowledges that genuinely novel regimes remain uncertain.

9. Autoregression scales only when the model sees enough tissue

  • OctoVC used masked autoencoding: divide each modality into tokens, hide most of them, and reconstruct gene expression, protein patches, or histology patches from what remains. The host recalls masking roughly 99%, and the response confirms that this is consistent with their approach.

  • Extreme masking is purposeful. At only 10%, a model can succeed through “boring behaviors” such as extending a nearby edge; hiding much more forces it to learn correlations between proteins and the holistic structure of tissue rather than local continuity.

  • TARiO changes the objective to autoregressive next-token prediction, effectively a structured form of masking that mirrors the loss behind LLM scaling. It was not Noetik’s first attempt, but it produced clearer gains from larger models and longer contexts on spatial-transcriptomic data.

  • The subtle result is that bigger models helped mainly when context was longer—meaning they could see more tissue at once. Low-expression but predictive genes might partly explain it, but controlled comparisons of the same context over smaller versus larger physical regions also favored larger regions.

10. GSK licenses a model as the asset, not a molecule

  • Noetik’s announced GSK agreement licenses OctoVC models trained on lung and colon cancer. The $50 million headline includes upfront payment and milestones, while an annual model-license fee sits separately; Tiwari also describes significant multimillion-dollar upfront and near-term economics.

  • GSK can use the models for simulation and therapeutic discovery, then fine-tune them on its own “mountains and mountains” of translational pathology and trial data. That turns a shared foundation model into something closer to GSK’s proprietary version without requiring the pharma to unify every silo and build the base model itself.

  • Tiwari sees pharma demand moving from bespoke, single-program collaborations toward access across dozens of pipeline programs. The deal resembles a traditional industry business-development transaction, but its breakthrough is structural: “The substrate of the deal is not a molecule. The substrate is actually a model.”

11. The moat required committing before any learning curve appeared

  • Tiwari says Noetik opened a laboratory, bought instruments, sourced human tumors, and ran two-week transcriptomics runs processing two slides at a time with “no prior indication that any of this would work.” His summary is blunter: “Big zero. Big crazy bet.”

  • The company lacked enough data to train a model for at least 18 months; individual transcriptomics runs took two weeks for two slides. Tiwari puts the investment “closer to the 10” million dollars than $50 million. There was no obvious off-the-shelf approach for spatial data, so the AI team had to explore an “alien landscape of data” from first principles.

  • Their advice to other startups is to begin with the machine-learning problem and design the dataset backward from it. “Any dataset” does not automatically have a valuable ML use case, and small pilots can fail below a critical scale even when the full experiment would work: “There’s no shortcut to it.”

  • Noetik’s closing bet is top-down abstraction. As simplified neural networks predicted real brain responses better than painstakingly stitched biophysical neuron models, functional tissue models may predict patient response sooner than bottom-up biochemical simulations. This may be the “first inkling of the ChatGPT moment for bio,” but literature-reading agents alone will not replace new data, new ML, and clinical translation.