Faster Science, Better Drugs
Summary
Arc’s wager is that science becomes materially faster when multidisciplinary teams and foundation models reduce both organizational and experimental latency. Patrick Hsu argues that biology still requires growing cells, tissues, and animals and moving things from tube to tube, while academic incentives discourage collaboration; Arc counters by putting neuroscience, immunology, machine learning, chemical biology, and genomics under one roof, hoping ultimately to enable experiments “at the speed of forward passes of a neural network.”
The investable promise of a virtual cell is not a perfect digital organism but a practical engine for selecting drug targets and experiments. Hsu defines the task as predicting which perturbations move a cell from disease state A to healthy state B, including deliberate, sequential combinations rather than accidental polypharmacology. Erik frames the analogous “AlphaFold moment” as roughly 90% reliability, while Hsu says today’s capability sits only “somewhere between GPT-1 and GPT-2.”
Biology’s missing data are real, yet Arc is betting that scalable measurements can serve as imperfect mirrors of layers that cannot yet be captured directly. RNA may be a “lower-resolution mirror” of protein signaling and metabolism, enriched over time with protein, spatial, and temporal tokens. Hsu concedes that “we’re not measuring many of the most important things in biology,” but argues a predictive model can still be useful without explaining every mechanism—much as a weather model predicts rain.
AI drug discovery has to solve two distinct causes of failure: choosing the wrong target and making the wrong drug for the right target. With roughly 90% of drugs failing in clinical trials, Hsu says the proportions attributable to each cause are unclear and both may contribute. Even a 90%-accurate virtual cell might recommend targeting a GPCR only in heart tissue—yet no suitable tissue-specific drug matter may exist. Hsu’s warning is that the vertically integrated AI-pharma pitch currently “precedes the fundamental research capability breakthroughs.”
Biotech’s economics improve only if technology raises success rates, lowers capital intensity, compresses timelines, and produces larger effect sizes. Jorge Conde says early investors currently absorb years of capital intensity without reliable valuation step-ups when scientific value inflects. GLP-1s show the upside: Hsu estimates Eli Lilly and Novo Nordisk added roughly $1 trillion of market cap, more than the market cap of all biotech companies combined over the last 40 years. Conde’s conclusion is that “the prize—the juice—needs to be worth the squeeze.”
Design may become computationally abundant, but physical testing and clinical development remain stubbornly serial. Even if models produce “a trillion binders in silico,” molecules still must be made, tested in animals, and advanced through humans; survival and longevity endpoints have irreducible timelines. In Hsu’s hype-versus-heft split, toxicity prediction and vague multimodal biology are overhyped, while protein design, pathology, radiology, and administrative automation already carry real heft.
Jorge Conde’s broader investment thesis targets biological, cognitive, and physical frontiers—and assumes AI agents will expand from software budgets into the services economy. He highlights synthetic biology, brain-computer interfaces, and robotics, while expecting computer-use agents to trail coding agents by perhaps a year and improve from minutes of error-free work to hours and days. Conde’s contrarian research opportunity is beyond the 2017 transformer architecture: “in 2025 we’re really overdue” for a new architecture, including old ideas newly tested at 1B, 7B, 35B, and 70B parameters.
Deep dive
1. Science is slow because physical work and incentives move in real time
Hsu’s starting point is physical constraint: outside AI research, science requires growing cells, tissues, and animals and moving things around and transferring liquids from tube to tube. Machine learning matters because it might massively parallelize that work, not because biology can simply become software.
The deeper problem is a “weird Gordian knot” of funding, training, career incentives, and the separation of basic from commercially viable work. Individual groups can be excellent at perhaps two domains—say computational biology and genomics—but modern problems increasingly demand five at once.
Arc is therefore an organizational experiment: neuroscience, immunology, machine learning, chemical biology, and genomics share one physical roof to increase “collision frequency.” Erik Torenberg’s pushback—that universities already combine disciplines—draws Hsu’s distinction: campus proximity is not enough when researchers must publish separately and own distinct discoveries.
Arc’s two flagship programs—finding Alzheimer’s disease targets and building virtual cells—require collaboration beyond any one laboratory. The moonshot is a tool trusted by skeptical experimentalists because it produces data and results, eventually enabling experiments “at the speed of forward passes of a neural network.”
2. Virtual cells turn drug discovery into navigation through cell states
Hsu uses AlphaFold as the adoption benchmark: it does not simulate molecular dynamics, but gives a sense of the protein’s end state with “90%+ accuracy” and becomes the default when no experimental structure exists. Arc wants perturbation prediction to become comparably routine for cell biology.
A virtual cell represents heart, blood, lung, and other cells across states such as inflammation, apoptosis, stress, starvation, or quiescence. Drug discovery then becomes a navigation problem: determine which perturbations “click and drag” a cell from a toxic, disease-causing state toward a homeostatic one.
Complex disease may require purposeful sequences—“these three changes” first, then two, then six—rather than a molecule accidentally hitting multiple targets. This reframes polypharmacology as designed, combinatorial control over time, with models proposing both targets and the drug compositions needed to produce the transition.
Practicality governs the scope: the model should help a wet-lab biologist choose perhaps 12 experiments in 12 conditions, then learn through a lab-in-the-loop cycle of prediction, measurement, and revision. Hsu wants in-silico target identification, but cautions that the AI-pharma company narrative currently outruns the underlying research capability.
3. Biology is only at GPT-1–2, so tangible benchmarks matter more than abstract scores
Asked how close virtual cells are to 90% reliability, Hsu places the field “somewhere between GPT-1 and GPT-2.” Arc’s Evo DNA foundation models, developed with Brian He, generate “blurry pictures of life”: Hsu does not think synthesized novel genomes would be alive, though he does not think that capability is “impossibly far away.”
Progress requires a full-stack hill climb—curating public data, generating massive private datasets, constructing benchmarks, and developing architectures together. The scale shift is already striking: early-2010s single-cell papers contained 20 or 40 cells; Arc expects to generate 1 billion perturbed single cells in a relatively short period.
A credible GPT-3 moment should rediscover famous biology: infer the four Yamanaka factors that reprogram fibroblasts into stem-like cells, identify factors such as Neurogenin-2, ASCL1, or MYOD for differentiation, or recover responses to HER2 inhibition. The benchmark must make sense to “an old professor who has never touched a terminal,” not merely improve mean absolute error.
Textbooks provide evaluation targets, not complete truth. Their A-signals-to-B-inhibits-C diagrams compress a multidimensional system full of exceptions; discovery often means finding those exceptions rather than treating the diagrams as literal simulators.
4. Scalable measurements may reveal what biology cannot yet observe directly
Language and image models advance faster because humans natively judge their outputs; “we don’t speak the language of biology,” except “with an incredibly thick accent.” Biological models require real experiments for ground truth, slowing every iteration.
Hsu expects capability to accumulate by layers: individual cells, cell pairs, tissues, and eventually physiologically intact animals, with spatial and temporal information added over time. His technology framework separates invention, engineering, and scaling—some measurements are ready to scale, while others still need fundamental invention.
RNA is the current scaling bet. It may be a low-resolution “mirror” of proteins, signaling, and even metabolism: unreliable for one cell, yet informative across enormous transcriptomic datasets, especially when supplemented with protein nodes. Arc does not want merely to industrialize today’s single-cell screens and look dated in three years.
Erik’s hardest objection—what if a core biological variable remains undiscovered—wins a direct concession: that is “almost certainly true.” Imaging and sequencing are biology’s two high-throughput tools, but a meteorological-style virtual cell could still predict the correct endpoint without explaining every causal detail, just as AlphaFold predicts a fold without reconstructing why it formed.
5. Better discovery improves biotech economics only if value survives the clinic
Software-first biotech companies initially competed for small SaaS budgets, then moved toward pharma R&D budgets and claims that biological agents could replace headcount. Hsu’s test is simpler: do they help build drugs more effectively?
Roughly 90% of drugs fail in clinical trials for two broad reasons—wrong target, or wrong drug matter for the right target—and Hsu says the proportions are unclear for any individual failure; both may contribute. Even a 90%-accurate virtual cell could surface technically unreachable instructions, such as inhibiting a GPCR in heart tissue but nowhere else, creating another demand for novel chemical biology.
Conde’s prescription has three parts: reduce capital intensity by improving success rates, compress discovery and clinical timelines where possible, and increase effect sizes so efficacy becomes obvious sooner. Clinical development has not yet seen AI-driven compression comparable to early discovery; cancer-survival and longevity studies retain naturally long clocks.
The investor consequence is acute: early capital bears the full development burden, yet scientific value inflections often do not produce valuation step-ups. Lower capital intensity and higher value creation would restore the reward for investing early rather than merely financing a long sequence of expensive proofs.
6. GLP-1 ambition meets the irreducible bottleneck of testing in humans
Hsu’s market-cap comparison is the episode’s sharpest industrial signal: Eli Lilly and Novo Nordisk gained roughly $1 trillion from GLP-1 development—more, he estimates, than the market cap of all biotech companies combined over the last 40 years. Large patient populations can justify the risk and ambition that the industry often avoided by clustering around validated mechanisms and small indications.
Conde agrees that “the prize, the juice needs to be worth the squeeze.” Small molecules, biologics, gene therapies, and gene editors should combine with better target understanding to produce larger effects against obesity, cardiometabolic disease, neurodegeneration, and cancers increasingly treated as chronic conditions—but “we have to deliver.”
Hsu’s constraint survives even maximal computational success: “a trillion binders in silico” still have to be manufactured and tested. Conde’s compressed path is “mice, then rats, then monkeys, and then man”; failure after that slow journey is precisely what better upstream models must prevent.
In a Dan Wang-inspired lawyers-versus-engineers exchange, Jorge Conde points to 10 of the first 13 American presidents having practiced law and all Democratic presidential candidates from 1980 to 2020 having attended law school. He sees echoes in FDA and regulatory bottlenecks and notes the idea of running Phase 1 trials overseas before bringing data back for domestic Phase 2 efficacy trials—but says that direction is “not enough” to solve making and testing.
7. The durable AI opportunity lies in verified work, new architectures, and open competition
Hsu puts toxicity prediction and loosely specified “multimodal biological models” in the hype bucket. Protein binding and design, pathology and radiology prediction, plus regulatory writing and reporting, already show heft; AI-designed drugs remain hard to isolate because AI will simply become native throughout designing, making, testing, and approvals.
Dario Amodei’s bullish scientific scenario depends on discoveries being sufficiently independent to parallelize across millions or billions of agents. Hsu finds that long-run compression plausible if models can reliably traverse target selection, molecular design, docking, off-target assessment, cell correction, and laboratory validation—but each interface still needs experimental feedback.
Conde targets synthetic biology, brain-computer interfaces, and robotics as ways to improve the human experience within one lifetime. The founding team must combine technical invention, product intuition, and commercial ability—the right “RPG dice roll”—and be funded when execution is possible within five or eight years, not merely imaginable in science fiction.
Agents matter because they attack services spending, not only software budgets; computer-use agents may trail coding agents by a year as reliable work expands from minutes to hours to days. Conde also sees a research opportunity beyond the 2017 transformer architecture: old machine-learning ideas could look different when scaled from 100M or 650M parameters to 1B, 7B, 35B, or 70B.
Arc’s open Virtual Cell Challenge offers $100,000 prizes sponsored by NVIDIA, 10x Genomics, Ultima Genomics, and others, creating a transparent capability track that can be followed today and in subsequent years toward the field’s “ChatGPT moment.”