🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub
Summary
Alex Rives’s central bet is that protein biology will obey the bitter lesson: scale a generic sequence predictor across evolution, and useful biological structure will emerge without hand-coded priors. Evolution constrains which amino acids can coexist, so predicting masked residues should force ESM to infer hidden variables for structure and function. After increasing models roughly an order of magnitude per generation since 2018, Rives says, “I believe in scaling laws.”
ESM-C’s decisive improvement over ESM-2 came from widening the evolutionary data distribution, especially with billions of noisy metagenomic sequences. ESM-2 showed diminishing returns despite larger models; ESM-C, at approximately the same parameter scale but with more diverse data, showed “no longer diminishing returns to scale.” The thesis-relevant bottleneck is therefore not clever architecture alone, but access to amino acids across as many evolutionary contexts as possible.
The newly MIT-licensed system packages a protein world model, ESMFold 2, mechanistic-interpretability features, and an atlas spanning 6.8 billion nonredundant proteins. The team predicted structures for 1.1 billion representatives clustered at 70% sequence identity, giving structural coverage of the larger set, while sparse autoencoders expose features from basic chemistry through abstract function. Rives calls it “the most comprehensive picture of protein structure and function that’s been created.”
The commercially consequential capability is search-based protein design, particularly scFv antibodies reaching affinity levels needed for therapeutic activity in a small number of trials. Rives says therapeutic design “basically emerges from that search” of a general sequence-structure-function model; he also claims significantly better antibody performance, where evolutionary information may be less useful. Full IgGs remain untested, although scFvs can be reformatted and he sees no reason the approach would not work.
The conversation treats static protein models as only the first rung; the larger prize is a virtual cell that predicts genuinely novel interventions in unseen biological contexts. Rives argues today’s virtual-cell models represent their training data well but have “a very limited ability” to answer new experimental questions. Reaching useful cellular oracles requires perturbational and spatial biology, multimodal measurements, experimental feedback, and models spanning molecules, genomes, cells, and ultimately physiology.
Biohub is committing $400 million internally and $100 million externally to build that missing data stack, while acknowledging this is only a fraction of what is needed. The near-term plan is to scale existing assays 10x–100x, then develop technology for another 10x or more while expanding interventions, measured modalities, and biological contexts. “We can’t wait decades”; Rives wants the enabling data created within a couple of years.
Neither protein data nor compute appears exhausted: ESM-C used roughly one billion sequences, while Rives estimates there may be on the order of 100 billion available sequences. Small variations should not be discarded as redundant because they may teach function—“a single mutation is enough to destroy the function of a protein.” A 100x compute increase would improve ESM-C, he says, but only if data scales in tandem; how long returns persist remains “truly an empirical question.”
Deep dive
1. Evolution supplies the supervision that protein models need
Rives traces the program to summer 2018, when his team at Meta AI trained an early transformer language model for proteins. Across subsequent generations, increasing scale roughly an order of magnitude repeatedly produced new capabilities.
The biological premise is concrete: residues touching in a folded protein cannot evolve independently. A change at one position requires compatible changes elsewhere, leaving statistical patterns in sequence databases that reflect underlying structure and function.
Rives’s task formulation is deliberately generic: mask amino acids and predict “the amino acids that evolution will choose.” Solving that across billions of sequences pressures the model to infer the otherwise-hidden constraints shaping proteins.
The host’s challenge—proteins are not natural language—draws Rives’s empirical answer: AI lacks a complete theory of when scaling transfers, but evolution has already generated an enormous training set through “four billion years of life running experiments in parallel.”
2. Metagenomics broke ESM-2’s data ceiling
ESM-2 improved from roughly the billion-parameter to 10-billion-parameter scale, yet its structure-representation curve showed diminishing returns. Rives now interprets that result as data limitation rather than evidence against scaling.
UniRef provided curated, clustered coverage of known sequence biology. ESM-C added metagenomic material collected indiscriminately from hydrothermal vents, polar environments, deep oceans, soil, human guts, and other ecosystems.
The metagenomic data is messy by design: researchers sequence environmental DNA, translate likely proteins from fragmented contigs, and often lack complete genomes, organism identities, or certainty that every inferred sequence is a full protein.
That noise bought diversity. With approximately the same parameter scale as ESM-2, somewhat more compute, and billions of additional sequences, ESM-C produced a clean scaling curve whose smaller-model extrapolation predicted the representational fidelity of larger models: “The data was really the critical thing here.”
3. ESM-C turns a language model into an open protein atlas
Rives describes ESM-C as the fourth generation, trained a little over a year before release and now fully open-sourced under an MIT license. The family contains 300-million-, 600-million-, and 6-billion-parameter models.
Around the language model, the team built ESMFold 2 for structure prediction and sparse-autoencoder tooling for exposing learned features. The result is meant to be a world model spanning protein sequence, structure, and function—not merely a next-token predictor.
The atlas combines major sequence databases into 6.8 billion nonredundant proteins. The team clustered them at 70% sequence identity and predicted structures for 1.1 billion cluster centers; related members should share the same fold, with smaller variations.
Those structures add hundreds of millions of entries to the accessible picture of protein diversity. Computed features also link distant proteins that share functional or structural patterns despite weak sequence similarity.
4. Interpretability reveals biology the model was never explicitly taught
Sparse autoencoders trained across the ESM-C layers reveal a hierarchy resembling biology’s experimentally developed reductionist picture: biochemical properties and structural building blocks at the bottom, then large functional themes and abstract concepts.
The nucleophilic elbow is Rives’s sharpest specimen. Protein families with different topologies may have evolved this structural motif independently, yet ESM-C uses “a single feature” for it across those evolutionarily distant families.
His explanation remains a hypothesis: prediction requires compression, and compression creates latent variables. Because every amino-acid choice is entangled with the rest of the sequence, reusable concepts such as a nucleophilic elbow help the model predict many otherwise unrelated contexts.
The same feature space clusters distantly related gene-editing systems and other proteins whose functions are unknown. Some might be undiscovered editing systems, though experimental validation is still required; Rives notes that Feng Zhang’s group used the first ESM Atlas to find a new gene-editing system.
5. Distributional structure offers a hypothesis for emergence
Rives invokes Zellig Harris’s 1954 “Distributional Structure”: the contexts in which a word appears are constrained by its meaning, so statistical structure can recover semantic structure without receiving explicit definitions.
His biological analogue is direct. The contexts available to an amino acid are determined by a protein’s structure, function, biological role, and relationships to other proteins; learning those context sets should therefore expose the hidden biological variables producing them.
This framing also explains why small sequence differences are valuable. Broad evolutionary diversity teaches structural abstractions, while dense variation within families may teach function at the resolution where one mutation can destroy function.
6. ESM3 and ESM-C offer two routes to programmable biology
Rives says ESM3 was consistent with the ESM philosophy and that both approaches have a place. ESM3’s goal was explicit programmability, using sequence, structure, and function tracks so biologists could prompt the model with the right biological information.
ESM-C approaches design as world-model search instead: specify desired criteria, then search its predictive landscape for molecules satisfying them. It has generated mini-protein binders and, more notably, scFv antibodies.
The hosts compare this with coding agents that begin with broad pretraining and acquire programmability through post-training or reinforcement learning. Rives calls conversion between these approaches promising, but says the right method is not yet understood: “We need both.”
7. Antibody results test whether general models beat specialized pipelines
An scFv combines one antibody heavy-chain component and one light-chain component into a single chain, allowing a complex binding interface. Rives estimates antibodies account for roughly a quarter of new drugs, making this a consequential therapeutic modality.
In a small number of trials, ESM-C search found scFvs at affinity levels needed for therapeutic function and activity. Rives emphasizes that this behavior came from a general protein model rather than a system trained solely to engineer antibodies.
The hosts’ pushback is important: mini-binders are increasingly routine, whereas nanobodies, scFvs, and especially antibodies become harder; antibody diversity also makes multiple-sequence alignments less naturally informative than for conserved proteins.
Rives has not tested full IgGs. scFvs can be reformatted as antibodies, and he sees no reason full-IgG design would not work, but that remains a prospective claim; his current, narrower assertion is that ESM-C performs “significantly better on antibodies.”
8. Fast structure prediction is an on-ramp, not a virtual cell
ESMFold 2 does not require multiple-sequence alignments, so it can produce atomic-resolution predictions directly from sequence in seconds. Rives says the system is state of the art for open models on multimer prediction.
A host proposes predicting every pairwise interaction in the human proteome as an initial interactome. Rives agrees this computational proxy could be valuable. The host also notes that static structures omit dynamics that are important to much of cellular biology.
At Biohub, the team is building cryo-electron tomography with higher cellular contrast. Rives hopes this kind of work will eventually enable structurally and empirically resolved interactomes, despite major technical hurdles.
First-principles simulation cannot yet bridge the gap: even physical folding simulation works only for a few fast-folding proteins. Rives also presents an information-theoretic view of the cell, linking genome, transcription, cellular programs, and phenotype.
Historically, the field expected protein-structure prediction to come from first-principles simulation; machine-learning pattern recognition instead made major progress. Rives argues that learning the underlying programs of cellular biology may provide the right abstraction available in the current era of information theory at scale.
9. Virtual biology requires a scaled experimental feedback machine
Rives’s standard for a virtual cell is generalization: it must predict an experiment absent from its training data. Current models are “good representations of the underlying data” but weak at novel interventions in novel contexts—the capability fundamental science actually needs.
Biohub’s initiative assigns $400 million to internal data generation and enabling technology and $100 million to outside efforts. Priorities include Perturb-seq, combined transcription and imaging, spatial biology, and simultaneous phenotype, transcriptomic, proteomic, genomic, and epigenetic measurements.
Existing programs may encompass roughly a billion cells by Rives’s estimate, but he wants multiple orders of magnitude more. Current technology might scale 10x–100x with reasonable investment; another 10x or more requires better assays, automation, flexible robotics, and cheaper multidimensional measurement.
Feedback completes the thesis: models could reason over thousands, millions, or even hundreds of millions of hypotheses, select a small number of experiments, observe outcomes, and update their representations—“something like RLVR” grounded in biology. Compute and data must expand together; whether ESM-C eventually hits diminishing returns remains empirical.