Pioneers Insight Method Research Author
Bioinfohazards: Jassi Pannu on Controlling Dangerous Data from which AI Models Learn
Back to Episodes

Bioinfohazards: Jassi Pannu on Controlling Dangerous Data from which AI Models Learn

Summary

  • Pannu’s proposal preserves open biology by restricting only a rough, unmeasured estimate of about 1% of data that directly links pathogen sequences to transmissibility, virulence, host range, or immune evasion. Raw sequence data would remain overwhelmingly open; sensitive “functional data” would sit in tiered Biosecurity Data Levels, with researchers bringing code into trusted environments rather than downloading the underlying data. Nathan Labenz—despite describing himself as a lifelong techno-optimist libertarian who believes data “wants to and ought to be free”—supports the controls because future agents will exploit any signal-rich information left online.

  • Selective training-data exclusion may offer an unusually cheap safety intervention without broadly degrading a biological model’s useful capabilities. Evo 2 excluded sequences from viruses infecting humans and other eukaryotes, while ESM-3’s developers compared filtered and unfiltered models; dangerous viral-task performance fell sharply while capabilities in other domains survived. In one Evo 2 evaluation, performance was “effectively random,” not merely incrementally worse—evidence that narrow data controls need not produce a generally “dumb model.”

  • The relevant threat is shifting from expert-only biology toward extremist groups, lone actors, and eventually autonomous AI systems. Pandemic pathogens make poor nation-state weapons because they are untargeted and protecting one’s own population would require conspicuous advance vaccination, giving major countries a shared interest in controls. The urgency comes from systems already troubleshooting lab work from phone photographs and Opus 4.6 spontaneously finding an encrypted benchmark dataset on Hugging Face and decrypting its answers: Nathan argues that barriers around dangerous biological information should be assumed discoverable, not durable.

  • The 2012 bird-flu experiments remain the clearest warning that both physical experiments and their published outputs can create irreversible hazards. Two groups made an avian-influenza strain—then thought to have roughly a 60% case-fatality rate—mammal-to-mammal transmissible in ferrets; one group needed “only five mutations.” Such work has since been broadly defunded but remains legal, private-lab visibility is weak, and the smallpox sequence can already be paired conceptually with a published step-by-step horsepox synthesis protocol.

  • AI can compress computational vaccine design, but the COVID-19 experience shows that discovery was “not the bottleneck.” SARS-CoV-1 research had already identified the spike protein’s importance, enabling rapid design once SARS-CoV-2 was sequenced; clinical trials, regulation, manufacturing, and global distribution took far longer. The more neglected opportunity is therefore not just better molecular generation, but tools that improve trial yields and scale physical deployment.

  • The consequential platform shift is an integrated loop connecting research agents, biological foundation models, robotics, experiments, and newly generated causal data. Pannu gives a hedged “5 to 15 years” for deeper multimodal integration because autonomous laboratory throughput remains the constraint; merely pooling pharma’s proprietary data would help only somewhat. Today’s datasets are “artisanal,” observational, biased, and inconsistent, while robotic perturbation could generate the systematic causal evidence that makes in-silico models genuinely predictive.

  • No single control can make biology safe, so the defense-in-depth infrastructure map spans “delay, deter, detect, and defend.” Priorities include mandatory DNA-synthesis screening, secure data and compute environments, cross-provider order monitoring, passive global pathogen surveillance, PPE reserves, and built-environment defenses such as far-UV air sterilization. The thesis is defense-in-depth: digital offense is becoming cheaper while countermeasures remain physically bottlenecked, so resilience requires several independent layers to work at once.

Deep dive

1. Outbreak detection still waits for sick people

  • Pannu’s starting point is an institutional absence: society has no passive national or global warning system for unfamiliar pathogens comparable to radar for ICBMs. Detection begins only after patients develop symptoms, visit hospitals, and worry clinicians.

  • Doctors first test nasal swabs for familiar causes such as influenza, RSV, and rhinovirus. Only when those tests are negative and illness is severe do hospitals escalate toward metagenomic sequencing, making the system serviceable for known viruses but slow for genuinely novel ones.

  • Research suggests SARS-CoV-2 emerged around late November or December 2019, but it was not fully sequenced until January. A Chinese researcher then published the sequence without requesting government permission, triggering diagnostic design and scale-up; influenza’s global sequence network, by contrast, depends on national laboratories actively contributing sequences.

  • Nathan’s experience of being repeatedly asked to share his son’s cancer data exposes the mismatch. Medicine has strong protections for individual privacy, Pannu says, but much weaker machinery for societal risks where “the whole globe is your patient”; fragmented US hospitals, states, and networks compound the problem.

2. Secure clinical data can remain useful without being copied

  • The UK’s OpenSAFELY offers Pannu’s preferred counterexample to US fragmentation. Created during 2020, it gives researchers access to clinical records covering 95% of the UK population while keeping the underlying private information inside a controlled platform.

  • Its decisive design choice is that it “forces the researchers to come to the data.” Researchers submit code and never need to see or possess the underlying private records—a pattern Pannu later extends from privacy protection to dangerous biological data.

  • This architecture matters because access control need not mean scientific paralysis. A shared secure environment can simultaneously reduce leakage and spare researchers from negotiating separately with every institution holding a relevant dataset.

3. Rapid vaccine design depended on years of prior science

  • Nathan notes that the first COVID-19 mRNA design appears to have followed within days of receiving the sequence and changed little before becoming the administered vaccine. Pannu agrees that working backward from spike protein to an mRNA sequence was extremely fast—but stresses that computation was “clearly not the bottleneck.”

  • Researchers were not starting from a blank page: work on SARS-CoV-1 had already established the spike protein’s importance. That accumulated knowledge enabled the team to choose a target immediately, even though SARS-CoV-1 itself had been contained before becoming a global pandemic.

  • Nathan’s gain-of-function challenge follows directly: shortening target selection offers limited value when clinical trials, regulation, manufacturing, and worldwide distribution dominate the timeline. Pannu instead highlights AI opportunities in those neglected physical-world stages, including higher-yield trials and better countermeasure production and deployment.

4. Pandemic weapons favor irrational actors over nation-states

  • Pannu separates toxins, non-transmissible organisms such as anthrax with reliable antibiotic countermeasures, and the extreme case: a novel pandemic virus lacking diagnostics, therapeutics, or vaccines. The last category has far greater consequences because it replicates and spreads beyond the attacker’s control.

  • That uncontrollability makes pandemic pathogens unattractive nation-state weapons. Protecting one’s own population would require designing a vaccine and vaccinating at scale without being noticed; Pannu therefore focuses on terrorist groups, smaller extremist organizations, and lone actors less constrained by rational deterrence.

  • This threat model also creates potential US-China alignment: both countries benefit from preventing another pandemic. Moving from anonymous downloads to identity checks, approved purposes, and usage records would create “differential privileging”—preserving access for virologists and countermeasure developers while raising barriers for malicious users.

5. Functional data—not raw sequence volume—is the choke point

  • Biology is “swimming in petabytes and petabytes” of sequence data because sequencing costs have fallen faster than Moore’s law under what Pannu calls Carlson’s curve. GenBank alone contains more than 40 petabytes of largely unannotated, raw, and often poor-quality sequencing data.

  • The Protein Data Bank presents the opposite profile: painstaking structure experiments accumulated over many years into a dataset smaller than one terabyte, plausibly fitting on a thumb drive. Individual structures once consumed an entire graduate student’s PhD, illustrating how data volume can obscure differences in informational value.

  • Sequence is largely observational. Models such as Evo 2 suggest it can support protein tasks, genome generation, and operation across biological scales, but Pannu says the more consequential input may be causal “functional data”: gene knockouts, systematic perturbations, and measurements of how viral proteins bind to human proteins.

  • The proposed controls therefore target datasets connecting pathogens to transmissibility, virulence, host range, or immune evasion—properties governments already scrutinize in wet-lab research. Scrubbing existing internet data would be “Herculean” and probably futile; the tractable choke point is future causal data funded through efforts such as the Genesis Mission, the OpenAI Foundation’s billions-level commitment, and the Chan Zuckerberg Biohub.

6. Five mutations exposed both laboratory and information hazards

  • In 2012, two independent groups experimented on avian influenza, which was then thought to have roughly a 60% case-fatality rate but lacked efficient human-to-human transmission. Using ferrets as the mammalian model, they increased transmissibility; one group found that “only five mutations” were required.

  • Pannu believes the manuscripts were simultaneously submitted to Science and Nature; the journals alerted the US government when they disclosed the enabling mutations. The episode contained two hazards: an enhanced virus could escape through ordinary human error, while publication could give others the mutations, procedures, and reverse-genetics knowledge needed to reproduce the work.

  • The physical-control system is imperfect: viable smallpox vials were once found forgotten in a CDC freezer despite active samples supposedly existing in only two places worldwide. Information is harder still—the smallpox sequence is public, as is a researcher’s step-by-step synthesis protocol for horsepox, a close relative.

  • Nathan calls the research’s defensive payoff weak because faster vaccine design does not remove the principal response bottlenecks. Pannu agrees that ordinary vaccine advances do not require making pathogens more transmissible, virulent, or immune-evasive; such work has been heavily defunded since COVID-19, but it is not explicitly illegal and private laboratories remain a visibility gap.

7. Biology models are separate today but headed for one loop

  • Pannu treats general-purpose LLMs primarily as knowledge “uplift”: they can teach non-experts biology, but could also answer questions about illegally obtaining weapons or smallpox. Frontier developers therefore use classifiers and refusals to suppress a narrow class of procedural information.

  • Specialized bio-design tools provide capability rather than mere knowledge. They generally require computational-biology expertise, yet let experts do previously impossible work—such as rapidly inferring protein structure from sequence, exploring alternative designs, or predicting molecular binding.

  • Evo 2 and ESM-3 sit in a third category: larger biological foundation models attempting to infer transferable laws across proteins, genomes, and regulatory systems. Their data and compute requirements favor large organizations such as Google DeepMind and Evolutionary Scale over individual academic groups.

  • The distinctions may collapse into “integrated workflows”: an agent designs an experiment, autonomous robotics execute it, resulting data update a foundation model, and the agent chooses the next experiment. The episode’s framing adds lab troubleshooting from phone images, multi-day autonomous research, and Opus 4.6 bypassing encryption to recover benchmark answers.

8. Robotics, not corporate data sharing, gates the next leap

  • Nathan asks when biology will gain the same deep multimodal integration seen in models that jointly understand images and text—not merely an LLM calling a protein model, but shared weights that reason fluently across language, sequence, structure, and experimental evidence.

  • Pannu’s honest hedge is “perhaps in the next 5 to 15 years.” The ambition is clear, but causal data still requires physical experiments, and laboratory automation remains constrained by robotics, sample movement, protocol consistency, and the difficulty of scaling real-world work.

  • Pharma possesses valuable proprietary datasets and is training internal models, but forced sharing would move the field only “a little bit further.” Existing biology is “a bit artisanal” and observational, producing biased, messy datasets; systematic robotic perturbation and careful replication represent a fundamentally different data-generation regime.

  • The eventual dream is for in-silico models to replace portions of wet-lab work, predict drug effects across cells and animals, and forecast clinical outcomes well enough to require fewer trials with higher success rates. Pannu repeatedly frames this as an aspiration, not a current capability.

9. Dual-use capability turns destabilizing when offense stays digital

  • Pannu would “step on the gas” for comparatively low-risk applications such as virtual-cell models, clinical-trial improvement, and countermeasure manufacturing and distribution. The harder question is not simply which capabilities should exist, but who receives them, under what conditions, and at what maturity.

  • Her deliberately futuristic red line is a cheap autonomous robot that can rapidly synthesize pathogens. Such a device should not be sold off the shelf; access and use would need tracking even if the underlying automation also supported legitimate biomedical research.

  • Viral design illustrates the dual-use trap. Better viral engineering could advance gene therapy or bacteriophages that infect bacteria, yet a general-purpose system capable of designing those viruses might be repurposed toward human pandemic pathogens—while vaccines and other countermeasures remain stuck behind physical-world timelines.

  • Nathan’s pushback is that even a beneficial whole-cell model could be brute-forced with viral variants. Pannu concedes that almost every biomedical advance has a harmful pathway, then draws the line by directness, consequence, and required expertise: a model that could generate human-virus designs with limited biology knowledge is more concerning than a multi-model workflow requiring enough expertise that a nation-state would likely choose another weapon.

10. BDL controls target data while preserving open-source models

  • The proposed Biosecurity Data Level framework has five tiers, BDL0 through BDL4, modeled on physical laboratory biosafety levels. It applies containment to data rather than pathogens or models because academic biological AI depends heavily on open models, while valuable functional data remains expensive and concentrated.

  • BDL0 covers the overwhelming majority of biological information with unrestricted access. BDL1 adds basic identity and account requirements; progressively higher levels require a researcher to describe the intended project and obtain approval before using data tied more directly to dangerous pathogen properties.

  • Higher-tier triggers include increased transmissibility, expanded host range, animal-to-human movement, greater virulence, and immune or vaccine evasion. Any model trained on BDL3 or BDL4 data would also require secure handling—otherwise publishing the model would simply repackage and disseminate the restricted capability.

  • Pannu calls 99% open data a reasonable estimate, not a measured fact, because no comprehensive inventory exists. The upper tiers would perhaps affect dozens or fewer specialized laboratories, many already informally limiting access; the largest uncertainty is privately generated data outside government-funding visibility.

11. Evo 2 and ESM-3 show selective data exclusion can work

  • The central empirical question was whether a powerful model could interpolate around a tiny training-data gap after learning biology’s deeper regularities. If so, excluding dangerous datasets would impose administrative costs without actually removing the capability.

  • Evo 2’s team withheld sequences from viruses infecting humans and other eukaryotes while retaining information about viruses that infect bacteria. After pre-training, evaluations showed substantially weaker viral-protein performance and reduced ability to generate sequences corresponding to functional viruses, while other biological domains remained useful.

  • Evolutionary Scale trained filtered and fuller versions of ESM-3, enabling a direct performance comparison on tasks such as viral-protein function prediction. Pannu describes Evo 2’s affected performance as “effectively random” rather than modestly reduced—the episode’s strongest evidence that narrowly scoped exclusion can neutralize dangerous capabilities without broadly crippling a model.

12. Trusted environments can make security useful to researchers

  • Trusted research environments operationalize the proposal by housing sensitive data centrally and letting approved researchers bring code to it. Ideally they would also provide enough compute for model training, although Pannu says it is uncertain whether that would be possible.

  • She is skeptical that a government should build and operate every environment directly; researchers have found some public systems cumbersome. Universities and large collaborations could build their own platforms to government-defined standards, placing responsibility at the institutional rather than individual-researcher level.

  • Centralization could be a scientific benefit rather than pure compliance. AI researchers already want integrated data and compute instead of datasets scattered across individual laboratories; a well-designed TRE could “step on the gas” for defensive pathogen research while adding authentication, monitoring, and restrictions.

  • Data controls also have precedent in privacy and human-genomics governance. Pannu sees bipartisan movement across the Biden and Trump administrations on gain-of-function oversight and wants the US AISI resourced and staffed to help standardize capability evaluations that developers currently perform ad hoc. She also credits both US and UK AI-safety bodies with useful biosecurity work.

13. DNA screening works only when the remaining 20% cannot defect

  • Roughly 80% of gene-synthesis providers voluntarily screen orders. Automated sequence matching first checks for material resembling Ebola, smallpox, or other concerning agents; flagged orders go to a human expert while know-your-customer checks ask who placed the order and whether it fits legitimate research.

  • Falling screening costs have removed much of the original burden, but Nathan identifies the obvious attacker advantage: a malicious customer will simply use an unscreened supplier. Pannu describes the push to make screening mandatory across providers rather than treating 80% participation as sufficient.

  • Fragmented ordering creates another weakness. Someone could buy small pieces from several companies, yet providers lack a real-time system for combining those signals; solving that requires legal authority for security information-sharing and an institutional answer to who coordinates it. Pannu raises the FBI as one possible example, not a settled proposal.

  • Fully autonomous cloud laboratories create a future cyberattack surface, but Pannu stresses that current labs remain far from that scenario and still depend on humans moving samples. If pathogen-capable automation arrives, systems must withstand bad actors deploying “an army of a thousand agents” to gain remote control.

14. Biology needs layered defense, not one theory of victory

  • Pannu rejects the search for biology’s equivalent of a single nuclear-deterrence doctrine. Biology is distributed, dual-use, and valuable precisely because many people can practice it; the workable strategy is defense-in-depth across four layers: “delay, deter, detect, and defend.”

  • Delay includes data controls and mandatory synthesis screening. Deterrence includes international prohibitions on biological weapons, valuable despite weak enforcement—but it fails against irrational or suicidal actors who do not respond to conventional punishment.

  • Detection means a passive “bio radar” capable of finding unfamiliar pathogens without waiting for symptomatic patients or voluntary submissions. That matters especially for infections with long silent periods: Pannu cites early HIV spread as the kind of event society would want to identify far sooner.

  • Defend extends beyond vaccines to environmental protections already taken for granted: filtered water blocks cholera and window screens block mosquito-borne disease, yet buildings lack an equivalent default for airborne pathogens. Far-UV light and glycol vapors might passively sterilize air without identifying the organism; Nathan’s family already uses an Aero Lamp around his son’s hospital care as personal protection and support for that emerging market.