Pioneers Insight Method Research Author
Gene Hunting with o1-pro: Reasoning about Rare Diseases with ChatGPT Pro Grantee Dr. Brownstein
Back to Episodes

Gene Hunting with o1-pro: Reasoning about Rare Diseases with ChatGPT Pro Grantee Dr. Brownstein

Summary

  • Labenz frames rare disease as an information-processing bottleneck rather than merely a sequencing problem. A rare disease affects fewer than 1 in 2,000 people or fewer than 200,000 people in the United States, yet collectively affects “more people with rare diseases in the United States than there are natural blondes.” Genome sequencing affordability has improved nearly 10,000×, from above $1 million in 2007 to a few hundred dollars, while the scarce asset has shifted toward interpretation.
  • Cheaper sequencing has made Boston Children’s referral cases harder, not easier. Early, rigorously selected cases produced diagnoses in roughly 80%; today, obvious findings are resolved elsewhere, leaving Brownstein’s team with previously negative cases and a roughly 10% diagnosis rate. The workflow increasingly depends on cohort formation, new assays and periodic reanalysis because “things get discovered all the time.”
  • AI’s immediate return is expert time reclaimed rather than autonomous diagnosis. Summarizing papers, genes, unfamiliar techniques and reviewer critiques has “changed my life,” Brownstein says, including narrowing 20 candidate genes to three in about 90 minutes. Labenz identifies a near-term workflow opportunity in reducing literature-search, command-line, fragmented-tool and repetitive-triage friction for scarce specialists.
  • Better models cannot compensate for inaccessible or unrepresentative data. A seemingly pathogenic variant may prove common in a rarely sequenced ancestry, leading Brownstein to argue that “we need to sequence the whole world” to separate disease from background variation. Investigator incentives, opt-in biobanks, siloed cohorts and weak sharing defaults remain major non-AI bottlenecks.
  • The current strongest operating model remains multidisciplinary humans plus AI. In a 2014 contest, 23 teams examined three apparently Mendelian families and two families received diagnoses; teams mixing clinicians, genetic counselors, researchers and bioinformaticians did better than homogeneous groups. Brownstein’s point is that “none of this exists in a vacuum”: AI is helping a multidisciplinary process rather than replacing it.
  • High-stakes adoption will be gated by verification and a harsher-than-human error standard. Brownstein still checks outputs, cannot reliably obtain citations and sometimes finds reasoning models making an impressive but unsupported extra leap; Labenz expects medical AI may face the self-driving threshold of needing to be roughly 10 times safer. Her stance has nevertheless shifted from skepticism toward conviction as “the reasoning is getting better.”
  • The OpenAI grant is testing the minimum data and compute required to move from triage toward case-solving. Brownstein is not yet feeding complete medical records and genomes into a model because of cost, energy use, irrelevant detail and tangent risk; she wants the smallest sufficient input and reliable workflows. Labenz predicts o3-class systems could return meaningful conclusions within a year, while Brownstein keeps the forecast properly hedged: “I would love to say that we’re using it to solve cases. I don’t know if that’ll be true, but I hope it is.”

Deep dive

1. Rare diseases are common in aggregate

  • Brownstein’s opening correction is the episode’s foundation: “Rare diseases are quite common, actually.” A disease can qualify by affecting fewer than 1 in 2,000 people or fewer than 200,000 Americans, while the combined population exceeds the number of natural blondes in the United States.

  • Classification depends on granularity. Autism is common, but autism caused by a de novo variant in KCNJ8 is rare; the definition is tricky and keeps evolving as researchers learn more.

  • Brownstein therefore describes herself less as a disease specialist than as “a gene hunter” trying to diagnose the undiagnosed.

2. Falling sequencing costs moved the bottleneck upstream to interpretation

  • When Brownstein began in 2011, an exome—the roughly 1% coding portion of the genome—cost about $3,800 to $4,000. She recently priced the same test at $160; Labenz adds that genome sequencing fell from above $1 million in 2007 to a few hundred dollars, nearly a 10,000× improvement in affordability.

  • Scarcity once forced rigorous patient selection. Clinicians had freezers of DNA awaiting affordable technology, and sequencing concentrated on cases where a child-parent trio was likely to reveal a clear de novo event, premature stop codon or major deletion.

  • That selection produced something like an 80% diagnostic rate: “shooting fish in a barrel.” As routine testing absorbed those cases, Boston Children’s increasingly received patients whose genomes had already returned negative, dropping Brownstein’s current referral yield toward 10%.

  • The economic inversion is central: generating a sequence has become ordinary, while explaining a negative sequence demands increasingly specialized labor, literature access and cross-case reasoning.

3. The diagnostic odyssey ends in a deliberately diverse team

  • Many families spend years moving from local providers to specialists before reaching Boston Children’s. The Manton Center for Orphan Disease Research can accept self-referrals, assemble records and prior genetics, then resequence or use RNA-seq, long-read sequencing and other newer methods.

  • In a 2014 initiative proposed by Zack Kohane, three apparently Mendelian families were released to 23 international teams. Two of the three families received diagnoses, and diverse teams did much better overall than homogeneous groups of bioinformaticians.

  • The successful mix included research assistants, genetic counselors, researchers and multiple kinds of clinicians. Brownstein’s durable lesson for AI deployment is that “having multidisciplinary teams with multiple strengths all working together means we do much better as a whole.”

4. Diagnosis is layered evidence, not one magical pipeline

  • A new case arrives with extensive medical records and, increasingly, a genome on a thumb drive. Brownstein runs multiple genomic pipelines because comprehensive systems can create false positives, while simpler black boxes may silently eliminate variants without exposing their reasoning.

  • Phenotypes are encoded through ontologies such as HPO, ICD-9/10 or SNOMED and combined with genetic data ranging from raw FASTQ files to aligned BAMs and processed VCF variant lists. The pipeline ranks variants by phenotype-gene fit and predicted pathogenicity.

  • When known disease genes fail, investigators look for unusual structural changes, translocations, deletions, duplications and disturbed gene expression. They may add epigenetic assays or consider multifactorial explanations through genome-wide analyses and rare-variant tests such as SCAT.

  • Even broadly, roughly 25% of genetic testing comes back positive, leaving about 66% to 75% negative. Brownstein’s team can exhaust this stack and still shelve the case for another year.

5. Good phenotype translation and relational biology determine what rises

  • A clue as specific as the absence of tears can point strongly toward one condition, but only if “alacrima,” “no tears” and culturally variable lay descriptions map correctly. Brownstein highlights the Monarch Initiative’s work connecting HPO terms across patient language and animal phenotypes.

  • More sophisticated pipelines use biological relationships: if gene A is related to a phenotype and interacts with gene B, a large variant in B can become a candidate even without an established disease association.

  • Brownstein’s first success followed that pattern. A patient with episodic ataxia became stiff and locked in position; sequencing surfaced KCNA1, not the expected gene but a related one, and it “rose to the absolute top of the list.”

6. Reanalysis can change a family’s future after medicine cannot change its past

  • Reanalysis matters because “you can’t be an expert in every gene, every condition, every structural variation.” A variant ignored one year may become the browser’s number-one answer after another group establishes the missing connection.

  • Brownstein recalls three siblings who died from a myopathy roughly 20 years earlier. After their DNA was exhausted, remaining RNA enabled RNA-seq and revealed a variant in “CFL2, I think,” giving surviving siblings information for carrier testing and family planning.

  • The team debated whether diagnosis after multiple deaths deserved to be called success. Brownstein’s resolution is modest but material: it cannot repair the past, yet lets the next generation proceed “with their eyes wide open.”

  • Another family received a rare-bone-disorder diagnosis spanning three generations, including a relative in their 90s. It could not undo unnecessary surgeries, but replaced a three-minute symptom narrative with a ten-second explanation.

7. Candidate findings now require cohorts, replication and population context

  • In 2011, a dramatic protein prediction could provoke immediate publication excitement. By 2024, “the waterline is rising”: one unusual variant may be random, so researchers seek additional families with variants in the same gene and comparable phenotypes.

  • Matchmaker Exchange and Beacon help investigators locate those families. A single case can join an existing 19-patient series, producing a more convincing gene-disease claim and connecting the family to specialists.

  • Prediction tools are evidence, not gospel. Brownstein uses CADD, SIFT, PolyPhen and protein-impact models, but warns that some known disease relationships would fail their filters because even a seemingly minor perturbation can matter in a particular gene.

  • An exciting conserved variant can also collapse when it proves common within a poorly sequenced ancestry. Hence her categorical data requirement: “We need to sequence the whole world” to distinguish causality from isolated background variation.

8. Biology’s missing interface layer wastes expert capacity

  • Protein folding and interaction tools such as AlphaFold and STRING-DB are “amazing,” but the biological interaction graph remains sparsely illuminated and many tools remain difficult to use. Brownstein is still surprised by how user-unfriendly cutting-edge protein analysis can be.

  • Her sharpest example is operational: a sequencing company returned data only through command-line instructions and offered no help. She needed a crash course on the Harvard/Boston Children’s supercomputer simply to transfer her own files—“a huge waste of time.”

  • Labenz identifies the opportunity as a UI and orchestration layer for biologists and doctors who are not programmers. Brownstein adds that restricting tools to people fluent in Unix also suppresses applications their creators never imagined.

9. Sharing incentives constrain discovery more than raw data scarcity

  • Brownstein understands why a young investigator might withhold a rare case, hoping to lead a career-making paper rather than become a middle author in someone else’s consortium. But the incentive can strand cases and prioritizes the investigator over patients.

  • Boston Children’s CRDC cohorts committee and hospital-wide GeneDx browser offer a better pattern: researchers can query de-identified genetic variation without receiving names, phenotypes or other identifiable information, then contact the responsible physician about a match.

  • One candidate-gene query found four additional patients; a physician then pointed Brownstein toward a Netherlands-led case series. The result was a more comprehensive condition description, publication and direct linkage between patients and experts.

  • Her “spitballing” estimate is that ideal sharing would have a huge effect. Cohorts remain “in the back of the freezer,” while enrolling in multiple programs and registries increases the chance that a future discovery will actually reach a family.

10. Administrative defaults can nullify a valid scientific result

  • Brownstein has watched biobank policy remain opt-in despite years of arguments for easier sharing of discarded tissue, urine and other samples. She sees enormous inertia around reforms where speed, rather than privacy alone, is often the affected family’s priority.

  • She also cites Mew, who made Milusen, as an N-of-1 drug story in which an opportunity and an extremely motivated family overcame major barriers. Such inspirational cases may represent the “dark matter” of many families that could not overcome them.

  • Research findings also require confirmation from a Clea-accredited laboratory and return through a physician or genetic counselor. Sometimes physicians “won’t play ball”—they see little value, avoid email or simply decline the work needed to confirm the diagnosis.

  • Her own family has been unable to obtain clinical confirmation of a research finding because a relative’s doctor would not cooperate. Multiplied across inadequate counseling and missed research enrollment, seemingly small failures become a systemic barrier.

  • Brownstein hopes AI can return some autonomy to families: help them recognize the right specialist, test or registry earlier and reduce dependence on infrastructure “that doesn’t work as well as it should.”

11. AI’s first breakthrough is eliminating scientific drudgery

  • Before the current case-triage work, Brownstein used an interactive model in a Picory grant to map patient phenotypes into HPO codes through a sequence of seven questions; that work was still being analyzed.

  • Brownstein’s largest realized benefit is deceptively mundane: article and gene summaries. Avoiding the cycle of finding a paper, encountering a paywall, logging into Harvard’s library and discovering the abstract is irrelevant has “given me hours back in a day.”

  • As a core-facility leader, she uses ChatGPT to learn newly launched sequencing methods, prepare for investigator consultations and summarize unfamiliar conditions among nearly 20,000 genes. It also helps decode reviewer number two’s vague claim that she overlooked “a whole body of literature.”

  • The common arrow is from tedious search toward expert judgment: “It’s all cutting down on this mundane, time-consuming, really tedious part of the job,” then returning her to gene discovery.

12. A 500-patient bladder cohort shows AI operating as a research filter

  • Brownstein studies severe interstitial cystitis/bladder pain syndrome, sometimes debilitating enough that patients cannot leave home. In a cohort of roughly 500, she and collaborators look for genes carrying more variation than expected and for enriched multi-gene pathways.

  • A pathway may contain 12 genes under a broad label such as small-molecule transport. ChatGPT can rapidly answer a first-pass question—whether each gene has a direct tie to bladder pain, bladder biology or bladder cancer—before Brownstein verifies and investigates the “yes” results.

  • One pathway had about 60% of its genes tied specifically to bladder cancer, suggesting bladder expression and known perturbations. Another gene’s link to urothelial issues nearly made her scream because it supplied a plausible mechanism worth discussing at a meeting the next morning.

  • In another case, the same approach helped narrow 20 plausible genes to three for presentation in about an hour and a half. Brownstein’s forecast is generational: future geneticists may not know how the work was done before these tools.

13. Reasoning models can overreach in ways that are either error or discovery

  • Brownstein is still experimenting with o1 and GPT-4o rather than following a settled model taxonomy. She is not using web search because “there’s a lot of garbage on the web,” and says checking outputs remains paramount even as hallucinations and reasoning improve.

  • Labenz tries to prompt neutrally because models can mirror a user’s preferred theory. Brownstein has the opposite control problem: asking whether a gene relates to a phenotype can produce a clever multi-step mechanistic argument when she only wants a paper directly linking the two—“No, too far, too far.”

  • Whether that extra inference is hallucination or an AlphaGo-style “move 37” remains unresolved. Brownstein’s honest answer is, “Maybe I’m not smart enough to understand it, and it’s right. I don’t know,” followed by her proposed method: keep testing it.

14. The grant targets minimum sufficient data, compute and integration

  • Brownstein’s OpenAI work asks how to diagnose faster, generate new hypotheses and identify the minimum dataset and compute required for an answer. She currently summarizes cases rather than uploading entire medical records, even behind Boston Children’s protected ChatGPT deployment, where she still cannot get citations to work properly.

  • Whole records and genomes invite tangents, irrelevant correlations and heavy compute use: “Not everyone who smokes gets cancer.” Every question takes energy, so the research problem is not merely maximizing context but selecting meaningful evidence.

  • A year from now, Brownstein envisions wrapped, guided interfaces that help patients ask the right questions and make researchers’ use cases clearer. She hopes case-solving arrives but refuses to promise it: “I don’t know if that’ll be true.”

  • Labenz is more aggressive, citing o3 FrontierMath performance near 25% with very high compute and 10% at low compute, versus a previous ceiling around 2%. Even with context limits, he expects summarization and filtering to make some rare-disease cases tractable; Brownstein proposes reconvening in a year to see.