Using AI to navigate son's cancer diagnosis
Summary
AI changed the trajectory of Nathan Labenz’s response to his six-year-old son’s rapidly advancing Burkitt leukemia. Human clinicians twice discounted an LDH reading around 1,500 because the samples showed hemolysis and said they had “ruled out anything life-threatening”; GPT-5 Pro, Claude 4.1 Opus, Gemini 2.5 and Grok 4 instead treated it as a red flag requiring action within 24–48 hours. Labenz’s strongest practical conclusion is categorical: in a serious medical crisis, “use both human and AI doctors aggressively.”
The outcome so far is close to the best case available for an otherwise terrifying disease. Burkitt leukemia can double in as little as 24 hours, and Ernie’s disease included a large abdominal mass, widespread lesions, bone-marrow involvement and possibly ambiguous spinal-fluid involvement. Roughly 10 days after treatment began, the main mass had shrunk about 80%, many lesions had disappeared, and the latest marrow and spinal-fluid tests were clear—making cure, rather than chronic management, the overwhelmingly likely outcome.
The product lesson is that model choice, tone and context presentation can alter real-world decisions even when the underlying data are identical. Claude 4.1 Opus was so alarmist that it helped trigger a panic response, while GPT-5 Pro delivered essentially the same warning in a longer, more clinical voice that Labenz preferred; yet he concedes Claude’s urgency “might have been what I needed.” Regular GPT-5 also repeated the doctors’ error when portal screenshots foregrounded the hemolysis warning, showing why users should seek “n opinions,” vary the prompt and inspect the totality rather than trust a single output.
At bedside, GPT-5 Pro functioned less like a diagnosis app than a continuously available clinical analyst. Labenz fed it nearly 20 rounds of labs and updates over nine or 10 days, using it to validate the treatment protocol, interpret changing vitals and question mineral replacement, blood-pressure medication and additional testing. It could not examine Ernie, prescribe drugs or perform procedures, but it gave the family enough understanding to advocate intelligently and enough agreement with clinicians to trust that the local team was delivering state-of-the-art care.
AI’s most differentiated contribution may be personalized synthesis across specialties and beyond standard workflows. When a basement flood created mold concerns during immunosuppressive chemotherapy, GPT-5 Pro combined the child’s oncology context with air-test and remediation data—an analysis that would otherwise require arranging a meeting between “your oncologist and your mold guy.” It supported a concrete plan involving remediation precautions and eight to 10 HEPA filters, reducing both health anxiety and the perceived need to leave the house.
Labenz intends to depart from standard follow-up by pursuing circulating tumor DNA, or ctDNA, to detect a relapse earlier than symptoms or scans might. The proposed workflow fingerprints stored tumor tissue, then looks for that signature in blood at sensitivity he believes may reach five cancer cells per million; his oncologists regard it as unproven or nonstandard. His reasoning is that a 24-hour doubling cancer should be attacked at the smallest detectable burden, both to reduce treatment-linked cytokine storm and to give the cancer fewer “at bats” to evolve drug resistance: “If ctDNA comes back positive on a Tuesday, you want to be infusing one of these next-generation drugs on a Thursday.”
The revealed willingness to pay makes healthcare one of the clearest demonstrations of AI consumer surplus, while exposing the regulatory bottleneck ahead. Labenz pays $200 a month for GPT-5 Pro against an estimated $500,000–$1.5 million six-month treatment bill and says he would willingly pay $10,000 a month in this crisis; in his phrase, the economics are both “GDP destroying” and “absolutely insane.” He expects AI-supported patients to challenge standard-of-care rules, clinical-trial access, liability doctrine and poor data portability—while still supporting a ban on uncontrolled superintelligence because faster medical progress and existential-risk governance are not, in his view, contradictory. The experience also left him feeling there are “no decels in the pediatric oncology unit”: delays in AI progress have real human costs.
Deep dive
1. A terrifying diagnosis quickly produced a best-case response
Nathan Labenz opens with the fact that reordered everything else: his previously healthy six-year-old son, Ernie, has Burkitt leukemia. Doctors initially called it Burkitt lymphoma, but bone-marrow involvement changed the classification as the family understood it.
The disease’s scale was frightening—a large abdominal mass, lesions throughout the abdomen and above the diaphragm, marrow involvement and possible, still-ambiguous spinal-fluid involvement. Burkitt can have a doubling time “as fast as 24 hours,” making delays unusually consequential.
The counterweight is equally important: this is among the cancers most responsive to chemotherapy. After the initial treatment wave, the central mass was down roughly 80%, many colonies had vanished, and the latest marrow and spinal-fluid tests were clear.
Labenz waited for that milestone before speaking publicly. With the response “about as well as one could hope,” the overwhelming expectation is cure: the cancer disappears, never returns and Ernie lives a long, normal life.
2. Cancer made the abstraction of exponential growth brutally concrete
Labenz’s analogy joins the subject of his podcast to the family crisis: both AI capability and cancer can look “flat behind you and vertical in front of you.” One malignant B cell began dividing out of control until seemingly minor symptoms crossed thresholds into a crisis.
Looking backward, the family cannot know whether a week of stomach pain in early September was the first “little murmur” of the disease. A doctor’s constipation explanation appeared to fit, prune juice seemed to help, and the complaint disappeared.
Knee pain after a trampoline-and-bounce-house weekend also had an obvious benign explanation. Amy worried sooner and even considered leukemia, partly through her own AI use; Nathan’s baseline response was closer to “Yeah, he’ll be fine,” as it had always been before.
By the October 22 vacation, fatigue, knee pain and the earlier stomach complaint still each had plausible explanations. Within days, however, Ernie was restless at night, moaning, intermittently screaming in agony and reporting pain in his tooth, head and multiple other places.
3. Episodic illness repeatedly disappeared inside the clinic
On October 24, after participating fairly normally in a Louisiana food festival, Ernie screamed in pain during the drive home and asked to go to the hospital—then chose to wait until morning. An October 25 urgent-care knee X-ray was normal, and staff could not complete a blood draw.
The mismatch mattered: clinicians saw a calm child with a tablet, not the child his parents had watched screaming. Labenz’s practical advice is to film the worst episodes because “there was a disconnect between what we had seen and the way he was presenting.”
He also recommends producing a written symptom chronology once a confusing pattern becomes concerning. Giving a clinician the entire history “in black and white” forces a survey-level view before any one symptom pulls the visit down another benign rabbit hole.
Labenz adds a third retrospective recommendation: record major appointments and use AI to transcribe them. Preliminary interpretations often arrived through short conversations without a formal report, leaving the parents with less precise information than the care team possessed.
4. One discounted lab value became the hinge of the case
At Children’s Hospital New Orleans, nearly every blood count and chemistry result looked normal except lactate dehydrogenase, or LDH, at roughly 1,500—around three times the cited reference range. LDH reflects cell turnover and can rise when cancer cells are proliferating and dying.
Both samples carried a hemolysis warning because cells had begun breaking down in the specimen. Clinicians treated the elevated LDH as an artifact and told Amy they had “ruled out anything life-threatening,” while recommending routine pediatric follow-up after the family returned home.
Back in Detroit, Labenz manually copied every portal result and assembled the parents’ full chronology. He trusted himself more than an agent for this one-off, high-stakes extraction: “This is N of one,” and accuracy mattered more than automating a repetitive task.
GPT-5 Pro, Claude 4.1 Opus, Gemini 2.5 and Grok 4 varied in tone but agreed on the substance: hemolysis might raise LDH, probably not that much, and the family should obtain a proper workup within 24–48 hours. Leukemia was not necessarily most likely, but it urgently needed exclusion.
5. The same evidence produced materially different AI behavior
Claude 4.1 Opus responded as though leukemia were the likely explanation and used language Labenz found emotionally overwhelming. It precipitated something like a mild panic attack and pushed him toward GPT-5 Pro’s longer, more clinical and less emotional style afterward.
His judgment remains deliberately unresolved: “Claude might have it right.” For a father inclined to assume things will pass, the model’s alarm may have been precisely what moved him out of denial quickly enough for a cancer capable of doubling daily.
Amy’s regular GPT-5 made the opposite error when she supplied screenshots containing the portal’s hemolysis disclaimer. It discounted the LDH much as the physicians had, whereas Nathan’s text prompt presented the same caveat in the parents’ narrative and elicited an urgent warning.
The resulting method is not “ask AI once.” Labenz recommends the best available models, Pro for anything that truly matters, multiple formulations of the evidence and comparison across outputs: AI makes “n opinions” practical, but presentation can still steer an individual answer badly.
6. Parent advocacy and a physical exam finally exposed the mass
Their pediatrician still considered a self-resolving childhood complaint most likely, but called Children’s Hospital of Michigan himself and secured an oncology appointment the next morning. The oncologist initially favored an autoimmune explanation over cancer.
The decisive moment came when she palpated Ernie’s lower-left abdomen and he recoiled. Amy and Nathan immediately reported that the pediatrician had elicited the same reaction the prior day, turning a detail that might have been forgotten into a repeated physical finding.
Labenz’s hindsight is that AI should have interviewed the parents for missing information and guided them through a parent-observed physical check. Models cannot replace an exam, but they could have helped identify where touch produced a uniquely abnormal response and made that evidence legible to clinicians.
Amy then pressed to obtain an ultrasound immediately rather than accept a later appointment. The October 31 scan found a roughly 5 cm by 6 cm by 4 cm abdominal mass, with possible liver, spleen and kidney involvement; once the report reached the portal, AI made the seriousness difficult to deny.
7. Most-likely thinking underweighted the fastest plausible disease
An MRI was scheduled for Monday under general anesthesia. Over the weekend, clinicians and AI both leaned toward neuroblastoma, described as slower-moving on a weeks-to-months scale, and advised that waiting a couple of controlled days was reasonable if pain and fever stayed manageable.
Labenz now thinks the system anchored too heavily on that leading hypothesis. AI estimates at the time put neuroblastoma around 30% and lymphoma around 25%; the smaller probability was not remotely small enough to ignore when its time sensitivity was radically greater.
Monday’s MRI led to a Tuesday biopsy, liver sampling, marrow aspiration and other procedures, followed by direct admission. The scheduling was impressively agile, yet Ernie’s decline showed how a system can move quickly by normal standards and still nearly lose a race against a 24-hour exponential.
The broader lesson is not that every uncertain case demands maximum intervention. It is that high uncertainty should be tested against the worst plausible trajectory, especially when the consequences of waiting are asymmetric and weekends materially slow the medical system.
8. Clinicians broke protocol when the trajectory outran pathology
By November 5, Ernie was swollen, profoundly weak and barely leaving bed while the team waited for definitive pathology. The doctors’ refrain—“We need to get to the diagnosis”—reflected the legitimate need to match treatment to cancer subtype.
Nathan’s pushback was temporal: if results slipped through the weekend, Monday was five days away, and “five days from now seems like a really long time.” When Ernie then said he could not move his legs, the bedside oncologist’s concern visibly jumped.
Surgical appearance had already made neuroblastoma unlikely and lymphoma increasingly probable. The oncologist returned ready to “break the rules a little bit”: candidate lymphoma regimens were similar enough to start treatment under imperfect information and refine it when the reports arrived.
A steroid began Wednesday, followed by conventional chemotherapy Thursday; the working diagnosis reported Friday, November 7, was Burkitt lymphoma. The family’s subsequent understanding was Burkitt leukemia because of marrow involvement, while further genetic and subtype testing remained pending. Labenz says he does not know the counterfactual timeline, though it seemed plausible that a few more days could have brought organ failure or death.
9. Fast growth created both the danger and the therapeutic opportunity
The pathology reframed the outlook: “super fast growing, super aggressive, but also super treatable.” Chemotherapy targets cells actively growing and dividing, so Burkitt’s defining liability also makes it unusually responsive compared with slower cancers that can be harder to kill.
Treatment still begins cautiously because rapid tumor destruction can itself overwhelm the body through tumor lysis syndrome. The prephase, or debulking phase, was intentionally less intense and nevertheless reduced Ernie’s cancer volume by about 80%.
The planned six-month roadmap then moves through two hard-hitting phases, two moderate “mop-up” rounds and two milder consolidation rounds. The goal is every last cell: residual disease can restart the exponential, mutate under selection and return in a much harder-to-treat form.
Labenz is equally careful not to generalize this pathway to all cancer. For this particular disease, diagnosis converted panic into a managed schedule because a highly effective, protocolized treatment exists; many slower-growing, rarer or surgery-dependent cancers pose entirely different problems.
10. GPT-5 Pro became a continuous bedside second opinion
The first parental lever after diagnosis was nonstop verification of the team’s work. Blood was drawn as often as every six hours, medications and vitals changed rapidly, and fluid treatment once caused a nearly 50-pound child to urinate close to a gallon within several hours.
Labenz maintained a long-running GPT-5 Pro thread containing symptoms, medications, reports and successive portal PDFs—close to 20 uploads over nine or 10 days. Each update asked what the results meant, what clinicians might be missing and whether the proposed response made sense.
The model repeatedly confirmed that the main protocol was state of the art and highly standardized. That helped answer whether the family should relocate to a more prestigious center: for this known disease and established regimen, the analysis said they were already receiving the care a competent leading center would provide.
Its value was therefore partly adversarial and partly reassuring. Rather than constant Googling and free-floating anxiety, Labenz gained enough mechanistic understanding to advocate when necessary and enough independent confirmation to believe the local team’s judgment was sound.
11. Productive disagreement improved questions without displacing doctors
GPT-5 Pro was somewhat more eager than clinicians to replenish calcium and magnesium when fluid dilution pushed them low. The family’s advocacy may have moved supplementation earlier on the margin, though Labenz characterizes the difference as small rather than a discovered medical error.
When Ernie’s blood pressure barely crossed the treatment threshold, GPT-5 Pro favored remeasurement and watchful waiting. The physicians preferred medication; armed with the model’s explanation, Labenz questioned them, found their reasoning compelling and said, “I trust your judgment. Let’s go for it.”
The model also proposed extra tests for abnormal liver enzymes that doctors considered predictable after a liver biopsy. Again, their disagreement was marginal: AI wanted more information, while clinicians could place the number inside a familiar pattern and avoid unnecessary investigation.
Labenz’s boundary is clear: AI cannot perform procedures or exams, order tests or prescribe. Its role was to make the parents competent participants in a human care system—and, when human explanations survived informed questioning, to strengthen rather than weaken trust.
12. Long context turned fragmented records into a usable case model
The recurring prompt grew toward 100 pages: parental history, MRI, CT and PET reports, spinal-fluid and marrow analyses, pathology, cellular-marker percentages and other PDFs. Labenz’s experience is that current models handle such long context well enough that old fears about “overwhelming” them no longer fit.
He also asked the bedside thread to generate a dense handoff document “for a new attending physician” capable of supporting cutting-edge care. Compressing the accumulated history into a few thousand structured words made it reusable in fresh model sessions without discarding the medical through-line.
Data acquisition remained the bottleneck. Important results arrived inconsistently through the portal, paper printouts or documents apparently degraded by faxing; phone photos of irregular tables demanded manual checking because a subtle transcription error could change which markers appeared on the cancer cells.
For crisp website screenshots, direct multimodal input generally worked. For difficult photographed documents, Labenz found Gemini 2.5 Pro better than GPT-5 Pro at transcription, though still not reliable enough to skip line-by-line verification in a high-stakes case.
13. AI connected oncology to a household mold problem no specialist owned
The second parental lever was preventing infection while chemotherapy suppresses Ernie’s immune system. Viral illness concerned the team, but bacterial and fungal infections could be more dangerous, including organisms already present in the mouth, skin or gut.
The family could adopt a temporary COVID-style exposure protocol because both parents work at home and only one other child attends school. Nathan suspended AI-event travel and on-site speaking, at least for several months of the six-month treatment roadmap.
A basement flood added a cross-domain problem: about two inches of water arrived just as the family returned from New Orleans, raising mold concerns for an immunocompromised child. GPT-5 Pro helped identify local remediation companies, interpret comparative indoor and outdoor air samples, and connect those findings to Ernie’s specific oncology risk.
The resulting plan was to isolate Ernie from active basement work, perhaps stay with family for several days, and run roughly eight to 10 new HEPA filters throughout the house. Slight basement elevation, clean results elsewhere and filtration effectiveness gave Labenz confidence that permanent relocation was unnecessary.
14. Relapse planning surfaced ctDNA before the doctors recommended it
The third lever is preparation for the smaller but devastating outcome in which the cancer returns. Relapse usually happens within six months and almost always within two years; if it occurs, Labenz says bluntly that most affected children die, though not all.
He is gathering the still-pending genetic subtype, remote second opinions from perhaps two leading centers and information about their trials. Formal reviews costing under $1,000 could both validate current care and establish relationships before an emergency requires CAR-T cells, bispecific immunotherapies or another experimental line.
The proposed nonstandard step is circulating tumor DNA. A laboratory would fingerprint stored tumor tissue, then test blood for that exact signature—at sensitivity Labenz believes may be around five cells per million—during treatment and surveillance.
Clinicians acknowledged ctDNA but called it unproven or questioned what they would do with a positive result before symptoms or imaging. Labenz’s answer is preparedness: ordinary monthly blood work and quarterly scans often detect relapse only after sickness or substantial growth, which is especially inadequate for a daily-doubling cancer.
15. Early detection links relapse biology, model competition and medical reform
The mechanistic case for ctDNA is that next-generation treatments can trigger cytokine storm, whose severity appears to rise with disease burden. Fewer cancer cells should also mean fewer “at bats” for evolution to discover another drug-escape mechanism; Labenz preserves the uncertainty but finds the causal argument compelling.
A model he believed to be a Gemini 3 checkpoint sharpened that plan with responses he found five to 10 times faster, shorter and more “pitch perfect” than GPT-5 Pro. He still used GPT-5 Pro as a checker and had to request citations explicitly, but within hours Gemini had become the output he most wanted to read.
The economics behind that preference are extreme: $200 a month for GPT-5 Pro beside an estimated $500,000–$1.5 million treatment bill. Labenz compares the subscription cost with somewhere between under a tenth of a percent and 2% of the total medical cost, calls the surplus “absolutely insane,” says he would pay $10,000 monthly in this situation, and argues that avoided complications, infections or unnecessary relocation can even make AI “GDP destroying.”
His policy conclusion cuts both ways. Delaying an AI oncology researcher costs lives, yet he still supports banning uncontrolled superintelligence; meanwhile, near-term medicine needs better data access, liability reform and a stronger right to try because AI-supported patients may increasingly bring well-reasoned ideas that could beat static standard-of-care rules.
Labenz’s preferred geopolitical race is therefore a “cancer cure Olympics” between the United States and China, extending eventually to healthy lifespans of 150 years. His closing claim is not that every AI proposal is right, but that “standard of care is not the end of history”—and he argues that the medical system’s current structures are becoming handcuffs. The experience also left him feeling there are “no decels in the pediatric oncology unit”: the costs of delaying useful AI progress are unusually vivid when a child’s life is at stake.