Pioneers Insight Method Research Author
E227 | The AI Battle for the U.S. Healthcare Market: Big Bets by Giants—Can Startups Win?
Back to Episodes

E227 | The AI Battle for the U.S. Healthcare Market: Big Bets by Giants—Can Startups Win?

Summary

  • The inflection point for healthcare AI in 2026 is that large pharmaceutical and healthcare companies have moved from debating “whether to integrate” to accepting that they “must integrate.” Eli Lilly and Nvidia announced a billion-dollar-scale strategic partnership, while some healthcare companies have even established “artificial intelligence universities” requiring employees, management and even boards to undergo training. This means AI demand is moving into organizational processes and infrastructure.

  • The first value healthcare AI delivers in the U.S. may not be replacing diagnosis, but repairing the resource mismatch between doctors and insurers and documentation systems. MGH generalists work an average of 61.8 hours a week, yet typically see only 15 to 25 patients a day for about 15 minutes each; only around 10% of denied insurance claims are appealed, but roughly 80% of appeals are overturned, showing that much of the waste comes from process rather than medical judgment. “Why have such expensive people doing low-value repetitive work?”

  • medical coding, medical billing and prior authorization combine huge markets, clear rules and low tolerance for error, making them the most realistic entry points for AI. Coding is essentially “translating” diagnoses and procedures into standardized codes such as I10 and CPT codes; errors can lead to denials, delayed reimbursement and even upcoding fraud risk. The work is highly structured, and many steps do not even require generative AI; Gen AI’s incremental value lies mainly in checking supporting documents and predicting denial risk.

  • HIPAA, data control and deployment models are not ancillary compliance items, but the entry barriers and competitive moats for healthcare AI. Claude for Healthcare is focused on infrastructure including billing, coding, HIPAA-compliant cloud deployment and connectivity APIs; federated computing allows more than 60 healthcare systems to collaborate without physically moving data. The priority order is “compliance, security and no hallucinations” ahead of general-purpose model capability, leaving room for vertical small models and on-premise deployment.

  • OpenAI is competing simultaneously for the consumer gateway and the hospital operating system, but control over hospital workflows remains undecided. ChatGPT Health addresses more than 200 million healthcare-related questions asked by users each week, with data isolation, no training use and explicit disclaimers; ChatGPT for Healthcare aims to connect to EHRs, summarize medical records, draft prior authorization materials and let hospitals build agents. The issue is that Microsoft products such as Office and Outlook are already embedded in hospital workflows, so “adding a feature” may be easier than OpenAI rebuilding system integrations from scratch.

  • OpenEvidence proves that a narrow use case can quickly create product advantages, while also exposing the risks of a high valuation, advertising conflicts and low switching costs. It constrains RAG with top journals, authoritative guidelines and traceable citations, and reportedly is used daily by 40% of U.S. doctors; roughly $100M in annual revenue corresponds to a $12B valuation, while doctors use it for free and revenue depends on pharmaceutical advertising and content promotion. 张璐’s judgment is blunt: “I don’t think it is a company with core AI capabilities”; its true moat is content licensing and physician penetration.

  • AI can become a tool for healthcare stratification and resource balancing, but the episode insists on “human in the loop,” with doctors remaining the decision-makers. For common colds, test results, emergency guidance and care navigation, AI may provide a “reliable first step,” with rural and underserved areas benefiting most; complex diseases still involve individual variation, unknown biomarkers, liability and treatment trade-offs. O3 scored 60% on HealthBench, with a top score of 32% in the hardest mode, which is closer to real clinical encounters than a high multiple-choice score and makes the capability boundary clearer.

  • Healthcare AI is unlikely to be winner-take-all: giants are better positioned for foundational platforms, while startups can still win through vertical data, customization and trusted relationships. Healthcare data is fragmented across hospitals, pharmaceutical companies and medical-device makers, and customers may not want to tie all their risk to a single technology giant; at the same time, To C opportunities are expanding from treating illness after it appears to continuous monitoring, preventive health and quality of life. Demand is real and durable, but the key variable is still who pays, while free substitutes will continue to pressure consumer subscriptions.

Deep dive

1. The Healthcare Industry Has Moved from Testing AI to Explicit Integration

  • 张璐 observed at the January 2026 J.P. Morgan Healthcare Conference that the mindset of pharmaceutical and healthcare companies had fundamentally changed: the previous year, they were still debating whether to combine with AI; this year, the answer was that they must integrate it, with substantially greater intensity, efficiency and speed.

  • The clearest industry signal was Eli Lilly and Nvidia announcing a deep strategic partnership, with an initial budget of $1B. Several Fusion Fund startups are also working with both companies; 张璐 said that once these companies launched new projects, the overall integration moved “very quickly and very dramatically.”

  • The deeper shift is taking place inside organizations: some large healthcare companies have already established institutions resembling “artificial intelligence universities,” requiring ordinary employees, management and even boards to take courses on using AI to process healthcare data, extract value and accelerate new projects and external partnerships.

  • 周叶斌 added a practical constraint for pharmaceutical companies: confidential materials generally cannot be uploaded to external ChatGPT, so companies are more likely to use tools available internally, such as Microsoft Copilot. Joining a platform operated by a large pharmaceutical company such as Eli Lilly may deepen collaboration, but it can also make smaller companies worry: “Am I handing over my information to you?”

2. Doctors Working 61.8 Hours a Week Expose a Massive Resource Mismatch

  • 周叶斌 cited a study tracking MGH generalists: they work an average of 61.8 hours a week, which could mean having only one day off and working more than 10 hours a day on the other six; yet doctors typically see only 15 to 25 patients per day, with visits of around 15 minutes, and encounters lasting more than 25 minutes are already quite rare.

  • Spread across total working hours, each patient represents roughly 1.7 hours of physician labor, most of it outside the consultation itself: reviewing records, documenting care, dealing with insurers and handling other administrative processes. The time doctors actually spend seeing patients may account for only one-third of their total work.

  • The result is not merely an efficiency problem, but also physician burnout. 周叶斌’s explanation is that many people do not attend medical school to process administrative paperwork or repeatedly call insurance companies; when workflow drains the satisfaction from the profession, the consequences become more serious.

  • 张璐 calls the EHR an unintuitive outcome: it was intended to automate and digitize care, yet turned high-cost doctors into “record clerks” who continuously enter and review information. “Why have such expensive people doing low-value repetitive work?”

3. The Core Friction in Insurance Denials Is Often Process, Not Medicine

  • Fragmentation in the U.S. healthcare system compounds the documentation burden: a new patient’s electronic records may come from different hospitals and systems, and even when the same software is used, the records may be structured differently; prior authorization then requires doctors to obtain insurer approval before treatment or the cost may not be reimbursed.

  • The denial data 张璐 cited is the clearest illustration: only around 10% of denied claims are appealed, but roughly 80% of appealed claims are overturned. That means large amounts of money that should have been paid are blocked from the payment system by documentation and process errors.

  • Hospitals have already delivered the care, but must wait for coding and appeals to be completed before collecting payment, lengthening the reimbursement cycle; the pressure then flows back to doctors, forcing them to minimize documentation errors. 张璐 believes this is a major source of the intense resentment Americans have accumulated toward healthcare companies.

  • The experience of one emergency physician turned entrepreneur makes the mismatch concrete: he “might be saving someone’s life one minute,” only to find himself standing beside a fax machine at 3 or 4 a.m. after the resuscitation, processing medical coding. That experience ultimately pushed him to leave medicine and start a company.

4. medical coding Is One of the Best Healthcare Tasks for Early Automation

  • 周叶斌 and 张璐 clarified that medical coding is not merely a matter of recording responsibility. It “translates” diagnoses and treatment actions into standardized codes: hypertension may correspond to I10, while procedures such as ECGs, MRIs and CT scans map to different CPT codes; insurers recognize the codes, not just the doctor’s natural-language description.

  • Coding errors can cause denials and delayed payments; coding too aggressively can be deemed upcoding, leading to fines or fraud risk. The task is “actually not complicated at all, but has an extremely low tolerance for error,” and directly affects hospital revenue and operational assessments.

  • For AI, these tasks have clear rules, a structured format, high repetition and standardized answers. Many coding companies do not even need generative AI; rule-based systems can extract diagnoses from records and match them to codes. Gen AI can then check whether supporting documents are complete or estimate denial risk.

  • 张璐 said medical billing and medical coding are themselves multibillion-dollar markets. There is no need to create new medical facts here; the commercial value comes from reducing errors, cutting administrative costs and accelerating reimbursement, not from demonstrating open-ended reasoning.

5. Claude for Healthcare Is Betting on the Compliance Layer, Not the Clinical Front Door

  • 张璐 believes the smart move behind Claude for Healthcare is its choice of a “less sexy” infrastructure route: covering medical billing, medical coding, HIPAA compliance, connectivity APIs and cloud deployment rather than going directly to patients and taking on the risks of complex diagnosis.

  • She said it appears to include basic functionality for diagnostic codes such as ICD-10 and emphasized connections to official databases. As an enterprise infrastructure solution, hospitals retain direct control of their data, while Claude for Healthcare provides the connective layer between existing medical, insurance and compliance systems.

  • The automation logic itself is relatively simple; the real time sink is implementation: where the data is hosted, how systems connect to hospitals and insurers, and how the product passes security reviews can all extend the deployment cycle. The infrastructure direction is therefore critical, but difficult to launch through rapid self-service onboarding like ordinary SaaS.

6. HIPAA Determines Which Healthcare AI Can Actually Enter Hospitals

  • 周叶斌 summarized HIPAA as a hard threshold for medical privacy: who doctors may disclose a patient’s name and medical records to, and how pharmaceutical employees handle health information, are subject to strict rules. Uploading a complete medical record directly to a noncompliant public model is “obviously not acceptable,” and the consequences of a violation can be severe.

  • 张璐 added that HIPAA violations can trigger fines in the millions of dollars and class-action lawsuits. Healthcare AI startups therefore cannot build only a model from day one; they must also establish security teams, compliance architecture, audit capabilities and legal expertise, while undergoing lengthy hospital security reviews.

  • OpenAI says more than 200 million users ask ChatGPT about healthcare topics every week, creating a major impetus for ChatGPT Health. Its public commitments include isolating health conversations from ordinary GPT data and not using them for model training; but the consumer product still explicitly tells users to consult doctors and does not assume diagnostic responsibility.

  • Anthropic has committed to HIPAA-compliant deployment on AWS, Google Cloud and Azure, while emphasizing that enterprises retain direct control of their data. The episode repeatedly notes that even with cloud compliance, hospitals and patients may still want core data to remain on-premise, leaving room for small language models and localized deployment.

7. OpenAI Is Using To C to Capture the Gateway and To B to Pursue High-Value Workflows

  • 张璐 framed OpenAI’s entry into healthcare as driven by two forces: consumer monetization has a ceiling, while healthcare accounts for more than 20% of U.S. GDP and represents a vast To B market; meanwhile, around 30% of the data generated by human society is related to healthcare, but less than 5% is actually used.

  • Healthcare data has already become highly digitized through EHRs and administrative processes, with doctors unknowingly completing data integration, generation and labeling. 张璐 believes that as continued increases in consumer-side data produce diminishing gains in model capabilities, high-quality data inside hospitals becomes strategically more valuable for improving models.

  • ChatGPT Health is designed for ordinary users and can connect to Apple Watch health records, medical records and exercise data to help explain everyday health issues; its disclaimer draws the line on responsibility. It can provide information and suggestions, but cannot be treated as an independent diagnostic authority.

  • ChatGPT for Healthcare covers both administrative and clinical knowledge: drafting prior authorization materials, summarizing the latest research and publicly available treatment protocols, and helping doctors design potential care pathways. 周叶斌 believes the value is clear whether the goal is lowering administrative costs or making knowledge updates more evenly available.

8. OpenAI’s Hospital Operating-System Ambition Will Collide Directly with Microsoft

  • OpenAI is not satisfied with being a standalone app. It wants to connect directly to EHR workflows, summarize medical records, fill in missing information and provide explanations, ultimately becoming a hospital-level AI operating system. Its further vision is to let hospitals build agents for medical coding, assisted communication and other tasks on the platform.

  • There are currently only six initial hospital partners, including Stanford Children’s Health; 张璐 has not yet seen enough feedback to judge integration speed or hospital adoption. She believes the ecosystem strategy may offer the greatest value, but says, “I don’t know how receptive hospitals will be to this ecosystem strategy.”

  • 泓君 noted that as model capabilities converge, existing distribution channels may matter more than a pure model lead. 张璐 then raised the Microsoft variable: hospitals already use systems such as Office and Outlook extensively, so automatic medical-record summarization may be merely “adding a feature” inside Microsoft, while OpenAI would have to build an entirely new system integration.

  • The capability gap between open- and closed-source models is also relatively small, while open-source solutions are cheaper and easier to deploy internally. Highly regulated customers may deploy sensitive applications themselves, and startups can iterate at lower cost; but healthcare has almost no tolerance for hallucinations, so human in the loop remains essential.

9. OpenEvidence Won Doctors with Constrained Evidence Retrieval

  • OpenEvidence reached a $12B valuation in January 2026 and reportedly is used daily by 40% of U.S. doctors. It does not try to cover all of healthcare; instead, it lets doctors quickly search top journals and the latest national treatment guidelines for clinical questions, receiving answers with citations.

  • 周叶斌 explained the value of the narrow use case: when ordinary ChatGPT produces a medical literature review, it may fabricate papers or mischaracterize research; if a doctor relies on that to treat a diabetic patient, the consequences are far more serious than an ordinary conversational hallucination. OpenEvidence prioritizes data and answer quality by restricting its sources.

  • 泓君 described it as “a more intelligent search library,” and 张璐 broadly agreed, calling it a highly optimized RAG architecture. The model is constrained to search, filter and organize answers within licensed content, must provide sources, and is not encouraged to improvise or engage in creative generative reasoning.

  • The product is also optimized for doctors’ professional conversational context. The target user is singular and the task is singular: “When you need the latest knowledge, I give you the latest and most authoritative knowledge.” That restraint improves stability, and usage is especially high among younger doctors.

10. OpenEvidence’s Content Moat and Advertising Model Have an Inherent Conflict

  • 张璐’s core judgment is that OpenEvidence “does not have core AI capabilities”; its competitive strength comes mainly from licensing high-quality medical content and penetrating the physician population at scale in a short period. In other words, the moat lies in data rights and distribution, not the underlying model.

  • The company generates roughly $100M in annual revenue, while doctors use it for free. It relies primarily on medical advertising and content promotion, and may launch an enterprise version in the future. Its budget is effectively drawn from spending that pharmaceutical companies previously allocated to drug representatives and physician promotion, and it may influence younger doctors online.

  • The business model also raises questions about objectivity: when a pharmaceutical company wants to promote a drug, will doctors see related research first when they ask a question? If advertising affects the ranking of evidence, the product’s most important positioning—authoritative and neutral—comes into conflict with its revenue source.

  • A $12B valuation is roughly 120 times current revenue, leaving 张璐 cautious. Doctors face very low switching costs; if Gemini, Claude or ChatGPT Health offers equally convenient medical search for free and without ads, it remains unclear whether OpenEvidence can sustain its current advantage.

11. Vertical Small Models Will Remain Necessary Even as General Models Improve

  • 周叶斌 warned that AI is also contaminating medical data: paper mills can use AI to generate low-value research. If a general-purpose model draws equally on top journals and wellness influencer content, it may confuse the hierarchy of evidence; the New England Journal of Medicine and popular internet opinion cannot be treated as equivalent.

  • 张璐 believes healthcare may look like one industry, but internally contains a huge number of diseases, data types, devices and use cases, with substantial individual variation. A vertical model can optimize a specific task with a limited but high-quality dataset, matching or even outperforming a much larger general-purpose model within a single domain.

  • The enormous number of medical devices and smart-hospital edge devices also creates demand for on-device AI: large models are constrained by compute, power consumption and model size, making it difficult to run them everywhere; small models are better suited to edge AI and local deployment, while reducing the need to upload sensitive data to the cloud.

12. Complex Medical Decisions Still Require Doctors to Bear Primary Responsibility

  • 张璐 used autonomous driving as an analogy: healthcare AI must not only answer accurately, but also address who is responsible when something goes wrong and who has final interpretive authority. In the end, a specific person—usually a doctor—must bear responsibility and make the judgment.

  • The human body is not a fully digitized, closed system. Whether a given severe illness calls for conservative or aggressive treatment, and how different individuals will respond to a disease or therapy, remain highly variable; many biomarkers are still unclear. AI can plan based on existing data, but cannot treat unknown variables as though they do not exist.

  • 张璐 cited a company using smart capsule-like miniature robots to collect microbiome data. The data may be associated not only with gastrointestinal diseases and cancer, but also become a biomarker for brain diseases such as Parkinson’s. Because key data is still incomplete, there remains a significant gap in fully personalized judgment.

  • She does not reject the direction: “Personalization, digitization, and the integration of healthcare and AI are inevitable.” The condition is that human in the loop must always remain, with doctors using the tools, incorporating the patient’s specific circumstances and making the final judgment.

13. The Most Realistic Role for Consumer AI Is First-Step Assessment and Care Stratification

  • More than 200 million users ask ChatGPT about healthcare every week, showing that the demand already exists. Users upload blood-test results and describe colds, insomnia or sudden discomfort; 周叶斌 uses it himself, but emphasizes that the key question is not merely whether to use it, but whether to keep asking how the conclusion was reached and what supports it.

  • AI’s average capability may exceed the average level available in areas with limited medical resources, while also easing the constraint of scarce time among top doctors. 张璐’s qualification is that users should not blindly believe AI can solve every problem; it is better suited to rapid feedback, data interpretation and helping determine whether further medical attention is needed.

  • 泓君 placed this capability in the crowded setting of China’s top-tier hospitals: simple questions also pour into major hospitals, and some people who genuinely need care abandon appointments or checkups because of the wait. Having AI handle basic questions first could reduce waste and give patients a clearer next step.

  • U.S. emergency departments face the same need for stratification: patients often wait for long periods while staff determine who requires immediate attention; uninsured patients may delay care until they become seriously ill, while limited reimbursement makes emergency departments heavily loss-making units. 周叶斌 cautiously said that adding AI “might” improve triage, service and operating efficiency, especially in remote areas.

14. HealthBench Moves from “Passing Tests” to “Seeing Patients,” but 60% Is Still No License to Practice

  • MedQA was traditionally used to test textbook clinical knowledge and PubMedQA to test understanding of scientific literature, but both remained close to multiple-choice exams. Even if a model scores higher than many doctors on the U.S. medical licensing exam, 泓君 noted that ChatGPT may have reached above 90% very early, which does not show that it can identify risks, ask follow-up questions and reassure patients in a real conversation.

  • HealthBench instead uses real medical dialogues, with scoring criteria built by 262 doctors from 60 countries and 26 specialties covering 49 languages. Answers must not only be medically correct, but also cover key action points, use appropriate language and provide safe advice in scenarios such as “a neighbor suddenly collapses.”

  • The result cited on the program was a 60% score for O3, falling to a maximum of 32% in the extremely difficult mode. Common scenarios such as first aid and hypertension performed relatively well, while complex diseases were harder to address comprehensively; dangerous errors also incur penalties, potentially taking the score to zero or even below zero.

  • 泓君’s summary is worth preserving: previously, AI was asked to answer multiple-choice questions; now it is effectively being “sent into the clinic.” That makes 60% more meaningful than a traditional 90% multiple-choice score, but it is first and foremost a capability map, not proof that AI can bypass doctors and practice independently.

15. Giants Will Not Own the Whole Market; the Long-Term Upside Remains in Continuous Health Management

  • 张璐 believes healthcare AI is “unlikely” to be covered entirely by giants: the use cases are extraordinarily diverse, while core data is distributed across hospitals, pharmaceutical companies and healthcare companies. These institutions may not want to hand over all their data to large technology companies, and may instead prefer startups with greater control and more vertical products.

  • 泓君 offered the counterargument: SaaS stocks plunged after Claude plugins were released, while giants have strong reputations, extensive compliance certifications and more capable models—shouldn’t they be more trustworthy? 张璐’s response is that healthcare compares compliance, security and freedom from hallucinations first, and general-purpose capability second; a vertical model can absolutely tie a general model on a single task.

  • Giants are more likely to hold advantages in EHRs, cloud and standardized infrastructure, but healthcare institutions’ cloud services, data architectures and workflows remain highly fragmented. Startups can take on the customized deployments that giants may find too expensive, while hospitals may spread risk across multiple vendors rather than bind themselves to a single platform.

  • The long-term To C opportunity extends from health into wellness, preventive health and quality of life: watches, rings and other devices turn one-off diagnosis into continuous monitoring, allowing disease progression, immune changes and daily condition to be observed over time. The goal is “not merely to live longer,” but to retain the quality of one’s body and brain at 80 or 90, or even 100.

  • 周叶斌 used 23andMe to illustrate the difference in business models: DNA is measured once and does not change, so repeat purchases are naturally limited; physical condition and health concerns continue to change, creating a theoretically broader market. The fundamental challenge remains payment—if OpenAI charges, users may immediately ask, “Can’t I just find a free alternative?”