王国鑫:请AI破解“看病难、看病贵”
Summary
JD Health has made medical foundation models a “core strategy,” betting on whether AI can cheaply expand supply in an industry constrained by limited capacity and extremely high service costs. Nico described the operating base as roughly 490K healthcare inquiries per day. JD Health generated about RMB35B in revenue and RMB3.5B in profit in the first half, but AI “has not yet fully worked out its business model”; the more important task now is to “answer for the future of medical services and the future of the company.”
The moat of 京医千寻 2.0 is not just medical knowledge, but the integration of synthetic doctor-patient dialogue, medical imaging, and verifiable evidence-based reasoning in one system. It supports CT, MRI, and X-ray, and can anchor lesions in images to explain its conclusions. Training data combines JD Health’s own consultations, internet data, synthetic data, data centers, and long-term cohorts from more than a dozen specialty hospitals. “Medical models often cannot be built entirely on existing data,” making Agent simulation an inevitable path.
As of September 2025, Nico believes AI is already changing the “hard and expensive access to healthcare” problem through three channels: information, services, and R&D. Chatbots first weaken the information mismatch created by paid search rankings, then triage mild, serious, and emergency cases, while extending common-disease support to “24/7”; the longer-term path to lower costs comes from doctors’ continuous learning, new therapies, and drug development, including faster R&D behind the rise in China biotech BD and outbound licensing.
The key difference between medical vertical models and foundation models is that the former must make a rapid judgment from a few high-value questions like a doctor, while specifically optimizing for small lesions, organ positioning, and left-right symmetry. When a general model sees “I have a fever,” it often interrogates the user exhaustively to rule out every possibility. Nico believes this “does not conform to medical ethics or medical practice.” Public benchmarks are not the final judge either, because “the fundamental holy grail in medicine is still diagnostic accuracy and effective treatment plans.”
The competition around AI Hospital 1.0 is not for an isolated chatbot, but for the “first entry point to future health” and the service loop behind it. JD Health can connect internet hospitals, doctors, drug supply chains, 30-minute delivery, home nursing, health-check centers, and physical hospitals. The model assesses the user’s needs, cost, and willingness to proceed, while back-end capabilities are broken down into callable atomic services. Competition may eventually shift from chat experience to “chatbot capability plus back-end service effectiveness.”
The pace of medical AI expansion will ultimately depend on trust engineering, not just model parameters. 京医千寻’s launch goes through three layers of human validation: JD Health’s own general practitioners, partner medical schools, and a quality-control committee of more than 100 specialists. The model, training code, and some data have been opened up so research institutions can reproduce and evaluate the work. Hospitals have also shifted from “you cannot make mistakes” to “I can allow you to make mistakes, but how do we coordinate and control them?”
Nico’s three filters for evaluating a vertical-model company are knowledge and data moats, a market of the right size, and a clear path to commercialization. The opportunity must be achievable in the near term without being so large that it immediately attracts foundation-model giants. The company must also be clear about whether revenue comes from APIs, products, or sales, and have partners capable of executing commercialization. For medical AI, the two most valuable outcomes are “three nines or four nines” substitution within a narrow domain and connecting consumers with physical services.
Deep dive
1. The core medical-AI thesis is expanding supply, not building another Q&A tool
Nico’s operating context is straightforward: JD Health generated about RMB35B in revenue and RMB3.5B in profit in the first half, but AI “has not yet fully worked out its business model.” Its current role is to find answers for the future, not immediately contribute standalone profit.
Healthcare’s structural contradiction is “limited supply and extremely high service costs,” while individual demand for health and longevity is nearly unlimited. AI can truly change the industry’s cost curve only if it can replicate part of the service provided by doctors.
His ultimate vision comes with a clear condition: if low-cost supply can give ordinary people access to expert-level services once available only to a small minority, “then each of us might live another three to five years.” That is the social value medical AI can create.
2. JD Health’s AI roadmap has moved through compliance, service extension, and lifelong companionship
The first phase came from real operating pressure: with roughly 490K healthcare inquiries every day, precise doctor-patient matching is difficult without AI triage, and online services cannot remain compliant without quality control. The task was to “put a business inside a compliance cage” while bringing down costs.
The second phase used digital therapeutics and even explored parts of brain-computer interfaces, extending the data horizon beyond a single consultation in both directions: understanding a patient’s lifestyle before illness and tracking rehabilitation after treatment, rather than reducing healthcare to a single video call or phone call.
The new imagination brought by foundation models comes from “a highly human-like level” and instruction-following ability: first generate low-cost services approaching a doctor’s standard, then gradually become a long-term health companion that keeps understanding a person’s condition, “like a family member.”
3. Medical data is “painful but rewarding”: highly digitized, yet not a complete truth ready for training
The fortunate part is that after years of informatization and intensive evaluation, Chinese hospitals have highly digitized medical records, cloud imaging, and quality-control standards. If the system were still based on paper reports and manual registration, “nobody could build a medical model.”
The difficulty is that existing records serve emergency care and hospital workflows; they are not designed to preserve a patient’s full state. Medical records are often conclusions distilled by doctors, and their quality may not meet the needs of model training.
The more important gap is the thought process. Doctors see symptoms and make judgments internally, but write only the conclusions into the medical record. Humans can learn through cases, symbolic abstraction, and word of mouth; models still depend on large volumes of raw or reasoning data. “When will it learn how to learn?” may be a major part of the next stage of AGI.
4. Fragmented data, unclear ownership, and equipment differences form a hard wall for medical models
A patient may visit multiple hospitals, so data is physically fragmented by default. Whether those records belong to the patient, doctor, or hospital creates ownership and privacy issues when the data is repurposed.
Even when two scans are both CTs, different equipment and instruments can produce different results. The historical difficulty of mutual recognition between tests and examinations was not necessarily hospitals’ attempt to charge more; it may have been a way to control medical risk.
Nico therefore calls healthcare “truly the hardest” vertical, while also seeing it as one of the most meaningful. These constraints apply across the industry and are difficult to eliminate through any single technical advantage.
5. A vertical model must clear both the data-ownership and commercial-appeal tests
To secure budget from management, Nico developed a two-dimensional framework: whether data is cheaply available or can be simulated, and whether the business model is sufficiently clear. A genuine vertical model must also possess unique, scarce data that a general model cannot easily absorb.
Education is one counterexample: language and basic knowledge are relatively explicit and easy to simulate, while AI can remove the psychological pressure of practicing a foreign language with strangers. General models can therefore cover a large share of the demand, and the industry’s knowledge barrier may not be deep enough.
Code sits at the other extreme. The data requires governance, but the commercial value is so clear that general-model companies cannot ignore it. Nico even sees coder models such as Claude as vertical models—“effectively, the company’s entire business model is riding on one vertical model.”
6. AI first changes how people obtain health information—and that is already shaking up search
Nico’s mechanism-level view is that search engines rely on paid rankings, and the business model itself may encourage information mismatch. At least in principle, foundation-model teams aim to generate high-quality, fact-based answers rather than sell a placement in the information-matching process.
The first layer of medical value is health education: improving awareness of the quality of physicals, screenings, and examinations, while reducing delays caused by misinformation. Gastrointestinal endoscopy was merely his example, not a recommendation that everyone undergo the same test unconditionally.
His more aggressive formulation is that Chatbots may, in China, “to some extent sweep the search-engine format into the trash.” Once a product directly provides knowledge and answers, users will ask: “Why do I need to look at information?”
7. Easier access requires triage and assisted care; lower costs require research and new therapies
The most practical task today is separating mild, serious, and emergency cases: mild cases receive standard solutions, while serious and emergency cases are quickly connected to medical resources. If AI continuously collects health data, matching costs and users’ psychological barriers could fall further.
Nico carefully distinguishes the boundary of the effect: this “may not be solving difficult access to healthcare,” but it does reduce the complexity of seeking care. The next step is to have highly trusted assisted diagnosis checked and reviewed by humans, extending common-disease services to “24/7.”
Doctors are lifelong learners, but the volume of medical papers is too large for any doctor to read in full. AI for doctors can raise the overall standard of healthcare only if it continuously filters knowledge and improves physician capability—“if doctors are not good enough, that is fundamental.”
The high cost of care is more closely associated with emergency and serious cases, especially the cost of treating severe disease. Nico sees AI as a core component of drug and new-therapy development. Pointing to China’s innovative-drug market, stronger outbound BD, and increasing licensing activity, he argues that these developments necessarily reflect faster and better R&D.
8. The rebuttal to “users cannot use AI” is an older man identifying an aircraft mid-flight
The host’s objection was that doctors are scarce, but ordinary people also need education to use AI. Nico disagreed that this barrier would permanently block adoption, citing something he observed during a flight delay.
The older man sitting next to him photographed the cabin and asked a Chatbot: “What kind of plane is this? How is this model? Which seat is the most comfortable?” The natural behavior of this nontechnical user made Nico realize that China’s AI penetration “has somewhat exceeded people’s actual expectations.”
Industry data also shows solid penetration among teenagers and people in their 40s and 50s or older. When the host posed the question of receiving RMB100M but being the only person unable to use AI, Nico seriously considered whether he would accept it and still found it extremely difficult. He described that refusal to step back as a “cognitive divide” that cannot be measured by money alone.
9. 京医千寻 2.0 earned its new version number through three upgrades
Version 1.0 relied mainly on papers, textbooks, academic articles, and real cases. Version 2.0 invested heavily in synthetic data and opened a medical doctor-patient dialogue simulation Agent to the industry for free. Users can ask it to play either the patient or the doctor and generate a consultation.
This approach does not abandon real data; it acknowledges the practical constraint that “medical models often cannot be built entirely on existing data.” JD Health’s roughly 490K daily consultations provide the operating base for judging and generating high-quality simulated dialogue.
The second upgrade is native medical multimodality: a single model supports CT, MRI, and X-ray. Nico’s rationale is simple—when a cough lasts more than a week, doctors may recommend screening, and a text-only system is “very far from the real world.”
The third is evidence-based reasoning: each step—A, B, and C—is tied to a specific type of evidence, with evidence graded by sources such as top journals and national guidelines. The model can also anchor a specific lesion in an image and explain why it reached its conclusion.
10. “Reasoning” here is not philosophical thought, but verifiable, compute-intensive format learning
Nico explicitly dislikes the ambiguity of the word “reasoning.” Current models do not associate and think like humans; they are closer to “format learning,” using more compute and tokens to improve answer accuracy.
Healthcare requires the intermediate process to be verifiable as well. That is why 京医千寻 2.0 connects to an evidence base: it does not merely display a string of plausible reasoning steps, but explains what evidence each step comes from and how strong that evidence is.
Multimodal reasoning further grounds textual conclusions in physical image entities. The model can point to a specific area or lesion in the lung and complete its subsequent judgment “based on the condition of that lesion,” linking the explanation to the actual diagnostic object.
11. Model launches still cannot bypass expensive, three-layer human validation
JD Health first uses internally built evaluation sets for R&D comparisons. A model launch generally then goes through three layers of human validation: JD Health’s own general-practitioner team, third-party evaluation by partner medical schools, and a quality-control committee of more than 100 experts from a broad range of departments.
Evaluation covers more than readability. It includes five core metrics such as faithfulness, professional accuracy, fluency, and consistency. Nico acknowledged: “This is very expensive.”
When OpenAI released HealthBench, Nico explained to the CEO: “It means OpenAI also has to use people to validate medical models.” The project likewise involved more than 60 doctors, including doctors from China. That shows high-risk models still cannot fully automate validation.
The solution for massive volumes of synthetic data is an iterative funnel: machines simulate lower-level doctors as much as possible, while simpler samples stay at the upper layers and severe or high-probability-error cases are passed to experts. Not every item is reviewed by a human; the objective is to improve the probability of correctness under the premise that “a foundation model is Bayesian—it is probabilistic.”
12. The advantage of a medical vertical model is not answering more, but asking fewer, higher-value questions like a doctor
The host used his own habit of asking ChatGPT first and questioned what an 80-person medical-model team had actually created. Nico’s answer was “quasi-expert capability”: doctors need to form a rapid judgment through a short dialogue and the core questions for a specific disease.
A general model learns every possible explanation for a symptom from textbooks and may interrogate the user exhaustively to rule out risk. A specialized model must focus on the key questions and make a rapid judgment from the full information, rather than list “three, four, or five questions.”
“I have a fever today” is the suggested comparison case. Nico believes that a general model’s exhaustive questioning of such a common symptom is not rational from a health-economics perspective. Ideally, a human doctor should review the questioning paths of both types of model.
The multimodal gap is even more direct: medical models specifically optimize for positioning, organ location, left-right symmetry, and small-lesion recognition. Foundation models lack comparable data and commercial incentives, so their efficiency falls sharply on specialist imaging.
13. Benchmarks drive R&D, but cannot replace expert judgment or individualized product experience
Nico calls the key post-consumer turning point an “experience inflection point.” Benchmarks tell technical teams what performance level to reach, but the actual experience is jointly determined by the model, product design, and interaction design. Leaderboards therefore have “no 100% correlation” with real-world use.
JD Health places greater weight on expert-style service and two outcomes that cannot be bypassed: diagnostic accuracy and effective treatment plans. Fluency, empathy, and “saying nice things” can be trained, but cannot override professional metrics.
According to a JDD product manager on site, the JD Health app will build patient profiles and incorporate past disease and chronic-condition histories into context. The host added that a health vertical product could also build family profiles, so the same question would produce different answers depending on the individual’s and family’s health situations.
Nico’s attitude toward public benchmarks is: “You can run them, or you can not run them.” When external scores do not match in-house expert evaluations, he still trusts real doctors more, arguing that different expert groups may define different evaluation dimensions. The host added that leaderboard dimensions are fixed.
14. Empathy can be trained orthogonally, but building a psychologist is harder than building an internist
The host cited the discussion around Baichuan’s goal of building a doctor, emphasizing that doctors must also reassure patients and families and help them rationally accept a treatment plan. Nico agreed that communication is “extremely important,” but clearly separates internal evaluation into professional and experience tracks. The two can be trained independently and combined later.
His automotive analogy is that communication experience resembles the infotainment system, while professional diagnosis and treatment resemble autonomous driving. Both matter, but their stability and logic requirements differ. “Professional issues cannot be compromised”; resource allocation must first protect the safety baseline.
Nico instead believes that “the difficulty of making a psychologist is far higher than that of making an internist.” A voice is easy to imitate, but something that truly feels like a person may reveal its flaws after two or three minutes. Empathy also lacks reliable metrics, and “once a metric can be measured, it can be optimized.”
JD Health previously worked with a leading domestic mental-health hospital on digital humans for anxiety and depression, developing both the front-end persona and the back-end model. Clinical trials are not complete. Performance is currently “quite positive,” but more resources remain focused on common and serious diseases.
15. The data system is layered by common disease, difficult disease, and compliance boundaries
The first layer involves data-center partnerships. Medical data undergoes strong de-identification and anonymization externally before entering the R&D process, reducing the risk of internal handling of private data. The team has also signed with a national-level data center, focusing on large multimodal models.
The second layer uses internet data, JD Health’s own data, and synthetic data as a baseline, then adds private data from large data centers to cover more provincial-level data units. Nico’s assumption is that this can cover the overwhelming majority of common diseases.
The third layer consists of long-term cohorts from more than a dozen top specialty hospitals, focused on difficult and serious cases. Partnerships typically involve research support, joint development, and third-party de-identification. The core assumption is that existing foundation models need only a small amount of specialty data to learn the corresponding doctors’ capabilities.
16. Hospitals have moved from zero tolerance toward defining the boundaries of human-machine collaboration
After visiting hospitals for three consecutive years, Nico has seen support expand from some academicians to hospital presidents and department heads. National medical centers have responsibilities for building disciplines, as well as incentives to extend expert experience and patient services. AI is becoming a tool for codifying that capability.
Nico was once surprised by the depth with which a young department director evaluated different models. His judgment is that the next generation of outstanding doctors will use AI extensively to improve productivity, but there remains enormous divergence within the medical profession.
From the end of last year through 2025, the industry’s discussion shifted from “you cannot make mistakes” to: “I can allow you to make mistakes, but how do we coordinate and control them?” The industry has begun looking for use cases together, while pushing the pressure back onto models’ ability to generalize.
Demand is concentrated in three areas: long-term patient services before and after consultations, research and personnel training, and higher efficiency when simply adding more people is no longer possible. A “doctor clone” merely combines long-term service and productivity tools in digital-human form.
17. AI Hospital 1.0 will aggregate multiple specialist Agents into the first entry point for health
JD Health has already developed different Agents for psychology, internal medicine, pharmacy, and nutrition. The simple logic of AI Hospital 1.0 is not to leave them scattered across different places, but to create a unified entry point that users will seek out whenever they feel unwell.
“AI Hospital” is both a product name and a statement of JD Health’s ambition to own the consumer mindshare around the future health entry point. On one end, it supports first-tier cities as they move from disease treatment toward health management; on the other, it gives regions with larger healthcare gaps access to screening, triage, and resource connections.
Nico sees AI as more of a B2B productivity tool. Large hospitals can continue radiating services into local areas, community doctors can use AI for general screening, consultations, and referrals, and major hospitals can extend services into rehabilitation. Population aging and regional disparities may drive the formation of this network.
Execution is not about telling a new story from scratch, but layering AI onto the existing internet-hospital capabilities of cross-region matching and “24/7” service. “AI-powered internet healthcare” is simply a further upgrade of the existing internet-healthcare model.
18. Open source in healthcare is not marketing; it is infrastructure for building collaborative trust
Nico believes healthcare is “trust-driven,” and not merely collaboration-driven. Allowing hospitals and research institutions to see, try, and test the model is essential to building a technology brand and ecosystem relationships.
京医千寻 has opened more than model outputs. It has also opened training code and portions of the training data, with the goal of allowing participants to reproduce the work rather than merely call a black-box result.
Actual feedback has come mainly from universities, hospitals, and research institutions, while partners have proactively evaluated smaller models as well. Nico believes continued specialty collaboration will be possible, with the “priceless” trust created by open source playing an important role.
19. Commercialization must deliver both reliable substitution and service connectivity
When evaluating AI’s commercial value, Nico first asks two questions: can it achieve “100% or three nines or four nines” substitution in a narrow domain, and can it become a new bridge connecting consumers with services? Healthcare offers both types of opportunity.
JD Health’s differentiated resources are not pure traffic, but physical services: internet hospitals, at-home services, drug supply chains, 30-minute delivery, health-check centers, and its own hospitals. AI must organize these capabilities into solutions for individuals.
Back-end services are broken down into atomic capabilities, such as a nurse making a home visit to perform an examination. The model first assesses medical need, cost, and the patient’s willingness to proceed, then calls the supply chain. The design seeks to prevent commercial services from contaminating medical information at the front end.
Large general-purpose Chatbots already have substantial health-inquiry traffic, but Nico currently prefers to view them as partners. The Joy AI App launched within the broader group may also create synergies. Ultimately, competition will be the sum of chatbot experience, professional capability, and effective back-end services.
20. Overseas cases cannot be copied directly; the payer determines what medical AI becomes
Nico cautions that medical AI has weak cross-border portability. The key variables are the payer and the healthcare system. China places greater emphasis on efficiency and equity, while US doctors earn more and have greater willingness to pay. The same product therefore produces completely different business models.
Products such as OpenEvidence can commercialize in the US through doctor subscriptions. JD Health, by contrast, opened its evidence base to the industry for free. Nico summarizes the difference as “an orange grown south of the Huai River tastes different from one grown north of it.”
Hims shows another path: an AI-enabled internet product acquires users, assesses their condition, and continuously provides recommendations at the front end, while a specialized drug supply chain at the back end delivers solutions to help users look and feel better.
China’s market also includes doctor tools, patient services, hospital AI informatization, and To G services for medical-insurance and health-administration bodies. Hospitals have long procurement cycles and high existing informatization penetration, so Nico believes the opportunities most compatible with local conditions may still lie in patient services.
21. The two closing checklists both emphasize investing early rather than repairing problems later
His advice to individuals is to use age 35 as an example: set aside an annual budget for personal and family health, then research how to buy the best examinations and medical services within that budget. Nico later stressed that the principle does not begin at 35, and the amount need not be large. The key is “consciously using economic means to adjust your own cognition.”
This does not mean everyone should undergo gastrointestinal endoscopy. Tests should be selected based on family history, high-incidence regions, and personal condition. Regardless of age, long-term monitoring of blood pressure and blood glucose can be valuable. Even with the same diabetes diagnosis, early detection and control produce very different outcomes from late detection and control.
Medical experience can transfer to education, law, and finance because all involve modeling complex situations: individualized learning paths, multistep reasoning across multiple legal information flows, and financial portfolio recommendations are fundamentally high-specialization data and integrated-reasoning problems of “the same origin in different forms.”
The checklist for vertical-model investors has three parts: whether the company truly has knowledge depth and a data moat; whether the market can be realized in the near term without being so large that it immediately attracts giants; and whether the payment path—API, product, or sales-driven—is clear, with partners on the team capable of completing commercialization.