Vol. 230 Industry Watch 43 | How Is AI for Science Being Deployed in Pharmaceutical Chemical Synthesis?
Vol. 230 Industry Watch 43 | How Is AI for Science Being Deployed in Pharmaceutical Chemical Synthesis?
Summary
- Xia Ning, founder of Zhihua Technology, puts AI chemical synthesis at the level of a chemist with 10 years of experience. AI can generally find any route a chemist can think of, and “AI can come up with an average of 5 to 6 different strategies, while people usually come up with only 1 or 2.” In live tests using real client cases, “a route that took them 1 to 2 months to design, we could run in 5 minutes”—an inflection point comparable to XtalPi’s 2016 crystal-form testing for Pfizer, which has now been reached or even surpassed.
- The hardest investment case is that synthesis is the true bottleneck in the DMTA cycle. Only by making synthesis 10 to 20 times faster can the throughput of new-drug and new-material R&D rise 10 to 20 times. The design side can produce hundreds, thousands or even tens of thousands of molecules a day, and testing has high-throughput capability; the constraint is that “one chemist can synthesize roughly 3 to 5 molecules a month, or 2 to 3 reactions a day.”
- The hard rule for serious scientific deployment is that a black box cannot be tuned, and without a white box, commercial adoption is difficult. Conclusions from black-box models must be explained, verified and constrained inside a white-box system to eliminate hallucinations. Xia says white-box systems are actually becoming more common in the foundation-model era. The biggest gains in explainability have come from years of customer iteration, not model upgrades alone: the request list has grown to more than 1,000 items, while the product has expanded from 4 modules to 9.
- GPT fills the “generalization gap,” pushing the commercial model from SaaS toward agents, with the two formats likely to coexist for the long term. Generalization plus specialized expertise covers “almost the entirety of a chemist’s work.” Agents may develop memory and continuous learning, ultimately working like employees; usage-based pricing and SaaS will coexist. Token costs are a key constraint, so Zhihua uses intelligent orchestration and dozens-of-billions-parameter local models to lower client costs and reserve the most expensive resources for core problems.
- Foundation-model companies struggle to move into verticals: Xia says internal tests show that a vertical model can quickly reach 50 or 60 points in one field, but moving higher is extremely difficult and it generally cannot reach a usable level. AI for Science has limited data, and its data “looks more like multimodal data”; some fields have only hundreds or thousands of data points. The likely outcome is cooperation and mutual dependence between vertical players and foundation-model companies. The moat is “long-term dirty, hard work—continually laying a deeper foundation.”
- The macro signal is powerful: WuXi AppTec’s strong first-half report coincides with $110B of China’s license-out deals in the first half of this year. China accounted for just over 30% of global licensing value in 2024 and half in 2025. Li Xiang estimates that China “might approach 70% this year—an almost unimaginable figure.” License-outs are forcing pipeline companies to produce assets that are “more numerous, faster and better,” but Xia notes that Zhihua added more than 30 Indian clients this year, and “the growth rate actually isn’t slower than China’s.”
- Xia’s brutal judgment for latecomers is that “the window to start a company in this area may already have passed.” The 8 years required to build the technology stack cannot be compressed, while major clients are already using and deeply customizing existing systems, making replacement “almost impossible.” For the industry as a whole, AI for Science penetration is “not even 0.1%,” and the trend is close to inevitable: Xia compares it with the industry-wide shift triggered by China’s electric-vehicle strategy, while Li compares it with the microscope’s role in biomedicine.
- The lesson for scientist-founders is to stick with what they believe is right, while investors commonly make two mistakes. “Stick to what you believe is right, and don’t let capital’s views completely dictate your direction—capital often doesn’t really understand the industry’s actual conditions.” Investors assume founders must make drugs or materials themselves for the market to be large enough; Xia argues that “this is a result, not a cause,” and says they also underestimate the opportunity created when agents complete entire workflows and deliver outcomes rather than tools.
Deep dive
1. Policy and industry converge as AI for Science becomes one of the hottest venture-market directions
- Li Xiang, partner at Frees Capital, set the tone at the outset: the communiqué from the Central Politburo’s economic work conference in late July specifically highlighted the development of basic research. Combined with “AI+,” the push to apply AI to research—including life sciences—to improve efficiency, translation, prediction and screening has made AI for Science one of the hottest areas in the venture market.
- Guest Xia Ning, founder of Zhihua Technology, holds a PhD in organic chemistry, began programming as a child and won national-level prizes in chemistry competitions. His original motivation for entering the field was: “Could I find a job where computers solve chemistry problems?” Scientific research was being done by hand, which limited throughput; “computers have been replicable from day one.” He has been building companies for nearly 18 years, dating back to 2008.
2. Where the two-discipline capability came from: “Computer science is a technology; chemistry is a profession”
- Li’s industry observation is that the hardest people to find in this sector are those with “a foot in both boats”—people whose capabilities span computer science and chemistry, two fields that do not naturally overlap. Which boat matters more is a separate question.
- Xia chose his field after an admissions teacher told him: “Computer science is a technology; chemistry is a profession. You should first learn a profession—you already have the technology.” “It later proved to be right. Everyone who chose computer science and later tried to cross over into chemistry found it extremely difficult.”
- He maintained both skill sets through what Li jokingly called a “pathological level of interest”: doing experiments during the day and writing code at night. During his PhD, Xia also tried to build programs that simulated artificial intelligence for applications such as games and chess, modeling human judgment. “It’s especially unfortunate that I didn’t build AlphaGo.”
3. Seven years in France starting from zero: no capital, no data, no infrastructure
- In 2008, Xia joined a Strasbourg startup whose website had only just been built and whose company had not yet been formally established, working on chemical prediction with data. “From 2008 to 2015, there was no capital behind us—it was all personal ideals and interest.” He developed whatever was needed and “explored the entire field.”
- All 3 routes to building a data set from scratch ran into obstacles: manually commissioning hundreds or thousands of experiments still produced too little data; asking people to structure the data in their PhD theses was too expensive and inefficient; and building electronic lab notebooks for universities in exchange for data sharing ran into the fact that “professors generally don’t have much money,” making commercialization dependent on financing.
- The company was eventually acquired by its parent, which is now France’s largest CRO; much of the underlying architecture remains in use. Xia returned to China in 2015 and joined Wanghua during the chemical B2B e-commerce boom, working on data architecture and search logic. He was also “among the earliest in the world to register the Chemical.AI domain,” and began commercializing AI retrosynthesis. The technology’s roots can actually be traced to 2012, when he was still in France.
4. The WuXi AppTec test: running molecules live in front of a room full of executives and apparently finishing first
- WuXi AppTec, whose core business is chemical-synthesis CRO services, solicited software for global synthetic-route prediction. Xia entered as an individual before the company had even been formed: “There was a whole room full of WuXi executives… They put various molecules through the system, and the results were extremely good.” By his recollection, the product “seems to have” finished first and secured the earliest commercial partnership.
- Xia took 2 signals from the event: the industry took the problem extremely seriously, and “the technology had genuinely reached its first inflection point—before that, no software could do this well.” There was no GPT or Transformer at the time. The system relied mainly on traditional machine learning and cheminformatics—“studying how the logic of chemical reactions can be written as algorithms.” He considers this foundation important, though it may be overlooked by many people.
5. The macro backdrop: license-outs surge, and China’s share could approach 70%
- Li traced the industry’s cycle: pandemic-related research produced a high point in 2020–2021, followed by a downturn from 2022 through mid-2025. Starting in 2024, license-outs surged, led by companies such as Akeso, amid further volatility from measures including the U.S. Biosecure Act. WuXi AppTec released a “very impressive” first-half report a few days ago, and its shares broke to a multi-year high, sending a “very clear, very large signal.”
- Chinese biopharma companies generated $110B in outbound licensing value in the first half of this year. China accounted for just over 30% of global licensing value in 2024 and half in 2025. Li said that if global licensing value this year remains at $30B, China’s full-year share “might approach 70%”—“an almost unimaginable figure.”
6. Black boxes cannot be tuned; white boxes are what build trust
- The product that once outperformed models or software from several well-known companies had “a very large white-box component.” Xia’s core argument is that serious scientific software built as a black box “has a hard time earning customer trust”: it may produce hallucinations or bugs unpredictably, with serious and unexplainable consequences. “The biggest problem is that it cannot be improved. You can’t tune it—when you find a problem, you don’t know where the problem is.”
- Li translated the problem this way: change A inside a black box and you may simultaneously change related B, C and D, without knowing what each component does or being able to stop them from moving in the wrong direction. Zhihua’s structural answer is that every conclusion from a foundation model or black-box model “must receive an explanation, verification and constraint inside our white-box model,” removing hallucinations and distinguishing ambiguity, while the black-box and white-box layers remain tightly integrated and operate together.
7. The first MNC client: a more-than-$1M starting point and an iteration loop of more than 1,000 requests
- WuXi AppTec introduced Zhihua to a large multinational pharmaceutical company that was building an AI drug-discovery system. Zhihua was given the entire AI retrosynthesis module. A partnership worth “more than $1M” forced the product through true productization to meet a major client’s requirements for safety, speed and quality, making it “a very important starting point.”
- The initial product was “definitely not a 90-point product—perhaps a 60- or 70-point product.” Iteration filled the gaps: the client’s request list accumulated more than 1,000 items, while the number of modules rose from 4 to 9. “Many investors and industry people think that as long as the model and technology are good, you can instantly become the best in the world. But products are often built through this kind of iteration.” The white-box architecture was critical: with a black box, “the client keeps giving feedback, but you can’t change it, and eventually they become desperate.”
8. One product, 3 departments, 3 different objectives
- The molecular-design side—CADD/AIDD—tends to be weak on synthesis and needs to know whether a molecule can be made and how difficult it will be to make. Otherwise, medicinal chemistry colleagues may respond, “I don’t want to make this at all.” Medicinal chemists want to get a few milligrams quickly and generally do not focus on cost; later, at clinical process scale-up, the priorities become minimum cost, scalability and safety.
- The needs split further as the use case becomes more specific. When medicinal chemists design 30 molecules, they may first synthesize a common core and attach different groups through subsequent steps, minimizing the total number of steps. In the process stage, by contrast, only 1 molecule is made, and the goal is not necessarily the easiest reaction but the lowest-cost one. A reaction that looks difficult may still be adopted, including one using an enzyme catalyst, if it cuts costs.
- The more clients use the product, the more stages it touches, generating requests such as “Would this function work better this way?” and “Would it be better if you connected that function to this one?” “Real-world use cases are actually different from what we originally imagined.” The feedback has helped Zhihua understand customer needs and add modules.
9. AI may replace much of graduate students’ junior cognitive labor; acquisition offers from a CRO giant were rejected at the peak
- Drawing on his own experience as a graduate student tormented by synthetic routes, Li asked how many problems the tool could have solved back then. Xia believes that, from a student’s perspective, AI “may already be able to do most of your work.” This kind of junior cognitive labor will probably be replaced first. He is more optimistic about the broader impact: chemistry still contains extensive innovation work in new materials, and reducing the share of routine tasks will allow people to handle more projects—effectively increasing everyone’s bandwidth.
- During the 2020–2021 peak, a famous and very large CRO made acquisition offers more than once. Zhihua had only a dozen or 20-odd employees at the time. Xia ultimately refused and has no regrets: “As long as this thing can keep moving forward, I consider that a form of success.” Even at the peak, founders need to ask whether “the money is genuinely helping move the project or ideal forward.”
- His lesson for science-trained founders in today’s AI boom is: “Stick to what you believe is right, and don’t let capital’s views completely dictate your direction, because capital often doesn’t really understand the industry’s actual conditions.” Focus is also essential: “抓住这件事情最本质的部分,不要被高光和低谷影响”—stay focused on the essence of the work and do not let peaks or troughs dictate the strategy. Changing direction to chase capital’s attention may produce poor long-term results.
10. GPT fills the “generalization gap”: generalization plus specialization covers almost all of a chemist’s work
- Xia sees GPT as “a very large opportunity,” not a threat. Vertical systems solve specialized problems, but real work is full of generalization problems—for example, “the lab may not have a particular reactor, so the reaction can’t be run.” That is not a specialized chemistry problem, and “with what we had accumulated before, we couldn’t solve this kind of generalization problem.” Foundation models fill the gap: “Generalization problems plus specialized problems are almost the entirety of a chemist’s work,” making it possible to use agents and specialized tools to solve the full workflow.
- The architectural impact is not a direct change to vertical algorithms but a change in the delivery model: the product shifts from a tool used by chemists to an agent accessed through an API, with the agent completing the entire workflow.
11. Agents are more than a friendlier interface: they may remember, learn and work like employees, and MNCs are exploring the model
- Xia added that agents may gradually become familiar with users, develop memory and keep learning, ultimately “possibly becoming like an employee of the company.” After a user teaches the agent a few times, there will be no need to repeat the lesson; the user can simply assign a task. The parts that once required human decision-making and selection can also be delegated to the agent.
- Li summarized that agents will interact simultaneously with the user layer, the lab’s overall conditions and software systems. Early on, people may worry about missing critical checkpoints, but Xia believes that once delivery efficiency and quality exceed human performance, “people will stop caring about it.”
- The first MNC is still using the product and discussing agent-based cooperation. Xia says most of the MNCs he currently works with have not formally deployed agents; they are still testing and exploring the model, while concerns about being left behind by AI are also pushing them toward broader AI adoption.
12. Counterintuitively, white boxes are becoming more common in the foundation-model era; explainability is a commercial requirement
- Asked whether the balance between black and white boxes had changed, Xia replied: “There are more and more white boxes.” When Li asked whether large language models had strong explainability, Xia said the models can explain why they made a particular decision.
- The biggest source of improvement in white-box capability has been “the information obtained through years of continuous interaction and iteration with customers.” In serious scientific settings, customers must know why a conclusion holds, so products need to become progressively more white-box and provide clear explanations for customer questions.
- For AI for Science companies, Xia believes that users of R&D-assistance tools must be able to trust the system. As long as a human element remains in delivery, explainability is necessary. Li said the dynamic resembles the opportunities and challenges facing autonomous driving as it continues to develop.
13. Capability benchmark: a chemist with 10 years of experience, 5 to 6 strategies, and 5 minutes versus 1 to 2 months
- Internal and external evaluations conclude that “the capability of chemical-synthesis AI today is roughly equivalent to that of a chemist with 10 years of work experience.” AI can generally find any synthetic route a chemist can think of, and can generate 5 to 6 different strategies on average, versus 1 to 2 for a person. On more difficult process-route designs, AI can also broadly reach the level of what humans can conceive.
- The most striking client test was: “They may spend 1 month or 2 months designing a route, only to find that we can run it in 5 minutes.” The system also produced routes that clients had tried but ultimately rejected. Li compared this with XtalPi’s Thanksgiving 2016 test at Pfizer: Pfizer tested 3 or 4 molecules whose crystal forms had not previously been published, with the results for 2 or 3 already known; the test results matched Pfizer’s predictions. Xia confirmed that chemical synthesis has reached or surpassed a similar stage.
14. Synthesis is the true bottleneck in the DMTA cycle; AI finds routes humans miss but does not create new reactions
- R&D is fundamentally a multiround Design-Make-Test-Analyze cycle, and developing one drug may require 7 to 8 rounds. The design side can easily produce hundreds, thousands or even tens of thousands of molecules a day, while testing has high-throughput capability. Synthesis is the constraint: “one chemist can only synthesize roughly 3 to 5 molecules a month, or 2 to 3 reactions a day. That is the industry-wide efficiency.”
- The core argument is that accelerating the DMTA cycle by multiples requires multiplying synthesis efficiency. A 2x or 3x improvement would already be significant; reaching 10x or 20x could raise the throughput of new-drug and new-material R&D by 10x or even 20x. Client feedback shows that synthesis is where the entire R&D process gets stuck.
- The boundary for novel routes is that AI can produce feasible, more elegant routes that even experienced chemists would not have thought of, “but it does not create new reactions.” Innovation in synthetic routes generally relies on known reactions; the relevant reactions may simply be buried somewhere in hundreds of thousands of papers. A route that takes 10 steps when the reaction is not visible may take 3 or 4 steps once it is found.
15. The main obstacle for founders from pure AI backgrounds: poorly defined problems and difficulty building trusted systems
- Xia’s experience with collaborators from pure AI backgrounds, including university professors, is that they may not understand the industry problem or may misunderstand it, fail to communicate in time, continue down the wrong path and ultimately blame the data when the result is poor. “It’s not that their capabilities are inadequate,” but without industry know-how they cannot define the problem clearly, build a trusted system or recognize which industry information—often impossible to express in language—is important to the model.
- Li added that AI-oriented teams can go astray in choosing data sources, identifying data types and judging whether results are valid because they do not understand industry requirements. Xia said that for retrosynthesis, the largest contribution may still come from large volumes of public literature and patent data, because the model is learning actual synthetic approaches; high-throughput experiments do not solve the same problem.
- People from AI backgrounds may be good at improving training efficiency when data is abundant. That is precisely what AI for Science lacks: some fields have only hundreds or thousands of data points. Compared with the more than 40 years of public internet and text data used to train general-purpose foundation models, the gap is at least several orders of magnitude.
16. Positive data matters now; negative data will matter later—a different view from XtalPi’s “50/50” framework
- Li relayed XtalPi’s internal view from closed-loop experimentation: positive and negative data each contribute roughly half of model improvement. Xia disagreed: “Positive data—the successful data—is extremely important at the current stage; failure data will be extremely important in the future.”
- At the current stage, reaching human judgment is enough for large-scale commercialization. “If a person runs an experiment and it fails, there is a high probability that before running it they did not think it would fail.” Once the system reaches the human level, “it is already good enough without needing to know that the experiment will fail.”
- Negative data becomes important in the next phase. To surpass the knowledge of a chemist with 10 years of experience and identify reactions that the chemist believes are feasible but are actually infeasible, negative data is critical. Li therefore suggested that negative feedback should theoretically matter more for black-box models. Xia agreed, but said that before a model reaches human-level performance, it may not need to explain failures that cannot be explained.
17. Token economics: every call has a cost, making intelligent orchestration mandatory engineering
- Li highlighted an easily overlooked difference: “The biggest difference between today’s AI and previous software and systems is that every interaction and call has a cost. Interactions with traditional software and databases were effectively free.” He also cited a secondary-market colleague’s analysis of a U.S. model, where input tokens accounted for a notably high share and pushed up costs. This means AI is more likely to first land in work with a clear price tag—for example, a chemist with 10 years of experience at a domestic CRO costs roughly RMB20K–30K per month, or RMB200K-plus to RMB300K per year.
- Clients “care a great deal” about cost and will not use a system that is too expensive. Zhihua’s engineering answer is to save clients money: rather than calling multiple models indiscriminately, it must “call the right model when necessary, using the minimum cost.” Simple questions can go to lower-cost or self-deployed models; only complex problems should call larger models, reserving the most expensive and capable resources for core issues. Large clients have already deployed foundation models internally and can use those directly.
- Xia expects model prices to keep falling over the long term, eventually becoming a low-cost resource “like electricity,” but says “that is a very long-term” outcome. In the medium term, optimization will rely more on the agent layer. Client fees may remain relatively fixed while token costs become a significant part of Zhihua’s cost base, requiring continuous engineering work to reduce expenses.
18. SaaS and agents will coexist; demand is overseas, supply is domestic, and India is a variable
- The commercial judgment is that “the 2 models will coexist for the long term.” For some problems, clients believe SaaS already works well and do not want to pay extra. For others, SaaS is insufficient and doing the work in-house is expensive, so they will use agents priced by call volume. Xia also believes the Matthew effect or flywheel will be more pronounced in the agent model because users can more easily see changes in the speed and cost of end-to-end delivery.
- Domestic clients are more numerous, but overseas customers already generate more total payment value. Xia sees the U.S. as the largest biopharma market, alongside Europe, India and others. Li compared China’s biopharma industry with the precision-manufacturing and Apple-supply-chain phase around 2008–2012: MNCs captured most of the profit, while license-outs forced pipeline companies to become “more numerous, faster and better,” just as competition in precision manufacturing drove demand for industrial robots. Xia does not rule out China becoming the “world’s pharmaceutical factory” in the coming years, much as it did in manufacturing.
- Xia also sees India moving quickly. Zhihua added more than 30 Indian clients this year, and “the growth rate actually isn’t slower than China’s.” The clients include CROs, CDMOs and drug companies, though Xia is not sure whether they are engaged in innovative-drug R&D.
- In materials, the current coverage includes organic small-molecule materials such as OLED optoelectronic materials, new-energy battery additives and photoresists, as well as monomers used before polymerization. Biosynthesis is still fundamentally a set of chemical reactions, so recommended routes can include enzyme-catalyzed steps and suggestions for enzyme types. Zhihua cannot yet create an enzyme that does not exist through directed evolution.
19. Foundation models struggle to enter verticals, investors misunderstand the opportunity, and the brutal answer is that “the window has passed”
- Foundation-model companies have difficulty going vertical. Xia says internal tests show that a vertical model can quickly reach 50 or 60 points in one field, but moving higher is “extremely difficult” and the model generally cannot reach a usable level. The reasons include limited underlying data and the fact that vertical models resemble multimodal data more than pure text. Using image models as an example, he said it “may not be possible to use one large language model to generate images and outperform a model specialized in images.” Foundation-model companies that want to go deep into specific verticals may still need to build small teams, and their advantage over startups is not necessarily large.
- The two sides may ultimately cooperate and even depend on each other: foundation-model companies solve general problems, while Zhihua continues investing in the “dirty, hard work” of building a deeper foundation for vertical fields.
- Investors make 2 major mistakes. First, they assume founders must make innovative drugs or new materials themselves for the market to be large enough. “That is a result, not a cause,” in Xia’s view. The priority should be solving tool-efficiency problems and allowing drug and materials R&D to gain multiples of efficiency; otherwise, the company may lean more toward storytelling and “not necessarily be a real AI drug-discovery company.” Second, investors use the market size of past tool software to imagine the future and fail to see how tools combined with foundation models and agents can complete an entire workflow and deliver an outcome rather than a tool.
- On the customer side, adoption follows a prisoner’s-dilemma logic: companies that do not adopt AI will fall behind; early adopters gain a short-term advantage, but that advantage shrinks once everyone gets on board. Xia hopes senior executives understand that “the era is undergoing a major transformation.” Companies that do not actively improve efficiency and accept the shift could eventually be eliminated by the industry.
- His judgment for challengers with similar backgrounds in the first half of 2026 was unsparing: the 8 years required to build the technology stack cannot be saved, while major clients are already using and deeply customizing existing systems, making replacement “almost impossible.” Therefore, “the window to start a company in this area may already have passed.” Li’s response: “That answer is even more brutal than my question.”
- Looking ahead, Xia believes AI for Science penetration is “not even 0.1%,” while the trend is close to inevitable. In 10 years, people may do very little of the work they do today, with AI handling most of it. He compares the opportunity with the industry-wide shift that began when China adopted its electric-vehicle strategy 13 years ago. Li closed by comparing AI for Science with the microscope in biomedicine: microscopes made bacteria and cells visible, driving broad breakthroughs in biomedicine; AI for Science may similarly become a tool for raising discovery and R&D efficiency across vertical fields.