Pioneers Insight Method Research Author
140: Deep Potential’s Zhang Linfeng and Sun Weijie: AI for Science, from the Beginning to Now
Back to Episodes

140: Deep Potential’s Zhang Linfeng and Sun Weijie: AI for Science, from the Beginning to Now

Summary

  • Deep Potential’s core breakthrough was to preserve first-principles accuracy while dramatically improving computational efficiency. Deep Potential models “atomic coordinates → system energy” with a neural-network surrogate that satisfies translation, rotation and permutation invariance, delivering “more than six orders of magnitude” in acceleration. Zhang Linfeng said a result that originally consumed roughly 200 million core hours could be reproduced on a laptop in under half an hour after training was complete, though training itself still took time: “At the time, you could actually beat a supercomputer with a laptop.”

  • Deep Potential ultimately bet on a self-iterating platform rather than concentrating resources on one or two drug pipelines. The 2020 goal was to become “Dassault Systèmes at the microscale,” using quantum mechanics to fill the atomic and molecular layer left uncovered by traditional CAD and CAE. By 2025, the founders believed the five-year plan had been “more or less achieved”: existing drug and materials-computing software was relatively mature, while the next growth areas pointed to Science Navigator, the research Agent SciMaster, and a closed loop of “read, compute and do.”

  • The company’s technology roadmap successively absorbed four paradigm shifts: machine learning, pretraining, LLMs and multi-agent systems. DeepMD, DeepKS and DeepWF addressed mathematical and physical modeling; atomic, genetic and molecular foundation models beginning in 2021 pushed single-task models toward general-purpose models; LLMs opened the gateway to literature and knowledge; and Agents began to restructure “people” as a factor of scientific production. Sun Weijie’s assessment: “The final step is that Agents represent a restructuring of productivity and the factors of production on another dimension.”

  • Resource constraints were both Deep Potential’s biggest execution risk and the reason it built a full-stack moat. The team’s early estimate was that simply building general-purpose molecular-simulation infrastructure would require roughly RMB1B, while it had only RMB200K on the balance sheet. It got started with RMB12M in three years of competition prize money and subsequent financing from Baidu Ventures. Facing overseas giants with 100x to 1,000x more resources, the company had to build out compute scheduling, scientific data, industrial software, interdisciplinary talent and an open-source community in parallel, and it went through a platform-building quagmire in 2022–2023.

  • The value of research Agents lies not only in efficiency gains, but in their potential to disrupt the existing market structure around papers, peer review, disciplinary specialization and scientific software. Once AI can conduct research, write and review papers, and directly call computational software to interpret experiments, production relationships built on professional silos and scarce collaboration will loosen. Zhang Linfeng’s view: “The old way is becoming less and less workable,” and the real opportunity is to become the new interface and infrastructure connecting the scientific community.

  • The founders are highly aggressive about scientific output after 2030, but conditional on the timing of specific breakthroughs. They expect AI for Science to produce many disruptive advances in batteries, alloys, polymers, drugs and aerospace materials; the probability of successful controlled nuclear fusion within five to ten years could rise materially; and by 2035–2040, cancer may become an intervention-treatable chronic disease. But they reject the idea that lifespans will double in five to ten years, and stress that specific paths such as solid-state batteries still face “a series of uncertainties.”

  • The hardest risk to regulate may not be AI-generated dangerous molecules, but a structural break in human development and scientific evaluation. Sun Weijie believes existing physical supply chains already have a regulatory foundation; the deeper risk is that AI first replaces interns and junior tasks, depriving a generation of opportunities to practice. If papers, professions and universities all lose their old coordinates, mishandling the transition could “perhaps represent a regression for human science.” The variable worth watching next year is precisely “what an innovator actually looks like.”

Deep dive

1. “The party is over” pushed Zhang Linfeng to search for new problems

  • Born in 1993, Zhang Linfeng explored mathematics, physics and computer science at Peking University’s Yuanpei College before settling on physics. Sun Weijie worked on the preparation, editing and fact-checking of the mathematics, physics and chemistry content for The Intellectual. By his third year, Zhang’s problem was no longer that he could not master the courses, but “what do you do after you’ve learned it all?” Yang Chen-Ning’s line that “the party is over” made him feel that young people might never again have the chance to make discoveries on the scale of relativity or quantum mechanics.

  • What attracted him most was the process by which Einstein and others opened up new fields in their twenties. Yet he found that even the strongest students could finish string theory and the necessary mathematics faster, while still lacking a sufficiently exciting big question. “Those of us in this generation who want to start from foundational science and foundational physics to do something interesting may simply lack a problem big enough to drive us.”

  • The uncertainty took him into the mathematics, physics and chemistry content work at The Intellectual, and later to MIT’s mechanical engineering department, where he studied the role of lubricating oil in engine operation. The real spark came from an early algorithmic breakthrough in electronic structure and ab initio calculations during his undergraduate years—the first time he experienced the pleasure of scientific exploration.

2. Textbooks teach derivations but erase the path by which discovery happened

  • Zhang completed general relativity in his second undergraduate year, while his roommate in mathematics did not study differential geometry until his third year. The mismatch made him uneasy: physics classes simply derive field equations from established rules, while mathematical training asks where the underlying structure comes from. “You just accept a series of rules, and then get tested on whether you can derive the result.”

  • After completing special relativity in 1905, Einstein spent roughly another decade trying to unify gravity, only discovering along the way that he needed Riemannian geometry. The classroom starts from the final answer and works backward, compressing ten years of difficulty into a smooth path. Zhang believes learning an answer and acquiring the ability to “derive it from scratch” are two different things.

  • This later became a precursor to his thinking about AI and innovation: if a model presents learners with all the answers directly, it may make knowledge acquisition more efficient while simultaneously taking away the chance to rediscover a problem.

3. The 2016 decision was to stop “learning it again” and turn to machine learning

  • After arriving at Princeton, Zhang still planned to take courses such as condensed-matter field theory. When E Weinan reviewed his course selections, he reminded him that he had already completed more than 230 undergraduate credits: taking quantum field theory with another master would amount to “learning it again.” Zhang should be doing machine learning instead. Zhang immediately dropped every course and never took another class at Princeton.

  • The second week after he dropped the courses—in October 2016—Hinton won the Nobel Prize. Zhang did not see it as the opening of a new field, but more as confirmation of the direction in which the field was developing. The force that continued to propel it was likely the AI wave made visible around the same time by AlphaGo.

  • E Weinan subsequently brought Zhang together with Roberto Car and asked directly: “Machine learning is going to overturn the ab initio molecular dynamics you’re working on within a few years. What makes you think it won’t?” That discussion became the practical starting point for the team’s research.

4. The Schrödinger equation contains most of the material world, but complexity makes it unusable

  • The Schrödinger equation was proposed in 1926. In 1929, Paul Dirac judged that the fundamental laws governing all chemistry and most of physics were already clear from it; the remaining difficulty was that “the mathematics is too complicated for us to solve the equation.”

  • It describes the wave functions and evolution of particles such as electrons and atoms. In principle, materials, electric currents and proteins can all be described within the same framework. Zhang limits the scope to the electronic and atomic scale; more microscopic string theory and field theory, as well as the ultimate unification of relativity and quantum mechanics, lie outside it.

  • The difficulty comes from dimensionality. A protein immersed in water may contain tens or even hundreds of thousands of atoms, each with three-dimensional coordinates. Even with approximate algorithms, computational complexity can still scale with the seventh power of the number of atoms.

5. DFT and molecular dynamics only turned the impossible into the barely computable

  • Density functional theory, or DFT, simplifies complex particle interactions into relationships between particles and external fields, but complexity still scales roughly with the cube of the number of atoms. It typically handles only dozens to hundreds of atoms. As a comparison, a drop of water contains roughly (10^{23}) atoms.

  • There is a time-scale problem in addition to the spatial one. A single atomic response may occur over (10^{-12}) to (10^{-15}) seconds, while researchers may care about phenomena lasting hundreds of nanoseconds, microseconds or even milliseconds. The simulations cannot become large enough or run long enough.

  • Molecular dynamics itself can be approximated as Newtonian mechanics: the negative gradient of energy with respect to each atom’s position gives the force, which moves the coordinates to the next frame. But if quantum mechanics must be used to recalculate the electronic structure every time an atom moves, every frame of the “movie” becomes prohibitively expensive.

6. Deep Potential compressed quantum accuracy into a surrogate model

  • Deep Potential takes atomic coordinates from (r_1) to (r_n) as inputs and system energy as the output; force is obtained from the negative gradient of energy with respect to the coordinates. AI is not replacing physical laws here, but learning a high-dimensional surrogate model.

  • Compared with conventional molecular dynamics, which further simplifies interactions between atoms, the team wanted the surrogate to preserve the accuracy of ab initio calculations—especially DFT—while running at much higher efficiency. The aim was to ease the accuracy-efficiency trade-off spanning the Schrödinger equation, DFT and molecular dynamics.

  • Zhang unified the problem as follows: can complex wave functions, density functionals and interatomic potentials be effectively represented, approximated and solved faster by AI? The rules of physics are as clear as the rules of Go; the real difficulty is getting from the rules to an effective answer.

7. The real algorithmic hurdle was making neural networks respect physical symmetry

  • Simply feeding all atomic coordinates into a standard neural network does not work. The physical energy of the same water molecule should not change after translation or rotation; it should also remain invariant when atoms are exchanged, while wave functions involve exchange antisymmetry.

  • One extreme approach at the time was to design separate descriptors for water, silicon and other systems, a form of heavy feature engineering. But “a system that handles water cannot handle silicon.” The team chose first to implement simple, general symmetry handling, and then use a general neural network to model the high-dimensional function.

  • This was initially a mathematical-modeling problem and only later an engineering implementation problem in frameworks such as TensorFlow. After a demo appeared in May 2017, Zhang and Wang Han spent roughly a month turning the method into a general-purpose solution—fast enough that their adviser believed it was already sufficient for a PhD thesis.

8. Two hundred million core hours were compressed into half an hour on a laptop

  • Roberto Car’s team already had an expensive ab initio result that had consumed roughly 200 million core hours. After Deep Potential was trained, Zhang could reproduce a highly consistent simulation on a laptop in under half an hour. At the time, estimating one core hour at roughly RMB0.10, the traditional result represented a very large resource cost.

  • Zhang summarized the result as “more than six orders of magnitude in acceleration,” even saying that “you could beat a supercomputer with a laptop.” One famous demonstration was run on an airplane, but the key point was not the setting: computation had shifted from a scarce research resource to a tool for rapid iteration.

  • The research cadence changed from one discussion a month to a discussion in the morning followed by new results in the afternoon. Zhang’s self-deprecating explanation was: “The professors thought I was extremely productive. Actually, the computation was just very fast—not me.”

9. Idle P100s gave the team an early view of compute infrastructure’s value

  • At the end of 2017, Princeton acquired more than 200 P100s, and Zhang noticed that utilization was low. TensorFlow compilation was still troublesome at the time. He clicked “yes” on the CUDA support option and, unexpectedly, the build succeeded; training speed then improved by roughly 10x.

  • Combining idle compute with open-sourced code allowed the team to explore problems at the electronic, atomic, mesoscopic and even macroscopic scales in parallel. By 2020, it could use more than 20,000 GPUs and continue pushing the limits on what was then the world’s largest supercomputer.

  • The experience led the team early on to treat TensorFlow, CUDA, GPUs, supercomputers, elastic cloud computing and local private data as one infrastructure problem. For an algorithm to reach production, different forms of compute and data had to be managed together.

10. “AI for Science” initially meant using machine learning to solve high-dimensional scientific problems

  • AI for Science was proposed by E Weinan, Teng Chao and others at a Peking University conference in the summer of 2018. Sun emphasized that the definition at the time was closer to machine learning for science: using machine learning to accelerate physical and mathematical modeling, rather than today’s broader concept encompassing LLMs and Agents.

  • Deep Potential offered a methodological example: many bottlenecks in scientific computing are fundamentally problems of modeling and solving high-dimensional functions, precisely where machine learning performs well. E Weinan even wrote that machine learning “may be the last missing piece of the puzzle in applied mathematics.”

  • This was not a model for one drug or material, but a general possibility extending from foundational physical methods into chemistry, materials, weather, automotive and aircraft simulation.

11. AlphaFold 2 and Deep Potential represent two different data routes

  • The problem addressed by AlphaFold 2 was more data-ready. The field already had roughly 2 billion protein sequences and 200,000 protein structures, with the objective of learning the mapping from sequence to structure—similar to ImageNet, with massive data and a clear benchmark.

  • Deep Potential operated in a field without naturally abundant experimental data but with computable first principles. The team used first-principles calculations to generate “atomic positions–energy” data and then trained a model. Sun summarized it as “putting the principle into the data chain.”

  • This expanded the boundaries of AI data sources. It is impossible to conduct 10,000 physical experiments on robot motion or rocket ignition, but one can first simulate them on a computer according to physical laws and then train a model on the generated data.

12. The decision to start a company came from three abrupt instructions: graduate, return home, start a company

  • Sun and Zhang had been teammates in basketball, badminton and the student union’s sports department since entering Peking University. They continued discussing technology, investing and entrepreneurship in graduate school. Sun initially saw Deep Potential as a potential investment opportunity, but after a quick review concluded it was “the only one of its kind in the world”—and the only investable target appeared to be the two of them.

  • In 2018, E Weinan brought the two together and gave Zhang three suggestions: graduate now, return to China and start a company. His reasoning was that he had not seen an opportunity like this in more than 30 years, while continuing to publish papers, maintain code and answer questions at a university would not be enough to achieve engineering, organizational development and eventual commercialization.

  • Zhang had never considered leaving Princeton, but accepted the logic of elimination after the conversation: “The whole opportunity is AI for Science, not simply getting a simulation to work and publishing a paper.” If the goal was to influence an entire field, a new organizational form was required.

13. Two entrepreneurial motivations converged on “originating in China, leading the world”

  • Sun had decided as early as 2016 that he would start a company when the right opportunity appeared. Zhang was not looking for a direction in order to start a company; he believed “it had to be done.” One had prepared to become an entrepreneur when the right opportunity emerged, while the other had been pushed onto the court by the problem. The contrast became a source of complementarity between the founders.

  • The host’s challenge is worth preserving: if only a handful of people in the world are doing something, that may indicate originality—or a lack of prospects. Sun’s answer was that the downside risk was manageable: “In the worst case, I can go back to my county and become a teacher.” The upside was changing drugs, materials and the scientific paradigm.

  • He explained the decision through basketball: “The opponent gave you an opening; the game gave you an opening. If you don’t take it, you will definitely be punished.” The mission the two ultimately formed was to build a technology company “originating in China and leading the world.”

14. Drugs became the first entry point because demand, payment and validation were all in place

  • In 2019, the team studied almost one industry every day, mapping its data, value chain and R&D logic. It eventually narrowed the broad direction to chemicals and chemical engineering, drugs, materials, semiconductors and energy, and prioritized drugs.

  • Drug development depends heavily on interactions between atoms, but also has strong payment capacity, clear division of labor and a mature CRO system. Any computational result could be handed over for experimental validation relatively quickly. Schrödinger and Accelrys had already demonstrated that a market existed for related software and services overseas.

  • Schrödinger at the time mainly used previous-generation empirical molecular dynamics to calculate drug–protein interactions, while Deep Potential sought to add machine learning. The core modules of drug-design software covered docking and FEP, or free-energy perturbation. Technical validation came relatively early; product and customer validation emerged mainly in 2020.

15. With RMB200K on the balance sheet, the team estimated infrastructure would cost RMB1B

  • Deep Potential was incorporated on November 29, 2018, and began formal operations in May 2019. The founders estimated that turning molecular simulation into general-purpose infrastructure would require roughly RMB1B, while the resources on hand from research projects, outsourced R&D and other sources totaled only about RMB200K.

  • The early breakthrough came from competition prize money totaling RMB12M over three years. The defense took place the day after Zhang’s wedding. After giving a speech at noon, he reminded his friends: “Don’t drink too much. We still have to revise the slides at 3 p.m., and defend tomorrow.” The prize money functioned like an angel round that did not require immediate fundraising.

  • The pandemic disrupted plans in 2020. More than 20 people were spread across multiple time zones, creating a relay-style “company that never sleeps.” Sun and Li Xinyu returned to Beijing to launch the first financing round, eventually led by Baidu Ventures.

16. Dassault Systèmes turned the commercial vision from a point algorithm into microscale industrial R&D

  • In 2020, Sun reached a conclusion: “Whenever the R&D paradigm in an industrial field matures, R&D must be something that can be done on a computer.” Renovation starts with renderings; aircraft and cars likewise begin with digital drawings, CAD, CAE and multiphysics validation.

  • Dassault Systèmes moved aircraft and automotive drawings from paper into computers, then used solid mechanics, fluid mechanics, electromagnetism and optics to determine whether they could operate and whether they were safe. Even a wind tunnel is not simply “blowing air at an airplane”; software is needed to process multidimensional experimental data.

  • By contrast, drug development still jokes that it is “alchemy,” while materials R&D resembles “cooking: add some salt, add some MSG; if there isn’t enough water, add flour; if there isn’t enough flour, add water.” Traditional industrial software has absorbed many macroscopic laws, but has not fully exploited quantum mechanics at the microscopic layer.

17. The 2020–2025 goal was to build “Dassault Systèmes at the microscale”

  • In 2020, Deep Potential set a five-year plan to build an industrial R&D platform at the microscale, connecting huge industrial applications with the deep scientific law of quantum mechanics. By 2025, the time of the program, both founders’ review was that “the goals we set were indeed more or less achieved.”

  • The path expanded from drug-design software into R&D services, then into materials and battery software, as well as Bohrium, a research platform serving universities and research institutions. The original atomic and molecular models became a general-purpose technical foundation across use cases.

  • Zhang added that the early results were never static products. Once an algorithm is published, others can use it. The company therefore had to keep searching for application scenarios while continuously absorbing new AI technologies, compute and data systems into the platform.

18. Industrialization meant rebuilding the entire computing and data system, not packaging a paper as a product

  • “Attention Is All You Need” appeared in 2017 and GPT-1 in 2018. Rapid progress at the AI foundation layer meant the company could not simply defend one Deep Potential paper. Every algorithm had to advance toward deeper problems in drugs and materials while being engineered for high-performance production use.

  • Batch scientific computing requires elastic cloud compute, while sensitive data requires local private deployment. Models, compute, storage and data management therefore had to be built together. Zhang saw this as another gain from DeepMD: the team was forced to push supercomputing, cloud and local systems to the frontier.

  • Engineering also created a “single-plank bridge.” An innovative algorithm can prove that it is more than a one-off demo only by continuously absorbing user feedback and handling real data and industrial boundary conditions.

19. With no ready-made talent pool, Deep Potential had to build its own education system and community

  • In the early days, no single discipline covered mathematics, physics, chemistry, AI, software engineering, products and commercial deployment at once. More senior talent often already had established agendas, so the team had to begin training students from their first or second undergraduate year.

  • Zhang initially explained the algorithmic background repeatedly on a blackboard. As the R&D grew more complex, that approach could not scale; new hires needed to understand the APIs and take over quickly. The company began building an AI for Science version of Colab, Kaggle, Coursera and an open-source community.

  • Zhang said DeepModeling should now be one of the world’s largest AI for Science open-source communities. It is used to identify talent, observe the new scenarios into which users take the models, and continuously accumulate data and demand.

  • Sun defined the internal growth model as “learning by doing” and “learning as needed.” As AI lowers the barriers to acquiring knowledge, the real difficulty is deciding what one should learn around the problem that needs to be solved.

20. The biggest trade-off was not failing to see opportunities, but seeing too many with too few resources

  • Zhang said the team had looked at nearly every important direction in the early years, including AlphaFold-style protein problems, AI for materials and fusion. Some were “perhaps Nobel-level opportunities,” but with limited resources, the team could only keep the spark alive rather than push every one of them to the limit.

  • He described the gap with unusual specificity: overseas, PhD graduates were competing to join major companies; at Deep Potential, the team had to start by training undergraduates, plugging in power to the machine room and deploying environments. Against DeepMind and industrial giants, its resources may have been 100x to 1,000x smaller.

  • The founders’ retrospective was nevertheless that they had captured all four major opportunities: machine learning for mathematical and physical modeling, pretraining for scientific models, LLMs for research efficiency, and multi-agent systems for restructuring scientific workflows. The real difficulty was deciding how deeply to invest in each wave.

21. Pretraining turned “one model for every substance” into a unified scientific model

  • The pretraining route began in 2021. Previously, calculating a light alloy made of aluminum, magnesium, copper and titanium required a separate model for specific elements and properties. With a general pretrained model, the same foundation can be reused to calculate aluminum alloys today, high-temperature alloys tomorrow or molecules the day after.

  • Sun said the team’s 2019–2020 presentations had already sketched the path from “small models → community data accumulation → database → pretraining → large model.” They simply did not know when an LLM-style discontinuous shift would arrive.

  • The model lineage includes DeepMD, DeepKS and DeepWF, as well as large models for atoms, genes, proteins and molecules. The team said Uni-RNA and Uni-Mol were released relatively early, followed by efforts such as NVIDIA’s Evo 2 and DeepMind’s AlphaGenome, along with comparable work by DeepMind and Meta in other related areas. Its claim is only that it remains in the first tier, not that it leads on every metric.

22. The gap created by LLMs is that scientific corpora are still not truly AI-ready

  • LLMs first changed information processing, knowledge acquisition and the interaction interface, allowing scientific software to be called through natural language. But Zhang emphasized that LLMs are accelerating the mining of existing literature and patents, not yet exhausting scientific knowledge.

  • Public papers and abstracts represent only part of a complete scientific corpus. Molecular formulas, experimental spectra, scientific charts and patent details still require dedicated processing, labeling and modeling.

  • When the host asked why Deep Potential was investing now if large-model companies might fill the gap in two or three years, Sun’s answer was the window of opportunity: “Others currently have no way to do this well, and that is what gives us the opportunity.”

23. Agents are beginning to restructure the fourth factor of scientific production—people

  • Zhang summarizes the factors of scientific production as “read, compute, do” and people. Read corresponds to databases and literature, compute to scientific software, do to experimental instruments and laboratories, with scientists organizing the entire process.

  • Machine learning and pretraining mainly enhance computation and experimental signal processing; LLMs enhance literature databases. Agents do not merely upgrade tools but change the role of people. The current stage is co-scientists, providing an assistant to every scientist; only later might Agents execute complete research tasks autonomously.

  • Databases, software and experimental equipment in the past were all “designed for people to use.” In the future, more of them will be “designed for AI to use,” with Agents autonomously calling tools, validating results and starting the next round of research.

24. Deep Potential is not trying to do everything; it has to complete the loop no one else can

  • The host challenged the company’s breadth: literature Agents, scientific computing and laboratory automation already have separate companies, so why touch all of them? Sun acknowledged that “we do not need to do everything.” The company generally does not duplicate mature laboratory-automation hardware, robotic arms and similar capabilities.

  • The boundary depends on whether capabilities can be accessed publicly. If existing capabilities cannot be called by Agents, and read, compute and do are all indispensable, Deep Potential has to fill the gap; otherwise it cannot complete the discovery loop.

  • Zhang stressed: “Accelerating discovery is not complete when you discover something in your head; it is complete when you make the thing.” Answers from literature, computational results and physical experiments must validate one another in the same cycle.

25. Research Agents will first weaken the scarcity of papers and specialist collaboration

  • Zhang expects AI to soon write papers, review papers and independently complete large amounts of research that is not sufficiently original. This will force the community to ask: “Then why are we publishing so many papers?”

  • In the past, after an experimental scientist discovered a phenomenon, they had to find a computational scientist who knew a particular piece of software. In the future, the scientist may ask an Agent directly, with the Agent calling the software, obtaining the result and explaining it interactively. The shift does not necessarily mean the technology will be mature immediately, but it will rapidly unsettle the existing division of labor.

  • The approximations and terminology inside professional silos will also face pressure. An approximation accepted within one field may not matter for a new problem. When a large model can easily complete an established argument, specialized disciplines will have to explain their distinctive value again.

26. Once interdisciplinary knowledge is flattened, research value must realign with real problems

  • Sun used lithium as an example. Materials researchers focus on power batteries, while the medical field uses lithium carbonate to treat depression. Knowledge that was previously isolated may now be connected rapidly through AI. Those connections are already destabilizing disciplinary boundaries.

  • The sharper question is how much value increasingly specialized R&D ultimately creates, and how much qualifies as a scientific discovery that genuinely expands the boundary of knowledge. Platform AI makes the gap between investment, papers and deployment harder to ignore.

  • Zhang’s conclusion is that the old way is “becoming less and less workable,” but the new paradigm still needs organizations, tools and evaluation systems capable of a smooth transition. Agents may become the unified interface connecting users, production relationships and the scientific community.

27. “Originating in China” is not only a registration choice, but a bet on open science

  • The founders revealed that the team once had the opportunity to move as a whole to the United States or elsewhere overseas, with a promise of resources far greater than it had at the time. It ultimately chose to stay in China. “Originating in China, leading the world” was no longer merely an early slogan.

  • Zhang believes the assumption that “science has no borders” is being questioned. As foundational AI models increasingly run into these constraints, and funding at some U.S. universities and major research centers faces interruption, what the scientific community can and cannot do is becoming more dependent on external conditions.

  • What he wants to promote is a value system in which everything from scientific research to industrial R&D is “open-source and open.” He believes this position is more firmly held in China. The goal is not to reduce science to national competition, but to preserve its ability to expand human knowledge, feed back into technology and improve living conditions.

28. The 2022–2023 quagmire came from depth, platform-building and new technology locking together

  • Zhang believes the tightest resource constraints were not necessarily in 2022, but that period was when the company was “deepest in the mud.” New algorithms had to prove themselves in real settings such as drug pipelines, while also completing the engineering, feedback and iteration required for industrial software.

  • At the same time, every new platform module required a complicated build-out. ChatGPT opened an entirely new direction at the end of 2022, making literature, patents, wet- and dry-lab loops and Agents all necessary investments. The organization, technology and application directions became mutually dependent and locked up.

  • By the end of 2023, the combination of deep exploration, infrastructure accumulation and market changes had pulled the team out of the trough. Zhang believes Deep Potential subsequently developed a distinctive organizational form relative to internet giants, industrial companies and universities.

29. The key decision in 2023 was to prioritize the platform over an owned drug pipeline

  • Sun believes the logic of a platform company and a pipeline company is “not very compatible.” Drug pipelines require bets on one or two assets and test judgment about clinical demand, project selection and subsequent clinical management—areas that are not the founders’ strongest.

  • More importantly, a platform serving only its own pipeline cannot iterate sufficiently. Serving 1,000 or even 10,000 customers allows models and software to become self-improving. In 2023, the team ultimately chose the general-purpose platform. It had experimented with pipelines before, but did not continue to double down.

  • The host asked about the platform’s upside by pointing to the market-cap differences between ARM, Synopsys and downstream chip companies. Sun believes a platform, pharmaceutical company or materials company can each reach the $100B level if built deeply enough. The choice is not about which has greater value, but which mission projects the larger impact “in one’s heart.”

30. The Nobel Prize is only “the first”; the platform wants to be everyone’s “last”

  • The founders believe it will be difficult over the next 10 to 20 years to have a Nobel Prize entirely unrelated to AI. The AI for Science infrastructure itself may win prizes, while new mechanisms, materials and physical phenomena discovered with that infrastructure will continue to be judged under existing standards.

  • Zhang distinguishes two goals. Nobel Prizes reward the “first” major achievement, while infrastructure seeks to be the “last”—the platform on which everyone ultimately works. The ideal platform would have the best PMF, be used by everyone and create a positive feedback loop.

  • AlphaFold’s Nobel recognition in October 2024 affected him because the team had once stood before a similar possibility. But he does not regret returning to China to start a company, arguing that AI for Science will ultimately challenge the old evaluation system represented by the Nobel Prize itself.

31. Software is the most mature area today; growth expectations are shifting toward knowledge and Agents

  • Sun said drug and materials-computing software is currently the most mature, but the market is relatively limited. Products with stronger technology leadership and more room to grow include Bohrium’s Science Navigator and the research Agent SciMaster, which is now being developed.

  • Science Navigator is mainly purchased by universities and research institutions on a subscription basis, helping faculty and students process literature and acquire knowledge. Scientific-computing platforms also exist, while intelligent laboratory platforms remain immature and may become a focus next year.

  • Read, compute and do can each serve as an independent product. Agents can connect any one, any two or all three, creating a closed loop for customers from a question through experimental validation.

32. By 2030, the tangible changes may be new materials and a professor-level scientific friend

  • Zhang expects many drugs, materials, batteries, alloys and polymers to be developed through AI for Science by 2030. Which technologies ultimately prevail will depend on exploration over the next five years, and paths such as solid-state batteries face “a series of uncertainties.”

  • Sun’s picture for ordinary people is that everyone will have a cross-disciplinary, professor-level AI scientist as a friend, able to answer questions about pink skies, dark matter or any other scientific topic. On the industrial side, phones that go 10 days or even a month without charging could emerge, while electric aircraft could achieve still longer range.

  • For more distant goals, both believe AI could materially increase the probability of controlled nuclear fusion making progress within five to 10 years and help develop the materials needed to reach the Moon and Mars. By 2035–2040, cancer may become an intervention-treatable, manageable chronic disease, but “doubling lifespans in five to 10 years” is too aggressive.

33. The biggest risk is a break in human development; the next stop is the Innovator who can truly innovate

  • Dangerous molecules, drugs and false knowledge are visible risks, but Sun believes physical synthesis and supply chains already have a regulatory foundation that can be upgraded around AI. The harder problem is that once AI replaces interns and junior tasks, people may lose the ladder on which they practice, leaving universities, professions and human scientists unsure how to evaluate themselves.

  • Both guests believe foundational training in mathematics, physics, computer science, political economy and philosophy—as well as literature, art and sports—will become more important. The professions carved out during the industrial era may return to forms of essential learning resembling “natural philosophy” and “political economy and philosophy.”

  • Agents can already call read, compute and do tools and work continuously for two or three days. But Zhang distinguishes between “knowing everything” and “being able to innovate.” A truly capable model must explore the unknown from existing knowledge, rather than merely optimize next-token performance.

  • Zhang believes scientific results in specific fields beyond the understanding of ordinary scientists are highly likely to appear next year, though he considers FutureHouse’s claims about new drugs to contain a substantial promotional element. The founders’ own constraint remains the same: “Only idealistic pragmatists can change the world”—with both “the compassion of a bodhisattva” and “the force of thunder.”

Verification Notes

  • The drug-design software is called “Permit” in one reference transcript and “Ornate” in another, making it impossible to establish a unique name from the available material. The text therefore consistently refers to it as drug-design software rather than using an unverified product name.