Pioneers Insight Method Research Author
A Conversation with Former DeepMind Scientist 曹原: AI for Science Is Breaking Out, and a New Era Has Arrived
Back to Episodes

A Conversation with Former DeepMind Scientist 曹原: AI for Science Is Breaking Out, and a New Era Has Arrived

Summary

  • Jeff Dean, Oriol, Sanjay and Quoc leaving to found Discovery Loop marks the breakout of the AI for Science field. Former DeepMind scientist 曹原 offers 3 readings: the organizational conflict created by Gemini’s “concentrate resources to accomplish major tasks” model; Gemini’s “underwhelming” performance over the past 6 months, with 3.5/3.6 Flash proving mediocre and 3.x Pro delayed; and the idea that, once math and coding are in place, knowledge discovery is the obvious “next topic that can unleash AI intelligence.” But he says claims that “a new era has begun and the Google era is over” go too far: “Google may be the one among the big technology companies that has every full-stack ingredient needed to do AI well.”
  • The technical logic behind AI4S breaking out now is that coding, math and the agent harness have matured enough to run the scientific loop of hypothesis generation, experimentation, data analysis and self-iteration. 曹原’s view is that “verification is clearly the hardest part of the entire AI for Science chain”; it is the bottleneck, and once broken, “you can iterate very quickly.” Ideally, “if you could get one experimental verification every minute, the problem would be solved.” OpenAI’s closed loop with GPT-5 and Ginkgo Bioworks’ automated lab is, in his view, “one of the most convincing and effective closed-loop experiments so far.”
  • Biopharma is the best entry point: the market is already large, workflows are standardized and data-rich, and new-drug development can cost several hundred million dollars, making the demand for AI-driven cost reduction most urgent. Materials also has a large market, but it is too fragmented, with value chains spread across applications, making it difficult for AI to capture the full opportunity at once. Chip design, batteries and quantum computing remain largely exploratory.
  • Across the 3 major labs, Google has the earliest and broadest footprint, OpenAI is the most aggressive but suffers from confused priorities, and Anthropic is catching up. OpenAI launched Astra, published a paper claiming the model derived 10 long-standing unsolved math problems and set a goal of having an autonomous AI scientist by 2027; after AI4S head Kevin Weil left, some of the work was folded into Codex. Anthropic launched Claude Science and recruited Nobel laureate John Jumper, potentially filling the only major vertical-model gap in biomedicine. Coding and commercialization remain the top priorities at all 3 labs—“big companies always face the innovator’s dilemma”—creating an opening for Neolab and startups.
  • The real ceiling for AI4S is not compute but the ability to create concepts: 曹原 says “the process of abstracting concepts may be AGI’s final last mile,” and that “this step may be impossible to cross.” It may not even be Turing-computable. Nobel laureate Jennifer Doudna’s verdict is the evidence: AI generates many proposals, but “not one was something we didn’t already know.” AlphaGo’s Move 37 “was definitely a discovery, but that doesn’t mean it created a new concept.”
  • An industrial pipeline for AI for math cannot avoid Lean: natural-language proofs at IMO level are feasible, but “if a language model generates a 100-plus-page mathematical proof entirely in natural language, there is no way to guarantee that it is correct.” 曹原 cites Peter Scholze, who “apparently” won the Fields Medal in 2018: even Scholze could not establish whether his proof was right, and manual formalization could take at least 1-2 years before a successful compile provided confidence. Tao Zhexuan’s description of math moving from an era of scarce proofs to one of excess proofs raises the trust problem. 曹原’s analogy: even if a robot plays piano with perfect tone, you may not want to listen; “a correct result is not automatically enough—you need to understand it, then judge it.”
  • 曹原’s startup is betting on a hybrid LLM-plus-symbolic route: because “the entire universe of a language model is determined by its training data,” while scientific discovery “by definition cannot be in existing knowledge,” an external symbolic layer is needed to generate low-probability new ideas. He does not expect symbolic AI to revive and replace connectionism: “ultimately, 98% may be connectionism and 2% symbolic.” Lean verification in AI for math is already a neuro-symbolic architecture. The current bottleneck is balancing novelty with feasibility.
  • The investment timeline is tiered: intermediate outputs such as targets and molecular structures can already be monetized, but AI autonomously producing Nobel-level results is “at least 20-30 years” away. The core argument is that this is not AI for Science but “AI and Science”: science is not merely an application; it forces AI to develop causal reasoning, long-term memory and continual learning. AI is a “meta technology”—once it solves the hardest problems, “it will definitely do better in other, different fields as well.”

Deep dive

1. Jeff Dean Leaves for Discovery Loop: A Logical Next Step

  • 曹原 offers 3 readings. First, Gemini, as “the biggest project,” requires the organization to “concentrate resources to accomplish major tasks,” creating management conflict inside a DeepMind team that had previously operated more independently: “Some internal conflict was inevitable.” Second, Gemini lost momentum over the past 6 months. Third, Discovery Loop’s mission—having AI conduct knowledge discovery and scientific research autonomously—matches the direction of his own startup. Once math and programming are solved, the obvious “next topic that can unleash AI intelligence” is extending those capabilities to scientific knowledge discovery.
  • The team’s pedigree is formidable. Jeff Dean went from MapReduce, Spanner and BigTable to the Brain Team, TensorFlow, TPU and Gemini, making him “a legendary figure across the entire industry.” Sanjay joined at roughly the same time; 曹原 often saw the 2 peer-coding in the office, with “1 day every week when they had to write code” and a deeply hands-on working style. “If they leave to start a startup, they must want to set a new benchmark in this field—to become a new Neolab.”

2. Diagnosing Gemini’s Loss of Momentum: 80 to 100 Is a Different Game

  • 曹原’s assessment is blunt. At the start of the year, Gemini 3.1 Pro was “arguably one of the better models” on the rankings. Over the past 3-4 months, Anthropic, OpenAI and Chinese open-source models have “advanced by leaps and bounds,” while “3.5 Flash and the latest 3.6 Flash have been fairly mediocre.” 3.x Pro has still not formally launched. “Overall, Gemini is very far behind schedule,” and shifting toward Anthropic’s coding-agent commercial model “will still be very difficult.”
  • He rejects the idea that Google is finished: “A company cannot become unable to operate because of 1 or 2 people, even if those people are changes in the leadership.” Google has the full stack—from chips, infrastructure, data centers and tooling to models, talent, data, products and distribution—and is “completely independent of the external ecosystem.” The issue is reprioritization and resource allocation.
  • Meta is the relevant comparison. MSL reaching its current level over the past year “is already quite difficult,” but “going from 0 to 80 can, relatively speaking, be caught up within a certain period; going from 80 to 100, or even above 100, is a much bigger challenge.” Google has to “reinvent itself” and make the Gemini team operate like a startup.

3. Why Google Could Not Retain Its Scientists: Priority Overload and 曹原’s Own Exit

  • John Jumper and others left because DeepMind’s highest priority is getting Gemini to the top of the industry. “A lot of the work and near-term priorities have to shift in that direction,” including the work of people who had been focused on AI4S. “With a team and a project of this size and priority, you cannot make every person completely satisfied.” He believes management is improving, but it is a continuous iteration.
  • 曹原’s own motivation is worth recording as first-hand testimony. During Gemini post-training, he kept asking a deeper question: “If you really want it to solve scientific problems, its ability to innovate, propose new hypotheses, verify new hypotheses and continuously update itself is still nowhere near the level needed to help humans increase scientific knowledge exponentially.” That is why he decided to leave and try.

4. Demis Moves to Chairman: Google Will Not Abandon AI4S

  • 曹原 confirms that Demis still oversees the AGI and science divisions, and that Google was “the earliest, broadest and deepest” of all AI companies in its science push. His conclusion: using AI to advance science and technology autonomously is “an enormously important topic—not just for an organization, a company or a society, but even for a country. Human society ultimately advances through science and technology.” At that level, Google cannot simply walk away.
  • Demis’s personal disposition is “more like a scientist than a CEO responsible for products or commercial operations.” From around 2015 until the merger, London DeepMind’s Alpha projects were primarily scientific research. Later product successes—AlphaEvolve driving automated algorithm updates and AlphaFold winning a Nobel Prize before being commercialized through Isomorphic Labs—grew out of those early efforts. “Google has both the ability and the responsibility to advance long-term technological research. Demis is extremely well suited to that role.”

5. Concept Class: The Relationship Between AI4S, AI4AI and RSI

  • 曹原 defines AI4S and AI4AI as both treating the AI model itself as a researcher; the difference is whether the research target is a scientific problem or AI itself. In AI4AI, for example, a model receives a fixed compute budget and optimizes training on its own: it inspects the codebase, proposes parameter changes, writes code, runs experiments in a sandbox, observes the results and proposes the next step. “After round after round of iteration, the model’s performance keeps improving step by step.” Some startups are using this to explore non-Transformer architectures and optimization algorithms that do not require Backpropagation.
  • RSI “is not a task; it is a method.” Recursive Self-Improvement can be applied to both AI4AI and AI4S.

6. Why August 2026: 3 Capabilities Arrive Together

  • The breakout requires a chain of capabilities. To run a biology experiment, AI first has to infer the next experiment from current observations; second, write code to analyze whether the data match expectations; and third, trace back and self-iterate when the data are anomalous. On top of that comes the agent loop and the harness for managing memory. “Once these capabilities are in place, you can do science problems better, because science is so broad and the problems are so complex—possibly more challenging than coding and math. So this is the right time, place and people coming together.”

7. Why Capital Is Flowing to Biopharma, Not Materials

  • Biopharma has 3 advantages. The market “is already large; that has nothing to do with AI.” Processes are standardized and structured, with abundant data, making them relatively easy for AI to handle. Drug development requires selecting lead compounds from huge numbers of candidates and then clearing clinical and regulatory hurdles; “the entire cycle can cost several hundred million dollars,” making AI-driven cost reduction “extremely attractive” to pharma.
  • The problem with materials is not market size but fragmentation. Metals, leather, biomaterials, semiconductors and rare earths are divided into highly specific categories, and development processes vary by application, making it difficult to apply one relatively standardized method. Even discovering a new material molecule is only the first step; downstream production, manufacturing and packaging create a fragmented value chain that AI cannot easily capture in full. Chip design, batteries and quantum computing remain “largely in the exploratory stage.”

8. Productivity or Discovery: AI’s 2 Roles in Research

  • The first role is the Co-Scientist: “helping scientists accelerate discovery, without necessarily proposing new methods itself.” Claude Science is essentially a research-version IDE or workbench that speeds experiment design and data analysis, but “the AI model itself does not necessarily produce new knowledge during the process.”
  • The second role is autonomous knowledge generation. AlphaEvolve discovered new algorithms on its own, while AlphaFold helped generate new protein structures. The role depends on whether the model is a domain specialist or an orchestrator inside a general-purpose agent loop.

9. Humans Must Define the Problem: AI Works Best on Computable, Verifiable Questions

  • 曹原 draws a clear boundary: “The problem definition has to be defined by humans. AI still does not have the taste to identify the right problem on its own.” That is where human scientists remain irreplaceable. AI can search documents to structure and improve a question, “but its role is mainly assistive.”
  • The best AI problems are “anything that can be computed, or that can be easily verified and calculated.” AlphaFold’s problem is exceptionally well defined: given an amino-acid sequence, output its tertiary structure. In drug development, simulations of drug properties, stability and side effects are also relatively suitable. A counterexample is a nuclear-fusion device, where the cost and consequences of running an experiment are extremely high.
  • The underlying logic is that the basic elements of the universe are information, matter and energy. AI for math and coding operate at the information layer, “but to do science, you have to reach the material layer—you have to break through the information limit and enter the physical world.” Wet-lab work is unavoidable.

10. The Verification Bottleneck and Automated Labs: The Ginkgo Loop and A-Lab

  • The practical route to accelerating wet-lab work is the AutoLab or Cloud Lab. Standardized procedures can be fully executed by robots, down to dispensing liquids into test tubes: “You just call an API, and it runs the experiment.” OpenAI connected GPT-5 to Ginkgo Bioworks, a robotic lab in Cambridge, Massachusetts, allowing GPT to choose the next most promising and cost-efficient formulation from thousands of recipes for generating new proteins. That completes the loop of fully automated experimentation, feedback and verification.
  • Host 陈茜 adds the example of DeepMind’s collaboration with Berkeley and Lawrence Berkeley National Laboratory on A-Lab: 353 experiments in 17 days, producing 36 successes among 57 targets. 曹原 confirms it was an earlier case, using a model that was not Gemini, and says the physical bottleneck remains the robot itself. “To accelerate experiments, the robot’s embodied intelligence, precision and speed all have to improve.” Non-physical in-silico simulations may develop faster because they are easier to measure.

11. The Representation Problem: A General-Purpose Lab Head Versus Domain Specialists

  • The 2 model types require 2 different forms of representation. General models such as GPT and Gemini act as orchestrators—“lab heads”—and can internalize multimodal representations of almost any format through pre-training and post-training, including language, code and JSON experiment data. Domain models such as AlphaFold, GNoME, MatterGen and weather-prediction models accept only inputs from their own fields, such as amino-acid sequences plus metadata. They “cannot reason beyond the domain” and can only be called as specialist tools.
  • What cannot be represented? “As long as you can convert the data into a string, some modality-specific representation or an embedding, it can be represented. If you cannot measure it, you cannot represent it. That is a problem with AI itself, not just with AI for Science.”

12. How AI Chooses Experiments: Translating Intuition into a Value Function

  • Selecting experiments is a tree-search problem. Each node is an explored approach, and the system scores the node to decide where to go next. 陈茜 highlights the exploitation-versus-exploration dilemma: “It is already difficult for humans to choose; how is a machine supposed to choose?”
  • 曹原’s answer is to translate human intuition into an algorithmic value function, analogous to the reinforcement-learning algorithms used in post-training. “I cannot simply keep pushing the currently best method upward, because it may be a local optimum.” The final score is always a combination of an intuitive score—the value function—and how many times the prior path has already been sampled.

13. The Closed-Loop Methodology: “One Verification a Minute” and AI and Science

  • Coding loops work because Claude’s coding agent is “the most successful commercial closed-loop model so far.” Code can be verified immediately after it is written; users do not need to wait for or accept intervention from the outside world, which builds trust. AI4S involves the physical world, and “it is not as easy to verify as code.”
  • The core quantitative intuition is simple: “If, ideally, you could obtain 1 experimental verification every minute, the problem would be solved. I could generate unlimited data and train directly on it.” When that is not possible, there are 2 alternatives: build automated labs, or make AI smarter so it needs fewer experiments. Instead of running 10 experiments to obtain a target protein structure, improve the intelligence so it needs only 2.
  • That leads to a point he repeatedly emphasizes: “AI4S is not about treating science as an application. It is AI and Science.” To solve science problems better, AI’s capabilities and its ability to work with science must improve together.

14. Causal Reasoning Is a Weakness for Language Models: Change the Wording and the Answer May Break

  • 曹原 characterizes LLM causal reasoning as “unreliable,” not impossible. Change the wording or word order of a reasoning problem without changing its meaning, and the model’s answer “may still be wrong,” because training data usually contain only one formulation. “If your phrasing differs materially from what appears in its training data, the model gets confused.”
  • Interventions—Do-Calculus-style questions such as “what happens if I raise this variable slightly?”—remain beyond the model in many cases. That is why Yann LeCun stresses world models: if a system learns the underlying laws of physics, those laws are independent of how the event is described. Causal judgments should rest on physical modeling rather than language formulation. 曹原 says startups are already exploring ways to help LLMs capture causal relationships more effectively.

15. The Hardest Part Is Not Verification but Innovation: Doudna’s Reality Check

  • 曹原 adds a meta-recursive layer. The system should not only improve the experimental process; it should also learn how to arrange the workflow and when to update its own code. The agent harness itself must iterate in response to observations. That is a meta-recursive process.
  • 陈茜 asks which step is hardest. The answer has 2 layers. Technically, “verification is definitely the hardest; it is the bottleneck, and once you break it, you can iterate very quickly.” The deeper problem is innovation. If the model remains “forever limited to the knowledge available in current papers,” the loop is merely tuning parameters, testing formulations and recombining existing elements. Its ceiling is necessarily limited.
  • The most damaging testimony comes from Nobel laureate and gene-editing pioneer Jennifer Doudna. Asked whether she uses AI for research, she said yes, but “it generates many proposals, and not one was something we didn’t already know.” Some may be ideas people had overlooked, but all remained within existing knowledge.

16. The 3 Stages of Discovery and AGI’s Last Mile: Conceptual Abstraction May Be Uncomputable

  • 曹原 breaks human knowledge discovery into 3 stages: perceptual input through observation and experiment; formation and symbolization of a concept; and consolidation into a theoretical system. Seeing a person push a cart leads to the abstraction of “force,” represented as F, which then becomes part of F=ma. “This entire discovery process is extremely difficult for AI today.” AI4S currently searches and recombines within an existing representation space—far more efficiently than humans, but without generating new theories.
  • On AlphaGo’s Move 37, 陈茜 asks whether the “divine move” counts as discovery. “It was definitely a discovery—it suddenly sampled a path that was not commonly seen before, so it was certainly new. But that does not mean it created the new concept.”
  • His strongest metaphysical judgment comes in a mathematical context. Given only sensory inputs such as 3 birds and 4 geese, he finds it “hard to imagine a program” extracting the concept of number. “The process of conceptual abstraction may be AGI’s final last mile. That mile may be impossible to cross. It may not even be a Turing-computable process.” That is also why he thinks the human brain may not be a Turing machine at the most fundamental level.

17. Unreasonable Labs’ Solution: Add a Symbolic Layer Outside the LLM

  • The starting point is that “the entire universe of a language model is determined by its training data.” Scientific discovery, by definition, “cannot be in existing knowledge.” New knowledge cannot appear frequently in the training set, so the model needs an external mechanism to generate new ideas. “It definitely cannot be another language model, otherwise it would have the same problem as the language model.”
  • Symbolic AI starts with prior rules—If-Then, or A>B and B>C therefore A>C. It is precise but brittle. “You cannot describe a cat purely in language” or capture the entire universe with rigid logic; the moment the cat lowers its head, the rule breaks. That is why symbolic AI failed to become mainstream, while connectionism lets a neural network learn the cat as a vector.
  • The proposed approach is to use symbolic logic to connect concepts and relationships across papers and fields, extracting “the skeleton of knowledge” and feeding it back as a new intuition for the model—“generating things outside its probability range, low-probability things.” With large numbers of new papers published every day, “you cannot train a new language model every day.”
  • 曹原 is candid about the bottleneck: “You have to balance novelty and feasibility. Making a model innovative is easy—throw it a word and it may generate a different idea. The problem is that the idea must be feasible and reasonable.” On a symbolic revival: “You cannot call it a revival. Ultimately, 98% may be connectionism and 2% symbolic.” AI for math is already neuro-symbolic: a neural network generates the proof, and Lean verifies it formally. “That part is entirely symbolic.”

18. From Bacon to Hegel: Today’s AI Debate Replays Classical Philosophy

  • 曹原 maps the debate onto philosophy. The British empiricists—Bacon and John Locke—held that cognition comes entirely from experience. “In today’s terms, that is the large model: it starts as a random model, then trillions of tokens of training determine intelligence entirely through sensory input.” Continental rationalists such as Descartes argued for innate structure. Kant combined the 2: experience provides the material, but a priori structures—roughly 12 categories covering time, space and causality—make the input meaningful. Otherwise, a sphere or a tire is merely random input.
  • Hegel went further: concepts cannot remain fixed; they must continually negate and update themselves. “Mapped into today’s context, that is continual learning.” AGI will not ultimately be a fixed checkpoint. It must keep learning from its interactions with users and task execution, recognizing its own errors and continuously updating itself.

19. The 3 Labs’ AI4S Positioning: Aggressive OpenAI, Catch-Up Anthropic

  • OpenAI is the most aggressive. Its GPT-5 launch included a white paper on solving scientific problems; it has publicly targeted an autonomous AI scientist by 2027; it runs closed-loop experiments with Ginkgo; it built GPT-Rosalind for biochemistry; and its latest Astra paper claims the model derived 10 long-standing unsolved math problems, though 曹原 acknowledges that “some scientists have some disagreements.” The concern is prioritization. After AI4S head Kevin Weil left around April or May, “some of the AI4S work was all folded into Codex,” effectively betting on general GPT plus agents to solve science. With ads, hardware and enterprise products also in the mix, “priorities may still be somewhat confused.”
  • Anthropic is “catching up.” Claude Science launched less than 1 month ago and is essentially a customized interface combining general-purpose Claude with scientific databases and tools. Its recent work includes vibe physics for helping physicists derive results and the BioMysteryBench benchmark, alongside its recruitment of John Jumper. 曹原 believes Jumper’s role may be to fill the biomedicine gap at the only one of the 3 labs without a dedicated domain model. But Anthropic is “definitely preparing to go public,” and coding remains the top priority; the rapid advance of Chinese open-source models is also a major threat.
  • The startup opportunity follows directly. “Big companies always face the innovator’s dilemma. Once they find a clear path to commercialization, they definitely do not want to let it go.” AI4S and RSI are strategic priorities that no major lab can afford to ignore over the long term, but in the short term commercial priorities will dominate. That gives Neolab and startups the first move.

20. The Industrial Reality of AI for Math: 3 Bottlenecks Around Lean

  • The workflow for difficult math is straightforward. An LLM first proposes a proof strategy, but it does not know whether it is correct. The proof must be formalized in Lean; “once it compiles, the mathematical result is definitely correct.” There are 3 bottlenecks: the LLM must be strong enough; translating natural language or mathematical notation into Lean “is not particularly easy,” which is where Axiom Math AI and Harmonic AI are working; and mathlib must contain enough of the mathematics humans have already developed, or the pieces cannot be assembled.
  • Lean is not necessary everywhere. IMO-level math uses relatively shallow knowledge and proofs of 1-2 pages; “if the model is trained well, natural-language generation is largely trustworthy,” and the result can be checked by humans quickly. But if AI generates a proof of more than 100 pages entirely in natural language, “there is no way to guarantee that it is correct.”
  • The strongest example is Peter Scholze, whom 曹原 says “apparently” won the Fields Medal in 2018. Scholze “was himself doubtful about his proof”; after repeated discussions, neither he nor his collaborators could establish its correctness. They eventually translated the relevant theorems, definitions and lemmas into Lean by hand, a process that could take at least 1-2 years. Only after compilation did they confirm the proof. 陈茜 recalls that when Google competed in last year’s International Mathematical Olympiad, contestants had 2 days, while translating the work into Lean reportedly took 3 days.

21. The Key Mathematical Challenge: Inventing Concepts, Not Just Proving Theorems

  • 陈茜 cites a 5-level taxonomy of mathematical problems, from doctoral training exercises to beyond-Fields-Medal problems, and asks where AI stands. 曹原 rejects the scale: “As long as a math problem does not require generating a new mathematical object or definition—if the representation space is fixed—then no matter how difficult it is, a sufficiently strong model with enough search intensity may be able to find the proof.” Proof remains generation plus search.
  • His view is that the most beautiful part of math is not proof but concept invention. “A more important ability of a good mathematician is defining an elegant mathematical object. If all you see is a pile of matrices and cannot abstract the concept of eigenvalues, you cannot precisely capture the internal structure of the matrix.” Can AI do this? “I do not think it can,” and the process may not even be computable.
  • Is math invented or discovered? He admits the question cannot be falsified: “You cannot ask God whether the universe was built with mathematics.” But if forced to choose, “math may have been created by humans.” His argument is Gödel’s incompleteness theorem: every sufficiently powerful formal system contains true but unprovable propositions. “The universe is a thing-in-itself; it does not need to be described. The reason mathematics contains contradictions that can never be reconciled is that it is a man-made logical system.” He stresses that these are personal thoughts and cannot be verified.

22. The Era of Excess Proofs: Robot Piano and the Cost of Trusting a Black Box

  • 陈茜 quotes Tao Zhexuan: “Math has moved from an era of scarce proofs into an era of excess proofs.” Online problems are buried under dozens of AI-generated solutions, but no human expert is willing to take responsibility for verifying them. 曹原’s response: “A correct result is not automatically enough. As a person, you need to understand it and then judge it.”
  • His analogy is worth preserving in full. A robot plays piano with perfect key pressure and timing, but “I simply would not want to listen.” What you want is not only the tone, but “the pianist’s spontaneous emotional state.” Likewise, if an AI proof contains no insight capable of moving a mathematician, its contribution to math may be limited. Discoveries humans cannot understand “can have only partial meaning.” The computer-assisted proof of the Four Color Theorem in the 1970s is an example: the conclusion is usable—4 colors are enough to draw a map—but the proof was a tedious enumeration and “may not have particularly great value from the perspective of mathematics itself.”
  • The economics of black boxes are difficult. Accepting unexplained AI “creates many problems.” Engineers cannot understand the code AI writes; they do not know where the bug is and can only ask AI to debug itself, after which the next iteration may further scramble the code. LLMs are stochastic machines, so the code generated for the same task may differ on the second run. That uncertainty creates a major trust problem. The extreme case is clear: “If AI can one day replace 80% of economically valuable activity, but every activity still needs human intervention for verification, we might as well not let AI do it.”

23. AI for Physics and Abductive Reasoning: Imagination After a Random Event

  • The divide between physics and math is that “math only needs to live in the logical world, the ideal world; physics has to leave the logical world and enter the material world—you have to verify it and run experiments.” AI can help solve partial differential equations and derive conclusions from formal logic. Startups such as PhysicsX are building world models for fluid mechanics and aerodynamics, but work in this area remains limited.
  • 陈茜 asks about randomness, invoking Newton’s apple and cyclosporine, the immunosuppressive approach discovered by chance from a fungus-related finding in Norwegian soil. 曹原 classifies the leap as abductive reasoning. The falling apple was random, but imagining the underlying cause and developing it into a physical system was a human capability. “AI is currently best at deductive reasoning and may also do some inductive reasoning. Given a phenomenon, imagining a plausible explanation and generalizing it to other phenomena requires abstraction and imagination—it is a world model. AI may not be able to do that for a long time.”
  • On the “beauty of uselessness,” his answer is direct: “Aesthetics, art and philosophy are all useless from the perspective of utility. But the best things all start from useless things—from human curiosity and the pursuit of beauty.” That is the “great use of the useless.” In practical terms, AI for math can already be used for chip verification, program verification and any other verifiable domain. Theoretical physicist Edward Witten recently wrote in a footnote to an arXiv paper that “this passage was something Claude helped me think of.” “Phenomena like this will only become more common.”

24. Timeline and Endgame: 20-30 Years to a Nobel Prize, AI as a Meta Technology

  • Commercialization is tiered. “AI4S already has many applications in production. Even if the final product is not an AI-generated drug, intermediate outputs such as targets and molecular structures can already be commercialized.” Nobel-level results are different. That requires not treating science as an application, but pursuing “AI and Science”—first filling gaps in abstraction, causal reasoning, long-term memory and continual learning. Solving the leading-edge topics could itself take “5-6 years.”
  • Asked how long “a long time” means, 曹原 answers: “I think it will take at least 20-30 years.” AlphaFold won a Nobel Prize, “but this was not an AI autonomous discovery.”
  • The endgame is that AI is a “meta technology” and a general-purpose technology. “It is a brain—you can combine AI with anything,” not only science but also business, including an AI acting as a CEO. “If AI4S reaches that level, it means the model has performed so well on such difficult problems. It will definitely do better in other, different fields as well.” Once the model becomes that powerful, “all other related economic activities will necessarily become more efficient.”
  • 曹原 closes with a methodology for practitioners. Most current applications take a bottom-up view: Transformers work, the scaling law has not failed, and the harness and infrastructure should be improved. But “in the current AI frenzy, a lot of superficial and shallow views are inevitable.” The alternative is to reason top-down from first principles: What is intelligence? How do humans understand the world? What properties can AI not reproduce? “Once you turn the question around, you find that there are many things AI cannot do.” 陈茜’s closing question is where human value sits once AI closes the loop from model to experiment: definition, verification, application—or final judgment.