4-Hour Interview with 姚顺宇: Anthropic, Gemini, and the End of Heroism
Summary
The frontier-model race has shifted from “Can AI do it?” to “Can humans define the problem well?” 姚舜宇认为,Gemini、OpenAI、Anthropic已没有谁真正担心追不上,公开 benchmark 又多聚集在约80%,一两个百分点主要是“噪声而不是信号”;但用户仍能感到差异:据他个人了解到的信息,Claude仍偏通用 tool use/agent,Codex在纯 coding 上缩小差距,Gemini可能在 reasoning 与日常使用中更好。 “现在更难的事情是想明白要去做什么。”
Most application moats still sit with the models, and a general-purpose “wrapper” that cannot capture user mindshare extremely quickly will struggle to escape the foundation-model companies’ orbit. 姚舜宇给出两条生存路径:像Cursor一样高速增长、随后自研Composer,但要承受与Claude Code从合作转竞争的赢家通吃风险;或者像Midjourney一样守住一个“小到模型公司懒得管”的市场。“第一步不能一步登天”,更现实的策略是先吃下一个有想象空间的小场景。
姚舜宇明确否认模型进步正在放缓,并判断预训练至少未来四个月仍看不到撞墙迹象。 benchmark逼近100%后增速必然变慢,却不能代表用户体验或模型学习能力放缓;他的直接体感是,“以前让模型学会干一件事情需要动很多脑筋”,现在只要问题和数据定义清楚,余下训练往往顺其自然。许多所谓 scaling wall 在他看来不是规律失效,而是科学假设、数据配比或工程实现中“有一个bug”。
He is betting that 2026 will bring “train with finite context, use as infinite context,” potentially pushing agents toward continuously running personal assistants. 模型以有限 context length 训练,却在使用时持续筛掉不重要的信息,维持近乎无限的历史,可长期跟踪用户与任务;他先称“今年有机会”,随后把判断加重为“无论如何会实现”。不确定性不在有没有技术路线,而在多条路线中哪条对真实用户最高效。
Coding is currently the clearest AI-native market: verifiable rewards, GitHub data, and relatively consistent standards for good code create a training loop that other industries still lack. 姚舜宇称自己“保守估计”超过90%的代码由模型生成,不保守则是99%甚至100%;研究想法的实验效率较一年至一年半前提升约20至50倍,但工作时长和密度反而上升。其劳动力终局判断极端而有意保留不确定性:“千分之一的人干了过去所有人的工作,拿着现在一百倍的工资”,同时强调千分之一只是虚数,也可能是百分之一。
AI will crack math, coding, and scientific research first not because they are easy, but because they come with objective measures; products without an evaluation function are harder to train. “一旦想清楚这个事儿怎么去评价,你就知道怎么训练”,解释了为何最理性的高智力工作率先被自动化;数学推导、论文归纳和数值实验已被大量交给模型,而“什么叫一个好的产品”只能在真实用户使用后知道。产品经理式判断因reward不清,可能比具体编码更晚被替代。
The US-China model gap has narrowed materially over the past 1 to 1.5 years, but compute constraints, distillation strategies, and multimodal paradigms still determine whether Chinese teams can truly close it. 姚舜宇区分“硬蒸馏”和“聪明的蒸馏”:直接抄取别家token既不道德也暴露出团队“不知道自己想干嘛”,而让不同模型参与数据生成或评价虽处于商业灰区,却可能构成真正的multi-agent训练。中国的明确优势还包括豆包语音——“不客气地说,我觉得就是全世界最好的”——以及低成本机器人硬件;但机器人软件“还没有到GPT-1的时刻”。
The edge at top labs increasingly looks like organizational engineering rather than the heroics of a handful of geniuses. 姚舜宇把Anthropic的优势归于反应快、技术负责人能服众且拥有公司决策权,因此能从Claude 3的coding信号迅速形成top-down下注;大公司则背负安全、法律、品牌和持续服务成本。他刻意给行业祛魅:“本质上是那个浪”,最重要的从业者特质不是聪明,而是“靠谱,就是做事儿细,然后对自己做的事儿负责任”。
Deep dive
1. A Q1 2026 Industry Snapshot That Needs a Date Stamp
The episode was recorded in March 2026. 张小珺 noted at the opening that after recording, Meta’s acquisition of Manus was called off, Cursor was rumored to be a potential SpaceX acquisition target, and xAI was set to end its standalone operation, merge into SpaceX, and be renamed SpaceX AI. The M&A judgments in the conversation should be read in the context of that moment.
The guest, 姚舜宇, earned his undergraduate degree at Tsinghua and is pursuing a PhD at Stanford. He started out researching condensed-matter theory, non-Hermitian systems, high-energy theory, and quantum information; he joined Anthropic in 2024 and moved to Google DeepMind between late September and early October 2025. The host introduced him as having worked on Claude 3.7, Claude 4.5, and Gemini 3.
Another 姚顺宇 in Silicon Valley has spent years in computer science and later moved from OpenAI to Tencent. The two were classmates at Tsinghua and still speak regularly, but their paths diverged: the other 姚顺宇 started in CS and thinks about human-computer interaction and products, while this guest describes himself as a theoretical physicist who came to the field “halfway through.”
2. The Bottleneck Has Shifted from Capability to Problem Definition
姚舜宇 rejects a simple “first half versus second half of AI” framing. The qualitative shift he sees is that people are “less worried about whether AI can do something, and more worried about whether the task has been well defined.” Once capability becomes trainable, defining the objective itself becomes the real bet.
A year ago, Anthropic still worried about whether OpenAI’s reasoning would catch up. By Q1 2026, he believed none of Gemini, OpenAI, or Anthropic was genuinely concerned about being permanently left behind. The harder questions are what to build and how to specify the desired behavior.
张小珺 followed up by asking whether models had become homogeneous commodities. 姚舜宇’s answer has two layers: scores on paper have converged, but real-world use still reveals clear differences in character. Converging benchmarks should not be mistaken for interchangeable products.
3. Benchmark Convergence Masks Divergent Claude, Codex, and Gemini Experiences
In the past, public metrics such as SWE-bench and AIME offered a rough read on which model was strongest at reasoning or coding. Now many scores cluster around 80%, and a lead of 1 or 2 percentage points is “mostly noise, not signal.”
Based on information he has personally gathered, 姚舜宇 says Claude remains the strongest general-purpose model for tool use and agent scenarios. In pure coding, Codex has recently gained ground and narrowed the gap. Gemini may be better at pure reasoning and more everyday use, while coding and agents are still closing in on the leaders.
Some of these differences come from deliberate prioritization. Claude has long emphasized tool use; OpenAI invested heavily in reasoning for a period and later put more weight on coding. Priorities determine whether a team builds the relevant infrastructure, environments, and time-intensive data systems early.
Once capabilities improve across the board, intent no longer explains everything. Even when internal test gaps narrow, model behavior may be shaped by data structures nobody anticipated in advance. The hard work becomes defining the problem rather than stacking training against known benchmarks.
4. High-Quality Code Became an Accidental Model Advantage Through the Internet’s Data Distribution
姚舜宇’s example of an unexpected advantage comes from early pretraining. Years ago, models had no agent-style coding capability, yet were already remarkably good at code completion; teams may not have understood why.
One retrospective explanation is that GitHub code was naturally higher quality than ordinary web pages in the unscreened web corpus. Code data therefore received a cleaner signal by accident within the original internet distribution.
The example carries his broader warning: model differences are not necessarily the result of conscious management choices. They may also come from latent structures in the data that nobody recognized before training. “Maybe after some time, when we look back,” the reason will become clear.
5. OpenClaw Was Capability Overflow, Not a Technology Breakthrough That Arrived in 2026
姚舜宇 observed that OpenClaw created more shock outside the industry than inside it. Labs had already run similar experiments or demos; they simply had not polished them into public products. Its early GitHub code was “not particularly clean,” but its value was showing the possibility to everyone.
He believes the relevant capability was already largely available when GPT-4.5 launched, while Claude’s tool use at the time was stronger than OpenAI and Gemini 3. OpenClaw did not go viral immediately after launch; it took time to spread. It was therefore more a “natural overflow of model capabilities.”
The genuinely new consensus is that an upper-layer system can control multiple models, aggregate different tasks, and continuously execute “very, very, very long” long-horizon work. Model labs and larger startups may move quickly to follow and productize the demo once they see it.
Asked why Manus could not do the same thing, 姚舜宇 gave an honest non-answer: “I don’t understand why Manus couldn’t do it. Maybe it just didn’t.” He did not invent a generational technology gap between the two.
6. Agent Wrappers Have Yet to Build a Data Flywheel; the Main Moat Still Sits with the Models
On the fact that both Manus and OpenClaw ultimately moved closer to the model companies, 姚舜宇’s view is that a product needs a moat to survive long term, and “at least for now, a lot of the moats are on the model side.”
The industry often talks about data flywheels, but he has not yet seen an AI application truly build one. Outside agent coding, there are few strictly AI-native success cases. Chatbots still extend search around large demand pools, adding follow-up questions, interaction, and information compression.
张小珺 summarized the problem as a wrapper that lacks enough “escape velocity.” 姚舜宇 did not say products can never develop their own moats; he only said that a product-side moat is “not certain” to emerge. The current fact is that generic wrappers are easy for foundation-model labs to absorb.
7. Startups Have Two Ways Out: Move Extremely Fast or Make the Market Narrow Enough
The first path is to capture user mindshare before the model company can react, then build an internal model. 姚舜宇 sees Cursor attempting this: it first built the product on Anthropic’s models, then trained Composer to reduce dependence on upstream providers.
That has turned Cursor and Anthropic from “inseparable partners” into uneasy competitors. Claude Code has succeeded, while Cursor has begun developing its own model. Coding is a professional productivity tool where winner-take-most dynamics are plausible; losing the competition would be painful for either side.
The second path is to target a market “so small that the model companies simply can’t be bothered to manage it.” Midjourney is his example. A major lab might be able to replicate the product by committing money, data, and people, but the market may not merit priority.
Asked how he would start a company, he admitted that part of him wants to “take one big shot.” His more realistic judgment is that “the first step can’t be a giant leap”: first capture a small use case, while making sure it still has room to expand.
8. Big Companies Have Demos; They Just Cannot Release Them as Irresponsibly as Individuals
The fact that OpenClaw and Manus first emerged from outside teams does not mean researchers at large labs could not imagine them. 姚舜宇 explained that once a big company releases a product, it must address permissions, security, system crashes, legal liability, brand damage, and ongoing compute supply.
Google cannot tell users, “Buy another computer to run this, or it may take every permission and crash your system.” An individual open-source project can put out rough code and invite the community to improve it. That difference in liability directly changes product velocity.
On the acquisition itself, he does not understand why Meta could not build a Manus product internally. Setting price aside, the clearest value may have been acquiring a strong Asian product team and using Singapore as a base to attract talent from China, Singapore, and East Asia.
Google is not a company that “acquires nobody”; the host and guest also mentioned its absorption of the Windsurf team. 姚舜宇 sees these transactions primarily as talent and execution-capability allocation, not purchases of irreplicable product technology.
9. Near-Infinite Context During Use Could Arrive in 2026
姚舜宇’s clearest technology forecast for the year is: “train with finite context, use as infinite context.” Models would still be trained with a finite context length while maintaining extremely long, potentially near-infinite effective memory at runtime.
The simplest implementation would have the model interact continuously with the user and accumulate information, then discard low-value content and retain key state based on the current context. That is what would make a long-term personal assistant possible, rather than a one-shot question-and-answer box.
He first said the technology had “a chance of arriving this year,” then strengthened the call: “Technically, it will happen this year no matter what.” There is no consensus yet on the route. Multiple ideas already exist; the next step is to test efficiency in real, high-frequency use cases rather than wait for a bolt-from-the-blue insight.
10. Model Progress Has Not Slowed; the Old Ruler Is Failing
Asked whether model progress had slowed in Q1 2026, 姚舜宇 answered “not at all” repeatedly. But he refuses to discuss speed without specifying the yardstick: for any benchmark capped at 100%, the monthly point gain must shrink as the score approaches perfect.
Score movements and user experience are not linearly linked. Moving from 50% to 60% may make users feel only a little improvement; moving from 70% to 75% may create a much larger usability jump. Conversely, going from 80% to 90% may feel unchanged or even worse. No monotonic relationship should be assumed.
His view is not based on a single leaderboard, but on direct research experience: “The model’s ability to learn is getting stronger and stronger.” In the past, teaching a model one behavior required many tricks. Now the key is often simply to construct the problem, environment, and data correctly.
11. Pretraining Is Not Dead; Many “Walls” Are Just Bugs
姚舜宇 believes pretraining continued to improve over the past several months, contradicting the then-popular narrative that pretraining scaling laws had run their course. The window he can see is continued progress over the next 4 months; beyond that, he insists that “nobody in AI can predict what happens 4 months from now.”
When a regularity appears to have hit a ceiling, its domain of applicability may genuinely have ended, or necessary conditions such as data may no longer be available. His sharper third explanation is: “There’s a bug somewhere in the work, and they haven’t found it.”
The bug may lie in the scientific assumptions behind a scaling-law experiment—how many tokens each model should use, where the data comes from, or how the training horizon should match the setup—or it may simply be an ordinary engineering error. In the industry, fixing a bug often produces progress “far greater than some magical technique.”
What determines whether a team can break through is not optimism but systematization: when results diverge from predictions, can it design sensible ablations and rule out data, scale, implementation, and assumptions one by one? 姚舜宇 said “we and 海涛” are relatively good at this kind of pretraining diagnosis.
12. Data and Compute Drive the Mature Paradigm; Algorithms Matter Most at Paradigm Shifts
姚舜宇 does not treat data and compute as separable drivers: “Once compute goes up, naturally you absorb more data; once data goes up, naturally you need more compute.” Within the mature pretraining and post-training frameworks, the two are the main growth engines.
Algorithms often have a phase-transition effect. Before the method is found, the system cannot scale at all; once the key principle is discovered, it suddenly moves from impossible to possible. Transformer was that kind of jump for language-model pretraining. Subsequent improvements have mostly been smoother gains in compute and data efficiency.
For natural-language generation, he believes the current route is “scientifically fairly clear” before it hits a wall, with much of the remaining work leaning toward engineering. Whether RL or supervised post-training, the industry already has relatively clear paradigms.
Multimodal generation remains a scientific problem, with no fixed technical route and potentially meaningful architectural differences across companies. The main future driver here may still be algorithmic breakthroughs rather than simply more data and compute.
13. Coding Leads Because It Has Verifiable Rewards, GitHub, and Consistent Quality Standards
Coding has been advancing rapidly since Claude 3.5 New—called Claude 3.6 by some observers—not just over the past few months. The first structural advantage is a clear reward signal: whether the input produces the specified output and whether a feature passes its tests are easy to judge.
The second is GitHub, which contains decades of code written by excellent programmers and can support a large number of training environments. Compared with data that offers no feedback, it gives models a natural and executable foundation.
The product standard is also relatively simple. Social products and games depend on personal taste, while coding has substantial consensus. Good code is usually concise, clearly structured, easy to extend, and sensibly abstracted; strong programmers tend to agree more closely on what counts as “clean.”
These three factors form a causal chain: the evaluation function is clear, the environment can be scaled, and the target style is relatively consistent. Coding is therefore both easier to train and easier to turn into a repeatable professional product.
14. Models Now Write More Than 90% of the Code, Yet Researchers Work Longer
Because of restrictions on Google’s internal tools, 姚舜宇 cannot use Claude Code. But on model-assisted programming, he says that “as a conservative estimate, maybe 90% of the code is generated by models.” The less conservative estimate is 99% or 100%; the remaining 10% is “to give myself a little dignity.”
Human work has shifted to designing code logic, selecting related files, supplying references and context, and reviewing whether the implementation makes sense. Models are already “far better than humans” at directly outputting code, while old weaknesses involving multiple files and deep class definitions are also becoming less common.
Measured by the efficiency of implementing an idea and running an experiment, he estimates a 20x to 50x improvement versus 1 to 1.5 years ago. Researchers can run multiple agents on multiple ideas at once and have models monitor experiments and results.
Efficiency has not turned into leisure. In the past, getting stuck meant scheduling time with a colleague and waiting several hours. Now he can ask Gemini or an internal model, receive an explanation in 5 seconds, and continue working. His hours, intensity, and backlog of ideas to test have all increased.
15. The First Productivity Dividend at Frontier AI Teams Is Higher Intensity, Not Shorter Hours
姚舜宇 typically starts reviewing email and messages from the previous night at 9 a.m. and gets to the office around 10. When he is alone in the US, he may work until 10 or 11 p.m.; when his family is around, he goes home earlier, but “I’m working there anyway.”
Asked whether Google is still the old “retirement home” Google, he said that nobody in AI can coast unless they have lost interest in the technology and in their own ambitions. The organization may not force people to work harder; researchers are more often self-driven and want to keep chasing progress.
The description also challenges a simple way of measuring automation gains by output per line of code. When the cost of trial and error falls, teams tend to expand the set of experiments rather than lock the task count and cut hours.
16. AI Is First Replacing Objective, Rational Work Once Considered the Hardest
Math, coding, AI research, and theoretical science have long been viewed as the most intellectually demanding professions. 姚舜宇 believes they are precisely the jobs models can handle: “Once you figure out how to evaluate the thing, you know how to train.”
A Claude Code-style shift is already visible in basic science. In the past, a physicist with a numerical experiment in mind might spend half a day compiling and coding. Now the program can be running 5 minutes later. Mathematical derivations, proofs, paper reading, and synthesis are also increasingly delegated to models.
After Gemini Deep Research launched, he saw more researchers outsource reasoning-heavy work to AI. The impact is already happening; basic science simply lacks the built-in attention of consumer products unless a model discovers something on the level of “Einstein’s theory.” He was explicit that this moment has not yet arrived.
Excellent product managers may be harder to train. A good product has no objective scale predetermined in advance; it must be built and used by customers before anyone knows whether it is good. Without a clear reward, there is no obvious training handle.
17. Programmers Will Not Disappear Overnight, but Traditional Task-Based Roles Are Already Shrinking
姚舜宇 believes the day programmers are replaced will come, but not all at once on a particular night. Layoffs at some companies are evidence that the gradual contraction has already begun.
He describes AI as a “centralized technology”: a small number of people will be massively amplified, while most people lose the distinctive value they once held. His extreme scenario is that “one in a thousand people now does the work of everyone in the past and earns 100x the current salary.”
When 张小珺 pressed him on “one in a thousand,” he immediately qualified it: the figure is fictional. It could be 1 in 10,000 or 1 in 100,000, or it could be 1%. He calls himself a “famous pessimist” and does not present the ratio as a precise forecast.
In the near term, survivors will need strong technical skills, an understanding of how work fits into the broader company, the ability to break complex projects across multiple AIs, and skill in collaborating with models. But he repeatedly adds a time qualifier: these abilities are still difficult today and may be absorbed by models again in 6 months.
18. Seedance Raised Competitive Pressure but Has Not Created a New Multimodal Paradigm
Seedance drew attention during the Lunar New Year, but 姚舜宇 did not feel the pressure inside his own team. The people responsible for multimodal generation may feel it; he simply did not see a “paradigm-level change.”
His judgment is that Seedance is more likely to have excelled at product quality, data, and engineering details. Because the generation paradigm is not yet fixed, algorithmic differences may still matter. If forced to guess at the biggest factor, he would guess “data,” while repeatedly stressing that he has never worked at ByteDance and was only making a hard guess.
Multimodal understanding is more systematic than generation, but still less standardized than text tokens. What can be said with confidence is that ByteDance and Google DeepMind are both performing well. A product going viral is not evidence by itself of a fundamental scientific breakthrough.
Discussing 吴永辉’s move from Google to ByteDance, 姚舜宇 refused to offer a condescending assessment. From older code and projects, he saw someone at a very senior level who had nevertheless retained exceptionally strong technical ability—“very, very rare” in a large organization.
19. Chinese Models Are Closing the Gap, While Compute Constraints Are Driving More Sophisticated Distillation
Over the past 1 to 1.5 years, 姚舜宇 believes the US-China model capability gap has “clearly been getting smaller.” But he is explicit that he does not know whether it can be fully closed or even reversed.
The practical disadvantage for Chinese teams is that their access to compute resources is substantially worse. That may also have pushed them toward methods such as distillation. Distillation is an “open secret” across the industry, but not every version should be treated the same way.
“Hard distillation” means taking large quantities of generated tokens from models such as Claude and directly training a model on them. He considers this commercially unethical and strategically foolish: the team is copying others and making the numbers look good without answering what it actually wants the model to do.
“Smart distillation” lets external models participate in a self-generated data pipeline or act as answer evaluators. It remains a commercial gray area, but technically may constitute genuine multi-agent training. Different companies’ models have different distributions; integrating them into one training system is more scientifically interesting than simply asking multiple LLMs to work together.
20. Doubao’s Differentiation Comes from Voice and Consumer Context, Not General-Intelligence Leadership
姚舜宇 believes Doubao is certainly not as smart as Gemini or Claude, but it has a distinct character. The clearest strength is voice generation: “To be polite, it may be one of the best in the world. To be impolite, I think it is the best in the world.”
He has not built a comparable system and cannot assess the difficulty. He can only say that it certainly involves model capability and may also include product optimization. Whether in data or other optimizations, it is clearly “a very labor-intensive thing.” Friends and family have also told him it is “fun to talk” to, but he treats that as a subjective impression.
Doubao answers everyday questions quickly and does not display lengthy chains of thought, which fits its user base. US labs have generally prioritized intelligence ceilings and work efficiency. 姚舜宇 does not see short answers as technically difficult; Gemini 3.1 is already faster and less verbose than Gemini 3.
He rarely uses Doubao in the US. Using a Chinese model there is complicated, and most of his own daily questions are still technical, so Claude and Gemini fit his needs better. This is not an “intelligence hierarchy,” but a difference in geography, product positioning, and the distribution of his tasks.
21. The Key Constraint on Phone Agents Is Not Whether They Can Book a Ticket, but the Cost per Task
姚舜宇 calls the Doubao phone “a very good idea,” and says it performs well on actual tasks. What he is genuinely uncertain about is the underlying optimization and per-task consumption.
If the inference cost of having a model book a high-speed rail ticket exceeds the ticket price, the product is unacceptable. The key constraint for phone agents is therefore task cost, not merely whether a demo successfully completes a series of clicks.
On Apple, 张小珺 speculated that the company might advance AI through partnerships. 姚舜宇 replied that if Apple appears highly concerned from the outside but still cannot deliver, it would look foolish. His own view is that Apple does care about AI strategy.
22. Chinese Robotics Hardware Is Surprising, but the Software Has Not Crossed the Generalization Divide
姚舜宇 watched robot performances during the Lunar New Year and checked prices on Amazon. He had expected mature humanoid-robot hardware to cost several million dollars; the actual prices were far lower, highlighting the strengths of China’s hardware supply chain.
On the software side, he says he “hasn’t really understood it.” Existing robots are closer to the feature-engineering era: given a specific environment and task, they can be optimized through RL, simulated environments, and specialized data, but their capabilities do not generalize well.
Before language models crossed the Transformer/GPT threshold, translation and semantic analysis could already handle individual tasks. The real dividing line was whether one scale-up could improve capabilities horizontally and abstract single-task training to related tasks. Robotics has not reached that point.
He believes robotics and multimodal generation have both “not yet reached the GPT-1 moment”—neither has found a scalable unified path. VLA and related efforts are using language models as a base, so robotics will move closer to the main-model route. But “important in the future” does not mean “the path has already been found.”
23. Robotics Labs Are More Exciting Than Language-Model Offices, but the Tasks Remain Highly Narrow
姚舜宇 has visited Google DeepMind’s own robotics lab and Dyna Robotics. The tasks included folding clothes, retrieving items from shelves, and pouring water—relatively well-defined actions.
A language-model lab looks from the outside like an ordinary office. Robotics teams have to control hardware, collect data, and watch machines perform physical tasks, making the environment “much more interesting.”
That visceral excitement has not changed his technical view. Performing a single task well does not mean general-purpose robots have arrived; the biggest missing piece remains a scale-up mechanism that transfers across tasks.
24. From a Small City in Ningxia to Tsinghua, 姚舜宇 Kept Choosing Paths That Were New and Harder
姚舜宇 was born in Dawukou, Ningxia, a small city built around coal mining. He moved to Shanghai with his parents during elementary school, completed middle and high school there, and then went to Beijing for Tsinghua.
He describes himself as “pretty mediocre” as a child and says neither his elementary nor middle school was a competition powerhouse. Because he had never entered an academic competition, he decided that he “had to take one shot before college,” finding the difficulty itself exciting.
Given the choice between an ordinary class at one of Shanghai’s top 4 schools and the competition class at Gezhi High School, he chose the latter. He saw himself as the “underdog,” someone with nothing to lose. The real motive was not a calculated admissions return, but the feeling that without doing competitions, he would be merely “the smoothest stone among ordinary stones.”
The path did help him enter Tsinghua, but not through the physics-competition recommendation route often described online. He made the provincial team but not the national training camp; what actually mattered was Tsinghua’s independent admissions process.
25. A Text Message to an Admissions Officer Shaped His Operating Principles
On the final day of Tsinghua’s summer camp, 姚舜宇 heard that an independent-admissions exam was mainly for Beijing students. He immediately texted the admissions office, asking: “Why can Beijing students take the exam, but Shanghai students can’t?”
Tsinghua ultimately allowed approximately 7 or 8 Shanghai students to sit for the exam, with 2 apparently receiving offers. Although he could not complete half of one problem, he realized others had completed even less and ultimately received an offer with the admission threshold lowered to the regular undergraduate cutoff.
He took from the episode what he calls “the most important lesson in life”: “You have to be bold. If you don’t ask, you will never get it; even if you ask, you may not get it, but if you don’t ask, you definitely won’t.”
He did not feel particularly bold at the time. He simply believed the opportunity would disappear the next day, so he had to pursue it that day. His attachment to Tsinghua also comes from this: the school was willing to create opportunities and give people an equal chance to compete.
26. Autonomy, Competitiveness, and Constant Direction Changes Come from the Same Temperament
姚舜宇 says his parents’ greatest strength was that they “didn’t control me much.” For choices involving high school, college, and independent admissions, he generally informed them rather than consulted them. His principle is: “When you have no way to understand what someone else is doing, not pointing fingers is the best thing you can do.”
He cares deeply about what he actually wants to do. Once he has figured that out, other people should not stop him, and he will give it everything he has. For something he does not want to do, coercion will not help.
He has a strong competitive drive, but mostly competes with himself. If both sides happen to decide that something matters, he will say plainly, “I’ll definitely do it better than you.”
On repeatedly taking on unfamiliar directions, he offers two descriptions: “Put negatively, I like torturing myself; put positively, I challenge myself.” Suffering for its own sake is a sickness; accepting difficulty to learn and expand one’s capabilities is worthwhile.
27. Non-Hermitian Research Taught Him That Theory and Numerics Are Most Valuable When They Disagree
In Tsinghua’s basic-science program, he entered the Institute for Advanced Study because he wanted to do theory, working with 王忠 on condensed-matter theory, topological insulators, and open quantum systems. The field had relatively low barriers to entry but demanded a deep grasp of quantum mechanics, statistical mechanics, and solid-state physics.
Non-Hermitian systems study open quantum systems that exchange information and matter with their environments. Isolated systems are described by Hermitian Hamiltonians; real systems are usually not isolated, so the conventional framework may no longer hold.
The team initially found that hand calculations under periodic boundary conditions stubbornly failed to match numerical results under open boundaries. Further work showed that the Bloch-wave assumption commonly used for Hermitian systems breaks down here, and eigenstates may become concentrated on one side of the system.
The team then developed a framework for describing eigenstates, time evolution, and dynamics in open-boundary non-Hermitian systems, triggering extensive follow-up research. 姚舜宇 compares the experience with AI: develop an understanding, design numerical experiments to test it, and then implement the idea through a training pipeline.
28. Physics Taught Him Not to Worship Theory or Authority
After a successful undergraduate research project—and at a time when he believed it might become an important contribution to the field—姚舜宇 lost interest in simply accumulating papers and citations. He moved into high-energy theory, an almost unrelated field, because he wanted to do something he did not know how to do.
High-energy theory is extremely difficult, but because experiments cannot reach the relevant energies and microscopic scales, the field relies heavily on mathematical consistency. Consistency is a scientific standard, but when multiple frameworks are self-consistent, judgments about which is better can fall back on the subjective views of senior figures in the field.
His assessment of his PhD years is that his papers met the standards of the external community and that he worked hard, but “had almost zero impact on the world—almost zero.” What he could not accept was spending a finite life “serving the old guard.”
The 5 years left him with 2 principles: “Reading is not about reading a lot, but reading deeply,” and “Don’t trust pure theory too much.” Either have experiments and objective evaluation criteria, or at least choose problems with a tangible impact on the world.
29. AI Resembles Early Thermodynamics: Black Boxes Do Not Stop Empirical Laws from Driving Progress
姚舜宇 believes “everything in the world is a black box.” Physics has not explained every macroscopic phenomenon from microscopic dynamics; it has built effective theories at different scales. The fact that language models lack a neuron-level explanation does not mean humans understand nothing about them.
Scaling laws at least describe empirical relationships among model scale, data, and perplexity. The laws of thermodynamics were also initially empirical and only later acquired a microscopic foundation. Once model technology stabilizes, scaling laws may likewise move from empirical regularities toward deeper scientific explanations.
He refuses to dress “emergent intelligence” in pseudo-scientific language because everyone defines it differently. The clearer change is technical emergence: researchers discovered how to train at scale and improve multiple capabilities horizontally as they scale up.
Quantum computing ultimately did not become his career-change direction because its main bottleneck lies in experimental implementation rather than theoretical algorithm design. AI relies on numerical experiments, making it closer to his interests. He compares AI to “18th-century physics”: theory and experiment have not yet separated, understanding is incomplete, but empirical laws continue to provide direction.
30. Anthropic Gave a Physicist with No Industry Experience a Window to Bet on Reinforcement Learning
Toward the end of his PhD, 姚舜宇 began choosing between quantum computing and AI, looking for the field with more opportunities for young researchers. After receiving a Berkeley postdoc offer, he spent 2 or 3 months there in advance and resigned only 2 weeks after formally starting. Berkeley knew he was in talks with Anthropic and still let him keep the postdoc as a fallback.
He contacted Anthropic, OpenAI, and Google DeepMind. DeepMind’s interview process was moving too slowly; OpenAI did not match him with the right people or work; Anthropic’s first manager from a theoretical-physics background proposed a large-scale RL problem.
The timing was around August or September 2024. o1 had not yet been released, and industrial-scale reinforcement learning for language models was far from mature. His understanding of the process was only academic and rudimentary, so he completed every course and assignment he could find and built a NanoGPT-style project from scratch following Andrej Karpathy.
At Anthropic, he chose the less certain RL track over evaluation. The Horizon team had only about 10 to 11 people, while Anthropic had roughly 700 to 800 employees. Its small scale let him see more of the model-training stack and go directly to the people in charge to learn.
31. Anthropic’s Edge Is Turning Market Signals into Company-Wide Technical Bets Quickly
姚舜宇’s most consistent impression of Anthropic is that its execution is “extremely strong.” The organization is relatively top-down and concentrates resources once it makes a decision. Employees share information rather than hoard it, and the company’s smaller scale was particularly conducive to learning the full language-model stack.
After Claude 3 launched, Twitter feedback suggested that its coding was stronger than GPT-4’s. At a time when GPT-4 still clearly led most models, surpassing it on any single capability was a strong signal. Anthropic captured that signal quickly and began prioritizing coding.
The initial advantage had a concrete technical cause—some team did something specific—but 姚舜宇 cannot disclose the details and is unsure whether it began as a deliberate choice or a lucky discovery. His leaning is that it started bottom-up and was later converted by the company into a top-down strategy.
The demanding condition for top-down execution is that technical decision-makers must command respect through capability while also holding company-level responsibility and authority. He mentioned Jared Kaplan and Sam McCandlish as technical leaders near the top of the company and involved in decisions; he explicitly said he does not know how Dario participates in specific discussions.
32. The Model-Company Advantage Belongs to Organization and Reliability, Not Lone Heroes
姚舜宇 believes OpenAI may have been able to build a similar mechanism while Ilya still had decision-making authority. Later, in his words, Ilya “seemed to lose the ability to make decisions”; he does not know why or what ultimately led to his departure. Gemini, as part of a large company, also finds this model harder to replicate. Startups must make bets; large companies operate under a different set of constraints and resource-allocation rules.
This does not mean Anthropic gets every bet right. It means the company can turn a signal into coordinated action quickly. For model companies, speed of commitment, problem definition, and infrastructure execution are now as important as any single algorithmic insight.
His attitude toward the individual-genius narrative is almost deliberately provocative: “Everyone today is a surfer. At bottom, it’s the wave, not the surfer, because AI doesn’t really require that much brainpower in the first place.”
Pressed on what the industry actually requires, he returned to reproducible organizational qualities: “The most important trait in this industry is being reliable—doing things carefully and taking responsibility for what you do.” As heroism fades, the gap between labs will come from whether they can execute every bet reliably.