Chinese Characteristics
Key Views & Dialogues
Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics
- 🗓️ Date:
2026-08-02| 🎙️ Show:The Cognitive Revolution
China’s deployed AI safeguards trail America’s largely because OpenAI and Anthropic dominate the US average, while Chinese universities and companies now produce roughly 50-60 safety papers monthly. Open weights remain the fault line: Beijing regulates services and believes it can reverse domestic releases, but Nathan warns of irreversible bio and cyber risk as capabilities converge within roughly nine months.
View Dialogue Notes & Key Takeaways
China’s deployed AI safeguards currently trail America’s, but the headline gap is largely an OpenAI-Anthropic effect rather than a civilization-wide divide. Nathan estimates roughly 10-12 near-frontier developers in each country; remove “Openthropic,” and the remaining US companies overlap much more closely with Chinese peers, with Gemini only modestly ahead. Concordia AI’s evaluations tell the same story: mostly American proprietary models sit near or above the “45-degree line,” while mostly Chinese open-weight models tend to fall below it.
The claim that China “doesn’t care” and “will never slow down” is contradicted by both policy and precedent. Beijing reportedly held up many domestic chatbot launches for roughly six months in 2023 while it created standards, and today the CAC can require local and national reviews before a service enters the registry. China has also imposed costly rules on recommendation algorithms, gig platforms, children’s gaming and AI companions—evidence that the state will subordinate company growth when it believes intervention is necessary.
Chinese AI safety is scaling rapidly through universities and companies rather than America’s permissionless nonprofit ecosystem. Concordia counts growth from only a few papers per month in 2023 to roughly 50-60 per month by mid-2026, versus an estimated 50 to a few hundred across the US or Anglosphere. The work spans self-replication, evaluation faking, deception, mechanistic interpretability, multimodal attacks and hazardous-capability isolation; Nathan’s conclusion is that “AI safety has taken root in China.”
The central Chinese risk conversation has moved beyond censorship to agents that can act in digital and eventually physical systems. One major technology company insisted it “really do[es] care about catastrophic risks like CBRN risks” and said an agent produces a daily report on American AI-safety discourse. Xi Jinping’s WIC language likewise called for faster safeguards against “loss of control,” prevention of malicious use and keeping AI “under human control”—more safety-forward rhetoric than Nathan can point to from a prominent US official.
The largest unresolved fault line is open weights: China regulates services and believes it can put the genie back in the bottle, while Western safety analysis emphasizes irreversible global release. Chinese interlocutors argued that a 2.88-trillion-parameter K3 cannot simply be run on a distressed person’s laptop; serious inference requires substantial hardware and infrastructure. That logic may hold inside China’s enforcement perimeter, but Nathan worries it understates external bio and cyber risk once weights reach jurisdictions Beijing cannot control.
The regulatory gap could narrow as Chinese capabilities catch up, because companies appear to expect standards to rise alongside models. Nathan puts the capability lag’s credible center near nine months and notes that Claude 4.5 Opus marked the point when agents really began working. Chinese labs are now releasing models where agents really work as well. His conditional forecast is that, as they encounter the failures now confronting OpenAI and Anthropic, Chinese regulators will tighten requirements and deployed safeguards will “significantly” converge.
China’s biggest conceptual absence may be alignment by character rather than compliance by rule. Chinese AIs told Nathan that Confucius’s descendants, reportedly 79 generations later, still identify as his descendants and perform rituals in his honor, inspiring him to ask whether a Confucian constitution could carry values through recursive AI generations. Yet researchers told him, “We’re all engineers”; the current ecosystem emphasizes explicit rules and reliable obedience, leaving wisdom-tradition-based alignment “pretty much greenfield.”
🔗 Original source & video: Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics