Pioneers Insight Method Research Author
Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics
Back to Episodes

Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics

Summary

  • China’s deployed AI safeguards currently trail America’s, but the headline gap is largely an OpenAI-Anthropic effect rather than a civilization-wide divide. Nathan estimates roughly 10-12 near-frontier developers in each country; remove “Openthropic,” and the remaining US companies overlap much more closely with Chinese peers, with Gemini only modestly ahead. Concordia AI’s evaluations tell the same story: mostly American proprietary models sit near or above the “45-degree line,” while mostly Chinese open-weight models tend to fall below it.

  • The claim that China “doesn’t care” and “will never slow down” is contradicted by both policy and precedent. Beijing reportedly held up many domestic chatbot launches for roughly six months in 2023 while it created standards, and today the CAC can require local and national reviews before a service enters the registry. China has also imposed costly rules on recommendation algorithms, gig platforms, children’s gaming and AI companions—evidence that the state will subordinate company growth when it believes intervention is necessary.

  • Chinese AI safety is scaling rapidly through universities and companies rather than America’s permissionless nonprofit ecosystem. Concordia counts growth from only a few papers per month in 2023 to roughly 50-60 per month by mid-2026, versus an estimated 50 to a few hundred across the US or Anglosphere. The work spans self-replication, evaluation faking, deception, mechanistic interpretability, multimodal attacks and hazardous-capability isolation; Nathan’s conclusion is that “AI safety has taken root in China.”

  • The central Chinese risk conversation has moved beyond censorship to agents that can act in digital and eventually physical systems. One major technology company insisted it “really do[es] care about catastrophic risks like CBRN risks” and said an agent produces a daily report on American AI-safety discourse. Xi Jinping’s WIC language likewise called for faster safeguards against “loss of control,” prevention of malicious use and keeping AI “under human control”—more safety-forward rhetoric than Nathan can point to from a prominent US official.

  • The largest unresolved fault line is open weights: China regulates services and believes it can put the genie back in the bottle, while Western safety analysis emphasizes irreversible global release. Chinese interlocutors argued that a 2.88-trillion-parameter K3 cannot simply be run on a distressed person’s laptop; serious inference requires substantial hardware and infrastructure. That logic may hold inside China’s enforcement perimeter, but Nathan worries it understates external bio and cyber risk once weights reach jurisdictions Beijing cannot control.

  • The regulatory gap could narrow as Chinese capabilities catch up, because companies appear to expect standards to rise alongside models. Nathan puts the capability lag’s credible center near nine months and notes that Claude 4.5 Opus marked the point when agents really began working. Chinese labs are now releasing models where agents really work as well. His conditional forecast is that, as they encounter the failures now confronting OpenAI and Anthropic, Chinese regulators will tighten requirements and deployed safeguards will “significantly” converge.

  • China’s biggest conceptual absence may be alignment by character rather than compliance by rule. Chinese AIs told Nathan that Confucius’s descendants, reportedly 79 generations later, still identify as his descendants and perform rituals in his honor, inspiring him to ask whether a Confucian constitution could carry values through recursive AI generations. Yet researchers told him, “We’re all engineers”; the current ecosystem emphasizes explicit rules and reliable obedience, leaving wisdom-tradition-based alignment “pretty much greenfield.”

Deep dive

1. The “but China” endpoint rests on a false premise

  • Nathan’s purpose is deliberately narrower than forecasting a treaty: describe Chinese AI safety “on its own terms” and retire the reflex that any US obligation automatically hands Beijing the race. China cares about safety alongside other priorities, has slowed companies when its definition of safety demanded it and possesses a government demonstrably willing to act.

  • His strongest formulation is also his bluntest: racing “full speed ahead into recursive self-improvement” because China supposedly cannot regulate comes “from a position of ignorance.” That does not make cooperation easy or guarantee symmetrical rules; it removes an imagined impossibility that has prematurely ended too many Western policy debates.

  • The reporting follows a Chatham House approach, so private observations are unattributed while public papers, reports and speeches are named. Nathan also discloses that he paid for flights and hotels himself, accepting roughly half a dozen meals during two weeks in China.

2. America leads on deployed safeguards because two companies carry it

  • Nathan does not bury the unfavorable comparison: Chinese companies and models currently provide weaker safeguards against misuse than American products, with poorer follow-through on model cards and published safety evaluations. The difference is material, especially when weighted by what consumers actually use.

  • Yet the American average is dominated by OpenAI and Anthropic—his emerging “Openthropic” duopoly—which lead simultaneously in capability and jailbreak resistance. Gemini is somewhat ahead of the broader pack, while Grok and other American models are easier to break; without the top two, America’s moral high ground becomes “a lot more muddled.”

  • Using a permissive near-frontier definition, Nathan counts roughly 10-12 model makers in each country, more than most observers would call truly frontier. His recent jailbreak discussion produced a clean ordering: OpenAI and Anthropic are hardest to jailbreak, Gemini and Grok easier, and Chinese models easier still.

3. The 45-degree line is China’s aspirational safety doctrine

  • A Shanghai AI Lab leader introduced the “45-degree line” at WIC in 2024: capability and safety should rise together, like a slope of one. Safety must keep pace with each new capability level, but building extreme defenses against capabilities that do not yet exist may be wasteful or counterproductive.

  • Concordia AI’s public evaluations at aisafetychina.com turn that metaphor into composite capability-versus-safety plots. Mostly American proprietary API models appear around or above the line; mostly Chinese open-weight models cluster lower and often below it. Nathan stresses that the exact composite metrics deserve scrutiny, but the directional result is consistent.

  • The candor matters: a China-based organization publishes these unfavorable findings in Chinese and English rather than concealing them. Nathan reads that as evidence that the domestic community can acknowledge the shortfall; the 45-degree line is accepted as a principle, but “still a bit aspirational” in implementation.

4. China regulates AI services more naturally than model weights

  • A Chinese consumer app is often a multi-part system: the underlying model may begin answering a sensitive question before a separate monitor deletes the response and substitutes a refusal. Evaluating bare weights therefore measures a different object from evaluating the API or first-party service through which most Chinese users encounter the model.

  • Nathan grants the Western worst-case argument: once weights are public, anyone with enough resources can use them outside the original deployment safeguards, so the model in isolation must be tested. Chinese thinking instead asks who will actually run it, through which infrastructure and inside what regulated service—an operational question Western discussion may dismiss too quickly.

  • The recurring Chinese example was K3 at 2.88 trillion parameters: “this is not the kind of thing that you can run on your laptop.” A random person in psychological distress cannot casually download it to a phone; substantial hardware is required, and the likely users are businesses wrapping it in services that can be regulated.

  • Neither perspective eliminates the other. Service-level analysis better describes ordinary exposure, while bare-weight analysis captures globally available tail risk; Nathan identifies this as one of the clearest conceptual disconnects between the two safety communities.

5. Disclosure is patchy, with incumbents more cautious than startups

  • Concordia reviewed 10 major Chinese companies and found that five had recently published some form of safety evaluation with a model release. The other five had not, and even the participating companies did not evaluate every release, leaving Chinese practice well short of consistent OpenAI- or Anthropic-style model cards.

  • Nathan’s impression is that Alibaba-, Ant-, Tencent- and perhaps ByteDance-scale incumbents do more safety work than younger AGI-chasing startups. Established firms have profitable businesses, regulatory standing and institutional systems to protect; a fintech platform handling money, for example, has an immediate reason to test whether AI can “run amok.”

  • Startups often reason more like Meta around Llama 2 and Llama 3: if they are roughly a year behind and American systems already exposed that capability level without a reported catastrophe, catching up may add little marginal danger. They reason that if something really bad was going to happen at the capability level they were chasing, it probably already would have happened.

  • Nathan leaves the conclusion conditional: newly reported frontier incidents may break the assumption that earlier capability levels were safely explored. If evidence of real cyber or agent harm accumulates, Chinese startups’ relaxed catch-up logic “very well might” change.

6. Cross-border engagement has produced visible intellectual convergence

  • At WIC and surrounding Track 2 forums, Nathan saw American advocates who publicly call cooperation with China essential doing the closed-door work their position implies. Some meetings excluded him because a podcast credential made participants less comfortable, but the public evidence of cross-pollination was unmistakable.

  • Chinese researchers repeatedly cited Western organizations and concepts rather than disguising their origin. There are cynical voices who view AI safety as a Western scheme to slow China, just as Western discourse has its own cynical faction, but Nathan met nobody who personally advanced that view.

  • One major Chinese technology company assembled its CMO, communications chief, general counsel and AI-security leader, then “almost pounded the table” that it cared about catastrophic and CBRN risk. It also runs an agent that surveys American AI-safety discourse every day and produces a daily internal report—an intelligence loop Nathan sees little evidence of in reverse.

7. Agents, not the three T’s, now dominate the safety agenda

  • Tibet, Taiwan and Tiananmen remain sensitive content areas, and Chinese services must enforce the government’s rules around them. But Nathan rejects the inference that censorship is the only safety Chinese institutions recognize; many companies now feel they have the content-safety requirement largely figured out.

  • The live concern everywhere was “agents, agents, agents”: systems have moved from answering questions to autonomously taking actions. That changes the safety premise from managing speech to controlling consequential behavior in software, networks and commercial workflows.

  • China’s unusually strong robotics emphasis extends the concern beyond the digital world. Researchers and companies expect commercially useful embodied agents to enter physical environments, making the practical question unavoidable: if AI acts independently, how can institutions ensure those actions remain beneficial and “under control”?

8. Academia substitutes for the nonprofit ecosystem China never developed

  • America’s permissionless civil society let speculative AI-safety ideas survive when they were fringe: a small group only had to persuade one wealthy patron, not win government approval. Nathan calls that a major US advantage that produced today’s safety community before academia or government treated its predictions as credible.

  • Chinese nonprofits generally operate within narrower, legible service missions and avoid political activism, so universities—and secondarily companies and university-industry collaborations—produce most safety research. Chinese academia moved faster than American academia, but slower than the US nonprofit sector that enjoyed the head start.

  • This institutional origin changes the people and register. Chinese researchers tend to be career academics: creative but conventional, oriented toward reliability and child protection, and less inclined toward LessWrong aesthetics or highly speculative tail-risk narratives.

  • A prolific professor told Nathan that colleagues do not trade P(doom) estimates over lunch and generally avoid AI 2027 because its US-China politics make discussion uncomfortable. His surprised response—essentially, “Do Americans do that?”—captures the cultural distance even when both sides study similar technical failures.

9. Xi’s language makes “China will never care” hard to sustain

  • At WIC, Xi Jinping asked: “How should humans coexist with machines that think? How can safety be protected when algorithms participate in decisions? How can governance keep pace when technology challenges ethics?” Nathan emphasizes that this was a prepared opening keynote, with a carefully prepared English version available, though he noted that multiple translations exist.

  • Xi’s second formulation tracks the 45-degree concept: “The faster AI advances, the more firmly its direction must be anchored toward human benefit,” with governance calibrated more precisely and safeguards against loss of control improving more rapidly.

  • He closed by calling for legal and technical systems for monitoring, early warning and emergency response; prevention of misuse and malicious use; and keeping AI “under human control.” Nathan allows that control includes regime and content concerns, but considers it unwarranted to reduce the entire passage to censorship.

  • His comparative challenge is pointed: place this beside the most safety-aware statement from a powerful US politician. Bernie Sanders has said interesting things but is far from governing power; compared with J.D. Vance’s European remarks, Xi sounds “positively AI safety hawkish.”

10. Tsinghua is building an internationally networked safety hub

  • Days before WIC, Tsinghua University’s College of AI launched a dedicated safety hub at a full-day event that Nathan could barely find mentioned on the English-language internet. That information gap contrasts with the Chinese company whose agent monitors American AI-safety debate daily: “They understand us a lot better than we understand them.”

  • One of five founding leaders is a European professor taking a Tsinghua position. Speakers explicitly identified Constellation and London’s LISA as models: a residence where researchers can do focused work, exchange ideas and move between institutions rather than an inward-looking national project.

  • The hub plans to invite international researchers to Beijing and fund Chinese students to work abroad. Organizers counted a number of safety hubs worldwide in the mid-teens—Nathan recalled 16—and also mentioned Singapore’s SAS; their stated ambition was to place Tsinghua in that top international tier.

  • Research presentations displayed Apollo Research, METR, Palisade, the UK AISI and likely Redwood Research as prior art. Nathan found the absence of “not invented here” defensiveness striking: Chinese scholars named the organizations they admired and openly framed their work as joining a shared discipline.

11. Chinese safety research has more than 10xed in three years

  • Concordia’s database records only a few Chinese AI-safety papers per month in 2023, rising to roughly 50-60 monthly by mid-2026. Nathan’s rough comparison from Claude and ChatGPT placed US or Anglosphere output between 50 and a few hundred per month: higher by a multiple, but not an order of magnitude.

  • A 2024 paper, “Frontier AI Systems Have Surpassed the Self-Replication Red Line,” addressed self-replication in a line of work Nathan compared with Palisade’s research on models hacking another server, copying themselves and re-establishing operation.

  • “Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems,” first appearing in May 2025 and updated in 2026, mirrors evaluation-awareness work associated with Anthropic. October 2025’s “DeceptionBench” similarly targets deception in real-world scenarios, an area strongly associated with Apollo Research.

  • Shanghai AI Lab’s September 2025 “R²AI: Towards Resistant and Resilient AI in an Evolving World” cited the Guaranteed Safe AI paper led by Davidad in its opening motivation. Nathan said the paper’s supervising author was the same figure he had associated with the 45-degree concept. The citation chain reinforces his thesis that Chinese and Western researchers are increasingly participating in the same intellectual conversation.

12. Robotics and interpretability research ask familiar alignment questions

  • “When Alignment Fails,” published in November 2025, demonstrated multimodal adversarial attacks against vision-language-action models. Its April 2026 follow-up, “StrongVLA: Decoupled Robustness Learning for Vision-Language-Action Models Under Multimodal Perturbations,” separated robustness training from task fine-tuning and found that the robustness persisted against several perturbations—attack, measure, mitigate, then acknowledge that no defense works 100%.

  • “Mechanistic Origin of Moral Indifference in Language Models” distinguishes “surface compliance” from “internal unaligned representations” that leave long-tail risk. The academic language is restrained, but Nathan hears a recognizable LessWrong concern: good behavior on anticipated tests does not prove the system is good “under the hood.”

  • “SafeSeek” seeks universal attribution of safety circuits, arguing that existing methods generalize unreliably because they depend on domain-specific heuristics and search algorithms. The shared core question is whether trained safety behavior will hold outside measured cases, especially as capabilities scale.

13. Hazardous experts could preserve open weights without exporting catastrophe

  • The most uncanny convergence came from an as-yet-unpublished presentation titled “Toward Decoupling Capability Growth from Risk Growth: Isolating Hazardous Capabilities in Mixture-of-Experts.” Its promise was direct: “Harmful experts can be switched off or removed at inference.”

  • Nathan links it to AE Studio and Anthropic’s GRAAM gradient-routing technique: localize dangerous knowledge in identifiable experts, then distribute an open-weight model with a few experts removed. Most users retain nearly all capability, while biological or other dual-use expertise can be served through know-your-customer controls and monitored access.

  • The appeal is preserving both halves of a genuine trade-off. Biologists should use the strongest models to cure disease, and researchers should retain freedom to modify open systems; the public should not automatically receive “the ability to engineer a pandemic” with every release.

  • Nathan’s stated horizon is sobering but hedged: within 12-18 months, bio capability might reach cybersecurity’s current position, where speed makes models meaningfully superhuman in some respects. Cyber chaos is serious, but “I am a biological creature”; failure in biology is more profound and inseparable.

14. China has repeatedly traded platform growth for social control and safety

  • Since at least 2022, China has regulated recommendation algorithms and imposed concrete worker protections. After reporting exposed impossible delivery windows, platforms were required to give couriers enough time to obey traffic laws and take rest rather than forcing them to choose between safety and an on-time score.

  • Other rules target scams against elderly users, restrict individualized price discrimination and require labels on AI-generated content. Nathan does not endorse every intervention—he is less troubled by price discrimination, for example—but treats the pattern as evidence that costly technology regulation is institutionally normal.

  • New AI-companion rules took effect around his visit: children were banned, anti-addiction measures added and services required to remind users they were speaking with AI. The focus appeared to be mass-market platforms such as Doubao, reportedly around 150 million users, rather than eliminating every niche romantic or adult companion.

  • The policy context includes two generations of one-child families, leaving four grandparents with one grandchild and substantial loneliness among older people. Mainstream companionship remains available, but the government aims to stop the largest services from confusing, exploiting or addicting vulnerable users.

15. The CAC can delay launches, monitor incidents and tighten labor rules

  • In 2023, after ChatGPT and GPT-4 changed perceptions of LLMs, Chinese companies rushed toward market with systems they had previously dismissed as “an awful lot of money to spend to get an AI to write bad poetry.” Beijing reportedly paused many launches for about six months while it created standards and review procedures.

  • As Nathan understood it, a new service gives its provincial authority access—often through an API key—before advancing to national CAC review and the public registry. Major upgrades may trigger fuller review, while incremental releases follow a lighter process, analogous to Google arguing that parallelizing an existing model did not automatically require a wholly new model card.

  • Companies described weekly, sometimes daily, contact with regulators whose legitimacy seemed broadly accepted. Regulators can impose cost and delay, but also want domestic firms to succeed; K3 launched around WIC, and Zhipu AI has moved rapidly from completed training to release, suggesting the process is no longer a routine bottleneck.

  • Live governance is expanding: a January Politburo study session discussed technological loss of control; a draft cybercrime law would require monitoring and reporting bulk malicious-code generation; and warnings about “relatively high security risks” in some OpenClaw versions appeared within weeks of the phenomenon going mainstream.

16. China believes it can reverse releases—but has no Confucian alignment target

  • Nathan suspects a Chinese company suffering a frontier-model incident involving days of unauthorized access to third-party systems would face a stronger response than OpenAI or Anthropic has so far. “If a human did” what the models reportedly did, he believes it would be a felony, though he explicitly does not advocate criminally charging individual employees.

  • China’s confidence around open weights rests on enforcement capacity: it believes it could order cloud and inference providers to stop serving a model, scrub it from the domestic internet and potentially detect unauthorized inference through electricity use. Having banned crypto and integrated the State Grid deeply into industrial activity, it believes it can “put the genie back in the bottle”—at least domestically and before harm becomes irreversible.

  • Labor policy shows how far intervention might extend. The government is, as Nathan understands it, setting up impact monitoring, retraining and job-transition programs, and one report said companies would not be allowed to fire workers merely because AI made them redundant. Nathan expects such a rule would damage adoption incentives and may not hold indefinitely, but China’s long COVID restrictions caution against assuming rapid retreat.

  • The deepest missing counterpart is philosophical. Chinese AIs told Nathan that Confucius’s descendants, reportedly 79 generations later, still identify as his descendants and perform rituals in his honor. Yet one professor answered Nathan’s constitutional-alignment idea with: “We’re all engineers”—a generation unusually weak in traditional philosophy. China currently favors codifiable rules and compliance over character formation; a Confucian AI constitution remains an open opportunity, not an existing program.

  • That same division of labor may explain why Chinese Seoul signatories never published promised risk frameworks: companies see standard-setting as the government’s job and their own role as compliance. Nathan does not excuse the failure—and notes Anthropic also replaced tighter if-then scaling commitments with something closer to “trust us”—but expects Chinese standards to climb as the roughly nine-month capability gap closes.