China's AI Upstarts: How Z.ai Builds, Benchmarks & Ships in Hours, from ChinaTalk
Summary
Z.AI has broken into the frontier conversation by deliberately pivoting from general chat to coding and agents. Nathan Labenz’s benchmark snapshot places GLM-4.6 at No. 19 on the LMArena text leaderboard—about 65 Elo points behind the leaders, still winning roughly two of five head-to-head comparisons—and No. 9 in web development, competitive with GPT-5.1. Zixuan Li says the company followed where users saw greater value: “Coding and agentic stuff are more useful,” while ordinary Z.AI Chat remains free.
Open weights are both a research contribution and Z.AI’s practical route around the fact that Western enterprises may not use Chinese-hosted APIs. Enterprises may run GLM through Fireworks, another provider, or their own chips even when they cannot use Z.AI’s API; without opening the model, Li says, “we’ll never have the opportunity to join this conversation.” DeepSeek-R1 demonstrated that visibility and commercial returns can coexist: “You need to expand the cake first and then take a bite of it.”
The monetization thesis does not require market leadership—a narrow slice of a rapidly growing coding market could be enough. Its GLM Coding Plan uses subscriptions to remove anxiety over agent loops that might consume a million tokens, creating stickier usage than metered APIs. Li’s illustration is blunt: “You don’t have to persuade 50% of people… Maybe you only need 5%,” and 5% of Claude Code users would already constitute “a huge market.”
Z.AI treats organizational speed as a competitive asset, releasing models only several hours after training and evaluation finish. GLM-4.5 combined three specialized teacher models—reasoning, agents, and coding—through a tightly coordinated 100-to-200-person core organization whose team leaders and founder still run experiments themselves. The operating instruction is “Get it fast,” sometimes giving integration partners only two or three hours’ warning because “the open-source itself is the biggest event.”
Li does not believe more data alone will carry the present architecture indefinitely: “There is a wall.” GLM-4.6 is a 355-billion-parameter model, but hypotheses must be tested on 9-billion- or 30-billion-parameter models because full-scale experimentation is impractical; he says roughly 90% of those experiments fail. Nominal million-token windows may work effectively only around 60K or 100K, so Z.AI is exploring new architecture, on-policy reinforcement learning, longer effective context, and multi-agent systems while compressing production context when necessary.
Chinese use cases create differentiated post-training priorities, particularly role-play, culturally fluent translation, and social-media interpretation. Z.AI’s role-play work trains models on long character instructions so they retain identity, emotion, and prescribed behavior. Translation work includes abbreviations, danmu, memes, and context-sensitive emoji—for example, turning a whale back into “DeepSeek” when the surrounding sentence is about AI. Li claims Chinese-English translation is on par with Gemini 2.5 Pro and says social applications demand nearly complete comprehension, not the “80% is enough” threshold of casual video translation.
Li portrays China’s near-term AI anxiety as centered more on employment than existential catastrophe, while Z.AI’s overseas economics remain heavily US-weighted. Developers experience replacement risk concretely through Claude Code and Codex, but Li says the broader public still sees hallucinating models as useful rather than terrifying: “We are not there yet.” India supplies the most overseas users, yet the US generates about 50% of overseas revenue because customers buy higher-priced Pro and Max plans; training stays in China while overseas services and data are hosted in Singapore.
Deep dive
1. Coding, not chat rankings, put GLM on the global map
Nathan Labenz’s benchmark snapshot places GLM-4.6 at No. 19 on the LMArena text leaderboard, next to Qwen 3 Max, Kimi K2 Thinking, and DeepSeek V3.2—the four leading Chinese open-weight models, though Mistral is not too far behind. Its roughly 65-point Elo deficit still implies winning two of five comparisons against the leaders.
Web development is the stronger showing: GLM-4.6 ranks No. 9, competitive with GPT-5.1 and meaningfully behind only Gemini 3 and Claude Opus 4.5. That performance explains why a company unfamiliar to many Western listeners suddenly became visible through coding products and integrations.
Li traces Z.AI, or Zhipu AI, back to 2019, when it pursued AGI through graph computing and built the scholar-mapping product AMiner. It shifted to large language models in 2020 and published GLM in 2021, one year before GPT-3.5; GLM-4.5 and GLM-4.6 finally made that long-running effort internationally legible.
The strategic break came after Z.AI ranked roughly sixth to ninth on Chatbot Arena in 2024. Seeing Manus and Claude Code gain traction in 2025, the company moved its top priority toward economically useful coding and tool use: “We need to follow the trend and also predict the future.”
2. A small, hands-on organization is designed around a unified model
GLM-4.5 emerged from three separate teacher models for reasoning, agents, and coding, subsequently distilled into one unified model. Li credits collaboration more than organizational novelty: pre-training and post-training teams “just sit next to each other,” working against the same target instead of optimizing isolated mandates.
Leadership cannot retreat into goal-setting because the frontier moves during the training run itself. “You need to feel the trend yourself,” Li argues; Z.AI’s founder reads papers and runs experiments, combining live results with competitor moves rather than asking subordinates to mediate the evidence.
The core research-and-engineering team numbers roughly 100 to 200 people, which Li considers sufficient and easier to keep focused. Inside larger companies, a domain team may still have only 10 to 20 core members, supported by perhaps 80 or 100 people handling training or data preparation.
Ongoing PhD students treat model training and academic work as compatible rather than competing tracks: building a unified agentic coding model might be “one of your greatest achievements ever.” Hiring rewards papers, GitHub work, competitions, GPU experience, and actual training; startups compensate for lower pay by seeking people with ambition to “fight together.”
3. Returnee credentials matter less than demonstrated work
Asked about claims that Chinese labs discriminate against overseas graduates, Li rejects the premise: firms want the best people, and returnees may simply perform worse under a particular interview standard. He joined Zhipu AI after studying at MIT, yet estimates that perhaps only 10 people at Z.AI know where he studied because “people don’t care.”
China’s larger companies—Baidu and Alibaba among them—usually select first because they can pay more. A startup instead needs self-directed employees attracted to young teams and the possibility of making something that “seems to come from nowhere” compete with established models.
Li’s own path reflects that market: other prominent AI companies rejected or ignored his applications amid heavy résumé volume. His MIT work on AI for science and alignment was not directly relevant to his initial domestic-chatbot strategy role, but it gave him a working picture of what OpenAI and Anthropic were attempting at the frontier.
4. Open weights function as distribution before they function as ideology
Li first presents openness as a contribution to research, placing Z.AI alongside Llama, Qwen, and Kimi. The commercial constraint is equally decisive: US organizations may not use Z.AI’s API, but they can still deploy GLM themselves, obtain it through Fireworks or Groq, or run it on their own chips.
Z.AI changed course after keeping its flagship model closed in 2024. DeepSeek-R1 showed that a company could become globally famous through open-sourcing while retaining API, partnership, and collaboration revenue: “You need to expand the cake first and then take a bite of it.”
Global technical acceptance also feeds domestic credibility. Chinese media rapidly recirculate the preferences of Andrej Karpathy, Sam Altman, Elon Musk, and other US technology figures; even Chinese enterprises evaluate a provider’s “global brand and global performance,” correcting Z.AI’s earlier assumption that domestic API sales alone would suffice.
That recognition remains incomplete. Z.AI tracked only about 20,000 followers on X versus roughly a million for DeepSeek, and Reddit users still asked where GLM had “come from.” Nathan Labenz pushed back that San Francisco circles now discuss GLM more than Mistral and arguably Llama, but Li maintained that branding and technical-community engagement lag model quality.
5. The business model layers services and subscriptions over open models
Chinese enterprise demand splits between self-hosters and API customers. Data-sensitive organizations ask integrators to deploy models such as DeepSeek on private chips, then add RAG, data storage, and workflows; technology and media companies accept APIs and choose primarily on performance and price, a market Li believes ByteDance currently dominates.
Qwen3-Max illustrates a hybrid strategy: some models are open while Qwen3-Max remains closed-source for API sales. Because Z.AI opened its foundation models, buyers continually ask why they should not host them themselves; the answer must be engineering rather than exclusivity: faster decoding, search, MCP capabilities, and other improvements around the same weights.
GLM Coding Plan turns that infrastructure into a subscription. Users need not calculate how every agent interaction consumes tokens—one Claude Code dialogue might use a million—and the subscription removes that metering concern. Li sees the resulting predictability as a source of stickiness rather than merely a different billing mechanism.
Labenz questioned whether GLM can gain paid adoption when Claude Code, Codex, Gemini, and basic Cursor are available to try. Li’s answer was market-share arithmetic, not dominance: converting even 5% of Claude Code users would be meaningful, with further niches available in agents, role-play, and potentially major platforms such as Meta.
6. Role-play and cultural translation require their own data strategy
Before GLM-4.6, GLM-4.5 was relatively weak at role-play because it had not been post-trained on the relevant data. Without long character specifications in its training, the model forgets who it is and falls back to generic conversation; with that data, it can retain instructions, emotion, and prescribed behavior across the interaction.
Nathan Labenz’s concrete analogy was a text RPG whose user supplies five pages of identity and history. Li offered a pop-cultural specimen: describe Stewie’s personality and backstory from Family Guy, and Z.AI can create the text persona; adding a suitable speech model could extend that performance into voice.
Translation is another deliberate specialty. Li claims GLM’s Chinese-English performance is on par with Gemini 2.5 Pro, including abbreviations, memes, and emoji: a whale in an AI sentence may mean DeepSeek, while the same symbol in an animal sentence should remain a whale. “We understand the culture” is the asserted advantage.
Z.AI cannot scrape private WeChat conversations, so it studies public Xiaohongshu, TikTok, and comment sections, where users are “very naughty,” plus danmu and image memes for vision training. The “TikTok refugees” episode increased translation demand; unlike YouTube, where 80% comprehension may suffice, apps such as X, Xiaohongshu, and WeChat require users to understand essentially every comment.
7. Safety debate follows capability, and job loss feels most concrete
Li had read OpenAI’s work on reducing unhealthy attachment—training ChatGPT to identify itself as AI rather than human—and said such issues are discussed internally when relevant. Yet Z.AI prioritizes capability while its models trail the best closed systems: “If we have a model that can perform like GPT-5, then we can move on to remove the addiction.”
His rationale is partly technical: as Z.AI changes data collection and post-training to chase capability, behavior can shift dramatically between versions, making interventions built for an older model obsolete. The claim is not that attachment risk is imaginary, but that the company believes performance remains the binding problem.
Software developers and data analysts feel fear most directly because coding agents can already complete concrete tasks, especially for junior developers. Writers and managers have long used SaaS and other assistance, so another brainstorming or polishing tool does not yet feel as substitutive; in Li’s account, job loss is the clearest concern he hears.
Labenz highlighted the contrast with America’s vocal minority focused on risks beyond employment, a culture Li encountered at MIT. Li estimates perhaps a million people closely follow frontier developments while a billion continue daily work largely unaffected: “The more you learn, the more fear you will have.”
8. Global inference, mixed chips, and local search shape deployment
Z.AI trains in China but serves overseas users through Singapore, where its international services are hosted. Li says overseas data residency is a strict requirement and the privacy policy is revised almost monthly; he explains that co-locating the GPU, CPU, and database in Singapore avoids the latency of routing requests back through mainland China.
Blackwell is attractive not only for the chip but for FP4, which could reduce costs substantially. Z.AI nevertheless matches GLM-4.6, the upcoming GLM-4.6 Air, and older models to domestic and NVIDIA hardware according to requirements: one customer may need 30 tokens per second, another 80.
India provides the largest overseas user count, with meaningful demand also from Indonesia, Norway, and Brazil. Discovery is driven mainly by X, Reddit, and some YouTube, which biases the mix; the US nevertheless generates about 50% of overseas revenue because its customers choose Pro and Max plans rather than the Lite plan.
Irene Jiang asked how search works when Chinese platforms are walled gardens. Li linked the problem to the US too, noting that Google lacks a search API and Bing is trying to stop its API. Z.AI can aggregate multiple resources or let an agent log into platforms and browse pages itself, preserving source access that a generic API may omit.
9. The next gains require architecture experiments—and most fail
Z.AI is exploring on-policy reinforcement learning after becoming relatively mature in off-policy RL, along with multi-agent systems. Its current product is effectively one GLM-4.6 actor that searches repeatedly and can generate slides, presentations, or posters while retaining the preceding context.
Multi-agent designs might improve speed or performance, but orchestration introduces sharp trade-offs. If several agents receive the same context, they may duplicate one another or fail to work together; one hallucinating agent can contaminate the entire research result.
Advertised context length is not effective context length: a model may claim one million tokens yet perform well only within 60K or 100K. Z.AI can compress customer workloads to 60K or 30K because most users do not need a million, but Li does not treat that engineering workaround as the underlying solution.
His categorical assessment is that “there is a wall” that data alone cannot cross. GLM-4.6 has 355 billion parameters, so Z.AI tests hypotheses on 9-billion- or 30-billion-parameter models; “90% of the time we just fail.” Progress will require better architecture, pre-training and post-training data, and potentially a new framework.
10. Shipping within hours makes release a coordination challenge
Li previewed a 30-billion-parameter next-generation model, calling it GLM-4.6 Air but adding that the name might be Mini. He expected GLM-4.6 Air, GLM-4.6 Mini, and a GLM-4.6 Vision model to be available by the time the podcast launched. He said smaller-model experiments planned for the next generation would not be put into practice in 2026, but would provide ideas for future training.
Release cadence is startlingly literal: training ends, evaluation runs, and the model can ship “several hours” later. Z.AI does not first send the endpoint to LMArena or Artificial Analysis for evaluation or run a media-buzz campaign; “if you want to open-source the model, the open-sourcing itself is the biggest event.”
Li would prefer roughly a week to coordinate inference providers, benchmarkers, and coding-agent partners. Instead, he may tell them, “We have a new model coming in two hours, maybe three hours, maybe you are sleeping,” then amplify their integrations after launch.
His global role can consume 18-hour days because US partner meetings occur at 2 or 3 a.m., but he rejects that schedule for researchers: a brain may produce only eight hours of serious paper-reading, experimentation, and coding. Z.AI still needs model quality and recognition to stand out; without a solid model, “only the most famous one gets all the attention.”