Investor Meng Xing on open source impacts, Chinese AI personalities, and the future of RL data
Investor Meng Xing on open source impacts, Chinese AI personalities, and the future of RL data
Summary
- Kimi K3 convinced Xing Meng that China’s open-source labs are now frontier-adjacent, not fast followers. Kimi remains a sub-400-person, idealistic lab focused on AGI. It integrated fundamental architecture innovations, including KDA, some tested only on 30–40B toy models, into a 2.8-trillion-parameter model on constrained compute. Meng recalls that K2, launched about a year ago, would have ranked sixth or seventh on the closed-model list; K3 is getting close to second or third. He expects Kimi and DeepSeek to alternate the number-one spot in open-source model rankings over time.
- The reported tradeable impact of Chinese open source is on token revenue, not frontier researchers’ attention. Frontier-lab researchers have not cared much about these models, and few have tried them, while many agent developers have switched to K3. Meng heard that this significantly affected frontier labs’ top-line revenue, especially Anthropic’s: July token consumption grew more slowly than expected—some say it stagnated—and if consumption is flat while users swap expensive models for cheaper ones, revenue falls. Users may combine models, such as using K3 or DeepSeek to write code and Claude to review it. Even among Anthropic’s Fable 5, Opus 5, Opus 4.8, and sometimes Opus 4.6, users cannot always decide which is better for a task.
- ByteDance’s reported decision not to distill frontier models is brave and suggestive, not simply an admission. Meng found the reported condition—accepting that Seed could fall behind Chinese peers—especially significant: if a company refuses distillation, it may already be behind or risk falling behind. He distinguishes blatant distillation from indirect distillation through synthetic data, arguing that everyone distills something to some degree. He says Seed and DeepSeek can afford to be behind for a few years because they do not have to raise funding, preserving the ambition to reach the top by doing things “the right way.”
- ByteDance’s Seedance emerged from a risky scale bet with little prior evidence that it would work. Meng says the model was scaled up by an order of magnitude over earlier multimedia models, requiring tens of thousands rather than thousands of GPUs, at a time when compute was increasingly needed for LLMs and coding. He credits a very young female ByteDance researcher. Revenue appears to be driven mainly by prosumers: video-canvas services such as Topview and LiblibAI reportedly grew from small or mid-single-digit millions to almost 100 million or more after Seedance’s release, with some buyers prepaying for API tokens. The host said short films had surpassed movie theaters in China this year; Meng accepted that comparison after clarifying it included both human-acted and AI short films.
- The new-lab concept is changing: Meng sees Cursor’s recipe as increasingly important. A vertical application with unique, “untouchable” user-trace data can post-train a model such as Composer and potentially outperform many pretrained models. The advantage is not a new methodology but access to data frontier labs lack. Vertical applications such as legal or medical tools may have a better chance of building a distinctive model than general-purpose apps. The deeper problem is that researchers’ data choices reflect their familiarity with coding and math, while the real test environments should be the working software people use—Slack, Zoom, Salesforce, and similar systems.
- The future of training data is environments, with difficulty ranging from stable optimization tasks to partially observed and dynamic worlds. Meng describes static kernel-optimization environments, Slack-like partial observation, and finance or social-media environments where every action changes the next state. Vending-Bench extends the idea toward the physical world by having models run virtual vending machines; its creators also set up a real shop or market nearby to test an LLM in a less predictable setting.
- Data provision is booming through two models. Headcount-based providers such as Scale and Mercor price recruited labor like consulting, while newer providers sell environments ranging from a few hundred dollars to roughly $100,000 or more; a Booking.com task might cost about $1,000 or less, while recreating much of an AWS-like DevOps environment could approach $500,000. Meng says the former currently has higher revenue, while the anticipatory researcher-led model has faster growth. At a recent conference in South Korea, young researchers told him their top aspiration was to build data-provider startups.
- AI for science’s binding constraint is verification, not simply modeling. AI-scientist agents can read literature, design experiments, sometimes control lab equipment, analyze results, and write papers, but scientific databases, tools, workflows, and wet-lab execution remain difficult. Protein-design models face severe data and multimodality problems, while virtual-cell models aim to test biological interactions without costly experiments. The true verifier is whether a drug works in a human body and cures disease, but full feedback may not arrive until an FDA-approved clinical-trial stage roughly five or six years into the process.
- Meng’s stated misconception in Silicon Valley is that China’s AI progress is entirely a product of government support. He argues that the US ecosystem has not given enough credit to the individuals, entrepreneurs, and researchers producing this work under difficult constraints, and believes they could do it anywhere if they can do it in China.
Deep dive
1. Kimi is a sub-400-person idealistic lab that just shipped a near-frontier model
- Meng’s portrait: “a hub of really smart young talent gathered together to achieve a very idealistic goal” — AGI, held through three years of peak media attention, “horrible times,” and a return to the peak, without diluting into applications or multimedia. Interns “almost in high school” get real compute and built important components of K3.
- The headcount is deliberate: Xing said Kimi is still fewer than 400 people; the host estimated “300 or so” after three years, not for lack of resources but so the team stays intact and avoids context loss.
- Did K3 surprise him? Yes: Xing recalled that K2, launched about a year ago, would have ranked sixth or seventh on the closed-model list; K3 is near second or third while open. Integrating fundamental architecture innovations — some validated only on 30–40B “toy models” — into a 2.8-trillion-parameter run would normally drag short-term performance down; they did it on limited compute and a short timetable. “I would be lying if I said I wasn’t surprised.”
2. DeepSeek’s founder, scarcity-bred innovation, and why each Chinese lab has a character
- On DeepSeek’s founder, whom he has never met: “He seems to have figured out the world at a very early stage” — how to become rich, make money, trade, and hire talent — but does not seem very interested in wealth or influence and does not capitalize on his ability to build them. The lab hires raw young talent out of Chinese schools rather than high-level, well-known researchers from top US labs; VC friends photographing office floors at 9–10 p.m. found DeepSeek and Kimi to be two of the hardest-working firms by the share of people still working. “A lot of this is really voluntary.”
- The status inversion he keeps encountering: infrastructure people at OpenAI, Anthropic, and Gemini “look up to the young guys at DeepSeek,” including engineers born in 1998 or later. The mechanism is constraint: “If you make the resources scarce enough, magic will automatically happen” when enough great people are put together. DeepSeek also releases unusually detailed technical papers and supporting work; researchers at Google and elsewhere say, “I never knew this could be done this way.”
- Kimi Delta Attention and DeepSeek Sparse Attention differ technically but share some principles. Features appear to leapfrog from one release to the next. Unlike the US, where senior researchers can move between firms frequently, Chinese researchers less often move between Chinese labs; each lab therefore develops a distinct character shaped by its founder’s beliefs and the people attracted by them.
3. Z.ai and MiniMax: two listed labs, different characters
- The host introduced Z.ai by pointing to the strong coding adoption of GLM-5.2. Xing said Z.ai is doing good work, has a larger team built around Professor Tang Jie and his students, and has expanded into computer use, multimodality, and other product suites. Its market value, he asked, had reached 1.23 trillion RMB — nearly $200 billion — despite very small 2025 revenue. It is relatively less focused than Moonshot and DeepSeek, more like a full-fledged AI company in the style of OpenAI.
- MiniMax’s stock has been “a roller coaster”: it rose alongside Z.ai, then fell significantly, leaving it around or slightly above its IPO price and roughly five or six times below its peak. It initially built well-known companion products, similar to Character.AI, and helped define how Chinese companies built such products. It later shifted toward multimodality models, including MiniMax M1 and the Hailuo series, but Xing says it has been less successful on the pure LLM side of reasoning and coding.
- The talent tell from early 2024: many of the “cool people” who might otherwise have joined ByteDance or TikTok joined MiniMax, including product managers, developers, researchers, and go-to-market staff. “The vibe is sort of like TikTok.”
4. ByteDance’s no-distillation decision: brave, suggestive, and affordable for only some firms
- Zhang Yiming reportedly told staff that Seed would refrain from distilling frontier models “even if it means we’re going to be behind our Chinese peers.” Meng fixated on the conditional: “If you don’t distill, that means you’re going to be behind or you’re already behind. That’s actually the bigger news for me.”
- His taxonomy: blatant distillation sends a prompt distribution to frontier models and trains on their returns, including Claude’s chain of thought where available. Indirect distillation occurs through synthetic data: because models generate much of that data, “by default the synthetic data distills the model which you used to create it.” Everyone therefore distills something to some degree. Meng says both synthetic data and blatant distillation have proved effective, while adding that he has no proof about who is doing what.
- Why refuse it: the shortcut “will give us a false sense of satisfaction, because we achieved this in the wrong way.” When the host asked whether anyone could afford to be behind for a few years, Meng answered: “If you don’t have to raise funding, yes.” He gave Seed and DeepSeek as examples with that luxury.
5. Open source’s reported bite is the token economy, not frontier researchers’ attention
- Meng’s surprise from the trip: despite the K3 debate and attention to open source, frontier-lab researchers “haven’t cared much about this”; few had tried the models. Agent developers, by contrast, had switched to K3 in substantial numbers, particularly according to news coverage.
- Xing heard that this had a significant effect on the top-line revenue of frontier labs, especially Anthropic. July total token consumption was growing more slowly than expected, and some people said it had stagnated. If consumption were flat, he said, swapping expensive models for cheaper ones would reduce revenue. The hybrid pattern is to use K3 or DeepSeek to write code while retaining Claude to review it.
- Differentiation is narrowing for ordinary tasks: “For the most part, I don’t think you’ll be able to tell the difference among the top models that easily.” Even among Anthropic’s Fable 5, Opus 5, Opus 4.8, and sometimes Opus 4.6, users cannot always decide which is better for a particular task. K3 is not especially cheap, but it is cheaper than comparable models.
- Why non-coding tokens stall: company-led AI adoption can take two to three years because of compliance, security, and privacy work. White-collar tasks also lack easily predefined goals — “it’s very hard to set this before you see the PowerPoint” — so humans remain in the loop and throughput is capped. His proposed lever is speed: creating a PowerPoint in 10 seconds rather than 10 minutes would allow more iterations. “This is not yet happening.”
6. Seedance’s risky scale bet built a prosumer short-film economy
- Looking back to 2023, Meng says the obvious guess for the best future multimedia model would have been ByteDance, Kuaishou, or Meta, given their data, distribution, platforms, and GPUs. Kuaishou’s Kling reached number one before being surpassed by ByteDance’s Seedance. ByteDance scaled Seedance up by an order of magnitude over prior multimedia models, requiring tens of thousands rather than thousands of GPUs, when there was little evidence the approach would work. He credits a very young female researcher at ByteDance. Meta, despite having relevant data and capabilities, has not yet shown the progress he expected.
- Revenue is mainly prosumer-driven, not individual-consumer-driven: video-canvas services such as Topview and LiblibAI reportedly went from small or mid-single-digit millions in revenue early in the year to almost 100 million or more after Seedance’s release. At one point, buyers had to prepay to secure Seedance API tokens.
- The host said short films had surpassed movie theaters in China this year. Xing first clarified that this included both human-acted and AI short films, then said it made sense. He said AI-generated short films had previously been strongest in science fiction, traditional Chinese science fiction, and ancient-swordsman genres, where special effects mattered more than subtle human emotion. Seedance’s demos showed smoother emotional transitions, such as laughing into crying, and more accurate high-dynamic motion, such as kung-fu fights without hands passing through arms. As a result, it had become almost the exclusive provider for many short films.
- On interactive media between television and games, Xing revised his earlier optimism: “I have to say that I’m a little disappointed.” The host cited Detroit: Become Human as an example of the half-movie/half-game format; Xing said such content has produced individual hits but has not yet reached high enough quality or popularity to take off. He still hopes it will happen and says, “I think this will be bigger than games themselves.”
7. New labs are being redefined: the Cursor recipe and the researcher-taste problem
- The earlier new-lab concept centered on young researchers or experienced researchers with academic backgrounds. Xing now thinks more about Cursor: an application-focused company with a super app and user traces that trained Composer through post-training rather than pretraining, allowing it to beat many pretrained models. “It’s not a new methodology — it’s a new set of data that the frontier labs don’t have.”
- Criteria: the data must be unique and “untouchable,” and a vertical is better suited than a general market. A legal or medical application could potentially tune a model to a distinctive task and outperform general models, whereas general-agent prompts resemble data already used to train Claude and OpenAI models. Xing cited Harvey and other vertical application agents in the US and expects similar development in China.
- The deeper argument is that data choices are “attributed to the key researchers’ own taste.” A fresh PhD may know coding and math but have “probably never done much PowerPoint work or accounting work.” The true measure of a model should be the working environments people actually use — Slack, Zoom, Salesforce, and other SaaS systems — rather than an imagined environment shaped by researchers’ preferences. Cursor came early partly because coding is familiar to researchers.
8. The future of training data is environments — graded by how much the world fights back
- Data has moved from static material, whether crawled or synthetic, toward environments where models attempt tasks and learn from rewards. Xing’s taxonomy ranges from static kernel/CUDA optimization — “you can rerun this a million times and it won’t change” — to partial observation, such as a new employee in Slack “trying to figure out what the boss wants,” and dynamic settings such as trading or social media, where every action changes the next state. A social-media post changes followers’ perceptions, so the same experiment cannot simply be repeated 1,000 times.
- Vending-Bench tests the virtual-to-physical transition by having an LLM operate three virtual vending machines: choosing suppliers, setting prices, and deciding on marketing. The creators told Xing they had also set up a nearby market or shop to test whether a model could manage a physical-world business with unpredictable conditions. “Those data are not on the internet… they have to be made on the spot.”
9. Data providers are booming on two very different business models
- Revenue has surged in the past six months, with many providers reaching $2–3 billion in ARR at that scale; Chinese providers have not reached that scale but are growing quickly. Model one: headcount pricing, used by Scale and Mercor, sells recruited labor like consulting and is relationship-driven B2B, involving substantial wining and dining of the researchers who choose the data.
- Model two sells environments, from a few hundred dollars upward. A Booking.com-style booking task might cost about $1,000 or less; recreating nearly the entire AWS experience for a difficult DevOps agent could approach $500,000. Researcher-founders anticipate lab needs three months ahead: create a benchmark, make it popular, and then sell the data that helps models rank highly on it. Xing thinks headcount models currently have higher revenue, while the anticipatory model has faster revenue growth.
- At the academic conference in South Korea last month, young researchers told him their top aspiration was founding startups that sell data: “They’re not building model companies, they’re not building agent companies — they’re building data companies.” The host described a broader shift from human labeling toward environment and benchmark building; Xing said Chinese providers similarly began with human-labeling businesses. He mentioned HLE, or Humanity’s Last Exam, as an example of high-end expert labeling.
10. AI for science: easy to build, hard to verify — plus Silicon Valley’s China blind spot
- AI-scientist agents can read literature, design experiments, sometimes control lab equipment, analyze results, and write or publish papers. Xing says they are “not that hard to build these days,” but it is hard to make one genuinely good and meaningful, measured in part by whether it can produce a well-received paper. The blockers include backbone OpenAI and Anthropic models that are not specifically trained on scientific databases and tools, plus wet-lab experiments where execution, data collection, and verification require time and human intervention.
- On the model side, since AlphaFold and especially AlphaFold 3, efforts from DeepMind and companies such as Chai and Boltz have pursued state-of-the-art protein design and complex-structure prediction. Xing also mentioned companies they had invested in building protein-design models. These models can attempt de novo design — creating proteins or antibodies that do not exist in nature — while virtual-cell world models could test whether a designed binder interacts as intended without a costly wet-lab experiment. The biggest problems are limited cell data and multimodality: enzymes, antibodies, antigens, and peptides must be represented together rather than handled only by disconnected small models.
- The structural verifier gap remains severe: intermediate verifiers exist, but “the true verifier is if this drug works on a human body and if it cures the disease.” Full feedback may not arrive until an FDA-approved clinical-trial stage roughly five or six years into the process. “You can’t imagine that happening with an LLM… but this is a reality for drug-design models today.”
- Xing’s stated misconception in Silicon Valley is that China’s AI progress is entirely tied to government support through funding, compute, and similar help. He argues that the US ecosystem has not given enough credit to the individuals, entrepreneurs, and researchers making progress under difficult constraints, and believes they could do it anywhere if they can do it in China.