E247 | A Conversation with 盛颖: xAI, the Romance of Infra, SGLang, Open Source, Equality, and “甄嬛传”
E247 | A Conversation with 盛颖: xAI, the Romance of Infra, SGLang, Open Source, Equality, and “甄嬛传”
Summary
- Inference is a game with no losers. 盛颖’s framing: inference providers keep multiplying, but “none has been diminished by the existence of competition; everyone is growing in sync.” The host runs through the market tape: a company reportedly formed by core members of vLLM, Inferact (name as heard), raised $150M at an $800M valuation; Vast (name as heard) raised $300M at a $5B valuation; Fireworks is valued at $4B; and Together is seeking $1B in new funding. He identifies 2 major shifts: AI now has to “make money and cash in,” moving market effort from training toward inference; meanwhile, AI-native companies are running into the limits of prompt engineering and need to customize with proprietary data, with some established AI companies expanding into RL.
- RadixArk announced a $100M seed round led by Axial in May, drawing an unusually broad group of backers. Nvidia’s investment arm, AMD, Intel CEO 陈立武, OpenAI co-founder John Schulman, and others all participated. Its underlying asset, SGLang, “has never been marketed or sold,” yet now runs on hundreds of thousands of GPUs worldwide and generates tens of trillions of tokens every day for Google, Microsoft, Nvidia, and xAI.
- The company’s lifeline is not SGLang. SGLang is fully open source and belongs to a community apparently associated with LMSYS rather than to the company: “There is no private fork. We never intended to create a difference here and extract value from it.” RadixArk’s end mission is “next-generation AI”—not necessarily a model, and “not something that already exists in the world today.” Inference is only the starting point and a prerequisite. Baseten and Fireworks are both users and potential rivals; his preference is that “all of them become our allies.”
- His account offers a rare inside view of xAI. Early xAI was “full of talent, with absolutely no politics,” a description that sounds like “two different companies” compared with the public narrative of its co-founders leaving en masse. His diagnosis applies to any fast-scaling AI company: xAI failed to navigate the transition from running on people to running on institutions. “If you don’t pay attention, it will be broken.” He left 2 months before his one-year cliff so the open-source community would not stall, giving up his first-year vesting.
- Day-zero compatibility is an exercise in market-driven economics. Importance is determined by market psychology—“you launch a product and I rush to buy it on day 1.” DeepSeek V4 introduced enough architectural innovation that “most of the code was basically rewritten,” and it was the first DeepSeek large model to support RL. As for latent-space inference, such as Meta’s possibly Coconut, he is unfazed: “The world changes and structures change; that is part of writing an inference engine. The real difficulty has always come back to the people and the team.”
- The differentiator is taste. “Infra is not a support role; in my eyes, infra itself is the product.” RadixArk reverses the priorities of big tech: a training job can pause, and benchmark scores are not urgent. Urgency belongs only to solving deep problems systematically. Taste means: “My system doesn’t merely run; I know how it runs—whether it is well-designed or a shoddy construction project.”
- Equality is the through line. He learned from CSDN and online problem banks, so open source feels like air; gender inequality is captured by the fact that “every win of mine has to be explained,” and the only real fix is for the top 0.001% to 1% to genuinely hold power. His end-state for AI is similar: closed source can exist, but it cannot be centralized—“everyone should have the equal ability to make their own AI, while the AI they make belongs to them.”
Deep dive
1. Opening portrait: Behind a $100M seed round, an open-source project that never did promotion
- 泓君’s opening fact pattern: RadixArk announced in May a $100M seed round led by Axial, with Nvidia’s investment arm, AMD, Intel CEO 陈立武, and OpenAI co-founder John Schulman among the investors. Underneath is the open-source inference engine SGLang, “a project that has never relied on marketing or sales promotion,” now running on hundreds of thousands of GPUs worldwide and generating tens of trillions of tokens each day for Google, Microsoft, Nvidia, and xAI.
- The timeline starts in 2023, during the final stage of 盛颖’s Stanford PhD, when he and researchers apparently associated with LMSYS built SGLang together—“the closing work of my PhD.” He then joined xAI to build Grok’s inference system before turning the open-source project into a company with 朱邦华. The episode was recorded in early June 2026, just after the company moved into a new office. It had more than 40 employees: “The expansion has been faster than I expected, and the scope has grown faster than I imagined. Every part of the company needs people.”
2. From Jiangxi to New York: Breaking out of the comfort zone twice
- By the time he finished his undergraduate degree, he had already accepted a PhD offer from CUHK and passed the point of honor-bound commitment. On day 2, Columbia sent him what he describes as “a very ordinary master’s offer.” He chose to pay the breach fee and walk away: “At that point, the center of academia was still in the US. I wanted to go see it.” His father made the decision in 1 minute over the phone: “Okay, go.”
- New York’s culture shock was liberating: “No one cared what I wore, what I thought, or what I did”—a move “from a constrained state to a completely free state.” He did not need to be controlled, nor did he need to control anyone else.
3. Princeton’s void: The moment he wanted to become a mathematician
- He had not planned to do research in graduate school. A course in computational complexity and a group of friends pursuing theory pulled him in. What truly hit him was a workshop at Princeton: “Everyone was above the fray. The things people cared about, the things you normally struggle and scramble over, were all profoundly meaningless.” The conclusion was paradoxical: “The fact that nothing matters is what matters—only what is real matters.”
- Mathematics drew him through the flow state created by certainty: “When the brain is fully engaged, you enter a kind of physiological pleasure.” His ambition at the time was absolute: “I thought I was going to become a mathematician. I only wanted to do mathematics.”
- He also observed a cultural split between the coasts. On the East Coast, “I hoped the world would forget me.” On the West Coast, “you hear more noise about whether you are succeeding or not. I began to feel that I wanted to participate in moving the world forward; I no longer wanted to be unrelated to it.”
4. Four unsuccessful advisor rotations at Stanford, then formal verification by default
- His first few Stanford rotations produced neither a problem that could put him into a flow state nor a professor with whom the excitement was mutual. On the 4th rotation, he met Clark Barrett and drifted into formal verification: “I needed to keep doing research and producing papers because I wanted to get the PhD. It was actually a fairly passive state.”
- His explanation for outsiders: take code written by a person, decompose it, and map it into first-order logic. A specification—preconditions, invariants, and postconditions—describes what the program should satisfy; mathematics then proves whether the code actually conforms to that logic. His own research focused on optimizing the SMT solver itself: efficiency, correctness, and “what kind of semantic space you define to express what a piece of code does.”
5. The truth behind a high-output CV: Thin stretches, a blank year, and rewriting from scratch
- “The real process is much thinner and much more boring than the CV.” He had not expected to win the 2020 IJCAR best paper award for his Politeness paper: “I was a complete beginner. My advisor gave me a problem and I solved it. I only learned that I had won when people started congratulating me.”
- Then came the pandemic. He spent “close to 1 year in a state of depression” and did nothing. Clark cared only about his mental state, telling him to take classes and serve as a TA. “Clark never cut off my funding.” He returned to China to recover for 6 months and emerged 1 year later.
- The sequence paper that followed required a proof of more than 20 pages to be revised 4 or 5 times, including 2 near-total rewrites: “We discovered at the very beginning that a foundational point was wrong, so every subsequent inference was wrong.” Clark called that proof “the highlight of his past year.” The paper came close to a best-paper nomination but was not selected. “If we are talking about my formal-methods work, that paper was the period I enjoyed most.”
6. Leaving formal methods: Too expensive, too narrow, too far from the world
- Formal methods gave him a new perspective: “I had spent 20 years working with programming languages, but I had never looked at them from a mathematical perspective.” Languages face a fundamental trade-off: the closer they are to mathematics, the easier they are to verify; the closer they are to hardware, the more flexible and efficient they are. “It is hard to find an anchor that puts a language close to both mathematics and hardware.”
- His decision to leave was stark: there were too few scenarios where formal verification could be used practically in the real world; verification was too expensive; and the set of programs that could be verified was too small to create meaningful impact. “I wanted to be a little closer to the world.”
- The pivot into AI came after another nearly 1-year blank stretch. “From the outside, that work looked like the product of my past year. In reality, it was the product of the final 3 months. For the first 9 months, I had no idea what I was doing.” He jokes that he published 3 formal-methods papers in 3 years, but each took only 3 months of genuine progress—“9 months every year spent not knowing what I was doing.”
7. “甄嬛传” and ADHD: Finding a key for using himself
- During his master’s, theoretical work “looked very relaxed”: 4 hours of thinking a day and plenty of video games. “You can only wait for inspiration to arrive.” The solution to one hard problem appeared midway through watching “甄嬛传”: a paper he had read suddenly echoed in his mind, and he realized that this was the breakthrough point. “There really was a phase transition. It happened in an instant. Pause. Thank you, pause.”
- Understanding ADHD helped him make peace with himself: “It is not a disease; it is a characteristic. It was as if I had found a key for how to use myself.” He should stop fighting his nature and work with it. His focus is bimodal: in flow, he could not hear people calling his name, lost track of time, and missed multiple meetings; otherwise, “even if something is important, I cannot focus.”
- “I am someone who has to be intense. I have been that way since childhood.” His pandemic depression came from being unable to maintain his mental health in an environment that had slowed down. Competition was not about winning: “I do not really care that much about winning or losing. A game can be lost, a competition can be lost—but I care about that feeling of intensity.”
8. Defining talent: Wanting to do it when nobody else cares
- He downplays his national silver medal in the informatics Olympiad and his regional New York championship. The regional event was a 3-person team, and “very strong teammates carried me.” He sides with talent in the talent-versus-effort debate, but defines it unusually: “When the outside world does not take research seriously, when it does not think research is a great thing, and you still want to do it—that is your talent.”
- Publishing papers, doing research, and being a scientist are 3 different things. “Publishing papers has a formula. Once you understand the rules of the system, you can keep publishing many papers.” Research “is like any ordinary person’s job; it is a very ordinary profession.” A real scientist is someone such that “if it were not for him, the thing would happen much later.” There are very, very few of them; most professors probably are not.
- Asked whether Hinton, LeCun, and Bengio belong in that 3rd category, he says they probably do, but he has not interacted with them directly. “I care deeply about first-hand information and truth-seeking. I do not want to judge people I have not encountered.”
9. Google’s 2022 internship: The shock that redirected him toward large models
- To move into AI, he took an AI for Code internship, an internally incubated Google project whose summer program was aggressive enough to hire as many interns as full-time employees while testing a wide range of directions. Coming from program analysis, he saw the opposite of an interdisciplinary dream: traditional tools were “extremely powerless, basically unable to help at all,” while “large models had already completely covered the technologies in traditional fields.”
- This was before ChatGPT, just after Google released PaLM, with a model apparently called LaMDA circulating internally and 1 engineer claiming it was conscious. “I realized that you did not need explicit reasoning at all. I was the person who was shocked.” The experience convinced him to enter the new field and work on large models.
- Choosing infra was not a plan but the result of selection. He tried algorithmic work without much success; infra worked. “My breakthroughs have all been on the infra side.” Infra was closer to his habits of thought: deterministic, grounded in mathematics and code, both of which he knew well.
10. SGLang’s starting point and its relationship with vLLM
- SGLang was the integrated culmination of his LLM-inference work during the final 2 years of his PhD. Its positioning was clear from day 1: “From the moment we finished it, we wanted it to be production-ready and usable at scale as an inference engine.”
- The history with what was possibly vLLM is more intertwined than outsiders realize. “The idea that was possibly vLLM’s at the time was similar to another idea we were thinking about. From my perspective, I was merged into that project.” He did not participate much in the maintenance phase after possibly vLLM became open source, instead continuing work on S-LoRA, fairness, and other research.
- The difference between the 2 engines exists only within a time slice. “If you remove the time dimension, they are similar. Good ideas will always be learned from and copied.” Possibly vLLM moved earlier into scale-up serving production at the 1,000-GPU and 10,000-GPU levels; SGLang moved earlier on community coverage and long-tail model support. Both sides are now learning from each other to close their gaps.
11. Possibly Ion Stoica: A lifetime mentor and a different value system
- The possibly Ion Stoica—co-founder of Databricks and a Berkeley professor—was his visiting mentor and remains a “lifetime mentor.” 盛颖’s observation: “He does many things and starts several companies at once, but at the core, his strongest self-identification is still professor and student. He cares deeply about making an impact, and he passes that spirit on to his students.”
- He draws a deliberate distinction around his own coordinates: “I care about impact, not making. If someone is willing to help move something forward, that’s great. But I gradually learned that no one will do the things I truly want to do for me. Only I can fight for one thing over a long period of time and make it happen.”
12. Why xAI: The only place that offered both support and freedom
- He spent 5 months at Databricks trying unsuccessfully to promote SGLang. Mature companies require sufficient justification for every initiative; as a new-grad researcher, he had little power, and he was not mature enough himself to drive a major effort inside an established enterprise. “I did not yet know how.”
- His move to xAI in October 2024 had a clear logic: his husband 连敏 was already there, and the company had reached the point where it needed production inference. Together they could build the entire inference stack. “It could give me 2 things: support and freedom. I had other opportunities, some offering only support and some only freedom. At the time, xAI may have been the only one that truly offered both.”
- He joined the engine group, connecting product-serving deployment and configuration, supporting research-team evaluation, and running RL rollouts on their inference engine. The small team was an advantage: “You could see how every layer and component worked with the others, and how the whole thing happened as a system.”
13. Early xAI: A golden era without politics
- “xAI was the most beautiful period I remember. Everyone was talented, mutually supportive, friendly, and humble.” It was the first company he experienced with “absolutely no people games and no politics.” When he joined, it had fewer than 100 employees. “I only had to think about how to get things done, and everyone around me was thinking about how to get things done. There was friction, but it was friction that did not affect unity.”
- He specifically liked the boundaryless culture shaped by Elon: “He has his problems, but inside xAI I could touch anything. I was not restricted to a very narrow scope.” There was little people management; the culture was intensely technical and engineering-driven.
- His theory of intense growth is physical: “Like working out, you have to feel pain before muscle starts to grow.” A year there could deliver the growth that might take several years elsewhere.
- His version history: V0.0, before he arrived, had “Elon being very nice to everyone, like a family.” V1.0 introduced growing pains, “but everyone was still overcoming difficulties together.” He did not sense a collapse in morale; they had achieved a degree of comradeship forged in battle.
14. Diagnosing xAI’s co-founder exodus: The transition from people to institutions was not absorbed
- 陈倩’s contrast is stark: the xAI he describes and the public narrative of a period when all the co-founders left, either fired or resigning, “sound like 2 different companies.” His explanation is direct: xAI did not manage the transition from a small, tightly unified startup that ran on people to a large company that had to run on institutions. “The growing pains were not smoothed over well, and that produced these huge changes.”
- He leaves room for a future reset: “Elon is still there, and that is the most important thing. It has also been absorbed by SpaceX, which is going public soon. One day it may find a new solution.” But in terms of personnel, “it may already be another company.”
- The lesson for his own company is institutional. A small team can operate through culture and personal management; as it scales, both culture and people become fragile, and the system must carry the load. “You have to think about this process intentionally. It does not happen naturally. If you do not pay attention, it will be broken.”
15. A community built by online acquaintances, and a departure that could not be delayed
- The core SGLang developers all met online: “We are all internet friends.” They worked together for 2 years without meeting in person and held countless meetings without turning on cameras. “We had an implicit understanding: if I only see your code, I know who you are. That is the only part I want to see.”
- Demand surged in 2025, and the model of staying up late, working weekends, and coming in during holidays stopped working. “After forcing it for several months, we realized we could not deliver. We were scraping by at the borderline, or failing to deliver. The team’s bandwidth could no longer be stretched.” By July, his anxiety had reached a point where he could no longer defer the decision: “If we did not fill this empty space, we were about to disappoint the outside world.”
- The cost can be measured precisely. He had been at xAI for 10 months, with 2 months remaining before his 1-year cliff and vesting. “I could not stay even 2 more months, because we could not wait. I had reached the point psychologically where I felt I no longer deserved to remain. It was not something I could control.”
- His attitude toward the money is similarly practical: “It was necessary, not important. Its meaning was that if one day I could not raise money, I would still have money of my own. That was the confidence it gave me.”
16. The Two Sigma interlude: The most organized company
- During the 6-month gap between Columbia and Stanford, he joined hedge fund Two Sigma. The appeal was curiosity about finance and a desire to talk to smart people, plus an on-air moment: “One of my ex-boyfriends studied finance. Cut that.” “I am definitely going to publish that.”
- His ranking of the companies where he has worked: xAI first, Two Sigma second. “It was at its best every time I was there. It really was at its best.” Two Sigma was then open-sourcing projects and embracing AI and big data; “everyone was just doing things, and it was very open.”
- Hindsight changed his view of its organization. “At the time, I did not think it was organized. It was only after going to other companies that I realized how chaotic everywhere else was.” Two Sigma was not a large company, but its infra was the most stable he had seen: a stable business model, few major upheavals, and enough continuity for people to do extensive maintenance. Finance was never his destination: “Playing with money every day is not very interesting. I knew from the beginning it was not my priority.”
17. What RadixArk is building: Foundation models are infra too
- His structural ambition is a foundation toolkit. Given different use cases, data formats, and product and serving objectives—“a model that solves simple problems but must be fast, or one that solves hard problems but can be slow”—the system should “continuously produce the best model for your particular objective.” The foundation spans more than inference, training, and post-training: “Foundation models are infra too.” It also includes the codebase, toolbox, sandbox, environment, and intermediate checkpoints.
- The current focus is inference and RL because they sit at the end of the production pipeline. Post-training and RL are fan-out points: “One model will not solve every problem. Multiple enterprises want their own models.” That makes them an ideal place to build infrastructure that shares technology with everyone. Inference sits at the furthest downstream point, closest to the user.
- Doing both is manageable because they are closely related. A large share of RL’s challenge comes from the rollout engine, which is the inference engine, plus some glue work for weight transfer and scheduling algorithms. “At the systems level, there is enormous overlap between them.” The RL framework is called MILES.
18. RadixAttention and the agent-first idea that got buried
- In plain English, when requests share a common prefix, the corresponding KV cache does not need to be recomputed. A radix tree indexes prefix relationships; after mapping those relationships to a KV memory pool, the cache can be reused. To the objection that every new user question is entirely new, his answer is simple: once a conversation becomes multi-turn, later exchanges reuse the earlier history. “Prefix reuse exists extensively in almost every scenario.”
- The original agent focus has been obscured by backend work. The SGLang paper from 2 or 3 years ago was designed around agents, with a frontend language serving as “the interface for interaction between people and AI.” Much of the development effort has since shifted to optimizing the backend runtime, but “one day we will revisit this part. Ultimately, it should still take the form of a language.”
19. Day-zero compatibility: The market decides what matters
- Why does day-zero support matter? “The market thinks it matters. We are still very practical.” It serves the mindset of “you launch a product and I rush to buy it on day 1.” That demand travels from the user to the inference provider, then to RadixArk, and ultimately to the model maker.
- Supporting DeepSeek V4 required an exceptional effort. Its architecture introduced many innovations and differed radically from existing models. “Many operator-level implementations had to be rewritten. Most of the code was basically rewritten.” Its feature set was also unusually complete, making it the 1st DeepSeek large model to support RL. Conversely, when the architecture is unchanged, adaptation may be unnecessary; with only minor updates, it is almost unnecessary.
- Would latent-space inference—Meta’s possibly Coconut or ByteDance’s Kola DLM, both names preserved as heard—upend inference engines? “If it matures, many things will indeed change. But the world changing and structures changing is already part of writing an inference engine. That is not the transformation itself.” With AI now writing code so quickly, the technical work can always be done. “The real difficulty has always come back to the people and the team.”
20. The inference market: A game with no losers
- His counterintuitive observation: inference providers have existed for years, starting with Fireworks and Together, while new companies continue to appear every year. Yet “this game has no losers.” None has been diminished by competition; everyone is growing in sync.
- 陈倩 adds the market tape: a company reportedly formed early this year by core members of vLLM, Inferact (name as heard), raised $150M at an $800M valuation; Vast (name as heard) recently completed a $300M round at a $5B valuation; Fireworks is valued at $4B; Together is seeking $1B in new funding; and cloud providers including AWS and Google are building managed inference offerings in parallel.
- He sees 2 major trends. As AI becomes more useful, it has to “make money and cash in,” shifting market effort from training to inference. At the same time, AI-native companies are feeling the limits of prompt engineering on existing models and need to customize with their own data. “You will see some established AI companies expand toward RL.”
21. “We are not in this segment”: The lifeline is next-generation AI
- Baseten and Fireworks are both users and potential rivals, but he rejects that framing: “We do not regard them as competitors. We hope they will all be our allies.” SGLang is fully open source and belongs to a community apparently associated with LMSYS rather than being a company-owned asset. “There is no private fork. We never intended to create a difference here and extract value from it. The company’s lifeline is not here.”
- He will offer only an outline of the real target: “RadixArk’s ultimate mission is to build next-generation AI.” It may not be a model, and “what we want to build in the future is not something that already exists in the world today.” Inference is a prerequisite and a starting point, not the long-term core. Asked by 陈倩 for more detail, he says no.
- The customer spectrum runs from the bottom up: data centers, neoclouds, and hyperscalers that need help selling GPUs; infra providers, currently users whom RadixArk hopes to convert into customers; AI companies that train and serve their own models; and traditional enterprises seeking to use AI to improve existing businesses.
22. The romance of infra: Infra itself is the product
- Asked about 朱邦华’s claim that infra taste differs from that of a pre-training researcher, he declines to settle the point: “I am not sure that is actually true. At its core, they may be connected.” His definition of taste is attention not just to the end goal but to what makes it up: “My system does not merely run; I know how it runs—whether it runs because it is well-designed or because it is a shoddy construction project.”
- His most personal thesis is that infra professionals have long been underestimated, forced to support other components and patch over problems with hacks. “They are always in a state of being pulled in different directions.” But because of his “romantic relationship” with infra, “infra is not a support role; in my eyes, infra itself is the product.” No one says they want infra to “touch the human spirit,” but from his perspective it does: its far-reaching effects propagate upward from the foundation.
- RadixArk reverses big tech’s priority order. It does not ask the infra team to keep training jobs on track at all costs. When a training job is blocked, the focus is to solve the problem systematically rather than patching it so the job must run. “A job can pause. We are not in a hurry to chase benchmark scores.” Benchmark performance is not the source of urgency. The cultural keywords are focus, humility, and continuous refinement: RadixArk’s people are unusually demanding of their own work.
23. Open source like air: Taught by the internet
- He first rejects the premise that he is “obsessed” with open source: “I am not obsessed with open source. It happened naturally.” In Jiangxi, there were few programming resources. He learned from CSDN posts and online judge problem banks, from people whose identities he still does not know. “I was taught by people somewhere on the internet, in places I do not know. The culture of open sharing feels like air to me. If you make something and do not share it, that is what feels strange.”
- The open-source community is changing in both directions. The positive is that appreciation from users and recognition from capital are rising. The negative is that “impure things are starting to mix into the enthusiasm. Some utilitarians have entered open source, making it harder for people who were not utilitarians to tell which parts are genuine.” They affect more than just themselves.
24. The end state for open and closed source: Closed source can exist, but cannot be centralized
- His position is a dual rejection. “I do not think AI should be completely exposed to everyone. It carries risk, and I do not deny that. But I also do not believe AI should be entirely held by a small group of people.” AI “should be held equally by different people.” The company’s role is to provide the toolchain for creating the best closed-source models: give everyone the equal ability to build their own AI, while the AI they build belongs to them.
- He doubts a call to stop all AI research—citing a major figure whose name was rendered in the audio as “Andre Big”—could succeed. “It is contrary to human nature.” If an all-powerful God came down and assigned everyone what to do, it could happen. He does not think a world in which everyone stopped and 1 company led would be an extremely bad outcome, but the world could not remain there. Someone would inevitably revolt.
- His long-range scenario has 2 stages. In the foreseeable future, AI will not be strong enough to fully control or threaten humanity; it will replace human functions and create social and power-structure upheaval—“more of an Industrial Revolution than the Industrial Revolution”—but the disruption will remain understandable. If AI can truly dominate humanity, “we cannot stop it. I think that day will come.” The answer is not to call for a halt, but to invent another path that allows AI to advance without humanity losing. The prerequisite is first to let all humans coexist equally and unite, then search for a solution for humans and AI to coexist.
25. LMSYS and equality: Lifting people without pedigree
- LMSYS, apparently the nonprofit he founded with 连敏, carries enormous weight for him. “If RadixArk does not make it all the way one day, LMSYS is the事业 I will work on for the rest of my life.” Its mission is to incubate exceptional talent even when people lack sufficient background, connections, or branding; direct credit to the people doing the work; and protect unestablished developers’ control over their projects.
- The motivation comes from a gap he experienced personally. The same work is much more likely to attract attention, be remembered, or even be exaggerated when he is at a high point. Some work done during low periods may be just as good but is much harder to see. “That difference does not need to exist, and it damages the structure of fair competition in society.”
- He acknowledges benefiting from identity-based advantages himself. Some papers he considers merely average were treated as standards because he wrote them or because they came from Stanford or Berkeley. “When I was a beneficiary, I finally had enough power to push this issue outward a little.” LM Arena began as a fun project during the apparently LMSYS period, attracted unexpected attention, and was later developed by 韦林 and others into a commercial-scale operation.
26. Gender inequality: Every win has to be explained
- His textbook example: “If I were a boy, people would say I was a genius. If I were a girl, they would say she was hardworking, obedient, or lucky.” He placed near the top in competitions year after year, but “every win of mine had to be explained”—the opponent was supposedly out of form, or she happened to guess the problem because she had seen it before, which she had not. Normally, a person’s win needs no explanation: “They are good.” Even when he worked with an established male partner, people wondered whether the idea was actually his.
- The damage comes from scale rather than any single incident. “Each example is tiny when isolated, but these ordinary little things happen every day, every minute, every second. Multiply the impact of that one small thing by 100 million, and that is the true impact on a person’s life.” There is no way to step outside the frame; once you do, people call you petty or narrow-minded.
- He sees only one solution: education alone cannot achieve the goal. The disadvantaged group—in this case women—must “actually hold power.” Equality at the top 0.001%, 0.01%, 0.1%, and 1% of humanity is what changes the system. The past 10 years have probably brought incremental improvement and society is moving forward, but values remain volatile.
- His methodology for transmitting change is worth isolating. Training more students to become teachers is not enough; you must train “teachers who want to train students to become teachers.” “Only building the recursive chain works.” It is slow, but it does not roll back: once you push it forward, it has moved forward. When someone whose name is unclear asked whether he thought he could change the world, he answered: “I cannot change it. But if I can change 1%, I am willing to spend my life changing that 1%.”
27. Raising money without knowing what a term sheet was, and the first lesson in authority and responsibility
- He started from zero: “I did not know what a term sheet was, what valuation was, or what a seed round was. I had never even heard the phrase pitch deck.” When someone asked whether they had received a term sheet, his response was: “What is a term sheet?” The process was circuitous and he declines to discuss the details, but afterward he had “more respect for investors.” Sand Hill has a set of rules, but ultimately they are about human nature: make something good and pursue something real. It echoes his view of paper publishing: all common activities have a playbook.
- Why did 3 major semiconductor companies invest together, an unusual outcome? He refuses to speak for them: “I also want to ask the investors that question. I will not answer for them.” His own pitch is that RadixArk is “a player that has built up speed in this segment,” able to enable more AI companies—and companies seeking to enter AI—to build AI capabilities. The 1st round was clearly about both people and the work, given the open-source foundation.
- His first management lesson came from Tony and 国栋 at xAI, who independently told him the same thing: “Your authority and responsibility have to match.” At xAI, “I could not decide things, so I could not be responsible for them.” Now he has to make the decisions: success belongs to the team, while failure sits entirely on him. He is candid that he is moving from seeing everything clearly as an individual to no longer seeing everything clearly as a company scales. “I have not found the answer. All I see right now are challenges.”
28. Closing: The books did not lie to me
- He added this segment himself. He had almost declined the interview, but was moved by the document the team prepared. “Many of the investors who backed me have never looked at my experience this way.” He later realized that “for someone simply to complete what their profession requires is already extremely rare. Many things that are supposed to happen by definition never happen.”
- The 3 turns in his story ended in unexpectedly ideal outcomes: difficulty finding a PhD advisor led to “a textbook-perfect advisor”; attacks on the open-source project led to “a team with so much energy and unity”; and a difficult fundraising process led to “textbook-perfect investors.” “The books describe beautiful things, while reality is much more complicated. But the books did not lie to us. Good people do exist. You need the patience to let the grains of sand that are not ideal fall away on their own. You have to keep believing even when no such people are around you, believing that one day someone who believes as you do will meet you.”
- Can he reconcile this with the idea that the world is a rough, improvised operation? “I am actually very emotional. Very small things can have a very large emotional effect on me.” But he can separate emotional impact from the construction of his values. He will continue to believe in things he has not seen but trusts. Positive feedback has lengthened his patience: the 1st time, validation took several months; the next time, several years; after that, 10 years. “I know it will always eventually be validated. Next time I can wait 20 years.”
- His final boundary is uncompromising. “Even if my entire life is fragmented, I do not want to compromise.” The only thing that matters is whether you understand and act according to your own thinking, or let others pull you along. “As long as you do not feel that you are compromising, even if you are actually wrong, then you are wrong. That is okay.”
Verification Notes
- The names LMSYS, vLLM, Ion Stoica, LaMDA, and Coconut were heard with varying degrees of uncertainty in the original audio; the English edition preserves the source’s qualifications where applicable.