Pioneers Insight Method Research Author
4小时对谈英伟达研究VP刘洺堉:Cosmos 3、世界模型与黄仁勋
Back to Episodes

4小时对谈英伟达研究VP刘洺堉:Cosmos 3、世界模型与黄仁勋

Summary

  • 刘洺堉’s core view is that frontier-model capabilities will gradually converge: “A world where one company has all the technology and nobody else can catch up has never existed, and I don’t think AI will be an exception.” Talent moves, know-how diffuses, and once models converge, the edge shifts to integration with ecosystems and products. That is the underlying logic behind Nvidia open-sourcing Cosmos: the goal is not only to chase SOTA, but to help companies built on its GPU ecosystem succeed. “Model competition shouldn’t be a Squid Game.”
  • Cosmos 3 brings language, video, audio, and action into one omni world model, with action treated as a first-class citizen—an important distinction between physical and digital world models. The team consolidated what had once been planned as 22 models into one, after 黄仁勋’s story about Louis Vuitton reducing SKUs and increasing sales made the point. It is now running at “10,000-GPU scale,” with daily compute and storage costs potentially reaching several million dollars. 黄仁勋’s instruction was: “Take Cosmos to 97”—the number was off the cuff, but the commitment was real.
  • He made a clear industry call: “The ChatGPT moment for physical AI is coming”—perhaps this year, perhaps next year, but it is inevitable. The defining moment will be generalization: a person demonstrates a task once, and the machine learns it. Humanoid robots will be an important direction, but wheeled or semi-humanoid systems can solve many problems too; the key variable remains operating cost. 谭杰 of DeepMind forecasts a humanoid inflection point in 2035–2036. 刘洺堉 hopes it comes sooner.
  • Nvidia’s strategy is to create markets rather than fight over an existing pool of demand: 黄仁勋 prefers businesses that start from zero and can reach $1B after they work, while a $100M-a-year business is “a distraction” for Nvidia. If every household had robots, “my family might buy two or three,” and every node would need compute. The risk lesson comes from CUDA: Nvidia effectively sold it at a loss in the early days. “Not doing it is even riskier. If you don’t, your product gradually becomes a commodity.”
  • DALL·E and Sora were two career-defining moments of self-criticism: “I realized my ambition wasn’t big enough. I wasn’t brave enough”—even though his job talk a decade ago had predicted that deep learning would replace traditional graphics for image generation. After DALL·E, he wrote to 黄仁勋 and secured 1,000 A100s. Sora was later built in part by Tim Brooks, who had interned in his group. His question became: “What did I get wrong in my decisions and judgment that kept me from doing the right thing at the right time?” Asked about OpenAI largely winding down Sora as a product, he pointed to the token economics of Anthropic’s coding agent: chat may finish in a minute, while a coding agent can run for 24 or 48 hours and generate huge token volumes. OpenAI therefore reallocated resources—“the more things you prepare, the thinner your advantage.”
  • 黄仁勋’s management system is unusually high-resolution: he says every day that the company has only 30 days of cash left, avoids layoffs and forced ranking, dislikes internal horse races, uses top-five emails to spread weak signals, and runs an organism-like organization where “mission is the boss.” His deepest personal lesson came from being asked face-to-face, “Are you a crybaby? Nothing in this world is fair.” “I stopped making excuses after that.” The operating model is trust paired with pressure—a willingness to “torture you to greatness.”
  • His view of China is broadly constructive: DeepSeek’s ability to iterate 3 times in a year and produce DeepSeek-V3 directly shaped Cosmos’s iteration methodology; Chinese companies hire large numbers of interns, accelerating knowledge diffusion, unlike US frontier labs, which hire relatively few; and vertical integration between robot hardware and manufacturing can enable things “that are difficult to do in the US.” His 2026 view came up in the context of IPOs: 刘洺堉 said “2026 will be a year of major change,” while stressing that technology will not stop there.
  • His personal philosophy has shifted from being highly competitive to seeing “low ego” as confidence: “Have the ability to be number one, but do the thing that contributes the most to the world.” Building the strongest model for a small group of users and concentrating power in too few hands “may not be the right thing.” His advice to young researchers is three words plus one sentence: “Don’t be nervous—this is a good time to enter physical AI.”

Deep dive

1. Positioning the opening: a GSD research VP and a 200–300-person Cosmos army

  • 张小珺’s opening framing is worth preserving: people inside Nvidia say 刘洺洧 “doesn’t look like a typical researcher; he looks more like an engineering leader.” 黄仁勋’s description is GSD—Getting Shit Done. He is now Nvidia’s VP of Research and VP of the Cosmos Lab. His direct team numbers 80–90 people; the broader Cosmos effort has 200–300, making it “one of the major projects inside Nvidia.”
  • He comes to China 2–3 times a year: “It’s part of my company mandate, and I’m personally interested in AI for robot dogs. Things are developing very vigorously here.” His observation this trip: dexterous hands continue to make breakthroughs in key technologies. “The pace is extremely fast. I learn a lot every time.”
  • The organization changed early this year. He left the Research organization, where he had reported to chief scientist Bill Dally, and now reports to Dwight Diercks, Nvidia’s first software engineer and a figure featured in 黄仁勋’s biography. The former Deep Imagination Research team was renamed Cosmos Lab because “we’ve reached a stage where success is mandatory, and the resources required are much larger.”

2. The job talk from a decade ago: deep learning would replace traditional graphics

  • His pitch when he joined Nvidia sounded aggressive in an era when movie effects still relied on traditional graphics pipelines: “One day, most video-generation compute will be deep learning, not traditional computer graphics.” You would even be able to “put a novel in as input and get a movie out.”
  • Nvidia “wasn’t a company everyone recognized as an AI company” at the time, and the research organization had no one working on this kind of applied research. He was among the first to establish a generative-AI direction at Nvidia and an early advocate of unpaired image translation—using an information bottleneck to learn “horse to zebra” without paired inputs and outputs.

3. The full GauGAN story: a magic brush, a canceled keynote, and the first mass audience

  • GauGAN, known in Chinese as “Ma Liang’s magic brush,” worked like a child’s paint program: each color represented a semantic category, and a few strokes could turn into a landscape painting. It helped Nvidia build the image that it could produce not just excellent GPUs, but excellent AI models. Collaborators included interns Taesung Park, Zhu Jun-Yan, and Wang Ting-Chun.
  • The underlying algorithm was AdaIN, or adaptive instance normalization. 黄仁勋 was excited by the demo but objected that AdaIN was “not user language.” The paper was accepted to CVPR but held off arXiv for the 2019 GTC. Marketing genius Greg Estes came up with GauGAN, linking the painter with the painting app in an instantly legible way.
  • It was originally slated for a GTC keynote, but too many products forced a last-minute switch to a press release and private media demos. After seeing it, the media concluded that GauGAN was one of the most important projects at GTC. Feedback from the online app made him realize, “I can influence more than just researchers who read my papers,” and pushed him toward applications.

4. DALL·E’s wake-up call: “My ambition wasn’t big enough,” followed by a request for 1,000 A100s

  • His reaction to OpenAI’s DALL·E launch was blunt: “I realized my ambition wasn’t big enough. I wasn’t brave enough. I hadn’t made up my mind to build this thing”—despite being at the place with the most compute. “I didn’t realize that once you had enough compute and data, the conditions were already there.”
  • He wrote to 黄仁勋: “The future is clear: image generation will be done with deep learning. Nvidia cannot fall behind,” and asked for 1,000 A100s. That was a huge request at the time, and efficiently training across 1,000 GPUs was not easy.
  • He also wrote to Getty Images seeking a partnership. The eventual proposal combined data, compute, and a text-to-image model.

5. Picasso and Getty: research becomes a product that helps customers win

  • With Getty Images and Shutterstock, he built Picasso, an AI service for companies with data and customers but no AI talent or infrastructure. The idea was to build a generative-AI model for them, hand it over, and let them serve their own customers—“so we could both benefit.”
  • The identity shift mattered: “My research wasn’t just producing a model or a paper. It was also producing a product that could help customers succeed.” He later shut down Picasso because Cosmos consumed his bandwidth, while helping customers find technology companies that could continue supporting them.

6. Sora came from a former intern: happiness, self-criticism, and why OpenAI reallocated resources

  • 余佳辉, now at Meta, and Tim Brooks both interned in his group. During his internship, Tim “kept wanting to do video generation, the thing that later became Sora,” but the technology was not mature. After graduation, 刘洺堉 failed to retain him; Tim went to OpenAI and helped build Sora. He was “very happy” when Sora launched, but kept asking: “What did I get wrong in my decisions and judgment that kept me from doing the right thing at the right time?”
  • On the host’s suggestion that OpenAI had essentially wound down Sora as a product, he said Sam Altman has an investor-like portfolio mindset. But over the past year, OpenAI came under pressure from the “huge economic effect” of Anthropic’s coding agent: a chat session might finish in a minute, while a coding agent can work for 24 or 48 hours and generate huge token volumes. OpenAI therefore concentrated resources—“as Sun Tzu said, the more things you prepare, the thinner your advantage.” He does not take that as evidence that Sora was bad work; the company simply had something more important to do at that moment.

7. Why a platform company does its own research: chip cycles are long, so it must walk the customer path first

  • 张小珺 put the question sharply: if an infrastructure company can simply sell chips to frontier labs, why do research itself? His answer was about cycle time: “The cycle for designing chips and computer architectures is very long. You can’t say, ‘I’ll change it right now.’” Nvidia has to walk the paths its customers may take so it can design chips and infrastructure that satisfy leading AI companies.
  • The deeper belief is that “if everyone with ambition and ideals had the tools to build what they want, the world would be a better place.”

8. The logic of open source: not competition, but a predictable signal to the ecosystem

  • Putting an open-source model online for free is only the beginning. “There is still a lot to do in deployment, especially in embodied AI, where many hardware and commercial problems remain.” Open source is a commitment signal: if customers see Nvidia investing through generations one, two, 2.5, and three, they gain confidence and can deploy resources elsewhere.
  • Is that competing with customers? “Embodied AI is so difficult, and there are so many problems to solve. These models take some of the pressure off them. I don’t see that as competition. I see it as helping everyone succeed.”

9. Why bet on embodied AI instead of LLMs—and Jensen’s $1B standard

  • The reason not to build an LLM is straightforward: LLMs were already highly developed, he came from computer vision, and Nvidia already had a strong LLM team. “I wouldn’t have contributed much there.” Sora was oriented toward creation; the team saw physical AI as addressing urgent needs: mobility loss among the elderly, dangerous and labor-intensive work, and household chores. World models could contribute directly to physical AI.
  • Jensen likes “$1B-scale businesses.” Later in the interview, 刘洺堉 described the target as a “zero-billion-dollar business”—something that does not exist today but can become a $1B business. It does not have to make money immediately, but when it works, it must be enormous and consequential. “Making $100M a year is a big deal for many companies, but for Nvidia it’s a distraction.” The challenge is that Nvidia is not a solutions company: “It’s a bit like scratching an itch through clothing. If you don’t understand the customer, you can’t optimize properly.” That is one reason he comes to China so often.

10. A Taipei teenager changes course: from wireless communications to computer vision at an Intel internship

  • He grew up in Taipei, where the hottest technology in high school was wireless communications—the era of Nokia’s giant phones. He studied wireless communications at National Chiao Tung University, then became unsettled in his junior year after comparing his research with US PhD programs. He was working on protocols to make transmission faster, an incremental improvement rather than a zero-to-one leap; “every successful project at a top US university seemed to be a zero-to-one change.” When he decided to study abroad, his parents were “shocked. In more than 20 years of knowing me, they had never heard me say I wanted to study in the US.”
  • Before leaving Taiwan, he spent a year at Intel and met a friend working on image recognition. It was his first exposure to AI: “Inferring what something is from pixels was amazing.” He originally applied to study quantum communications, but after arriving at the University of Maryland found computer vision more interesting. He joined the group of field pioneer Rama Chellappa, with a first project that inferred a person’s identity from their gait in video.

11. The gift of traditional computer vision: nothing worked, so learn everything

  • “In the traditional computer-vision era, nothing worked, so you had to learn everything”: SVMs, continuous and discrete optimization, graphical models, energy minimization, and 3D calibration. “It was a very solid foundation.” AlexNet arrived in his final year and “changed the whole world.” What he had learned began diverging from the future, but he was not disappointed: “Because nothing worked, I had already developed the habit of learning everything.”
  • In the final year of his PhD, he was the senior student and organized deep-learning seminars for the junior students. One said, “Only a few people know this stuff. There’s no reason to study it.” “I ignored him and kept organizing the seminars.”

12. MERL and the origins of generative AI: “If I cannot create, I cannot understand”

  • His first job was at Mitsubishi Electric Research Laboratories, or MERL, in Cambridge next to MIT—a lab that in the 1990s competed with Microsoft Research. Several MIT professors, including Bill Freeman, had done research there. The environment was freewheeling, and he taught himself deep learning. After seeing Ian Goodfellow’s work on GANs, he invested heavily in generation and became one of the earliest researchers in the field. “If you could generate a 32×32-pixel face that didn’t look like a ghost, you felt very successful.”
  • Why does generation equal understanding? He cites Feynman: “If I cannot create, I cannot understand.” A discriminative model can distinguish A from B; a generative model also carries reconstruction and description signals. “When you can describe something clearly, people think you understand it better.” His advisor Chellappa believed in Bayesian methods and in incorporating generation into understanding, planting the seed.

13. Three research lessons: the trap of beautiful math, saying 10 after doing 100, and N² communication

  • Lesson one was an early fixation on elegant mathematics. He repeatedly added assumptions to prove a bound—an answer within 10% or 20% of the optimum. “Assumptions make the math clean, but they also take you farther from reality. AI is an application. I increasingly prefer getting close to the essence of the problem and not being distracted by fancy math.”
  • Lesson two was communication: “I did 100 points of work, but my explanation was worth only 10 points, so others saw only 10.” From then on, making his work understandable to ordinary people became part of the research.
  • Lesson three is the complexity of collaboration: communication between people is N² and “doesn’t scale easily.” Major contributions require collective effort, but also extremely clear communication so everyone is pursuing the same objective.

14. The GAN lineage and why text is hardest: climbing from CoGAN to language

  • The starting point was CoGAN, his only major GAN work outside Nvidia. Two coupled generative networks formed an information bottleneck, forcing the networks to identify the shared essence of horses and zebras and enabling unpaired translation. The difficulty ladder for control signals ran from image to edge to segmentation to class, with text last: “The structure of text is most different from image, so it is the hardest to learn and needs more data.” He climbed from easy to hard, which also explains why DALL·E hurt: “Because I didn’t do it, I couldn’t build it.”
  • At an early NeurIPS, he and a friend discussed where AI might break through first. “The image signal has too much noise and is farther from the essence of knowledge. Text is already very clean, close to symbols.” He knew that, but stayed with vision: “I simply liked this kind of visual signal more.”

15. The failed Scale GAN experiment: 余佳辉, Jeremy Bernstein, and the shift to diffusion

  • His decision rule predates the popular bitter lesson: “When making choices, look for what scales better.” GANs were difficult to train, and he worried about that for years. During his internship, 余佳辉 ran remarkable experiments in which each GPU hosted a discriminator with different weights—“as if many people were fighting the generator at once.” Jeremy Bernstein studied ways to stabilize GAN training and later developed the Muon optimizer. “Every result pointed to the same conclusion: GANs don’t scale.”
  • The shift in 2021 was blunt. The team had assembled to work on GANs, but “that year I told everyone we couldn’t do GANs anymore.” Diffusion was “extremely stable,” so they moved to the Edify series for image, 3D, and video. He admits he was not an early Transformer adopter; researchers gradually brought Transformers into the generative direction later.

16. A mindset for paradigm shifts: as things get simpler, solid fundamentals win

  • After 3 paradigm shifts—traditional computer vision to GANs to diffusion—he concluded: “I realized things were getting simpler and simpler… I gradually understood that this was the essence of the problem.” Cosmos likewise consolidated several models into one.
  • Does he feel the earlier work was wasted? “I spent 5 to 10 years on each direction. You can’t call it a mistake. Given the compute and data available then, it was reasonable to do those things, and what I learned then has been very useful to my understanding now.”
  • His current state resembles “a heavy sword without an edge, great skill without flourish”: build the foundations carefully, step by step, and write the thing out.

17. Interns and the flow of ideas: how to measure the vitality of research

  • His signature analogy: “The vitality of an economy is measured by how quickly money moves from one person’s pocket to another. The vitality of research should be measured by how quickly ideas move from one person’s head to another.” In industry, talking only to the same colleagues every day can reduce idea diversity, so Nvidia brings in interns from different backgrounds.
  • The attitude is genuine: “Many of my interns are better than me, and I think that’s a very happy thing.”

18. Joining Nvidia: if GPUs are scarce, why not go where GPUs are made?

  • There were two triggers. MERL emphasized algorithms and paid relatively little attention to compute, so he was constantly asking his manager to buy GPUs. Then his long-term mentor Oncel Tuzel left for Apple to work on autonomous driving, removing “the main reason I stayed at Mitsubishi Electric.” The logic was first principles: unsupervised methods such as CoGAN did not need labels; “the most important things were compute and data. Data didn’t need labeling, so I should go somewhere with compute. Why not join a company that makes GPUs?”
  • His father did not understand: “Apple is good, Google is good. Why would you join this Nvidia?” Nvidia was still obscure. The punchline came later, when his former MERL manager retired and admitted that every time 刘洺堉 successfully won budget for GPUs, the manager bought Nvidia stock.

19. The early research organization and Tero Karras: the pressure of writing one paper a year

  • When he joined, NVIDIA Research had fewer than 100 people, most working on graphics. At an annual offsite near Monterey, he first met Tero Karras, Timo Aila, and other researchers. They wanted to solve GANs; he thought, “They may not yet know how hard this problem is.” StyleGAN then broke through repeatedly from version one to three. “They may have underestimated the difficulty, but they did it.”
  • From Karras he learned the opposite of a high-volume publishing strategy. While some people write 20 or 40 papers a year, “Tero wanted to write only one paper a year, putting all his energy into one idea, drilling deeper and deeper into the essence of the problem to produce a major contribution. The pressure was greater. I learned a lot from him about staying focused.”

20. MERL’s historical lesson: the cycle of blue-sky research

  • MERL’s 1990s golden age eventually changed. The company began questioning the impact of blue-sky research, and “many important scholars were laid off or left.” Some went to Microsoft, some became professors, and others moved to industry. Those who remained continued similar work, but it became harder to explain how it related to the company’s direction. Research tied to cameras, industrial cameras, and robotic arms was easier to support.
  • “I was lucky to see that process early in my career.” Google and Meta have gone through similar cycles: during upswings, researchers receive broad freedom; during contractions, companies retain people whose work is more closely tied to corporate priorities.
  • That is why he aligned early at Nvidia: deep learning is expensive, so “you have to explain why this is an important investment for the company.” Blue-sky research and corporate interests can align.

21. The convergence thesis: one company taking everything is “not a stable social state”

  • He is cautious about the scaling narrative. “Sometimes one model is better this month, another is better next month, and a new model catches up the month after that.” With multiple Chinese companies continuing to break through, “frontier-model capabilities will get closer and closer. Model quality alone won’t be enough.” To the pitch that exponential growth means anyone who does not invest will miss out, he responds with history: “A world where one company has all the technology and nobody else can catch up has never existed. I don’t think AI will be an exception. That is not a stable social state.” Talent moves and know-how naturally diffuses.
  • What matters after convergence? “Integration—with your own ecosystem and your own products.” Google is the example: its model is integrated across Docs, Search, Gmail, and the broader platform. More feedback makes it easier for the model to contribute to the company.
  • He considers himself lucky to have believed early that models would converge, which is why he insisted that models be integrated with the rest of the company’s ecosystem.

22. The innovator’s dilemma and the Anthropic template: strategy is deciding what to abandon

  • The opportunity for startups lies precisely in the blind spots of giants. Nvidia is not interested in businesses below $1B, “but there are many worthwhile things to do in that space.” The Innovator’s Dilemma describes how small companies enter neglected markets; Nvidia itself grew that way.
  • “Starting a company today to do the same thing as OpenAI or Anthropic is dangerous. They have more capital, more customers, and can make the product cheaper. You need to solve pain points they cannot solve—not because they don’t want to, but because they cannot justify it.”
  • Anthropic is the template. It was initially viewed as behind OpenAI. “When OpenAI pursued diversified investment and didn’t put all its eggs in one basket, Anthropic made a concentrated bet and believed in coding from the beginning.” As China and the US now crowd into the same coding market, he says it may be forced competition: “Nobody wants core technology controlled by someone else. But I don’t think that was their entire plan. Ultimately they still need to differentiate.”

23. The first person to go from researcher to VP: still reading code and staying grounded

  • He may be Nvidia’s first researcher to rise all the way to VP, moving up one level each year for several years. The reasons were that he joined before Nvidia had many AI researchers and liked shipping work; trust accumulated over time. His team may now generate several million dollars in compute and storage costs every day.
  • Management and research instincts naturally conflict—“researchers have high egos; scholars looking down on one another has existed since antiquity.” His answer is refusing to disconnect: “I may be one of the few VPs who still reads code.” He reviews code, reads papers, and challenges everyone’s ideas. “If I don’t do those things, I won’t understand where the pain points are, and I won’t see clearly who contributed what. I want to stay grounded.”

24. 黄仁勋’s operating system: first principles, prioritization, and computer science’s laziness

  • Early on, 黄仁勋 often forwarded him papers and asked for summaries. “He dared to build GPUs without knowing computer graphics, and dared to enter AI without knowing deep learning.” His decisions came from first principles: “Clarify the essence of the problem, ignore external signals, and move toward the direction that makes the most sense. Early on, he dared to go all in on deep learning and CUDA.” Before meeting 黄仁勋, 刘洺堉’s decisions were more short-term. Watching him make long-term decisions and persist had a major influence.
  • Prioritization is a daily discipline. 黄仁勋 reads huge volumes of email and forces himself to rapidly decide what matters, then executes without personal attachment. “My real goal is to take care of the entire company. As long as you stay in the company, I’ll take care of you.” 刘洺堉 changed from finding it exhausting to run 2 or 3 projects simultaneously to switching topics almost every 30 minutes.
  • His accompanying philosophy of computer science: “The most important thing in computer science is laziness. Don’t compute what you don’t need to compute; compute later what can wait. We only have so much time in life. Knowing what not to compute makes your life much more efficient.”

25. Martial arts, ice baths, and 4 a.m. starts: gentle in person, direct about problems

  • He has practiced Chinese martial arts for years since age 18. His teacher was 徐纪, who studied under 刘云樵, in a lineage tracing back to “God Spear” 李书文’s Bajiquan and Pigua Palm. He started because of Jin Yong novels, then discovered that practice was a form of self-cultivation: the maximum force of a punch comes from coordinating the entire body, transmitting power from the feet to the hand. “It requires constant practice and constant self-correction, which has a deep connection with my tendency to keep examining myself.” His latest hobby is ice plunging. His routine: exercise at 4–5 a.m., arrive at the office for meetings at 9–10 a.m., and sleep at 10–11 p.m.
  • Though he comes across as mild, he points out problems directly at work. “After doing this for a long time, you get a little frustrated that people aren’t living up to their potential. Not everyone is willing to tell you where you can improve further.” People who take it positively are grateful; those who see it as nitpicking think he is harsh.

26. Cosmos 97: an offhand number and a serious commitment

  • After Cosmos 1, he asked 黄仁勋 whether they should continue. The answer was: “Take it to Cosmos 97.” His interpretation: “The number was off the cuff, but his commitment was expressed in that number. This thing will keep moving forward.” If each generation took 6 months, 40 years would be required. “AI is developing extremely fast, so perhaps it won’t take 40 years.”
  • Jensen’s discipline matters too. Even for important projects, “he’ll tell you to pace, pace your steps, and not rush.” He is good at running a marathon. Timing also matters: before Sora, it was difficult to persuade people; after Sora, it was not. 黄仁勋 saw it once and immediately understood its impact on the company.

27. Building the Matrix for robots: the 3-part generalization stack and the pain of algorithmic intelligence

  • His personal homepage says he wants to “build a Matrix for robots.” The underlying problem for any AI is generalization—performing well outside the training data. Cosmos helps in 3 ways: better data, by generating data “that you need but the real world has not collected”; a better starting point, where a model that understands the world and predicts the future can serve as a backbone for policies; and a better environment, by generating multiple highly realistic environments in parallel so robots can learn through interaction faster. “The most challenging part is better environments.”
  • On the possibility of robots waking up and escaping the Matrix, his philosophical answer is cautious: “Pain and suffering are important drivers of human progress. I’m not sure whether intelligence made of data and algorithms has the same pain and suffering. I still have doubts.” Guardrails are also necessary to ensure technology does not disrupt the normal development of the world.

28. Inside the launch: March 2024, a group email, and being forced to name the project

  • The timing was around March 2024, after Sora. A group working on generative models and media generation—including him, Sanja Fidler, and Bryan Catanzaro, who runs the NeMo team—wrote to 黄仁勋: “We have to build this.” The meeting’s broad consensus was already in place; the discussion focused on how to execute and who would own what. 黄仁勋 asked him and Sanja Fidler to lead.
  • The naming episode carried a product lesson. He knew 黄仁勋 might eventually come up with a better name and wanted to ask him directly, but was told to think for himself. “The name is closely tied to the product’s positioning. If you call it Cosmos, the goal is not to generate content or create things, but to solve problems in the physical world.” Nvidia is more focused on the developer market; it does not directly operate social networks or creative tools.

29. The debate over what a world model is: purpose determines form; Cosmos is the base layer

  • His practical breakdown is that the purpose determines the form. A world model can predict the future, understand why a factory line has stalled—perhaps a worker forgot to turn on a machine—or reconstruct a 3D world for route planning and virtual exploration. “Ultimately, we live in the same world. These models describe the same thing through different observations.” He now focuses on understanding and generation for the physical world, essentially prediction, rather than reconstruction.
  • He is blunt about the definitional debate: world models may repeat the AGI experience. “After so many years, there is still no common definition. What impact does defining it have on humanity? I don’t think it’s very important.” LLMs are precise because “large, language, and model are all clear terms.” Cosmos calls itself a World Foundation Model: “We want to be something more fundamental. Even a company building its own world model should be able to build on our foundation. Not every company pursuing physical AI has the resources to build its own world model from scratch.”

30. No competitors: World Labs and LeCun’s AMI are investment and collaboration

  • Nvidia has invested in 李飞飞’s World Labs and 杨立昆’s AMI. “Each company sees a different point and wants to solve a different problem. We help all of them and want them to succeed. Companies with less capital than Nvidia need more help and can leverage Cosmos.”
  • “When I built Cosmos, I didn’t want competitors. If companies using our compute platform in the Nvidia ecosystem succeed, Nvidia succeeds. I’ve moved past the period when I needed to publish an algorithm better than everyone else to graduate.”
  • His message to peers includes a sales joke: “A world model is a tool. Ultimately, a world model alone is not enough to solve the problem. If you use our open-source work to build a better model than ours, and that model is open source too, that’s a good thing.” If someone says, “Nvidia is coming down to sell me GPUs again,” his answer is: “You need compute, right? Why not? What matters most is how much profit you make. A cloud company may buy a lot of GPUs, but it makes even more money.”

31. From 22 models to one: the Louis Vuitton lesson and the inevitability of an omni model

  • During the Cosmos 1 era, he calculated that meeting every partner’s needs would require 22 models. He told 黄仁勋, who first replied, “Very good, very hard,” then told the Louis Vuitton story: after reducing the number of products on display, the company’s sales rose. “You don’t want to confuse your customers with too many things.” 刘洺堉 admitted, “I couldn’t even tell whether I should use A or B.” Every model also had to be maintained. “The more things you prepare, the thinner your advantage,” and iteration slowed.
  • There were 2 theoretical reasons to unify the models. First, they describe the same world and learn the same internal representation. Second, a foundation model should absorb everything: combine Reason, Predict, Transfer, and Policy data so the model’s ceiling rises. After Cosmos 2.5, the team moved toward Cosmos 3’s Omni model. “The move from one to two was a data upgrade. The move from two to three was a huge gap.” Work began around October last year and continued for more than half a year.
  • The iteration doctrine came directly from DeepSeek: “DeepSeek is an amazing company. It iterated 3 times in a year and produced DeepSeek-V3. I’ve always emphasized that we need to iterate quickly, with each generation better than the last.”

32. The debate over Cosmos 3’s “limited innovation” and full open source

  • A Chinese world-model founder spent a week reading the papers and concluded that there was “not much major technical innovation, but many details were validated through engineering.” He does not dispute the characterization: “We were the first to scale a model this far and combine audio, video, and action. Combining them conceptually is easy; actually making it work is not.” All major language models may use the Transformer architecture, but details still determine why one model outperforms another.
  • He increasingly cares less about whether work is revolutionary or original. “The important questions are whether it is useful, whether it creates impact, and whether it solves the problems people want solved.”
  • Nvidia open-sourced the model, training framework, and parts of the data. Frontier labs are moving toward closed systems and do not disclose key details even when they publish papers. “We decided to take a different path. We explain everything we can explain.” He is not worried about someone reproducing the exact same thing. The connection with DeepSeek is the desire to create impact and disclose details, but the goals differ: “We want users to succeed and physical AI to enter the real world faster.”

33. Action as a first-class citizen: weighing 杨植麟 against 谢赛宁

  • The key input-layer decision was to take language, video, audio, and action simultaneously. “There are basically 2 ways to think about a foundation model: make language the first-class citizen, as in coding and chat, or make visual signals the first-class citizen.” Agents in the physical world take action and change the state of the world. “A good physical-world foundation model has to make action a first-class citizen too.” He would like to add touch, but “touch sensors are not good enough yet”; dexterous hands and better data have to mature first.
  • 张小珺 raised the opposing views of 2 researchers: 杨植麟 worries that vision could reduce an LLM’s intelligence, while 谢赛宁 worries about “language contaminating vision, with language serving only as a scaffold for humans.” 刘洺堉’s answer is teleological. “Video captures the physical operation of the world more precisely; language makes that difficult. But the model ultimately serves people. Without language, it cannot communicate; without action, it cannot change the world. My goal is not AGI. Perhaps intelligence could emerge directly from visual signals and develop another civilization, in which case language would be interference. It depends on what you want. I know clearly what I want.”

34. Dual towers, modality fusion, and “deep learning is like Chinese medicine”

  • The evidence for fusion is empirical. Cosmos Policy found that predicting the next observation alongside an action helps training converge—video helps action. A “click” can calibrate the moment of contact—audio helps video. More detailed text descriptions make video easier to generate—text helps video. What is the secret sauce? “There is no secret sauce in building a foundation model. Run rigorous experiments, control variables, accumulate knowledge, and build carefully step by step.” Deep learning “is a lot like Chinese medicine. You try things, keep detailed experimental records, and understand the intermediate patterns along the way.”
  • The dual-tower design—one discrete tower for understanding and one continuous tower for generation—is a practical choice for now. Customers can separate the two during post-training. If everything is tightly connected, changing the generation distribution can also affect understanding. “It is relatively inelegant. Ideally, the data should determine the architecture. The next goal is simplicity. Why do we need 2 towers? One should be enough.” Work such as Transfusion is exploring unification. Ultimately, “the model that is easiest to deploy, in terms of cost, is most likely to succeed.”
  • Could the team also beat Claude at coding? “The more things you prepare, the thinner your advantage. Strategy is often about what you decide to give up.” Robots do not need to solve advanced mathematics. “You care more about whether they can assemble things in a factory and clean a home. Will their architecture really be the same as an agent model? Looking across human history, the answer should be no. Tools for farming are different from tools for weaving. The wisdom of our ancestors is still pretty useful.”

35. Data, evaluation, 10,000-GPU-scale compute, and Nvidia’s product matrix

  • Nvidia divides physical data into navigation, which is easier to collect because displacement and rotation generalize across body types, and manipulation, which is harder. The trend is clear: “The world was built for humans. Most tools are suited to human hands, and humans are the ones collecting the data, so everyone is moving toward human hands.” Cosmos 3 specifically added egocentric data: how people see from a first-person perspective and how human hands manipulate objects.
  • Evaluation has 3 steps: benchmarks, arena-style comparisons, and solving real problems with partners who have genuine pain points. The third is especially important in physical AI. “In coding, everyone is solving the same problem. In physical AI, every problem is different. This is an important step in guiding model development.”
  • Cosmos has reached “10,000-GPU scale,” and he hopes to go further. Nvidia has 2 model tracks: Agent AI, including the NeMo team’s Llama Nemotron Ultra, which uses a hybrid architecture combining some Transformer layers with Mamba-style layers to reduce inference costs; and Physical AI, where Cosmos is the base and autonomous driving’s Alpamayo and robotics’ GR00T are specializations. “The internal autonomous-driving and robotics teams are also my customers.”
  • The hard advantages are relatively low GPU compute costs, the Thor chip, and the Isaac Sim simulator. The software advantage is trust: “Customers are willing to work with us because we won’t build the exact same thing and eliminate you.”

36. How 300 people make decisions together: aligned goals plus distributed authority

  • The organizational challenge is precise: “You want everyone aligned on the goal, but you also want everyone to be able to make decisions independently.” If only a few people decide, the process is slow. “Suppose only I can decide the data mix and architecture. The model won’t succeed because my knowledge and time are limited.” If everyone decides independently without shared goals, the organization fragments and horse races emerge. The answer is transparency and a clear vision drawn in advance. Before the Cosmos 3 paper was released, he had already described the finished product to each team, and it ended up “perhaps 90% similar.” It is like a basketball player visualizing the ball dropping cleanly through the hoop before taking a free throw. The organization also needs tolerance for error: “People afraid of making the wrong decision won’t make decisions. But when a decision is wrong, you have to correct it quickly.”
  • The meeting load is intense: more than 10 meetings a week, with 2 or 3 recurring sessions each for video and understanding. “You can hold no meetings and let everyone work separately; that may not succeed. You can also hold meetings constantly and accomplish nothing; that won’t succeed either. You need a balance, depending on the task and the moment. There is no fixed formula.” Is Nvidia bottom-up like OpenAI or top-down like Anthropic? “It’s both. Extremes reverse themselves. The world divides and reunites. You have to understand what is right for the team at that moment, rather than blindly following one model.”
  • What pleased him most and what he regrets: he told the team they could combine every past model and still outperform each standalone model, and “many people didn’t believe it, but it happened.” The regret is time. The model had not fully converged when they had to release it, so 3.1 and 3.2 will continue to add unfinished capabilities and experiments.

37. The ChatGPT moment, humanoids, and the US-China robotics gap

  • He first distinguishes the terms: “The GPT moment tells everyone that something can scale. The ChatGPT moment makes it clear how a model will change people’s lives.” His physical-AI version is concrete: “A person demonstrates how to do something once and the machine immediately learns it. You give physical AI an instruction manual and it knows how to operate the instrument. If I see a product like that, that is the ChatGPT moment for physical AI.” Timing: “Maybe this year, maybe next year. I don’t know, but I think it will inevitably happen.” Capital, entrepreneurs, and the other conditions are gradually converging.
  • The humanoid debate starts with data. Humanoids can use large amounts of human data; a wheeled base with a semi-humanoid upper body can solve many problems. Full humanoids will be important because the human environment was designed for human bodies, but the key is still operating cost. “That is beyond my expertise.”
  • 谭杰 of DeepMind forecasts a humanoid inflection point in 2035–2036. 刘洺堉 says, “I hope it comes earlier, but I’m not sure.”
  • Sim-to-real is also a generalization problem. Better Environment “has not fully arrived”; for now, Better Data and Better Starting Points work better. A model that captures physical patterns and predicts what happens next can help build a more generalizable policy. The core issue is “generalizing the observation-action pair—learning an operation from limited signals and applying it to a scene it has never seen.”
  • The US has relatively few companies building robot hardware, mainly Tesla and Figure AI. China has many more. “They can control more of the stack, with more room for integrated hardware-software operation. China’s strong manufacturing base lets these companies do things that are difficult to do in the US.” Is the US software-first approach wrong? “It’s hard to say. But if hardware-software integration is the direction of robotics, then having no hardware leaves you in a passive position.”

38. Nvidia culture: 30 days of cash, no easy layoffs, top-five emails, and “mission is the boss”

  • 黄仁勋’s self-pressure has become a company ritual: “Every day he feels the company has only 30 days of cash left and will go bankrupt after 30 days.” He uses every meeting to remind himself to make the decision best for the company. During the pandemic downturn, Nvidia considered cutting travel, asked whether it could operate without Slack—“we have email anyway”—and considered eliminating a cloud service to save money. “Some companies reduce personnel costs through layoffs. Nvidia wanted to eliminate unnecessary spending and get through the difficult period together with employees.”
  • During his time at Nvidia, 刘洺堉 says he never heard of layoffs or forced ranking. The logic is explicit: employees with identification and security are more willing to raise concerns they would hide if they feared being cut. In a place where everyone is fighting to stand out, cooperation is harder. The cost is also clear: when technology matures and people need to transition, retraining is slower than laying off staff and hiring experienced talent. “But you gain employee trust.” Nvidia also avoids horse races: “If you take on this mission, we trust you—but we set a very high bar. In 老黄’s words, we torture you to greatness.”
  • The company has 2 distinctive mechanisms. Top-five emails allow “small signals from every corner of the company to spread”; if 2 teams are doing the same thing, it tends to surface naturally. “Mission is the boss” means Nvidia behaves like an organism rather than a fixed hierarchy. When the company moves in a direction, teams with relevant skills align themselves automatically. Cosmos is one example. “Your boss is not your boss. Mission is the boss.”
  • Looking back over 10 years, TPU once appeared and Nvidia’s stock fell sharply; some colleagues lost confidence and left. Nvidia later beat TPU repeatedly in MLPerf. “These changes are difficult to predict.” Another offer when he joined came from an investment firm: “Whatever they pay you, I’ll give you 4 times that.” He still chose research. His description of 黄仁勋: “A hexagonal warrior. In meetings he can drill down to register-level details, then teach marketing, pricing, and customer relationships. He gives everything he has to help the company succeed.”

39. “Are you a crybaby?”: no more excuses

  • The famous exchange came during an internal comparison of 2 solutions. 刘洺堉 explained that the other side’s settings were unfair, which was why his team was behind. 黄仁勋 replied: “Are you a crybaby? Nothing in this world is fair. If your goal is to solve the problem, go get the conditions you need to solve it.” “I had heard ‘no excuses’ from West Point, but hearing him say it made a deeper impression. I stopped making excuses after that.”
  • The worst suffering he has experienced was not a single incident but a constant condition. After GPT, “everyone in AI felt enormous pressure. We consumed huge amounts of GPU resources and young people’s time, and the results failed to meet expectations. All of that created invisible pressure.”

40. The number-one philosophy: have the ability to be first, but do the most valuable thing

  • His arc runs from being highly competitive to his current view. Researchers graduate by achieving SOTA and “beating everyone,” but “you are SOTA for a while, and then the next one appears.” His current formula: “Have the ability to be number one, but do the thing that contributes the most to the world. You can build the strongest model and make it available only to a small number of people, concentrating power in too few hands. I don’t think that is necessarily right.” There is also a simple addendum: “Live a long life.”
  • Is low ego real or just hidden? “It’s a form of confidence. I set a very high bar for what I build, but my goal is not to beat competitors. It is to propose a better foundation so everyone can move faster.” He has also shed the obsession with external metrics: there will always be someone with more papers, more citations, or more money. “The point is whether you are proud of what you are building. Maybe what I care about now is whether enough people are being helped.”
  • Asked about an S-class team versus an A-class team, he uses a game metaphor: “Don’t make the entire team archers. You need people to defend and people to heal. Diversity is better. A group of very smart people who don’t trust one another, who are all guarding against each other and trying to prove they are smarter, will struggle to make one plus one greater than two.”

41. China and 2026: interns spread know-how, the LLM hypnosis debate, and IPOs

  • Chinese models are “very impressive.” He names DeepSeek, Qwen, Doubao, MiniMax, Kimi, and MiMo. The structural observation is more interesting: “US frontier labs now hire very few interns, so know-how stays inside. Chinese companies hire many interns, helping transfer know-how. That is good for the diffusion of knowledge across humanity.”
  • Does he agree that Silicon Valley has been hypnotized by LLMs? No. “You can’t deny that the leading companies all started from large language models. I don’t know how I would get through the day without a coding agent now.” Once something is fully understood, people will find new areas for breakthroughs. “It’s hard to imagine people this smart spending their entire lives on LLMs. For them, working on LLMs now may be exactly the right thing to do.”
  • Against the backdrop of possible SpaceX, OpenAI, and Anthropic listings, 刘洺堉 said “2026 will be a year of major change,” and discussed the wealth and recognition that IPOs can bring. But technology will not stop there. “The companies that were powerful 100 years ago are completely different from those that are powerful today.”
  • Can new players still catch up? “Kimi still has a chance. MiMo came later, but there is still a chance. Compute will gradually come online, and there are many ambitious, determined people who know how to move resources and build a large-model team.”

42. The endgame: models become commodities, market creation is not about killing rivals, and CUDA’s risk lesson

  • His endgame view is explicitly controversial: “Model capabilities will gradually converge. You will need other things organized around them. Everyone can write software now, and the software companies that survive have their own way to differentiate. There is no reason models should be different. I treat models as commodities. So many people are trying different things; eventually it becomes a body of knowledge that many people possess. The question is what you do with it.” Does every company need to train its own foundation model? “If it’s a commodity, do you need to build your own power plant?”
  • The host later used the relationship between Dario and Sam to ask whether teams could split apart. 刘洺堉 said anyone can become convinced that their own model is the right one, and people can move between companies and take ideas with them. Data and compute, by contrast, can be obtained through resources.
  • Is Cosmos a war Nvidia cannot afford to lose? “If other companies are also building open-source models, then no. What does war mean? I’m not trying to kill anyone.” The mindset is market creation. “Right now no household has a robot. If households did, my family might buy 2 or 3, and every node would need compute.”
  • The next CUDA? “The core software of the future will be AI. NeMo and Cosmos could both be future CUDAs.” But he rejects a one-to-one analogy: “Many stacks combined may become the new CUDA. It won’t be a single model.” CUDA’s path also cannot simply be copied. The key is to use first principles and current conditions to make a contribution, not to carry a hammer around looking for a nail.
  • The risk lesson comes from CUDA itself: “Building CUDA was effectively selling at a loss. For the same FLOPS, our price was higher than competitors’. Not doing it would have been even riskier. If you don’t, your product gradually becomes a commodity.”
  • His advice to young researchers in the closing rapid-fire segment: “Don’t be nervous. Young researchers often worry that AI will end next year or become self-training. Have confidence in yourself and calmly identify where it is worth investing. This is a good time to enter physical AI.” He also mentioned AI’s potential impact on materials, drugs, biology, and other fields.
  • His answer to the Squid Game question was direct: “If everyone wants to solve exactly the same thing, then it is a Squid Game. Model competition shouldn’t be a Squid Game. There are so many problems in the world to solve and so many ways to create value.”