A 4-Hour Interview with 柯丽一鸣 of Physical Intelligence on Robotics
A 4-Hour Interview with 柯丽一鸣 of Physical Intelligence on Robotics
Summary
- Pi (Physical Intelligence) was founded 2 years ago and now commands a post-money valuation above $5B, while its 3 core papers each have a defining keyword: π0 is capability, π0.5 is generalization, and π0.6 is performance.* The highest-leverage finding came from π0.6*: once a model is roughly capable of doing a task, let it operate on a real robot, collect “experience data,” and feed that data back into training. On the first job, it “clearly outperformed the best data collector.” Once deployment begins, data costs could fall: “You’re deployed and can collect real-robot data—why wouldn’t you?”
- 柯丽一鸣 is a true believer in real-robot data. As of the 2026 snapshot, simulating the physics of folding clothes—including fabric, adhesion, and friction—may be beyond what anyone can build; sim-to-real remains “half-frontier.” “It’s very possible that all 10M hours of data are garbage data.” π0.5’s curve flattened after data collection across roughly 100 Airbnbs, suggesting that a finite amount of in-distribution data may be enough—but she added: “We still have far too little data, so we definitely need to be more aggressive.”
- Not building a humanoid was a condition of her joining Pi. Her logic: a general-purpose brain should manipulate cars, excavators, and every other morphology the way the human brain does, rather than being tied to a human-shaped body. If a simple morphology can handle a complex task, the odds of transferring it to a humanoid later are “very high.” The wild-card argument is that wheels do not exist in nature, yet roads everywhere are built for cars—the world will reshape itself around useful forms.
- The industry’s genealogy is the underlying map of the robotics race. The CMU traditionalists—Matt Mason’s “dexterity is in the brain, not the hand,” leading to Boston Dynamics—stand opposite Berkeley’s learning school, from 吴恩达 to Abbeel to Sergey Levine and Chelsea Finn to Pi. Google’s 2023 decision to cut back robotics for large-model work became a “Whampoa Military Academy,” spawning Pi, Scale, Generalist, Sunday, and others. Covariant’s early focus on logistics and warehousing taught Pi to treat commercialization as a “distraction”: it supplies models to deploying companies only as a partner.
- China’s supply chain is the variable the US cannot catch. “It’s hard to imagine a robot with not a single component from China… To this day, I don’t know how Americans would catch up if they tried.” Chinese robotic arms cost about $10K versus tens of thousands of dollars for Franka; the industry joke is that “Tesla’s best factory is in Shanghai.” The open-source catch-up paradox is equally real: following Pi’s public papers is like following Google. “To catch Google, you first have to train a Gemini-sized model.”
- Robotics has no clearly defined frontier, and evaluation is the biggest bottleneck. Infinite initial conditions mean there is no standard racetrack or leaderboard, while each company’s internal metrics are incomparable. π0.6* proposed throughput—the number of successful tasks completed in a fixed period—as an answer. Commercialization is “closer than I initially thought”: controlled environments such as factories are the natural first market, while robots entering homes may arrive “more as exploration than as a product, and it’s not that far away.”
- The softer signals are worth recording too. Claude Code let her rebuild her entire workflow in 2 weeks and improved her efficiency by roughly 3-4x, while making this introvert rethink the value of the L in VLA. She rebutted AI extinction arguments through permissions, responsibility, and capability. Her ultimate romantic vision: “A robot assembling and building itself is a form of reproduction,” and perhaps “you build something that can see the end of the universe.”
Deep dive
1. Founded 2 Years Ago, Valued Above $5B and Touted as the OpenAI of Robotics
- 张小珺’s framing: Pi is a Silicon Valley star focused on “the brains of robots,” founded just 2 years ago with a post-money valuation already above $5B. Guest 柯丽一鸣 (K) leads the reinforcement-learning work and is a core paper author; she also writes science-fiction web novels in her spare time.
- Timing matters: this episode was recorded relatively early and does not include the latest π0.7 model. The technical discussion stops at π0, π0.5, and π0.6*.
- The cold open delivers the episode’s 3 hardest lines: no humanoids, robots building themselves as “reproduction,” and the 3-paper arc defined by capability, generalization, and performance.
2. Writing Fiction and Building Robots Share the Same Source: Creativity Plus Execution
- Her analogy is that when a robot cannot complete a task, “you need to come up with a new method”; likewise, the story you really want to write has not yet been written, so you have to create it.
- The second half is all execution: research means turning an idea into code, while writing a novel means “typing it out one word at a time.”
- She refuses to disclose her pen name: “Help, the alias—the alias must be protected.” Work and personal life are “kept very separate.”
3. Science Fiction’s 2 Questions: What Do People Do After Productivity Changes?
- The science fiction she obsesses over revolves around 2 questions: how people’s lives change after productivity changes, and how people relate to one another in future societies. The first is directly tied to work: “Once we automate what people do, what do people do? It’s a question I think about every day.”
- On the second, she sees a constant: “There should be aspects of human nature that survive and circulate even in a technologically advanced society with dramatically higher productivity.”
4. Claude Code Became an Intermediary Between People
- Joining Pi taught her that “some things can only be done through team collaboration,” but Claude Code is loosening that collaboration from the other direction: instead of asking the person responsible for a module, “why don’t I just ask Claude Code?” She now has “one person controlling 3-4 agents.”
- That led to a novel premise: everyone has a bot that crawls, digests, and feeds them the internet. “Without this bot, it will become very hard to survive in the vast ocean of the internet.”
- Her end-state vision for solo operators is to hand robot repair to “the eternal problem of American society”—fixing the plumbing. Or discover a small planet, settle there, and deploy a fleet of robots to build the infrastructure; 1,000 years later, “maybe they’ll look back at us as a primitive society.”
- A generational aside: the next generation may date only robots. “I can actually accept that, because I’m a longtime otaku.” She already knows people who do not date humans and can understand those choices.
5. Why the AI Extinction Argument Fails: 3 Constraints—Permissions, Responsibility, Capability
- She cannot make sense of the logic by which a creation meant to help humanity becomes its enemy, so she offers a 2026 snapshot: the worst thing an agent has done is, after receiving permission, “delete most of the user’s mailbox.” The task was executed, but it had no overlap with the user’s actual intent.
- Permission: everyone prompts, “If you’re not sure, don’t act—check with a person.” Responsibility: after injuring herself while rock climbing, she used AI to answer medical questions, but “there’s no way I would believe whatever the AI says”; it must be taken to a human expert. “Trust is built by whoever can bear responsibility.”
- Capability: the industry’s self-deprecating joke remains, “Teach a robot to put the red block on the blue block; a 3-year-old can do it—why not just have a child?”
6. Wuhu, 江涛’s Lost Hands and Learning Logo at 8
- She grew up in Wuhu, Anhui, where educational resources lagged the major cities. But the city had an excellent informatics-competition teacher, 江涛, who lost both hands in an accident as a student, returned home, and built a systematic pipeline for training Olympiad competitors. They never met, but she considers herself one of his beneficiaries.
- She entered an after-school program at 8 for a simple reason: “I could play with computers. I was so happy.” She started by drawing with graphical Logo.
- The only programming textbook she ever read carefully was C Language for Middle School Students. It explained recursion through a genie granting wishes, “bringing it into everyday life very naturally” until she internalized it.
7. From Crying on the First Day to “Sister Ke”: Obedient-Looking but Rebellious
- She showed up to the first day of elementary school with a buzz cut and was mocked by the entire class, “crying all the way to school.” By graduation, she was the school’s big sister: “If you have an idea, we can sit down and talk; if that doesn’t work, I’ll beat you into submission.” She specialized in dealing with bullies.
- Her own summary: “A child who looks very well-behaved but is actually very rebellious.” She has a clear view of fairness: “No one has the right to demand that another person act according to their wishes.” She maintains justice only within her own capacity.
8. 3-4 Tsinghua or Peking University Admits per 10,000: Anhui’s Educational Pain
- Her painful accounting after moving abroad: in Anhui, only 3-4 out of every 10,000 students can get into Tsinghua or Peking University; in major cities, the figure may be 80 per 10,000. “In Anhui, you wouldn’t even dare imagine that number.”
- She does not see herself as a winner of high-pressure education: “I was average in the competition class. I studied more of what I liked and less of what I didn’t.” She regards the students who can grind through it with “the highest respect.”
- Today’s solution operates at a different level: “Create newer forms of productivity and, at some level, flatten these inequalities.”
9. A Solo Southeast Asia Backpacking Trip in Her Second-to-Last Year of High School
- Her top-3 act of rebellion was backpacking alone through Southeast Asia for 2 weeks at the end of her second-to-last year of high school. She had become fascinated by Indian mythology novels and religious architecture. “There was no clear purpose, and I didn’t know what impact it would have on my life. I just really wanted to do it.” Her parents urged her to come home every day, then eventually came to find her.
- She cut her hair into a buzz cut and disguised herself as a boy as preventive self-protection, after seeing social news suggesting women are “more likely to be placed in positions where they could be harmed.”
- The host noted that some people would simply choose not to go. Her answer is a character note for the entire episode: “If there’s something I want to do, I have to find a way to make it happen.”
10. Bloodstained Rain, Gu Long and the Sudden Decision to Study Abroad
- The reason her conventional academic path stalled was disarmingly honest: “I can’t do things I don’t believe in or agree with.” She had to memorize biology and chemistry and shore up her weak subjects. “Why should I work harder? Why should I make one more push? Isn’t it okay if I just don’t do it?”
- The turning point was a game. She encountered Bloodstained Rain and was “blown away,” seeing it as a perfect interactive form of Gu Long’s novels. Its founder seemed to have graduated from Tsinghua before going to Yale: “For the first time, as I was walking down a rebellious road, I realized that this person was a formally trained genius.”
- A competition-based early admission would have sent her to a place like USTC—“a good university, but too close to home, so I couldn’t go.” She abruptly pivoted toward studying abroad midway through her second-to-last year of high school, following her idol’s path.
- Games gave her a humanistic foundation: “People are born to experience things.” She was drawn to collisions of pure emotion: “You carry a blood feud. How would most people in modern society carry a blood feud? You can experience that in a game.”
11. Psychology to Economics to Computer Science: Asking Why Until the Bottom Falls Out
- She changed majors twice in college. Psychology was “very interesting, but I really wasn’t suited to it”: she wanted to keep asking why, but the path was too indirect. Economics helped her understand the world, “but it couldn’t help me understand what I wanted to do in the world.”
- Her self-diagnosis for not making games: fiction is catharsis, while AI research is “an expression of my pursuit of efficiency.” The satisfaction comes from “doing something nobody else can do, or doing something better than everyone else.”
12. A First-Author Paper in Freshman Year: 李博, Adversarial ML and Game Theory in GANs
- American-style competition started immediately: “I began sending out résumés everywhere in my second month of college.” Her résumé included software she and competition classmates had built and sold in high school. That got her into a lab, where she met her first academic mentor, 李博, who later joined the University of Chicago faculty.
- The topic was a line of fate: she had studied economics and loved game theory, while 李博 worked on adversarial machine learning—the two were structurally isomorphic. She published a first-author paper in her sophomore and junior years, laying the groundwork for PhD applications.
- Her GAN explainer: the generator makes counterfeit money, while the discriminator determines whether it is real. When the discriminator can no longer tell, it can only guess. Through repeated competition, the images approach reality. “Game theory’s application in machine learning continues to this day.”
13. The PhD Arc: From Imitation Learning to Reinforcement Learning
- She entered the program expecting to do theoretical ML, but gradually preferred applications and moved into robotics. She started with imitation learning: “Someone gives you an example, and you figure out how to surpass it.”
- The more she worked on it, the less satisfied she became: “If you’re always copying other people, you can’t innovate.” She moved to reinforcement learning, which emphasizes exploration and pushing the ceiling higher. “That also fits my personality to some extent.”
14. The Robotics World in 2017: The Lab’s Only Learning Researcher
- At the time, putting machine learning on robots was “emerging but not mainstream.” Her advisor came from the traditional full-stack CMU school, working on planning and control and arguing that, in the most extreme case, one person should be able to complete almost any task. She was the only person in the lab working on learning.
- The traditionalist intuition was to connect the first, second, and third planning steps into a complete plan, with control ensuring that every step executed correctly. Her question started from data: “Could a robot’s intelligence and ingenuity come directly from data instead of from rules I specify?”
- The debate mirrored NLP’s 2016-17 shift from semantic structure to end-to-end systems: “Robotics may need to take a similar path.”
15. A Clash of Faith with Her Advisor: Elegant Guarantees vs. Black Cat, White Cat
- They argued frequently. Her advisor saw ML as “inelegant and without guarantees”: “You don’t even understand how to model friction, so how can you do this task?” Her response: “Maybe the task doesn’t require you to understand those things. Our goal is to let people who don’t know the prerequisites complete the task.”
- She describes ML’s mission in stark terms: “Make experts disappear from the middle of machine learning. Let the machine learn by itself.”
- She also credits her advisor with giving her a crucial standard: “So what if it performs best in your simulator? Everything we traditionalists build can run on a real robot. You need to run on a real robot.” Real-world deployment and performance became part of her research taste from that point.
16. Why Robotics Was Niche: $500K for One Hand
- The 2017 cost structure explains the niche. The neighboring Yejin Choi group spent $200K on an NLP dataset, already “more money than anyone had ever burned.” In robotics, a single Shadow Hand cost $500K-$1M, and when it broke, the PhD student had to repair it. “Sending it back to the factory meant losing tens of thousands of dollars.” A 3-finger Barrett Hand also cost over $100K.
- Even Franka, supposedly “the cheapest” option at the time, cost tens of thousands of dollars. Franka has since been acquired by a Chinese company; “maybe a Chinese version will come out for a few thousand dollars.” Chinese arms now cost around $10K.
17. 50 Years of Robotics: CMU and DARPA Ignite Autonomous Driving
- The lineage begins at Carnegie Mellon. Its robotics institute dates to around 79 and was established in the 1980s; the integration of perception—looking—and decision-making—moving—became influential there, later branching into disaster response and other fields.
- The 2004-05 DARPA autonomous-driving races made CMU “an overnight sensation” and helped trigger the founding of a wave of autonomous-driving companies from around 2006. It was the traditional school’s high point.
18. Founding Father Matt Mason: Dexterity Is in the Brain, Not the Hand
- Matt Mason, the founding figure in manipulation and the root of her academic genealogy, is best known for saying: “The key to dexterity is not the hand itself, but the brain.” Even a simple structure can perform very complex tasks. Her chopstick robot was directly inspired by this idea.
- His academic descendants spread widely: her advisor worked on path planning and built a “robot bartender”; another student, 桑北 (likely Sangbae Kim), went to MIT and built Mini Cheetah, “possibly the founder of the modern robotic dog.”
- Marc Raibert went on to Boston Dynamics and pushed the traditional school to its limit, building the first backflipping robot in the world with custom motors and showing everyone that it could actually be done.
- The Hawaiian shirts are an industry Easter egg. After colleagues criticized his unserious clothes at conferences when he was young, he decided, “From now on, this is what I’m going to wear.” Even the Boston Dynamics AI research institute’s team uniforms are Hawaiian shirts. “There’s a rebellious spirit in him.”
19. The ML Genealogy: 吴恩达 → Abbeel → Sergey and Chelsea
- The entry point for the learning school was straightforward: people entering the field in 2014-15 all watched 吴恩达’s online courses. His PhD student Peter Abbeel was “one of the earliest pioneers to put machine learning on robots”; he now teaches at Berkeley and leads research at Amazon.
- Sergey Levine—her reporting line at Pi and a co-founder—actually did his PhD in graphics and switched fields during a postdoc with Abbeel. “When I was doing my PhD, there weren’t many machine-learning textbooks. I learned by watching Sergey’s videos, reading his papers, and listening to conversations in their group.” He and Chelsea Finn were “very much on the same wavelength” and together advanced RL in robotics.
- The internal joke was “Sergey GPT.” Before ChatGPT, Berkeley researchers would send Sergey a paper draft and receive “an incredibly polished version, fast and good.” Even after ChatGPT arrived, it “couldn’t compare.” All of Pi’s science-fiction citations in its papers were written by Sergey.
- Chelsea’s profile: she swims for 1 hour starting at 4 a.m. every day, is extremely disciplined, and has “an animal instinct.” She drove many of π0’s task choices; when nobody knew whether they were feasible, she thought they were.
20. More Schools: 李飞飞’s Lineage, 宋舒然 and Scale’s 2 Founders
- 李飞飞’s students 朱玉可 and 金范 lead Nvidia’s GEAR Lab. 宋舒然 is a female scholar she “respects enormously.” Diffusion Policy came from 宋舒然’s group, while ACT and ALOHA came from Chelsea’s group—“milestone papers.”
- The CMU lineage also produced Deepak Pathak, known for curiosity-driven RL, who co-founded Scale with 阿比纳姆. Scale is explicitly the company name. Both are building brains for embodied intelligence.
- 阿比纳姆 is the kind of freewheeling thinker who is “impossible to predict from one second to the next.” He advocated early for robots to collect data outdoors and built a gripper mounted on a long pole, costing “$5 on Amazon,” to gather data—a precursor with “a lot in common with UMI.” His student 王小龙 now teaches at UCSD.
21. “Robotics Is Far Too Important to Be Left to Roboticists”
- Berkeley computer-vision heavyweight Jitendra’s line is: “Robotics is far too important to be left for robotists.” The hostile reading is, “These people can’t do it, so we’ll take over.” She agrees with the charitable version: “People from different fields have to work together.”
- She makes the contrast concrete. Given one data point, a large-model researcher thinks, “Collect 1M more so the numbers get better.” A traditional roboticist first asks what the data looks like—the action and hardware environment. The latter focuses on physical motion; the former can step outside the frame and ask what information is missing.
22. The Input-Output Curve: Why the Backflip Does Not Scale
- Her curve for the 2 schools is simple. Traditional methods have a clear input-output relationship, but the backflip required “many engineers continuously tuning parameters for this one task”; a new task demands the same level of investment. ML lets the algorithm do the work. “The amount of work behind today’s parkour videos is not remotely comparable to that era.”
- The end-state is democratized interaction. Robots currently “work better when an expert knows how to use them.” The direction of travel is to let an ordinary person interact naturally and get the desired result.
23. Why the Traditional School’s Ultimate Morphology Was a Dog
- The host asked why everyone built dogs in that era. Her physics-based answer: manipulation requires modeling friction. “Microphones, clothes, folds—millions of things are hard to model.” A leg’s contact with the ground is momentary, so “you don’t need to fully understand friction in detail.”
- Bottom line: dogs require less physics knowledge, while manipulators require more. A dog as a hardware platform may also be easier to obtain than a robotic arm.
24. The Chopstick Robot: One Pair of Chopsticks Solves 90% of the Problem
- In a conversation with her advisor, she offered the provocative claim: “I can solve 90% of problems with chopsticks”—grasping and placing. Her research philosophy was to attack the hardest case first: “If the algorithm can succeed on chopsticks, it should be easier on other methods. Everything left is engineering.”
- Her advisor bet on the martial-arts fantasy of catching a fly with chopsticks in an American movie: “If you can build it, you can graduate.” She scavenged parts from the lab warehouse and built the robot from scratch. The full system ran in 3 months; teleoperation was established in 6 months. She did it all alone.
- Her argument for the advisor was elegant. It did not matter that the hand-built robot was inaccurate: “If I can manipulate such an inaccurate robot and use chopsticks to catch a ball, then as long as the algorithm is as smart as I am, it should be able to do it even with joint backlash.” The first paper became a working system in 2018-19.
25. From Catching Balls to Catching Glass Balls in Midair: Imitation Was Not Enough
- The imitation version left her unsatisfied: “I was adjusting the motion every second and spending huge amounts of effort providing the robot with high-quality data. It was exhausting. How can lazy people like us do this? It has to explore on its own and exceed the data I give it.”
- After switching to RL, the robot eventually caught a glass ball swinging through the air—“even a person would find it somewhat difficult.” Her analogy is Olympic training: find the strategy best suited to your body and repeatedly train until it is muscle memory.
- The “master problem A first” approach directly answered the traditional school’s most serious objection: “If you can do everything but do nothing well, what’s the point of the algorithm?”
- She admits the challenge remains. Traditional systems for loading glass in factories offer guarantees on success rate, stability, and speed. “Nobody can say machine learning is equally robust yet. It remains a frontier problem to this day.”
26. Geography and Migration: Learning on the West Coast, Traditionalists on the East Coast, Everyone Headed to Silicon Valley
- Stanford and Berkeley give the West Coast naturally high ML density. The traditionalist schools at CMU and MIT are “more deeply rooted” on the East Coast; many of their newer scholars were hired only in the past 5 years.
- The signal from the past 2 years: “No matter where people are working, they’re moving toward Silicon Valley.” Entrepreneurs and professors in Boston have taken roles at companies whose offices are in Silicon Valley.
- Traditionalists are also “actively embracing the new toolbox,” while expert intuition still matters. The night before, she had been reading a motor blog written by someone from ML who thanked a traditional CMU professor.
27. The Silicon Valley Startup Map: Pi, Scale, Figure, One X and Dyna
- Pi and Scale are both “academics starting companies” to build general-purpose brains, but their wedges differ. Pi is betting on bimanual manipulation: the provocative thesis is that “if you solve pick-and-place, you can eliminate every household chore in life.” Scale, judging from its videos, is more focused on legs, humanoids, and dogs, and is closer to commercial deployment.
- In humanoids, Elon Musk’s 2021-22 comments about Tesla ignited the hardware community. Figure was founded in 2022 by a founder without a technical background but with a record of successful entrepreneurship: “It’s interesting to enter an unfamiliar field simply because you believe in it.” One X has gone deep on motors and pursues a distinctive cable-driven design.
- Dyna’s founder combines entrepreneurial and technical experience, with a stronger emphasis on commercial deployment. Its folding-clothes demonstrations belong to this category. Its 2024-25 publications are testing the proposition that ML can gradually move into deployment and create commercial value, making it one of the earlier groups to cross from research into deployment.
28. Google’s Robotics Whampoa Military Academy and the 2023 Startup Explosion
- ChatGPT was “so good that it shook many people’s assumptions.” Google’s decision to “effectively cut back” its robotics unit and pull people into large-model work instead released an entire generation. Google Research became robotics’ Whampoa Military Academy: all of Pi’s founders came from Google, while Figure and One X also absorbed people from the wave. Robotics startups exploded in 2023-25.
- The UMI twins came from the same wave: Generalist AI and Sunday Robotics. Their founders worked together on the UMI paper, using handheld grippers to collect data from people’s everyday lives and demonstrate transferability to robots. Judging from videos, Generalist leans industrial while Sunday leans household; Sunday currently has “the cuter-looking robot.”
29. The Big-Company Bets: Tesla Bets on Humanoids, Google on Multimodality, Nvidia on Selling Cards
- Tesla is “the most aggressive company on humanoids.” She personally thinks hands “would be very useful if you have them, but it’s fine not to have them now; they aren’t necessary.” Tesla wants the ultimate form, with “every aspect broadly humanoid” and a strong humanoid conviction. Its automotive experience leads people to assume it has formidable hardware capabilities.
- Google’s robotics effort may be part of Gemini’s multimodal map: giving large models spatial perception, planning, and control capabilities.
- The industry joke about Nvidia is that “everything it does is to sell cards,” which is why it works on card-intensive projects such as world models. “But history has repeatedly shown that sometimes, more data and more cards really do improve performance. There’s a beauty to brute force.”
30. The Humanoid Debate: Practicality Plus Wildness
- She describes herself as “a practical person plus a wild one.” The practical case is that capability, performance, and manipulation matter most right now; “you don’t need a humanoid to do these things, and you can already deploy them to get work done.” Humanoids may be solved by others first; “we can catch up later.”
- The wild argument is more memorable: wheels do not exist in nature, but roads everywhere are designed for cars. “If a form is truly useful, people will eventually reshape their living environment around it.” Why insist on 2 hands? “Could it be that the forms natural evolution could never produce are actually the best forms?”
- Her personal end-state is dynamic reconfigurability: a mature robot should be able to swap out body parts anytime and anywhere for more suitable ones, “like its tools.”
31. Defining a General-Purpose Brain: One Brain Drives a Car and an Excavator
- In response to the claim that a humanoid is the most general morphology, she splits “general” into 2 parts. Pi is building “a brain that can benefit many morphologies”: trained on data from different forms and used on robots with different forms. “That doesn’t mean it can’t work on a humanoid.”
- Her analogy closes the argument: “I drive a car, operate an excavator, and manipulate a robotic arm with the same brain controlling different morphologies. That is the most fundamental definition of a general-purpose brain. Building a brain for a humanoid does not make it a general-purpose brain.”
- Task complexity and morphology are separable. Humanoids have many difficult-to-control joints, but “if you can use a simple morphology to perform a more complex task, the odds of transferring it to a humanoid later are very high.” Building a microphone with 2 arms matters more than making a humanoid run better.
32. The Biggest Frontier Trap Is Evaluation
- Robotics has no leaderboard. Unlike other fields, there is no giant public dataset where every model runs the same race to see who comes first. The reason is that evaluation must happen on a real robot and the initial conditions are effectively infinite: cup position, lighting, background, table height, and angle can all affect performance. You want to evaluate everything, but reality cannot be controlled.
- The result is an invisible frontier. Every company has its own evaluation suite and priorities, so the field is fragmented and full of different approaches. The macro question is consistent: how to make a mechanical body perform many tasks well in real life. The wedge each group chooses is simply different.
33. How a Robot Is Born: From Motors to a 2-Layer Brain
- At the atomic level, a motor is “a joint that moves.” Connect 6 or 7 joints in arrangements corresponding to a person’s elbow, wrist, and shoulder, then fix them with 3D-printed or metal parts, and you have a primitive robotic arm. Low-level control translates a position signal such as “90 degrees to 180 degrees” into physical motion.
- Above that sits a 2-layer brain. The high level observes the task and makes decisions—“to pick something up from the table, first raise the arm”—then sends the command to the joints for execution. A new scene enters the next cycle. The goal is “something that can understand, execute, and also be easy to debug.”
34. Pi’s 3 Core Papers: π0 for Capability, π0.5 for Generalization, π0.6* for Performance
- To understand Pi, follow its core publications. “The core thesis is something that can push the entire field forward.” The keywords are capability for π0, generalization for π0.5, and performance for π0.6*.
- π0, released in November 2024, did 3 things: folding clothes at a level “we had never seen before”; folding cardboard boxes, asking “Can this thing really do it? Let’s find out”; and desktop busing, which required handling unprecedented object diversity.
- When Pi was founded in early 2024, “everyone was excited to apply the large-model approach to robotics, but nobody knew what it could become.” π0.6* returned to the old challenge: if a model can do everything but only halfway, what is it actually for?
35. π0.5: 100 Airbnbs and a Curve That Flattened
- The fundamental generalization problem is in-domain versus out-of-domain. “If training only makes you good inside the data, it’s useless outside the data and the impact is limited.” π0.5 left the office and collected data in about 100 Airbnbs—“messy, unfamiliar homes”—performing everyday tasks before testing on new homes.
- The most encouraging result was a flattening curve: perhaps the robot did not need data from every home in the world to work. “In-distribution data has a finite size; reaching that size may be enough.”
- The title Open-World Generalization Model may have come from Chelsea. As usual, the science-fiction citation in the paper was written by Sergey. “Only Sergey remembers.”
36. π0.6*: Experience Data Beats the Best Data Collector
- The asterisk means that data generated by the model itself is put back into the training pool. The core claim is that once an agent can almost do a task, having it collect its own “experience data” and return that data to training can outperform fixed data collection. This was the first publication from her RL team and emphasized methodological simplicity.
- The project was a bet: could a robot be faster and better than the data collector most accustomed to operating robots? The result “clearly met that bar, at least on the first job,” surpassing the starting point set by the best data collector.
- The evaluation metric followed: throughput, or the number of successes in a fixed period. Even defining success and failure required extensive work. “How to ensure a large model fully understands what a person means is itself an open research question.”
37. Data Costs Could Fall: Deployment Is Data Collection
- She breaks down the cost of real-robot data: the platform, bringing it into homes for maintenance, hiring operators, and managing them. “The biggest cost right now is managing the person who is collecting the data you want.” Replace the operator with a trained large model, and “the cost of data should fall a lot.”
- The same applies to correction data. The old imitation-learning problem is error accumulation, which pushes the robot into bad states it would never encounter while a human was collecting data. So teams collect repair data specifically. “But corrective data can also be collected by having the robot run by itself. That is something fundamental to reinforcement learning.”
- Her provocative conclusion: robot prices should eventually become cheap and convenient. “Many of the concerns people have about data today will be unimportant over the long arc of time. The data collected during deployment should all be available for my use.”
38. A Real-Robot Data Believer: Simulated Clothes-Folding Is Still “Half-Frontier”
- Her position is pragmatic: “Black cat or white cat, the good cat is the one that catches mice. Believe whichever data really works.” With the same $100K budget, real robots and simulation produce different quantities of data. The question is the quality and meaning of the data—who can make the model perform better.
- The hard constraint as of 2026 is fabric flexibility, friction, and adhesion. “You may not even be able to build a simulator for that in simulation.” Sim-to-real clothes folding remains “half-frontier”: nobody has done it extremely well, but some believe it is achievable, perhaps within 2-3 years.
- The warning is worth preserving: do not look only at “we have 10M hours of data.” “It is entirely possible that all 10M hours are garbage data.” If the goal is every scene and every task, real-robot data remains irreplaceable for now.
39. Does Clothes Folding Generalize? Infinite States Are Generalization
- In response to the criticism that the demo is impressive but may not generalize, she decomposes the dimensions. Clothes have countless states: one additional fold or a slight bend creates a different state. “The states you collect will always be finite, but you still have to guarantee completion.” Different garments are another dimension, and folding more categories in π0.6* is yet another.
- π0’s generalization showcase was desktop cleanup: hundreds of props could be picked up and swapped freely, with the data collectors using their own imagination. She helped design the task strategy.
40. Side Projects: FAST, Hi-Robot, Olympics and Partners
- Between π0 and π0.5 came 2 important side projects. FAST studied the optimal representation space for actions: the model predicts in a learned action representation, improving performance with methodological guarantees. Hi-Robot formally introduced hierarchical outputs within one model: a high-level policy turns instructions into shorter executable subtasks, solving the problem of a model getting lost during a 10-minute task.
- Olympics began as an internal event. Researchers teamed up with data collectors to teleoperate robots through tasks they “wouldn’t even dare imagine,” asking: “If a person can do it by hand, can a robot?” The exercise taught large models to unlock doors and open doors—actions that sound easy but absolutely are not. The first time she teleoperated a door opening, she spent more than 20 minutes thinking through it.
- The latest partner releases involve one clothes-folding company and one packaging company running models trained by or using Pi. Pi itself does not currently have many commercialization ideas, but partners can “indirectly push the research frontier.”
- Pi is not merely a hardware spectator. π0.6* uses proprietary hardware, including a low-cost swappable gripper designed for making coffee. “With this kind of R&D, you start to feel what the fundamental problem is.” Hardware is optimized less for software than for the final task and its performance.
41. Compared with Which GPT? Not Yet GPT-1, but Getting Closer
- Her direct answer: “I don’t think we’re at GPT-1 yet, but we’re getting closer.” What she sees every day is imperfection: this performance is not good enough, that one is not good enough; perhaps a major transformation is required to reach the desired result.
- Her brief 2017 stint in NLP provides the best self-deprecating footnote. Models trained to write Shakespeare produced daily “tsk-tsk-tsk” reactions. When she heard someone was using RL for language models, she asked: “Can you combine an extremely difficult A with an extremely difficult B and succeed? I had no idea. By 1920, someone had done it.”
- Robotics has one unusual advantage. After VLA emerged, “despite having nowhere near NLP’s data volume, it had already produced a lot.” Robots can complete concrete tasks with less data, something that is harder in natural language.
42. Scaling Laws and Compute: There Is Still Too Little Data, So Be More Aggressive
- She is wary of oversimplifying scaling laws: “Every time I see an oversimplified statement, I feel it lacks detail.” π0.5’s house curve was one scaling experiment—how many homes are needed to produce results in a new home, and after how many homes the improvement becomes negligible.
- But the conclusion is not to stop. “We still have far too little data. We definitely need to be more aggressive. I think it will be fun.”
- On compute, her answer became a quote: whether there is enough compute depends on whether the people involved enjoy working. If “everyone really likes working, compute will never be enough.”
43. Zero-Shot and Concrete Performance: Stepping on One Foot with the Other to See If You Can Fly
- Some say generalist teams care about zero-shot generalization while application teams care about making specific scenarios work, implying completely different needs. She sees them as complementary. Robotics has 2 central themes: how to make performance complete, and how to produce something credible in a new domain.
- Her intuition is that perfecting one scenario may improve results across more scenarios. “Research into performance on specific tasks can improve the model’s overall capability. That’s something I strongly believe.” Conversely, once a large model is roughly stable in a new environment, specific improvement techniques can be layered on immediately. “Step on one foot with the other and see if we can fly.”
- Zero-shot improvement comes from architecture and pretraining research, as well as data. Some believe the next-generation architecture will be a world model, but “there is no consensus yet on what it must look like. Everyone is exploring.”
44. 2024 to 2026: From Uncertainty to Visible Success
- In early 2024, she “really wanted to fold clothes but didn’t know whether it could be done.” By year-end, uncertainty had gradually become visible success. “You can’t guarantee 100%, but there is a possibility of success. That progress matters enormously.”
- In 2025, π0.5 entered Airbnbs and π0.6* pushed deployment performance to a level “purely collecting data would have struggled to reach.” Reinforcement learning came only after clothes folding had reached the point where “there was nothing more to do by collecting more data.” Gemini’s spatial understanding also impressed her and is “clearly very important for robotics.”
- Her commercial instincts have changed: “Some of what we do may already have commercial value. It used to feel distant; now it feels quite close.” The most natural environment is controlled settings, with factories as the obvious example: control the environment to fit what the system can do today.
- Her 2026 expectations include more astonishing demos—“seeing it is one thing; understanding the progress underneath is another”—humanoid manipulation becoming mature enough to explore, and major changes in model architectures.
45. Robotics vs. Autonomous Driving vs. Large Models: Error Tolerance Sets the Difficulty
- Autonomous driving is harder than robotics because of the stakes: “When you make a decision, a human life is involved.” The performance bar must be extremely high. It is easier in that control is nearly solved and the decision space is smaller: going from A to B does not require an endless sequence of C-to-D transitions.
- Robot manipulation is difficult at both ends. Payload and friction can prevent an arm from reaching the intended position, while folding clothes offers countless ways to rotate the wrist. Large models “can say something wrong, and everyone is used to it.” Robotics sits between the two, though “at its current stage, it aligns more closely with large models.”
- The lesson from autonomous driving is that leading companies all have their own simulators, with 1-2 public paid simulation companies as well. “We still don’t have an equivalent simulator for manipulation. Is there something to learn from that?”
46. The Hidden Race to Enter the Home: Whoever Ships First Must Be Safe First
- “The hidden metric every robotics company is competing for is who can build the first robot to enter the home.” The home is the implicit direction of competition: when you feel you have lost the plot, think about which household problems robots still cannot solve.
- There are 3 layers of difficulty. Environment: “No 2 people’s homes or bedrooms are exactly alike”; Pi has folded quilts, but not in a way that works in every new home. Actions: the list of household chores is endless. Hardware and safety: “I don’t know whose hardware could run in someone’s house for a month and perform all kinds of tasks.”
- Her most vivid fear: “What I’m most afraid of with Optimus is it falling in my house. How much would I have to pay to repair the floor? If it fell on me, wouldn’t I be, well, finished?”
- Smaller form factors involve trade-offs. A small machine like those from Unitree is less threatening, lighter, more energy-efficient, and instantly safer. But once human-scale tasks become smaller, “some things can no longer be done. So how do you use human data?”
47. China’s Supply-Chain Dominance: “I Don’t Know How the US Would Catch Up”
- Watching Unitree perform at the Spring Festival Gala, she thought that seeing a demo at that level 1-2 years ago was unimaginable. The result reflects both algorithmic progress and hardware strength. Company chats immediately turned to how it was built, and a reading group considered inviting a locomotion expert to speak.
- Her strongest industrial judgment: “It’s hard to imagine a robotics company assembling a robot with not a single component from China.” The manufacturing advantage of China’s industrial chain is so large that “to this day, I don’t know how Americans would catch up if they tried.” Rapid hardware iteration also gives Chinese teams an advantage in taking existing algorithms toward commercialization.
- Technically, she pointed to recent work on Omni Target from CMU professor 石冠亚, a friend from his postdoc days at BGI. Legged robots use high-level gait planning plus low-level execution; reinforcement learning has “achieved substantial success in locomotion.”
48. Covariant’s Lesson: Early Commercialization Is a Distraction
- Pi is “not considering commercialization at all for now” for historical reasons. Peter Abbeel founded Covariant in 2015 or 16 to build a general machine-learning solution for robots, then focused deeply on logistics and warehousing. “Looking back from the perspective of large-model development, that was a distraction: early commercialization cost it generality and transferability, and it failed to return to the underlying problem.”
- Pi took the lesson to heart: do not do things that do not help research just to make money or close a commercial loop. “Just maximize the research performance and think about commercialization later.” She offers a possible contrast: Chinese companies may emphasize pragmatism and payback more strongly for commercial reasons.
- She leaves the end-state product structure open: perhaps some companies specialize in brains while others commercialize specific scenarios, eventually inserting someone else’s brain into their products.
49. The Open-Source Catch-Up Paradox and America’s Hardware Problem
- Asked whether Chinese teams can simply follow Pi, she separates the question into 2 layers. Pi does publish its papers, unlike Google and Tesla in some cases. But “even if Google published everything, to catch Google you would still have to train a Gemini-sized model first. The same logic applies to Pi: even if we tell you how we did it, you may not be able to match it.” Competition remains, however, and someone will eventually match it.
- America’s hardware problem is close to intractable: “Looking at the US industrial chain today, all you can hope for is innovation.” The lack of industry has weakened education pipelines, reduced interest in relevant majors, and created a talent gap that cannot be repaired in 1-2 years. Immigration may help, “but it takes so many years to build a subway station in the US.”
- The industry joke: “Tesla’s best factory is the one it opened in Shanghai. The output is simply higher.”
50. The Essence of Reinforcement Learning: Exploration, Attribution and the Mario Bug
- She breaks RL into modules. Exploration means that “to do something better, you have to do something you have never done before,” from microscopic muscle adjustments to choosing a research direction. “Large models still do not possess a highly active exploratory ability.”
- Credit assignment means that if a pet performs a sequence of actions and then receives a reward, it needs to know which action was decisive. π0.6* helps the agent sift through large volumes of deployment data: “Everything else is garbage; this is the essence. Do more of this essence in the future.”
- Reward design is a foundational problem that textbooks skip. In Super Mario, RL discovered a bug, repeatedly exploited it, and cleared the game by maximizing the reward function: “I’ve done what you wanted me to do. What more do you have to say?” But the reward function itself was wrong from the start.
- The conclusion is about communicating intent rather than writing a function: “The problem is not writing a reward function; it is communicating to the agent what you want it to do.” NLP is most useful on verifiable tasks, such as code that runs. “Good rewards can be designed for the context.” She never folds clothes, so she can design from a pure robotics-task perspective; Chelsea has “an extremely sharp intuition” for which folding method suits a robot.
51. The Future of VLA: Claude Code Changed Her View of the L
- She once doubted the value of language interaction—“I’m introverted. Wouldn’t it be better if many things didn’t need to be expressed verbally?”—but became a heavy user of Claude Code and changed her mind. “You can use speech to provide more context and let it search or make plans on its own. Logical reasoning is important for robots; the L in VLA still has a major role.”
- Today’s VLA is “still a relatively primitive architecture.” Language operates at the level of “fold this piece of clothing,” without exploring the fine-grained link to how 2 hands should perform each step. The future may add video and auxiliary losses at the output, and more context at the input—“take in much context, like a large language model.”
- She sees 2 possible end states: a dedicated, pretrained, self-contained brain for robots, or “everything folded into a model that has not yet been created—language, image editing, video generation, and robot control all inside it. By then, robots would already be part of the training process.”
52. Organizational Slice: Reading Groups, CO2 Alerts and 70 People with No Time Clock
- The detail density is high. Roughly 70 people—“that’s my feel; I don’t count every day”—were packed into the first office. The meeting-room CO2 monitor turned red with more than 3-4 people, so they had to leave the door open and keep working. There was no clock-in system. As a night owl, she arrived around noon and went home at 1 or 2 a.m.
- The early 10 a.m. daily meeting drew her protest. The CEO was stunned: “It used to be 8, then we moved it to 9, then 10. How can 10 still be too early?” Overtime was unpaid but voluntary: after dinner, a discussion would begin and she would want to push it one step further.
- Reading groups are part of the research machinery. π0.6* was presented twice: once at kickoff to explain what the prior art had achieved, and again before maturity to collect another round of objections. There were also sessions on “whoever can use Claude Code best, get up and present.” After one, she was shocked and rebuilt her entire workflow in 2 weeks; “the effect should be around 3-4x.” But “research is not something you can hit by firing wildly. You still need human input.”
- The governance structure: CEO Carl Houseman oversees infrastructure and keeps research moving efficiently, in a role “similar to Sam Altman.” Sergey and Chelsea run the research.
53. An Interview She Did Not Know Was an Interview: Quitting the Cambridge Faculty Job at the Last Minute
- She had been determined to become a professor. When ChatGPT appeared, she slapped her thigh and thought, “I can just have it write my funding proposal.” She got the Cambridge faculty job she wanted: Cambridge’s humanistic character fit her personality, allowing her to conduct research and write novels. Her advisor even announced at her thesis defense that she was going to become a professor.
- The pivot came through a student of Sergey’s—likely Abhishek—who joined BGI and became her “junior boss.” The interaction was completely different from the critic-style dialogue with her advisor: “I have this idea, and I have one too. Combine them and 3 or 4 new ideas emerge.” Her conclusion: “Having both approaches is best.”
- Within a week of Pi’s founding, she was pulled in for a chat. Nobody told her it was an interview, and she spent the entire conversation turning the tables and asking what they were doing. “They said they felt enormous resistance while interviewing me.” But after the conversation, she thought the group was “quite truth-seeking,” unlike academic work that sometimes exists only “to make the story internally consistent.”
- Her decision to leave academia came down to a student’s 3-5 years: “Is that enough time to let them do genuinely pioneering work?” Compute, data, and hardware also require industry. After 7 years in a PhD and building a robot from zero herself—3 months to get the system running, 6 months to establish teleoperation—she accepted that “some things can only be done through team collaboration.” Her view of the PhD years is warmer: she never worried about graduating; it was a coherent and happy process, with free overnight parking, a 24-hour fried-chicken shop, and 2 project ideas that came from Christmas vacations.
54. Consciousness, Robot Species and the Ship of Theseus
- On the idea that awakened consciousness would destroy humanity, she wants an explanation first: “I really want someone to explain how consciousness emerges—from this pile of physics and chemistry, how does consciousness suddenly appear?” If the definition is human-like consciousness, “I’m somewhat pessimistic.” If it is broad enough that chatting with an AI feels like talking to a person, “you can’t say it has no consciousness.”
- A robot species is her dream research topic: “A robot assembling and building itself is a form of reproduction.” She said exactly that when applying for faculty positions. Being repairable and able to change form would be “a very free and unconstrained way to exist,” especially after spending a long time at home recovering from a climbing injury.
- The Ship of Theseus is already playing out at Pi. A robot retains the same number while every internal component is replaced through repairs and modifications. “Should we give it a new name?” She jokes that a future robot ethics review board will check whether every part from head to toe has been replaced.
55. Love and Death: Robots Will Not Be Humanity’s Final Creation
- Asked whether robots would be the last thing humans ever make, she answers from a humanistic premise: “I don’t think we are alive to increase productivity. We are alive more for love and death”—the timeless subjects of literature. Even with extremely high productivity, people will still want to make different things.
- She accepts the romance of continuation. A human life lasts only 100 years and is infinitesimal against cosmic history. “Could you build something that can see the end of the universe?”
- She gives the black box a philosophical home: ML’s process may be incomprehensible, but it is “bounded by the objective that you define” and explainable mathematically. Her final provocation is whether laziness is a fundamental drive of human civilization: “Improving productivity and being lazy are the same thing.” When robots achieve that, she wants to go home, write novels, make a game, and take a run at the Nobel Prize in Literature.
56. Distinction and Greedy Algorithms: A Gentle Rebellion Against Silicon Valley’s One-Way Street
- Her book of life is Pierre Bourdieu’s Distinction. Reading it helped her ask how much of what she likes is her own and how much was shaped by society. “Silicon Valley is full of worship of productivity—efficiency, simplification, creation. Isn’t that also part of this society’s value system?”
- She describes the rebellion in algorithmic terms. A greedy algorithm chooses the local optimum at every step, while dynamic programming shows that “making the best decision at every step does not necessarily produce the global optimum.” No one has a God’s-eye view, so “it’s okay to do 9 points and leave 1 point undone.”
- The detours where she “never took things all the way to the end” are what brought her to this path. Robotics may be the first time she has felt she must finish what she started, fundamentally different from people in Silicon Valley who choose the same path early and compete along it for years.
57. Quickfire: Tomato-and-Egg Stir-Fry, Edinburgh and the Brain in a Vat
- Her favorite food is tomato-and-egg stir-fry: “I hope that when artificial intelligence is about to arrive, it leaves behind some marks of everyday human life.” Her favorite place is Edinburgh, where “the weight of history is placed in front of you” and creates a sense of stability. Her obscure fact is a piece of anatomy: what is commonly called the hymen is medically the vaginal corona, and first intercourse does not necessarily cause bleeding. “We have so many gaps in our knowledge of our own bodies.”
- Too many papers have influenced robotics to list. ChatGPT “also had a huge impact on the trajectory of robotics,” along with Diffusion Policy, Transformer/ACT, and earlier foundational work in imitation and reinforcement learning. Her key fact based on current understanding: “Robots entering homes may not come in product form; they may come more as exploration. It’s closer than I initially imagined.”
- Hearing “language is the world,” her first reaction was the brain in a vat. All the senses shape our perception of the world, so “you can completely imagine that we are living in a giant simulator.” Earlier language-only models could not ground a “water cup” in the physical world; that reflection indirectly drove multimodal research. As for whether we live in a simulation: “If the simulator is the world, then whether it simulates makes no difference. That’s idealism.”