Which Embodied Model Is Best? 范浩强 and 高阳 on RoboChallenge
Summary
RoboChallenge’s primary value is turning embodied capability from a handpicked demo into a repeatable generational curve. Traditional papers often show only 3 or 4 tasks, with inconsistent definitions; the new platform uses a standardized Table 30, real-robot environments at a centralized site, and hundreds of executions to reduce variance. 范浩强 is not trying to debate 99% versus 99.5%; the first task is to reliably distinguish “it’s zero, I’m not zero.”
The jump from π0 to π0.5 shows that, although embodied models are still at a very early stage, they have already made leaps visible to the naked eye. π0 averaged just over 20% success on 30 tasks humans can perform with near-zero error; π0.5 reached roughly 42%, while fine-tuning pushed several simple tasks to 100%. Standing beside the machine, 范浩强 felt that “it’s a little sharper than the previous model.” The host called it a short-term leap; he cautioned that the two generations were actually nearly a year apart, but the doubling remains a clear scaling signal.
On the latest leaderboard, 千寻智能’s Spirit V1.5 overtook π0.5 to rank first, but even the top-tier models remain below 60 points overall. That does not mean the leaderboard has failed: the team deliberately placed tasks in the models’ “sensitive zone”—shredding paper tests precise insertion under occlusion, flower arranging involves object-to-object contact, and scanning a QR code exposes the lack of memory in single-frame VLAs. For investors, the key is the generational gap and rate of improvement, not a seemingly precise decimal that cannot be reproduced.
“Scaling data may be the most important theme of 2026,” while the race over approaches remains unresolved. 高阳 divides the data into simulation, human video, wearable devices, and teleoperation; 千寻 is currently betting on every real-world source except simulation. 范浩强 takes the view that “all roads lead to Rome”: the ceiling will ultimately be determined by how much physical-world knowledge a team collects. Generalist AI claims to have 270,000 hours, adding 10,000 hours a week; another wearable-data team has proposed collecting 10 million hours in a year. Both figures should be understood within the conditions stated by the guests.
The embodied field has yet to find a high-value, scalable wedge comparable to AI coding. Coding is concentrated and high value per unit, while robot services are fragmented across home services, delivery, maintenance, and countless other verticals; traditional industrial arms have also raised the commercial bar. 范浩强 treats “automatically folding my bedspread every day” as a personal validation task, but the bedspread simultaneously tests the model, payload, precision, and whole-machine stability—evidence that the breakthrough will not be determined by the model alone.
Both guests compare current capability to the GPT-2 or CIFAR stage, not the near-consumer-product stage the public imagines from demos. 高阳 expects a GPT-4-style breakout that ordinary people can truly feel may still be 3–4 years away; industry insiders will see the signal earlier in imperfect models. The host added that a miracle in 2027 would of course be welcome. The breakout ultimately has to happen on the customer side: only when customers believe will demand and deployment begin reinforcing each other.
The two defining questions for 2026 are whether China can produce a DeepSeek moment in embodied AI and whether foundation models can approach GPT-3 or even GPT-3.5. 范浩强 wants to test whether domestic teams can move from looking up to overseas models to catching them; 高阳 says, “I’m extremely confident today,” while allowing that his confidence could rise or fall after a few more months. Within roughly 1–2 months of launch, RoboChallenge has accumulated more than 10,000 runs, and its roughly 9 machines once had a queue of 1–2 days—evidence of real R&D demand, but also of an evaluation infrastructure still constrained by supply.
Deep dive
1. Two technical careers jointly underpin the optimism about embodied progress
范浩强 joined 旷视 in 2012 after being admitted directly from his senior year of high school, and witnessed computer vision move from SIFT and SVM into deep-learning commercialization. He remembers facial recognition advancing from the low 80% range to “twelve nines” in total, building a strong conviction: “As long as the algorithmic principle is right,” sustained investment of intelligence, effort, and labor will ultimately pay off in technology.
高阳 began shifting from computer vision into robotics research around 2016 while pursuing his PhD at UC Berkeley, working with Sergey Levine and others. He returned to Tsinghua to teach in 2020, and the ChatGPT wave convinced him that humanity had found a path toward AGI. When he founded 千寻智能 around 2023, the goal was to build the strongest model in embodied AI and put it into robots capable of entering everyday life.
Their complementarity also explains the division of labor on the show: 范浩强 focuses more on turning hard-to-reproduce capabilities into operable, iterative engineering systems; 高阳 explains why large-scale evaluation has finally become possible through the history of robot learning, model paradigms, and data structures.
2. Carefully selected demos have long concealed robots’ real boundaries
范浩强 compares past robot research to graphics papers: every paper needed several cool demo videos, but usually it was “filmed for a long time, and in the end there was one good take.” That can prove only that something can happen at least once—not its success rate, generalization, or stability.
Real-robot papers often test only 3 or 4 tasks, and even papers that both call a task table bussing may use entirely different environments, objects, and action definitions. Simulation is easy to test and keeps noise under control, but cannot directly answer how a system will perform in deployment. The lack of a common leaderboard also slows R&D by leaving teams unable to tell whether a model change was an improvement or a regression.
3. Large-scale real-robot execution reveals a clear generational gap
RoboChallenge began with a simple question: can models stop being tested 3 or 5 times and instead execute at least several hundred trials, using average success rates to suppress randomness? Once the team committed to large-scale evaluation, task formats, textual descriptions, and testing protocols had to be standardized and made compatible with different robot configurations, including single-arm and dual-arm systems.
The team initially selected 30 tabletop tasks that humans could “absolutely perform perfectly,” yet π0 averaged only slightly above 20% success—roughly one success in 4 attempts. 范浩强 admits he was nervous at the time: if the results were made public, would the outside world lose confidence in the entire sector?
When π0.5’s weights were released, the team tested them immediately. Average success rose to roughly 42%. The shock was not just the doubling in score; standing beside the machine, it was “a little sharper than the previous model,” and fine-tuning had already made several of the simplest tasks reliably achievable at 100%.
The host described this as a rapid improvement over a short period. 范浩强 added an important qualification: “It actually wasn’t that short”—π0 and π0.5 were separated by nearly a year. Even so, the difference between the two models on the same real-robot tasks convinced him that the industry is making very concrete progress.
4. Physical-world noise means evaluation should focus on trends before decimals
高阳 points out that embodied models must interact with the physical environment in a closed loop. The handles on 2 cups, object placement, lighting, and even mechanical error can all vary. When a model fails, it is difficult to tell immediately whether the model is weak, the handle is different, or some environmental detail has moved outside the training distribution.
The host’s counterpoint is worth preserving: doesn’t this simply show that the model generalizes poorly? 高阳 acknowledges that early models did require nearly identical environments, while today’s models are much better. But research innovations often produce only 0.x% improvements, and even a major improvement may be around 5%; environmental noise is enough to drown the signal completely.
RoboChallenge therefore does not try to prove that a number is accurate “down to one percentage point.” It aims to identify obvious generational gaps. The current top score remains below 60, placing the benchmark in a “sensitive zone” that amplifies model differences; the industry is still comparing “it’s zero, I’m not zero,” far from the stage of 99% versus 99.5%.
5. Academia had the idea of remote real-robot testing, but lacked the resources to operate it continuously
高阳 recalls that Abhinav Gupta at CMU had experimented, even before 高阳 began his PhD, with opening a fixed laboratory setup and robot to external users through remote connections. Similar projects later appeared in China and abroad. The concept was close to RoboChallenge: lock down the machine and environment, then let others run their algorithms remotely.
The problem is that university labs have limited space and often can maintain only 1 or 2 machines and 2 or 3 tasks, while also requiring substantial labor to reset and supervise them. Once the students responsible graduate, the entire setup can easily be left without maintenance. Academia had “the idea,” but struggled to turn it into continuously operating infrastructure that broad users would trust and accept.
The technical baseline has also changed. Older robot algorithms could sew a handkerchief, cut a specific object, or perform a handful of fixed tasks, but could not run at all when the task changed. VLAs at least make it theoretically possible to perform any task; success rates remain low, but for the first time there is a real need for a general benchmark.
6. Simulation can reproduce code, but still does not stand in for the real world
Simulation’s advantage is straightforward: given the same simulator and code, anyone who downloads them should obtain the same result. But many simulated trajectories come from prewritten programs—move above the object, descend, grasp—so their state distribution is far narrower than human teleoperation and does not fully reflect the deviations a real robot will encounter.
高阳 says reviewers at robotics conferences generally will not readily accept a paper based only on simulation results, because improvement in a simulator does not imply the same improvement on a real robot. 李飞飞, MIT’s Russ Tedrick, and others are still advancing the area; the conclusion is not that simulation is useless, but that “it is a very difficult thing.”
Discussing World Labs’ Marble, 高阳 says it currently looks more like touring a static 3D scene. The underlying 3D or 4D Gaussian technology is still not well suited to dynamic interaction, so it may first serve games or 3D vision rather than directly function as an embodied-training world.
7. RoboChallenge and RoboArena test problems at different levels of maturity
Physical Intelligence’s RoboArena uses a distributed model: laboratories around the world each contribute some compute and real-robot time, with 2 models run on the same machine in the same lab for each comparison. Enough paired observations can ultimately be aggregated into a relative ranking; the labs do not need perfectly identical environments.
RoboArena tests zero-shot capability: a task is specified on the spot, and the model must understand the text and execute it without data for that task. 高阳 uses language models as an analogy: during the GPT-3 stage, a small amount of expert Q&A was often needed for fine-tuning—that was few-shot. Only when a model is strong enough to answer reasonably without any examples does it qualify as zero-shot.
RoboChallenge currently uses a few-shot or fine-tuned setting. Each Table 30 task provides roughly 1,000 human demonstrations; participants fine-tune using their own base model and training method, then control the on-site robot through a standardized API. The data itself defines the test distribution, while the centralized environment supplies comparable absolute success rates.
The host asks whether RoboArena is simply too difficult for today’s models. Both guests acknowledge that success rates on most tasks are very low; sometimes the only conclusion is that neither model has grasped the task, though one action looks “more promising.” If every score is close to zero, discriminative power also falls, making open-ended pairwise testing better suited to future zero-shot models.
8. Table 30’s discriminative power comes from real desires, not tidy task taxonomy
Table 30 was not derived from a theoretical classification system. The team asked the entire company to list “what robots must be able to do next,” producing a wish list of thousands of items. The data-collection team then assessed feasibility one by one, and one researcher initially selected 30 tasks.
The initial selection was fairly ad hoc. Only after testing did the team realize it covered flexible objects, millimeter-level positioning, contextual memory, actuator-object contact, and object-object contact. It did not count picking up a cup and picking up a ball as 2 nearly identical tasks; each item was designed to have a different failure mode.
范浩强’s retrospective summary is pointed: “The real needs of ordinary people are, in fact, the most representative and highest-quality benchmark.” Tasks derived from everyday desires may not look elegant, but they are more likely than an artificially ordered academic taxonomy to hit the model’s actual blind spots.
高阳 emphasizes that a foundation-model team does not design a separate trick for each task. 千寻 uses the same base model and standardized fine-tuning process across the full task set. The point of a foundation model is precisely that “whatever kind of thing you want to do,” it can learn through the same process rather than requiring a bespoke solution for every task.
9. The hardest tasks expose contact, occlusion, and flexible-object problems at once
Shredding paper requires the robot to insert paper into the shredder’s narrow slot. During insertion, the paper also blocks the camera mounted on the hand, so the model must maintain precise motion as visual information deteriorates. This is not ordinary pick-and-place; it combines occlusion, flexible material, and narrow-slot positioning.
Flower arranging expands interaction from actuator-object contact to object-object contact. The model must locate its own hand and understand when the flower stem it is holding touches another object. Tasks such as placing a cup into a hole have tolerances of only a few millimeters and similarly produce low success rates even for the best current models.
Wiping a table with a napkin looks simple, but the real difficulty is that after wiping, the napkin has become a shape that training data is unlikely to cover well, and the robot must still put it back in the bin. 范浩强 believes flexible objects are difficult to model, and demonstration data cannot enumerate every possible deformation case.
10. QR-code scanning shows that many VLAs still live in a single-frame world with no past
The scanning task requires one hand to hold an object and the other to hold a scanner, then put the object back after scanning. The motion itself is not complex; the difficulty is that the images before and after scanning look almost identical. From the current image alone, the model cannot tell whether it has just scanned, so some stop as soon as they pick up the scanner while others scan repeatedly.
高阳 reduces the problem to memory. Many open-source VLAs still take single-frame inputs: after viewing several camera frames, they output an action, with no mechanism for putting whether the scan has already happened into context. 范浩强’s metaphor is: “It forgets once every frame; every 0.7 seconds, it forgets.”
The industry already has patches, such as feeding in more historical frames or using an agent to record events. But both guests draw a clear distinction: this kind of memory is often an engineering layer bolted on externally, not something genuinely trained into the foundation model. Table 30 therefore ended up measuring a cognitive gap the team had not originally set out to design for.
11. Spirit V1.5 reaches the top as platform demand outruns machine supply
千寻 tested a unified base model, and Spirit V1.5’s latest result came in above π0.5, making it the number-one public leaderboard entry at the time. RoboChallenge requires only that runs be tied to a real individual; organizations can remain anonymous. Usually a result appears first and the company claims it later, preventing the leaderboard from becoming a commercial PR contest from day one.
From its launch announcement in October to the time of the interview, roughly 1–2 months later, the platform had accumulated more than 10,000 runs. The farthest submission came from London: code crossed the globe to connect to the lab and still made the robot complete the action, which the team found “pretty rewarding.”
Testing capacity was roughly 9 robots across 4 types, including 1 dual-arm and 3 single-arm configurations. Submitters at one point had to wait 1–2 days. 范浩强 says the enthusiasm was “far beyond expectations”; had they known demand would be this strong, they would certainly have prepared more machines at the outset.
When the leaderboard launched, there were no external results, so the team first found volunteers to reproduce roughly 6 open-source models as baselines. Later, more companies decided they were not comfortable having the platform tune the models for them and began fine-tuning and submitting independently. That shift itself shows the community moving from looking up to π0 toward mastering comparable model technology and challenging π0.5.
12. An open control API expands room for innovation while leaving credibility to the community
The platform does not require participants to upload a fixed model file or Docker image; it provides only an API for controlling the robot. Participants can use any compute architecture, parameters, and control logic. The reason is that fine-tuning involves many “craft” decisions, and even changing the GPU can affect the result; standardizing training at the platform level would remove a key part of the participants’ innovation.
The trade-off is that the platform cannot directly verify how many episodes a participant used, or even determine from the interface whether the system behind the result is a model or a human. For now, the approach relies primarily on academic self-discipline and encourages leaderboard entrants to publish their models and code so third parties can resubmit them. Results need not match perfectly, but the broad trend must be reproducible.
The simplest hack would be to submit the same task 100 times and keep only the best score. The rules therefore use the final run, while limited capacity also prevents unlimited retries. Since participants cannot access the physical environment, common forms of cheating such as repositioning objects are difficult; the biggest potential attack is instead using human teleoperation to impersonate a model.
13. RoboChallenge is evolving from a single leaderboard into shared real-robot infrastructure
范浩强 emphasizes that the project is no longer operated independently by his company; it is being advanced through a partnership and community, with more than 10 companies already involved. Hugging Face provides international collaboration, robot manufacturers can donate machines, and teams with data-collection sites or laboratories can contribute their own benchmarks.
Table 30 is only the first content block. Future sets could include Kitchen or Restaurant environments, or new tasks built around open-source dexterous hands. 高阳 also recalls that a company called Tesollo wanted to participate in designing related tasks. Only when different organizations contribute machines, environments, and data can models receive a more complete “physical” than a single tabletop setting can provide.
For 范浩强’s company, serving external users is only a small fraction of its internal model-evaluation cost, so the platform remains affordable for now. But the growing number of tasks, hardware units, and submissions will require more institutions to share the burden. The platform’s biggest short-term constraint is not demand; features, machines, and operating capacity can only come online in a queue.
14. Both guests see data scale as the clearest theme for 2026
高阳 believes data remains the main bottleneck for embodied models. If the field had nearly unlimited data comparable to language models, the problem could “most likely be solved.” He therefore expects very large model improvements in 2026 and the following year, rather than a linear sequence of minor fixes.
His proposed path follows the language-model “three-part progression”: pretrain on large-scale general data, perform supervised fine-tuning with high-quality expert data, then use RL or RLHF-like post-training to improve task success. The specific embodied architecture may differ, but the learning paradigm—pretraining, fine-tuning, reinforcement—is already fairly clear.
范浩强 jokes that Table 30 would need to reach at least 90% before he could be reasonably confident the algorithm was not missing “major parts.” 高阳 believes the field may approach that level quickly by following the current scaling-data path. The real challenge is getting algorithms to absorb large, heterogeneous datasets and organizing the social division of labor needed for real-world collection—not an insurmountable theoretical obstacle.
15. VLA describes the input-output interface, not a single architectural answer
高阳 cautions that VLA only describes a model receiving vision and language and outputting action; it does not require an LLM first, a VLM next, and an action head at the end. The model could also become a VLTA, with T standing for tactile. Whether a video-generation model is a better base is likewise still an open question.
千寻 currently starts from a VLM base and adds video, human teleoperation, and wearable-device data, but does not organize them as an explicit world model. Reinforcement learning can be added at the end. The goal is to use different data at the stages where each is most suitable, rather than insist on a single source.
范浩强 believes the broad learning framework has largely converged: first let the VLM absorb internet knowledge and invariances, then use large amounts of robot data to build the mapping between action and perception, and add the post-training process before release. What truly differentiates teams is “which building blocks they use, what data they add, and in what proportions.”
16. Four data streams trade off scale, quality, and diversity
高阳 divides the data into simulation, human video, wearable devices, and teleoperation. 千寻 is currently betting on the latter 3 real-world sources: large volumes of lower-quality data suit pretraining, while smaller volumes of high-quality data aligned with robot actions are better for post-training.
He is not currently betting on simulation, mainly because of the cost of diversity. Each scene still requires a professional artist to build it in the simulator; the tools are not easy enough to use, production is slow, and the work cannot be delegated to large numbers of ordinary operators as easily as teleoperation or wearable collection. This is 千寻’s empirical judgment, not a claim that other companies’ simulation efforts are necessarily wrong.
高阳 leaves room for a clear change of view: if simulation tools improve to the point where “you can casually make something like Black Myth: Wukong,” simulation data will of course become “really attractive.” The issue is not whether the data is nominally synthetic, but whether it can cheaply and quickly generate enough rich, physically valid interaction.
Static 3D demos such as Marble have not yet solved dynamic interaction, showing that visual realism is not the same as embodied usefulness. Robot training also needs continuous states after objects are grasped, pushed, occluded, or deformed—not just a 3D backdrop that can be toured.
17. Every collection route ultimately pays the cost of acquiring world knowledge
范浩强 is relaxed about the debate over approaches: “All roads lead to Rome.” Simulation teams must produce vast quantities of 3D assets and close the sim-to-real gap; real-robot teams must industrialize and reduce the cost of operators, equipment, resets, and quality control. Further down the road, both approaches may run into the same problem: how to enumerate everyday products and environments.
His team chose a primarily real-robot route because it had previously built a large-scale offline facial-data collection system. That does not mean real robots are the only correct path; it reflects the production method the organization is best at. The decisive variable is not the name of the data representation, but “how much effort you put into collecting knowledge about this world.”
Generalist AI’s off-body collection has people perform actions while holding a gripper, without requiring a robot to hold it, making equipment easier to ship and collection easier to scale. 范浩强 relays its public figures: 270,000 hours already collected, with 10,000 hours added each week; “if what they say is correct,” the total could approach but remain below 1 million hours by the end of 2026.
The scale claimed for wearable collection is even more aggressive: he has heard of a company proposing to collect 10 million hours in a year. Neither guest endorses these figures, but the plans show the industry shifting the focus of competition from model names to knowledge volume, collection efficiency, and the organization of labor.
18. Data sharing does not automatically eliminate cleaning and engineering labor
Open X Embodiment is an important academic data-sharing effort that organizes datasets that were previously public but scattered. At the company level, however, nothing comparable in scale has yet emerged. Embodied trajectories collected with corporate money are still more likely to remain private assets in the short term.
范浩强 pushes back on the simplified story that internet data was “free” for language-model companies, calling it a “fairy tale.” Web pages may be accessible, but large numbers of excellent engineers still have to clean, denoise, and organize them; that human labor is what “made this model great.”
Embodied AI may initially be able to use massive volumes of public video, but only if algorithms can extract actions and physical knowledge from it. 范浩强 believes the industry may still need to solve problems that come before scaling data. Table 30’s low success rate is evidence that models have not yet converted existing data into robust basic manipulation capabilities.
19. The embodied industry is still looking for its own AI coding
高阳 separates capability from commercialization: foundation models determine what robots can do, while business logic determines what is worth doing. Deployment can begin, but current foundation models are still not strong enough, so improving capability is more urgent than rushing to find a single universal application.
The host notes that roughly half of language-model traffic may reportedly come from AI programming, with search replacement and Q&A among the other major uses. 高阳’s response is that the transformation is still developing: turning model capability into phone assistants, automatic food ordering, and similar products takes time, and the ultimate boundary cannot be inferred from current traffic.
The central question both leave open is: “What is the coding of the robot era?” Coding consumes high-value intellectual labor, has high value per unit, and is concentrated in form. Robots’ potential work—parcel pickup, room cleaning, domestic services, and maintenance—is highly fragmented, with no equally clear large market to pull models, data, and products forward at the same time.
高阳 therefore expects the robot industry to form many vertical companies: some focused on domestic services, others on roof repair. China’s relatively abundant labor supply makes the urgency of some services harder to feel; expensive labor in Europe and the US may lead to economically necessary robot applications emerging earlier in certain niches.
20. “Folding the bedspread” tests the model, hardware, and product loop simultaneously
When he founded the company, 范浩强 used a personal task to convince himself: every morning, he should be able to get up without thinking about it, and have the bedspread folded automatically so that he comes home to a tidy bed at night. Robot vacuums and traditional appliance companies may not view this as their job, while classical detection and segmentation algorithms clearly cannot do it, making the task representative of the moment embodied AI moves from impossible to possible.
The signal is already emerging, but the field remains far from that goal. 范浩强 mentions that π0.6 can fold cardboard boxes in a factory; the host then cites demos from several US companies involving crumpling socks and folding towels. But “crumpling” a sock does not mean being able to pick it open and turn it inside out, which requires a higher level of manipulation.
A bedspread is heavy: the robot needs enough payload to lift it without toppling as its center of gravity shifts. Socks and clothing likewise expose the precision and payload limits of the hardware itself. Many failures that look like algorithmic failures may first require the body, dexterous hand, and control system to clear their own thresholds.
Household tasks are popular in academia because they are easy to stage, require relatively low precision, and visibly demonstrate intelligence. Commercialization may still be highly vertical. What 范浩强 hopes for is a few “mind-blowing demos” that persuade customers more tasks are feasible—“the breakout cannot be just a breakout among the public and investors; more importantly, one day the customer has to believe.”
21. The public either overestimates or dismisses robots, while demo engineering amplifies both extremes
范浩强 sees 2 opposite misconceptions. One side assumes robot companies are all editing videos; the other thinks robots have already arrived because bionic robots could run and jump in the past. When the public sees clothes being folded, it naturally infers that crumpling socks and putting them in a closet must also be solved, without realizing that the technical boundaries between actions do not follow human intuition.
高阳 lists 4 risks to guard against when watching demos: cherry-picking, video editing, teleoperation, and AIGC. Running a task 100 times and showing only the one success, or fixing every object’s position before recording, can create an impression of generalization far beyond reality. The only strong evidence remains having an observer on-site interact with the system.
There is still no complete answer for online certification. The industry may place an iPad clock in the frame to prevent speed-up editing; others require participants to obtain a unique random hash for each run, place it on-site, film it, and submit immediately, or even imagine using blockchain. These methods raise the cost of fraud but still cannot conclusively prove that the model completed the task autonomously.
22. Industry insiders may see the inflection point 3–4 years before the public
高阳 compares current embodied foundation models to roughly GPT-2. 范浩强 prefers an analogy from computer vision, arguing that the industry is still in the CIFAR period and ImageNet has yet to appear. CIFAR has only 10 classes and 32×32 images, making it obviously a toy example; once ImageNet expanded to 1,000 classes and full-resolution images, people began to believe models could “really do work.”
高阳 expects a GPT-4-style breakout for ordinary people in roughly 3–4 years. He does not consider GPT-3.5 a true breakout yet. Industry insiders will identify the signal earlier: even if a model does not follow instructions well and its post-training alignment is poor, experts will know that if it is “very good at everything,” adding post-training may bring it close to usable.
The host added that a miracle in 2027 would of course be welcome; 高阳 did not offer a shorter, definite timeline. Both emphasized that intelligence gains may be nonlinear, and today’s near-zero success rate cannot be directly extrapolated into a 3-year forecast.
范浩强 summarizes the information gap with a joke: “The best way to catch the LM express is to go heavily into Nvidia after the GPT-2 paper comes out.” This is not investment advice from the show; it underscores the inevitable time lag between insiders seeing early CIFAR-like signals and the public seeing a usable product.
23. 2026 will test both catch-up speed and the ceiling of foundation models
What 范浩强 most wants to test is whether embodied AI can produce a Chinese “DeepSeek moment”: after years of watching overseas work with envy, can 2026 bring a domestic model that genuinely catches up? Using computer vision as a reference, he says the felt time from Google’s work seeming like “something from outer space” to China and the rest of the world jointly reaching human performance was about 3 years; tools such as Codex and Cursor have further improved experimental productivity today.
高阳 spends less time on the China-US narrative and is more interested in whether foundation models can move from GPT-2 to GPT-3 and even GPT-3.5. His wording is optimistic but falsifiable: “I’m extremely confident today,” while acknowledging that after working for a few more months he may become more confident—or lose confidence. For now, the trend is fewer questions and greater certainty.
The closing discussion frames the industry around 3 axes: data, models, and hardware. Data and models directly determine intelligence, while hardware must also solve for lifespan, durability, reliability, precision, payload, and battery life. The brain, cerebellum, and body can theoretically be decoupled, but vertical software-hardware integration may accelerate R&D and deployment, and Chinese teams can draw on local supply chains.
The talent base is equally diverse: autonomous driving, long-running academic robotics research, computer vision, and traditional robotics each contribute different capabilities. The host’s final view is that embodied AI is chaotic precisely because it is interdisciplinary—and therefore unlike a field that one giant can quickly consolidate simply by spending money. Evaluation, data production, model architecture, hardware, and applications all still contain substantial uncertainty.