155: Jia Peng of Simplicity Dynamics: From Nvidia to Li Auto, the Hexagonal Warrior of Embodied Intelligence
Summary
Jia Peng argues that the first gate in embodied intelligence’s shift from To A to To B is not the model, but hardware that has yet to reach mature mass production. The industry is still largely operating at the level of a few thousand units, hand assembly and inconsistent behavior across identical units; in his observation, the long-term repair rate of existing robot bodies is “basically 100%,” with hardware-consistency noise potentially outweighing the benefit of additional data. He expects hardware to start converging around the end of 2026, but repeatedly emphasized: “Hardware is nowhere near that point.”
Simplicity Dynamics chose to develop its entire stack in-house—from the robot body, data-collection hardware, AI Infra, models and applications—because commercial scale, real-world data and model iteration must form a single flywheel. Jia defines the company as a “hexagonal warrior” with no weak side across technology, strategy, product, brand, organization and commercial execution; the more products it sells, the more real long-tail data it generates, the faster the models improve, and the better the products and repurchase rates become. “The best autonomous-driving technology company must also be the best commercially and have the best product.”
Its core technical bet is the Unified Model: bringing the world model, VLA, fast and slow systems, understanding, generation, reasoning and Action into a single model as far as possible. Jia believes dual systems inherently face different frame rates, arbitration during conflicts and difficult joint training. The ideal model delivered in 2025 already attempted to decide for itself when to use CoT while predicting both future frames and actions. Simplicity Dynamics will continue to back “Unification” while seeking a Self-evolution loop in which Action and generative models provide feedback to each other.
Commercialization will not begin by replacing one station on a mature production line; it will prioritize end-to-end, standardized and replicable flexible tasks. Second-level takt times, near-zero tolerance for errors and the durability requirements of high-frequency actions mean robots are currently ill-suited to replace line workers. Simplicity is instead focused on the complete workflow of retrieving materials from a warehouse, loading, unloading, inspecting, deburring and returning goods to the warehouse. Its current deliverable capability is defined as “general-purpose mobility + simple manipulation + object generalization,” with dexterous manipulation, safe interaction and household scenarios pushed back in sequence.
Jia expects the industry to complete the shift from To A to To B between the end of 2026 and early 2027, reach PMF around 2028, and put robots into homes at scale around 2029-2030. The embodied-intelligence window will last longer than the foundation-model window because it has neither an OpenAI-style validated paradigm nor naturally available internet data, and users cannot simply take over whenever the system fails. But the window will not stay open forever: once PMF is proven, large companies with software and hardware capabilities, including Xiaomi, Huawei and Li Auto, could accelerate their entry. “The first to do it in this industry will definitely die, but the ones that ultimately succeed will definitely be from the first batch.”
Li Auto’s autonomous-driving team, which went from a 20-30-person group considered to have almost nothing to a first-tier player, is the clearest case study for this entrepreneurial playbook. In 2021, roughly 80 people delivered highway NOA in 100 days, starting with code that was essentially built from scratch; in 2023, the team switched to mapless driving amid layoffs and product deterioration, repaired the issues one by one within weeks and rolled it out to all users in June. By the end of 2024, the team was validating the product through daily active users, the share of autonomous-driving mileage and NPS. The Max mix rose from the low teens to above 70%, while users paid about RMB30,000 more for autonomous driving, proving that the team had shifted from a cost center to a revenue contributor.
From Nvidia to Li Auto, Jia’s biggest change in approach was moving from “ecosystem tools” to directly facing users, standard hardware and data-driven development. He credits Jensen Huang with making an early, decade-long bet on CUDA, Jetson and in-vehicle edge computing, but believes Nvidia’s L4-, map- and Toolkit-led approach made it difficult to close the loop with real users. Tesla, by contrast, committed early to abundant compute, dual chips and pre-installed hardware even for users who did not subscribe. “You have to let users curse at you,” because only real-world deployment exposes long-tail problems and reveals the direction for iteration.
The capital boom has left openings for new entrants while turning survival into a systems battle. Simplicity Dynamics currently has more than 80 people and about 10,000 hours of data, while GPUs account for more than one-third of R&D expenses. Its financing is being staged across top-tier financial investors, strategic investors such as Alibaba and Tencent, and supply-chain and application-side industrial partners. Jia confirmed that the company has reached its fifth round, but did not confirm a post-money valuation above RMB10B. The listing of the first embodied-intelligence companies will lift industry enthusiasm while giving leaders stronger capital and financing advantages; companies that survive the subsequent bubble must have capital, full-stack products, organizational firepower and the ability to generate cash internally.
Deep dive
1. High-performance computing pulled Jia Peng into Nvidia’s automotive bet
Jia traces his technical beginnings to 2008-2012, when he bought dozens of nodes and assembled his own high-performance computing cluster. He worked with C2050 and CUDA 1.0 and 2.0, accelerating GPU algorithms for nuclear explosions, protein folding, weather and ocean forecasting.
When Nvidia recruited him at the end of 2015, it was his HPC and parallel-acceleration background—not automotive experience—that mattered. He acknowledged that he was “completely at zero” on edge computing at the time, but edge computing was fundamentally still an acceleration problem.
Huang had previously tried putting chips into phones, only to run into heat and power constraints. The next category of smart terminal shifted toward electric vehicles, whose larger batteries and better thermal conditions made them more suitable for absorbing ever-growing compute demand.
2. China’s autonomous-driving team started with 1 person and 2 boards
Jia became Nvidia China’s first dedicated autonomous-driving employee. The team initially still sat within the data-center business and reported to a U.S. manager; his mandate was to push Nvidia chips and low-level SDKs into China’s automotive and autonomous-driving markets.
The early work was highly hands-on: he would “carry the boards himself,” taking PX 1 and PX 2 to automakers and internet companies for demonstrations. Baidu, Pony.ai, WeRide, Horizon Robotics and other teams that emerged during the same period initially wanted to use Nvidia chips and its toolchain.
As customers began accepting GPUs, his role shifted from selling hardware to algorithms, software acceleration and low-level SDKs. The local team also had to cover the full chain, including sensor calibration, perception, mapping and self-localization.
3. Nvidia was still a “small company” in 2015
Jia recalls that Nvidia’s market capitalization was about $10B at the end of 2015, with only 5,000-6,000 to 7,000-8,000 employees globally. Beijing had roughly 100 people and Shanghai about 800—nowhere near the company’s current scale.
He lived through the company’s climb from roughly $10B to $100B and described the subsequent change as “the same as Bitcoin’s gains over the past decade—5,000x.” This was his oral summary, not a figure independently verified by the program.
Performance grew so quickly in those years that it created a trading habit: by Jia’s recollection, the stock often jumped about 30% after earnings, so people bought before the report and sold the next day. “This company was growing in leaps every day.”
4. Autonomous driving moved from a research market to mass production
Jia uses To A, To B and To C to describe technology cycles. From 2015 to 2019, the market was mainly To A: customers were doing research and demos, and getting a system to run in a small area was already enough. Sales and commercial revenue should not have been the primary yardsticks.
Huang kept asking, “Why can’t we sell it?” The China team’s answer was that the industry had not yet entered mass production. Real growth had to wait until after 2019, when Nio, XPeng and Li Auto began mass production and platforms such as Orin were formally adopted.
When Jia left Nvidia in 2020, automotive still accounted for less than 2% of Nvidia’s revenue. That explains why the company could invest years ahead of demand but still fail to generate data-center-like returns from autonomous driving for a long time.
5. Nvidia tried to recreate CUDA’s ecosystem moat in cars
Nvidia did not want to deliver the final product for customers. It wanted to build another CUDA-like infrastructure layer for vehicles: tools and SDKs on top of the hardware, with customers developing algorithms and applications directly on the platform.
The China-U.S. gap was substantial, so the local team filled in capabilities spanning calibration, perception, localization and mapping. That gave Jia his first view from the chip layer all the way up to applications.
The model’s strength was horizontal coverage across customers; its weakness was distance from end users. Jia later put the problem bluntly: “You have to let users curse at you.” Otherwise, the company would get neither real data nor a clear understanding of why the product was difficult to use.
6. An L4-leaning approach left Nvidia without a product loop
Jia believes Nvidia’s early approach was heavily influenced by traditional L4 thinking: acquiring a high-definition mapping team, investing in maps and planning and control, and writing rules for different cases. It looked more like the Waymo path than the data-driven path for mass-market vehicles.
The choice also fit Nvidia’s corporate character—it provided ecosystems and toolkits rather than directly facing users. Even when it later took on more delivery responsibility in projects such as Mercedes-Benz, Jia speculated that one purpose was to validate the approach and work backward to define hardware requirements.
程曼祺 asked why Huang had not moved onto the Tesla path earlier. Jia’s answer was not that the technology judgment had been wrong, but that the commercial positioning dictated it: “He didn’t want to face users directly.”
7. Tesla linked abundant compute, standard hardware and data early
Jia recalls that Tesla’s first mass-production solution used Mobileye, but problems around openness and cooperation pushed Tesla to develop in-house. Before its own chip was ready, Huang provided PX 2 as transitional compute and opened up the underlying stack.
Musk made clear that the next generation would not use Nvidia again: autonomous driving was ultimately an AI problem, AI was ultimately a compute problem, and Tesla needed its own chip with sufficient compute. Hardware also had to be pre-installed even if the user did not subscribe to FSD.
The dual-chip setup and the vehicle-wide electrical and electronic architecture reduced service costs and enabled continuous data collection as part of the same design. “I’m willing to pre-install the hardware for this,” which Jia sees as an early implementation of Scaling Law thinking in 2016-2017.
8. A flat organization turned Jia from algorithm engineer into architect
Shortly after joining, Jia found himself only 5 levels away from Huang. The team was small, but Nvidia’s code and resources were highly open, allowing him to see from drivers, chip design, SDKs and acceleration all the way to upper-layer applications.
Over 5 years, he evolved from a pure algorithm-acceleration specialist into an architect comfortable across software and hardware: “I can talk to people about hardware, chips, software and algorithms.” That became the foundation for his later full-stack robotics work.
Nvidia’s general-purpose computing platform also kept engineers in a constant process of “opening up new territory.” The underlying tools stayed in place while applications expanded from gaming, Bitcoin and data centers to autonomous driving and Robotics, without requiring the infrastructure to be rebuilt each time.
9. Huang sustained long-term bets through energy, detail and communication
At GTC, Jia would drink with Huang and visit his home to talk. His strongest impression was that Huang had “far too much energy”: Huang woke at 4 or 5 a.m. every day to read emails or Papers, and once said his forehead was 1 degree warmer than his wife’s because his brain was running at high speed.
Nvidia employees wrote a weekly Top 5, which could cover both business progress and personal events. When the company was still small, Huang would respond personally, asking why Chinese customers had not yet reached mass production and whether the problem was hardware or software.
He might not immediately connect every face to a name, but hearing “Jia Peng, the China autonomous-driving team” was enough for him to continue discussing specific customer feedback. That attention to detail made long-term Vision more than a slogan.
10. A decade-long bet requires founder credibility and a cash cow
CUDA was also questioned early on. Jia remembers Nvidia’s stock falling as low as about $7; Jetson and automotive edge chips began receiving investment around 2014 and contributed only fractions of a percent of revenue for years, yet remained in the portfolio.
Jia identifies 3 prerequisites: the founder must be sufficiently committed to the judgment, engineers must believe in the founder’s logic and character, and the core gaming-GPU business must provide the cash base. Without any one of them, a small business is unlikely to survive inside a large company for 10 years.
Huang repeatedly said, “We are not a hardware company; we are a software company,” later upgrading that to an AI company. Jia believes that clearly sharing the Vision enables engineers to believe they are working on something “increasingly important.”
11. Jia left Nvidia to put technology in users’ hands
His departure in 2020 was not driven by a loss of faith in Nvidia. Jia wanted to see real users using products he had built; as a supplier too far from users, he could not complete the data loop he had come to believe in.
He also did not think 2020 was the right time to start another autonomous-driving company. 2016-2017 had been the startup and bubble peak, 2018-2019 brought an industry downturn, and the eventual recovery came through Tesla and the mass production of Nio, XPeng and Li Auto.
The more rational move was therefore to join an automaker and build the final product. “Make your own philosophy and your own product” mattered more than launching another algorithm supplier during a cyclical trough.
12. Li Auto won him over with a technology-product-commercial flywheel
Two things Li Xiang said particularly impressed Jia. First, “technology must ultimately serve the product.” Second, “data is everything.” Technology has no meaning without a user product and a feedback loop.
Jia was especially struck that Li Xiang did not come from an AI background but recognized early that autonomous driving would ultimately be a data problem, not a contest over whose algorithm description sounded better. That aligned closely with the conclusion Jia had drawn from Tesla.
Li Xiang viewed autonomous driving as the “lifeline” of the second half of the new-energy vehicle cycle: electrification would eventually converge, while intelligence would determine differentiation. The car would first become a mobile robot, then a second home for the user, before similar technology extended into household service robots.
13. Li Auto was still an Underdog without a technology halo
When Jia joined in September 2020, Li Auto had just opened up its position through range extenders after the failure of its first vehicle, while autonomous driving still relied mainly on suppliers. The team had only 20-30 people, and the outside view was that it had “basically nothing.”
He chose Li Auto over more mature teams because his own impact could be large, the philosophies aligned and he could fully apply his prior experience. The imminent IPO was not the key variable; role and team Fit were.
Chasing into the first tier from “a very poor state” over 4-5 years was the part Jia enjoyed most. Compared with taking over an existing system inside a mature organization, he preferred starting from zero in a headwind.
14. Li Auto unified the team’s thinking, then fought through projects
Jia believes that if Li Xiang, 王凯, 郎博, 王佳佳 and he had each believed in the Waymo, Tesla and Nvidia approaches, the project “would absolutely have died.” The first consensus was to learn from Tesla, make hardware standard and build the data loop first.
Li Auto’s middle management largely did not emphasize a fixed organizational structure. People were pulled from across the company by project and sent to hotels for closed-door development. Each task had a campaign name, with the organization reconfigured around the objective.
The team often watched the documentary 《全营一杆枪》. Faced with U-2 reconnaissance aircraft, its subjects compressed the time from radar detection to launch from minutes to roughly 8 seconds. Jia treated it as an organizational metaphor for unifying thinking, aligning action and improving time utilization.
15. Even the most urgent campaign must reserve people for the next technology
Li Auto’s autonomous-driving team consistently kept some R&D staff outside the current delivery machine. They explored immature directions including end-to-end systems, VLA, dual systems and world models.
Jia says the team could produce 20-30 top-conference Papers a year in recent years. The point was not the number of Papers, but that once research proved effective, it could quickly move into product development and mass production.
This explains why later iteration accelerated so suddenly. What outsiders saw as a delivery milestone had already been explored internally by a small team. “To truly widen the gap, you transfer this thing directly into mass production.”
16. The first existential test was 80 people and 100 days for highway NOA
In 2021, Li Auto began serious in-house development. Roughly 80 people had about 100 days to deliver a complete highway NOA, with code that was essentially built from scratch, and put it directly on users’ vehicles. Failure could have meant dissolving the internal startup team.
The solution was not simply to compress sequential work further, but to “split the time into 3 parts”: maximize parallelism, increase effective working hours and improve collaboration efficiency, making 100 days deliver utilization approaching 300 days.
“If others can do it, why can’t I?” formed the team’s emotional foundation. Jia believes existential pressure, a unified target and collective execution mattered more to completing the first battle than any individual algorithmic advantage.
17. The supply chain also needed partners who would “die if they failed”
Li Auto was first to mass-produce Horizon Robotics’ Journey 3, and later first to deploy Journey 5 and Journey 6. New suppliers including Hesai’s lidar and SenseTime’s millimeter-wave radar also took on the risk alongside the team during high-pressure development cycles.
Jia’s supply-chain selection rule is counterintuitive: “Don’t choose the industry leader.” Instead, choose the second- or third-ranked player, because failure would also threaten their survival, giving both sides the incentive to send their strongest people and form a single team.
Li Xiang, 王凯 and other senior executives aligned multiple companies on targets and resources, while the team received the trust that “whatever resources you want, I’ll give you.” Li Auto’s Lark, project-based operating model and strategic-analysis methods even influenced the organization of some suppliers.
18. Mapless driving was a product-fairness requirement, not a technology stunt
When urban NOA depended on high-definition maps, competition became a race over who could open cities faster. But automakers did not hold mapping licenses, leaving eventual coverage dependent on suppliers such as Amap and Baidu.
Jia recalls that maps at the time might cover only about 30 cities and their key arterial roads. Li Xiang’s question was simple: if users paid the same amount, why could people in first- and second-tier cities use the product while those in third- and fourth-tier cities could not? That was unacceptable for the product.
A small Mapless team had already validated feasibility in Shenzhen. Shenzhen’s elevated roads, side roads and tidal lanes were highly complex; although the performance was worse than the mapped solution, it showed that “there was something here.” After the first mapped version, the entire team quickly switched to mapless driving.
19. The 2023 mapless switch collided with product deterioration and layoffs
Opening up side roads, industrial parks, arterial roads and highways all at once exposed a lack of data and caused product performance to deteriorate sharply. The transition also coincided with Li Auto’s layoffs in April and May 2023, putting both team morale and external confidence under pressure.
Li Xiang had “zero tolerance” for product flaws. New versions went directly onto his car, and feedback could be: “If this thing doesn’t work soon, the team is disbanded and you’re fired.” Jia and delivery lead 王佳佳 bore most of the pressure.
The team did not simply repeat the long-term direction. It explained each problem’s data, root cause, solution and repair timeline. After 2-3 to 4 weeks, most issues were closed and Li Xiang began to recognize the mapless experience.
By June, the mapless solution had reached all users. Removing map boundaries allowed side roads, industrial parks, urban roads and highways to be used continuously, lifting the product to a new level; some employees who had previously been laid off were also recalled.
20. Real organizational capability means falling into the pit and climbing out together
The deepest line Li Xiang left Jia was: “No one can avoid the pit. In the end, what matters is who can climb out faster.” Mistakes are unavoidable; the key is not to wallow in the pit or rebuild the old solution.
Another rule was: “If we fall into the pit, we fall together.” Nobody could stand outside making snide comments—“See, I told you you were wrong.” Misaligned views turn retrospectives into blame allocation and prevent the team from reaching consensus on what comes next.
Li Auto’s autonomous-driving team had relatively little historical baggage. Once research showed that a new direction was better, it could overturn old code and old products. Jia sees that speed of climbing out as the core capability behind the later catch-up.
21. Autonomous-driving success is measured by usage, monetization and industry standing
By the end of 2024, after the end-to-end dual system had matured, Jia believed Li Auto’s autonomous driving had truly “been built.” Internally, the first metric was daily active usage: whether users used it every day, rather than whether the team thought the technology was advanced.
The second was the share of autonomous-driving mileage in users’ total trips, preventing users from merely trying it briefly. The third was NPS, measuring whether owners would recommend the product to friends.
Commercial metrics were equally direct. Max cost about RMB30,000 more than Pro; its share was initially in the low teens and later rose above 70%. That meant autonomous driving had shifted from a pure cash-burning team to a product capability users were willing to pay for.
Industry recognition was the third layer. Jia placed Huawei, Li Auto, XPeng and some other teams in the public’s commonly recognized first tier, with the key point being that Li Auto had caught up from a clearly lagging position in a short period.
22. Autonomous driving’s technical paradigm began to look “close to final” at the end of 2024
After end-to-end systems, dual systems, VLA and world models moved into mass production one after another, the channel from R&D to product had begun working smoothly. Future projects would be rapid iteration within an established system, rather than a life-or-death paradigm rebuild like the mapless switch.
In discussions with Tesla’s team, Jia found that the direction from V13 to V14 was close to Li Auto’s: a more integrated VLA, world model and language reasoning. He was pleased that the thinking aligned, but disappointed that Tesla had added some System 2 capabilities “even later than us.”
His view is that in-vehicle compute and models will continue to grow, but the “data-driven + integrated model” paradigm will not fundamentally change. The focus will shift to efficiency, operations and user-level L4, which is no longer the new problem he most wants to solve.
23. His startup choice came from eliminating several candidate tracks
By the end of 2024, stronger performance from VLA, π0, end-to-end systems and dual systems had increased Jia’s conviction. In early 2025, he was already discussing with Li Auto colleagues that “autonomous driving is more or less done” and that it was time to begin the next 10 years.
Foundation models had become a capital game that only large companies could afford. Agent, in his view, looked like building iOS apps 10 years ago and would soon attract a flood of teams. Consumer AI hardware leaned more toward product, operations and Shenzhen supply-chain speed than toward his strengths in algorithms and systems.
The remaining track that fit best was embodied intelligence: a large end market, AI as the core driver and a need to combine algorithms, software and hardware tightly. He wanted to find a venture of his own that he could “work on with peace of mind for 10 or 15 years,” rather than take another similar job.
24. The embodied-intelligence window is longer than the foundation-model window, but nearing closure
Foundation-model startups have 3 natural advantages: OpenAI has already validated the paradigm, the internet provides massive data and users are relatively tolerant of early mistakes in purely virtual products. Autonomous driving likewise has a Tesla paradigm, vehicle data and a driver who can take over.
Embodied intelligence has none of the 3: the technical paradigm has not converged, there is no naturally available dataset for cold start, and robots must complete physical tasks independently. A 60-70% success rate has no practical value.
The first entrants’ window will therefore last several years. Jia expects the industry to shift from To A, driven by demos and financing, to To B, driven by commercial results, around the end of 2026 or early 2027. After that, the opportunity for new entrants will narrow materially.
Li Xiang’s judgment was: “The first to do it in this industry will definitely die, but the ones that ultimately succeed will definitely be from the first batch.” The direction and timing can be backed early; the real risks lie in organization, control and execution.
25. Three long-term collaborators formed a complementary startup triangle
Jia left Li Auto at the end of June 2025. Initially, to minimize the impact on his former employer, he contacted only a small group of Li Auto’s existing investors. At Yuan Capital, he found 王凯, then a partner, and the two quickly decided to start a company together.
王佳佳 joined after completing Li Auto’s subsequent VLA delivery. The division of labor extended years of tacit coordination: Jia handles algorithms, models and technical direction; 王佳佳 handles hardware, low-level software, products and delivery; 王凯 handles strategy, management, IR and financing.
Jia was not attached to being the No. 1 executive, but in the early stage of embodied intelligence, the technical paradigm still determines how data and GPUs are deployed. The 3 ultimately decided that he should be CEO and make the calls. “You have to be willing to spend the money, but once you find the paradigm, you have to spend aggressively.”
He took 3 principles from Li Xiang: technology serves the product, the product serves the business; organizational culture sustains alignment during difficult periods; and final decision-making authority must clearly rest with one person.
26. “Simplicity” is a principle for the model, product and organization
The company’s English name is Simplicity and its Chinese name is 至简动力, derived from “the great way is simple.” Its slogan is “Simple Scaling”: simpler systems are, in theory, easier to scale.
At the model level, more human-designed modules and rules mean poorer scalability. At the product level, the goal is “out of the box,” reducing customers’ on-site debugging burden and learning curve as far as possible.
The organization likewise aims to stay flat, small, efficient and project-driven, without rushing to build large departments and layers of management. The name is therefore not an aesthetic statement, but a shared constraint across 3 layers of methodology.
27. The “hexagonal warrior” cannot stop at a polished Demo
Simplicity Dynamics defines itself as a “hexagonal warrior” that must be strong across technology, strategy, product, brand, organization and business, because the industry is shifting from To A to To B and point-model performance is no longer enough to ensure survival.
The company intends to control the full stack from the robot body and data-collection hardware through algorithms, training frameworks and inference engines. Jia’s autonomous-driving retrospective is that the early market had many algorithm, mapping and data suppliers, but the more durable teams were generally integrated across software and hardware and capable of delivering products.
Full-stack capability is not about doing everything. It ensures that the model can define the body, the body can generate data, inference can drive continuous iteration and products can generate sales and repurchases. “These links build on one another and are tightly connected.”
28. Demographics underpin Jia’s most personal market thesis
Jia has 2 daughters, born in 2018 and 2021. By the figures he cited on the program, China’s annual births fell from about 17M to more than 9M, then to more than 7M in 2025.
That led him to ask who would provide care and daily services when he is old. By the time his daughters become adults, the labor dividend and low-cost services will be disappearing year by year. Household robots may therefore be more than a consumption upgrade; they could be one response to the demographic structure.
At the end state, he considers it reasonable to imagine “a system for every household, even every person.” That is also why he is convinced the total addressable market is large enough to justify investing 10-15 years.
29. The embodied end state is more likely to be several vertical oligopolies
Jia rejects the idea that humanoids will unify everything. The upper body’s eyes, arms and hands may converge toward a humanoid form because that is a general-purpose manipulation structure. The lower body, however, may continue to support bipedal, quadrupedal, wheeled and tracked designs in parallel.
Humanoids will capture the largest general-purpose market, but many specialized environments have more efficient dedicated configurations. The end state may not be a handful of companies taking 80-90% of the market; instead, several software-hardware leaders may emerge in each vertical.
It would be a form of “distributed monopoly”: scale, data and the product loop would still govern each vertical, but there would be no inevitability that a single body design could cover every task.
30. Open Foundation Models and vertical closed loops will coexist
Jia expects 2 ecosystems to emerge, much like iOS and Android. In one, large companies such as ByteDance, Alibaba and Tencent provide general-purpose multimodal Foundation Models, while vertical teams add their own data and Workflow. The other resembles Tesla, with as much of the stack closed-loop and self-developed as possible.
Startup boundaries will not lie in training a foundation model on the entire internet, but in owning the body, Action data, vertical workflows and deployment capability. Commercial implementation is “dirty and exhausting work,” and internet companies may not want to complete it industry by industry.
Large companies may eventually cover general-purpose 3D understanding, spatial reasoning and generation, while startups handle Conditional Training, body adaptation and Action generation. The two will overlap, but this will not be as simple as “a large model plus an action head.”
31. Simplicity is building the bottom layer of a 3-tier technology stack first
Jia divides R&D into 3 layers. The bottom is AI Infra, including vision, robot bodies, components, training frameworks and inference engines. The middle is the embodied Foundation Model. Only at the top come specific applications in factories, supermarkets, logistics and services.
The first 6 months of the company focused mainly on the bottom layer because long-term competition is fundamentally about iteration efficiency. Data loops, GPU utilization, training speed and edge inference determine whether the same data can be converted into product improvements faster.
The team is also keeping some R&D focused on exploring middle-layer models, but is not rushing to deploy across a large number of scenarios. Jia is carrying over the Li Auto playbook: establish Infra early, and application-conversion speed can suddenly create separation later.
32. The integrated model is designed to eliminate dual-system conflicts
After Li Auto delivered an end-to-end dual system in 2024, the team saw the natural problems of 2 models: different fast and slow frame rates, difficult arbitration when opinions conflicted and the need to construct large amounts of paired data for joint training.
For its 2025 delivery, Jia says the team integrated the world model and VLA into a single model: it generated language CoT, predicted the next few frames and then output Action.
“Fast” and “slow” were no longer fixed responsibilities assigned to 2 independent models. Instead, the model decided for itself when to think and when to act directly. Simplicity Dynamics continues to develop along the unified VLA path.
He views generation, understanding, reasoning, world prediction and Action as mutually reinforcing capabilities, rather than modules that can be optimized separately over the long term.
33. OpenAI, Gemini and Tesla are all Unification examples in his view
Jia’s retrospective is that OpenAI initially advanced language understanding, Thinking and non-Thinking, visual understanding and visual generation separately. GPT-4o introduced multimodal understanding for the first time, followed by the gradual addition of multimodal generation; after GPT-5, Thinking models also began to be integrated.
Gemini 1.0 emphasized native multimodality from the start, but its early results failed to win broad recognition. By what the program called Gemini 3, the performance of understanding, generation and upper-layer applications made the “grand unification” route more persuasive.
Tesla V14 put 3D world-model prediction, language CoT and Action into a single system. Jia concluded that “Unification is definitely the direction of travel,” even though the specific architecture has not become an industry consensus.
34. Simplicity will not train an internet-scale foundation model from scratch
The team currently uses Janus-Pro partly because open-source models that unify generation and understanding are rare. Its 1B and 7B versions are also relatively small, making them better suited to embodied edge deployment and rapid experimentation.
Jia does not believe that directly adding an Action Head to a native VLM is enough. Internet images and text contain large amounts of knowledge irrelevant to robots, while lacking 3D understanding, spatial reasoning and future-3D generation. That path “could even be a dead end.”
The ideal division of labor is for Alibaba, Tencent and others to train underlying native multimodal models, with Simplicity adding body and vertical data to turn them into embodied Foundation Models before adapting them to specific tasks. The program said the team has maintained ongoing discussions with relevant multimodal teams.
Large companies may eventually build robots themselves, but general-purpose models and commercial deployment across individual scenarios remain different capabilities. Simplicity is betting that the latter will not be eliminated by a single general-purpose API.
35. There is consensus on data cold start, but Simplicity is behind its original schedule
Early in the company’s life, Jia’s biggest concern was that embodied intelligence would not naturally generate data in the way cars and the internet do. Discussions with teams including Sunday Robotics in July and August 2025 showed that wearable devices and semi-real-machine collection were beginning to produce some Scaling effects, strengthening his confidence.
Simplicity currently has about 10,000 hours of data, mainly from semi-real-machine collection but also including real-machine data. That is below Jia’s initial expectation. The bottleneck is not only collection scale, but also the reliability of wearable devices.
Gloves, for example, can suffer electromagnetic interference, causing position data to jump suddenly. Roughly one-third of early data was unusable; after iteration, the usable rate rose to about 90%. Hardware quality directly determines whether data can enter training—it does not become valuable merely because it has been collected.
36. Commercialization is the most valuable data-collection system
Jia rejects the separation of “don’t discuss commercialization yet; just build the model and data.” Tesla taught him that scaled sales create real user scenarios, which in turn generate long-tail data that semi-real-machine systems, synthetic data and collection facilities cannot imagine.
Autonomous-driving teams often spend “90% of their effort solving the final 10%.” Accidents and Corner Cases are difficult to enumerate from an office; the most complex problems and the best data are in real-world scenarios.
Whether a product has entered real use must also be measured through usage, repurchase and repair rates, not by whether it sits at a customer site after shipment. Commercial outcomes test product value and determine the scale of the next data cycle.
37. Standard dual-chip hardware reserves room for Shadow Mode iteration
Simplicity plans to make dual-chip hardware standard: one chip runs the stable production version, while the other collects and analyzes data and runs new models in Shadow Mode without affecting users.
The new model can be compared with the old version on the additional chip, with deviation data sent back to the cloud to judge which model is better. At sufficient scale, different models can also be distributed to different devices for large-scale A/B testing.
The design inherits Tesla’s logic: standard compute is both a data gateway and a test bed. Jia therefore places edge inference at high priority rather than focusing only on training-side parameter scale.
38. Synthetic data can extend the long tail but cannot replace real manipulation data
程曼祺 asked about the industry view that synthetic data should be emphasized and that reinforcement learning with real data in real environments is difficult. Jia did not call that view “non-mainstream,” but said people who have shipped mass-produced products generally believe synthetic data is “useful, but not the main force.”
Even with a relatively simple vehicle body and obstacle-avoidance problem, autonomous driving did not have synthetic data account for 90% of training. Its most effective use is to take a rare Corner Case and expand it to 100 examples through a Pipeline.
In embodied intelligence, synthetic data can support Locomotion, closed-loop testing and validation. But using it as the primary source of manipulation data offers “no visible hope at present.” The long tail of real-world scenarios still has a huge Gap.
On Nvidia’s enthusiasm for simulation, Jia says the explanation is the logic of an ecosystem company: if customers have a need, it makes the tool more complete and faster. That does not mean simulation has become the primary data source for all manipulation models.
39. Hardware consistency is the real problem most easily waved away
What the industry calls mass production currently often means only a few thousand units, far from the several million or tens of millions of vehicles produced annually. Many robot bodies are still “handmade,” without shared standards for process, consistency, stability or reliability testing.
Jia’s direct test is this: take data from the same manufacturer and the same dexterous hand, replay it on B using data collected from A, and the behavior may differ. The resulting noise can outweigh the gains from additional samples, depriving model training of a reliable foundation.
On repair rates, he declined to give figures for specific manufacturers, but said that by his observation they are “basically 100%.” Few devices can work for an extended period without repairs, making high-frequency commercial tasks difficult for now.
Unitree’s demonstrations of multiple robots performing difficult synchronized movements in large-scale shows gave him hope for consistency and process-based manufacturing. But he still believes the industry is at least 1-2 years away and cannot skip manufacturing fundamentals because of model excitement.
40. Simplicity uses a “semi-humanoid” body to match current model capability
The company’s current body has a dual-arm upper body and a wheeled lower body. Jia believes this is better suited to near-term scenarios than a fully bipedal form: it moves efficiently while retaining the core manipulation capability.
In just over 6 months, the hardware has already gone through 2-3 rapid iterations. The core principle is not to maximize human likeness or degrees of freedom, but to define the body around what the model can use at the time and prevent the hardware team from “running wild and entertaining itself.”
Tesla’s insistence on a fully humanoid form may force the supply chain to mature from the end state, benefiting later entrants. Simplicity instead uses existing integrated joints and the experience of industrial partners to avoid repeating past mistakes.
Production is in Suzhou, close to the field, allowing problems to be solved the same day. Jia views this proximity to manufacturing as a capability software teams were previously unfamiliar with but must now build.
41. Moving from 60-70 points to near-100% determines whether To B has truly arrived
A general-purpose Foundation Model may score only 60-70 on each task. That can already be valuable in a chat product, but for a physical robot it means continuous errors. A single commercial task must reach 99% or even close to 100%.
Jia says the industry broadly agrees on a 2-step process: first use similar-scenario data for SFT/Domain Adaptation, then use real-machine I/O or Offline I/O to reinforce a single task. The debate is not whether the 2 steps exist, but how efficiently they can be executed.
Simplicity hopes that after switching to a new task, it can bring baseline capability to a deliverable level in 20 minutes or half an hour. This remains a target, not a completed result. If every customer requires prolonged customization, the business model will revert to non-standard automation.
To B therefore requires 2 things at once: hardware that can be mass-produced reliably and a model that can move from a general-purpose score of 60 to a vertical score of 100% at low cost. Jia’s timing has a basis, but he did not present 2026-2027 as unconditional inevitabilities.
42. The first scenarios must cover complete workflows, not replace one station
Workers on mature production lines can complete an action in a few seconds, which robots currently cannot match. An error can also halt the entire line, while high-frequency actions amplify heat, service-life and repair problems.
Customizing a single fixed station would leave Simplicity no different in substance from China’s more than 7,000 non-standard automation companies. The team therefore starts by asking about the customer’s understanding and workflow, rather than rushing to the site for a Demo whenever it hears about a scenario.
More suitable tasks run from retrieving materials from a warehouse through loading, unloading, inspection and deburring, then returning goods to the warehouse. A robot can reconfigure the complete process, cover multiple links with one Skill Set and replicate the solution more easily at a similar customer.
The current capability boundary is “general-purpose mobility + simple manipulation + object generalization.” Grabbing, taking, retrieving, placing and transporting remain relatively simple, but pretrained visual capability can recognize objects that continue to change.
43. Flexible manufacturing, European factories and supermarkets form the early demand pool
Many factories do not produce a single product year-round, while changing economic conditions have made orders smaller and more fragmented. Traditional robotic arms with fixed trajectories are highly sensitive to changes in objects and positions and have almost no Zero-shot generalization.
王凯 has nearly 20 years of overseas experience and visited Europe repeatedly after the company was founded. Jia believes ordinary service workers there can cost about €3,000, giving flexible small-batch factories a stronger incentive to buy robots that can switch tasks quickly.
Supermarkets also have large amounts of structured but non-fixed work. By Jia’s description, 80%-90% of supermarket staff handle shelf organization, replenishment, transport and placement, making the environment suitable for systems that can recognize varied objects and perform basic handling.
The common feature of these markets is not spectacular robot movement, but end-to-end workflows, continuously changing objects, high labor costs and the ability to replicate one solution across customers.
44. Household robots must wait for safe interaction to become a general capability
Jia divides capability evolution into general-purpose mobility, general-purpose manipulation and general-purpose interaction, followed only much later by cross-body generalization. The realistic near-term capabilities are mobility, simple manipulation and object generalization.
Manipulation generalization requires high-degree-of-freedom dexterous hands covering 40-50 atomic actions such as grasping, pulling, inserting and twisting. But dexterous-hand service life, price and reliability have not converged; Jia expects this could still take 2-3 years.
Household use must wait for safe interaction to mature. A robot weighing dozens of kilograms falling onto a child or a robotic arm flinging something once could be fatal. Industrial robotic arms are fenced off for precisely this reason: bringing robots into homes cannot be judged by Demo success rates alone.
Systems such as Figure and π0.5 can already let a small number of Geeks test tasks such as folding blankets at home. Mass-market household products will need to work their way through factories, supermarkets, hotels and restaurants, and other closed or semi-closed environments before their boundaries are understood. Jia puts the timing around 2029-2030.
45. Organization, capital and the CEO transition all serve the same systems battle
Simplicity currently has more than 80 people and no second layer of management. Hardware, low-level software, motion control and product report directly to 王佳佳; algorithms, models and inference to Jia; and strategy, IR, GR, PR, finance, legal and fundraising to 王凯. The 3 resolve disagreements through a weekly strategy meeting.
Hiring starts with whether candidates believe in embodied intelligence, followed by whether they can accept “wasteland.” Roles are not fixed, and each person may be opening 2 or 3 areas at once. Jia wants the roughly 200-person Tesla Autopilot team that built a top-tier product as the organizational benchmark, rather than erecting heavy processes too early.
Financing is being staged across top-tier financial investors, strategic investors such as Alibaba and Tencent, and supply-chain and application-side industrial partners. When the host mentioned a fifth round and a post-money valuation above RMB10B, Jia confirmed the fifth round but did not confirm the valuation. The speed of fundraising brings higher expectations: the team must deliver something new almost every month.
Data, talent, GPUs and hardware are the 4 major funding needs, with GPUs already accounting for more than one-third of R&D expenses. Jia worries that excessive caution could cause the company to miss the To B expansion window, but still insists on first completing the minimum viable loop before deciding whether to scale out.
For 2026, he is relatively confident that the model and data paradigm will continue to converge; hardware remains the biggest uncertainty. For 2027, he is watching whether multimodal reasoning, Visual CoT, Unification and Self-evolution produce a signal comparable to a “GPT-3/GPT-4 moment.”
The IPOs of the first embodied-intelligence companies will give the sector a shot in the arm while giving leaders stronger capital and financing efficiency. Around 2028, when PMF emerges, large companies may accelerate their entry, and the 1-2 years after the “hundred schools of thought contending” phase could also bring a bubble rupture.
Entrepreneurship has turned Jia from an engineer focused only on technology into a No. 1 executive more like a parent. Beyond technology, he must care about employees’ desks, drinking water, sick children and career prospects. Strategy now means being explicit about what to pursue and what to leave aside: the long-term Vision cannot be wrong, near-term actions must be precise, and the pits in between must be climbed out of quickly.
The floor for 2-3 years from now is for Simplicity to catch up immediately when others find PMF and the technical paradigm. The higher goal is to be the discoverer itself. “I would rather take something from a state where nobody believes in it—or from zero—and bring it to a very good state.”