96. A Conversation with 郎咸朋 on 10 Years of Autonomous Driving, Key Technical Details, and Tesla
Summary
郎咸朋 compresses autonomous driving’s 10-year evolution into 3 paradigm shifts: HD maps and rules, BEV + Transformer, and end-to-end. The first generation treated roads like “tram tracks laid over the real world,” but highways cover only more than 300,000 km while ordinary roads cover more than 9 million km, making nationwide map coverage and updates impossible; the second used BEV to unify 360-degree visual understanding; the third shifted from coding every scenario to training driving capability—“light-map, mapless, all of that is really a minor detail.”
Tesla’s early advantage was not a single algorithm but a closed loop linking pure vision, BEV, ASICs, and mass-production economics. 郎咸朋 cites these figures: an early 64-line lidar cost RMB500,000-600,000 per unit, while the total sensor cost of one Baidu vehicle was about RMB5M; Tesla’s full sensor-and-chip package cost about $1,000. Its two 72-TOPS chips delivered 144 TOPS in total, far below the 500+ TOPS of two Li Auto Orin X chips on paper, but he argues that effective performance was at least as strong: “You can’t look only at the TOPS number.”
End-to-end turned autonomous driving from “software 1.0” into data-driven “software 2.0,” moving the moat to high-quality data and training compute. The traditional stack had to enumerate combinations of weather, traffic, lighting, and roads while compounding errors across modules; end-to-end generates driving trajectories directly from sensor inputs, shrinking Li Auto’s roughly 2M lines of mapless or light-map code to fewer than 200,000. 郎咸朋’s conclusion is blunt: “Training compute and data will determine the ceiling of autonomous driving.”
Li Auto’s 2024 autonomous-driving turnaround was a high-pressure organizational experiment, not incremental optimization. At a March strategy meeting, 李想 demanded a reversal by September; on April 15, about 180 people entered a closed development sprint in Zhongguancun, and around the May Day holiday the first model could drive from Zhongguancun to Beijing Jiaotong University. The initial model had fewer than 1M KPS, yet outperformed the rule-based system 郎咸朋 had driven on longitudinal braking; the team told him, “Dr. Lang, we really didn’t tune it.” His recollection of 李想’s first test drive varies slightly between late May and June, but he took over only once or twice in more than an hour, far exceeding expectations.
Li Auto’s vision for L3 is not a natural upgrade from L2 with more features, but a lead product for L4. The route is to first connect the full door-to-door experience, with humans supervising continuously when capability is insufficient, then gradually convert reliable road segments to unsupervised driving; the threshold is roughly one takeover every 200 km overall, including more than 350 km on highways and more than 50 km in cities. Based on Li Auto’s historical city/highway mix, an ordinary user would take over about once a week, while the separate safety target is “10x safer than human driving.”
The complete architecture is not just end-to-end, but System 1, VLM System 2, a world model, and a reinforcement-learning loop. System 1 learns common driving behavior; the VLM handles unfamiliar scenarios requiring reading and reasoning, such as variable lanes and time-based bus lanes; the world model uses 3DGS to reconstruct real exam questions and Transformer Diffusion to generate simulated ones, testing every model release across 5 dimensions: safety, compliance, comfort, navigation, and efficiency. “Using a Qing-dynasty sword to execute Ming-dynasty officials” is 郎咸朋’s verdict on the old feature-testing system.
郎咸朋 sees 2024 as the year intelligent driving began to materially affect vehicle sales, while capital shifted from headcount to training resources. He acknowledged that Li Auto lost deals in the first half as consumers switched to compare Aito’s intelligent driving; laggards that merely “do about as well as everyone else” still cannot reverse the competitive position. The team had about 800 people at the time of the interview, down from roughly 1,000, but training-card investment had multiplied. His competitively charged assessment: at the time, only 2 companies in the world were actually delivering end-to-end products—Tesla and Li Auto.
Li Auto sees autonomous driving as the first proving ground for a broader foundation-model strategy, not an isolated feature. 郎咸朋 advocates building its own multimodal Mind GPT: long-term dependence on GPT, Qwen, ERNIE Bot, and other general models creates supply risk, and their optimization targets may not match Li Auto’s needs; autonomous driving, Li Auto Tongxue, intelligent industry, and intelligent commerce should all grow from the same foundation. VLA has entered preliminary research, with a visible window of “within 1 to 3 years,” but he explicitly made no claim that it is the final destination.
Deep dive
1. Autonomous Driving Saw Only 3 Genuine Paradigm Shifts in 10 Years
郎咸朋 describes Baidu ADU in 2014-2015 as the rule-based era: HD maps, lidar covering the entire vehicle, and engineers writing “tens of thousands of lines of if-else.” The Wuzhen demonstration embodied the prevailing conviction of the time: if maps were accurate enough and rules complete enough, a car could drive itself along virtual tracks.
The second shift came around 2018. Tesla was the first to explicitly reject HD maps and lidar, adopt pure vision, and use BEV + Transformer to place multiple camera feeds into a unified spatial representation; 郎咸朋 believes it “saw correctly, saw accurately, and saw it very early.”
The third shift was the end-to-end breakout of 2024: instead of separately optimizing perception, decision-making, planning, and control, the system learned driving trajectories directly from raw sensor inputs. That is why 郎咸朋 puts map-based, light-map, and mapless systems in the old paradigm: “The rest, I think, is all a minor detail.”
2. Virtual Tracks from HD Maps Cannot Cover the Real World
郎咸朋 explains the early approach through the metaphor of a tram: first lay a sufficiently precise virtual track for every road, then use 360-degree lidar to detect dynamic people and vehicles. Highways, railways, and subways are easier to automate precisely because their tracks and operating boundaries are relatively stable.
Scale breaks the model first. His figures are more than 300,000 km of expressways in China versus more than 9 million km of ordinary roads; even if every road were mapped in HD, collecting and updating the full network at sufficient frequency would be impossible.
Ordinary roads are constantly dug up, rerouted, and put under temporary construction, while highways are comparatively stable. HD maps can support highways but struggle to become the common base layer for urban roads nationwide. “Dig a hole today, change the road tomorrow” turns map freshness into more than an engineering-budget problem.
Asked how long the industry stayed on this path, 郎咸朋 places the turning point around the emergence of BEV Transformer. After Tesla publicly embraced pure vision around 2018, the industry gradually understood that the old map-heavy, lidar-heavy route was not only expensive but difficult to scale.
3. Vision-Based Driving Was Not Sudden; Sensors and Compute Were the Constraints
郎咸朋 traces the history back 20 or 30 years: researchers in the 1980s and 1990s were already trying to make cars recognize their environment and control themselves, but cameras might have only tens of thousands of pixels, roughly 600×800, while computers could not process high-resolution images in real time.
Early experiments therefore relied more on one-dimensional signals such as millimeter-wave radar. The compute burden was smaller, but the car was effectively “driving blind,” or like a bat or dolphin, navigating through echoes. Once roads filled with vehicles and complex road users, the information was plainly insufficient.
Around 2013, ImageNet, convolutional neural networks, powerful chips, and better automotive cameras reopened the vision route. At the same time, successful entries in U.S. autonomous-driving challenges used lidar, directing industry capital and talent toward sensor-heavy systems. The 2 approaches then developed in parallel.
4. Tesla Turned Pure Vision, Compute, and ASICs into a Mass-Production System
郎咸朋 considers Tesla’s “humans have only 2 eyes and can still drive” a first-principles explanation aimed at the mass market. The technical substance is that images contain color, category, and texture, while pinhole geometry, camera parallax, and continuous video implicitly encode 3D position, speed, and motion trends.
Pure vision cannot process only the roughly 120-degree field in front. Tesla placed about 6 or 7 cameras around the sides and rear, using Bird’s Eye View to create a unified 360-degree view around the vehicle; the compute requirement consequently rose to 6 or 7 times that of a single front-facing camera.
Xavier-generation chips delivered about 30 TOPS, insufficient for Tesla’s envisioned BEV network. Tesla therefore began preparing its own chip in 2016 and reached 144 TOPS in 2019 with two 72-TOPS chips. 郎咸朋 stresses that although this seems ordinary today, it was far above the industry’s usable 30 TOPS at the time.
A proprietary algorithm and ASIC are like a tailored suit: they do not need to reserve generic redundancy for other customers or network structures. 郎咸朋 puts Tesla’s full sensor-and-chip cost at about $1,000; Li Auto used two Orin X chips with more than 500 TOPS on paper, but TOPS on a general-purpose chip cannot be equated directly with effective compute.
5. BEV’s Value Is Early Fusion, Not Stitching Together 6 or 7 Recognition Outputs
Before BEV, front cameras recognized people, vehicles, and lane lines, while the sides and rear relied more heavily on millimeter-wave radar. Radar from Tier 1 suppliers such as Bosch and ZF could output an object’s approximate position and speed—enough for blind-spot warnings—but was essentially a 2D point, inadequate for fine-grained planning.
The most basic multi-camera setup was to recognize objects in each image separately and then fuse the results. The same car might be missed by one camera and only half-recognized by another; the system would not know which source to trust or whether the reconstructed vehicle was 1.5 meters, 2 meters, or 3 meters wide.
Tesla’s approach is to extract features from all images, reconstruct surrounding objects once in a unified space, and then project the result back into each camera. 郎咸朋’s simplified explanation is: “First stitch all the images into one big image,” then recognize people, cars, and spatial position from the same source.
The old approach synthesized upward from multiple flawed outputs; the new one projects downward from a unified representation of the world, so consistency is inherently stronger. 郎咸朋 calls this Tesla’s recurring “dimensionality lift”: instead of eliminating post-fusion errors one by one, change the level at which the problem is represented.
6. Lidar Gives Direct Distance; Vision Offers Orders of Magnitude More Information Density
郎咸朋 compares the systems by resolution: a camera has roughly 8M pixels, close to 4K×2K; vehicles generally use an AT128 and another 128-line lidar, while early Velodyne units had only 64 lines. A person 1.8 meters tall at 150 meters might generate only a handful of points, with even worse results in black, low-reflectivity clothing.
Lidar’s clear advantage is that “if there is a point, it tells you the distance.” Vision must infer distance from parallax and temporal information. But a camera’s full frame is filled with color and texture rather than the large holes found in point clouds, making it better suited to reconstructing a complete world.
The cost gap was once extreme: a Velodyne 64-line lidar cost about RMB500,000-600,000, while early Baidu and Cruise test cars might carry 7 or 8 units. 郎咸朋 recalls that the total sensor cost of one Baidu vehicle was about RMB5M; it even mounted a 16-line lidar horizontally to search specifically for traffic-light poles.
张小珺’s counterquestion is worth preserving: if lidar later became cheaper, why continue insisting on pure vision, and why did so many people consider removing lidar unimaginable? 郎咸朋 did not reduce the answer to cost, instead pointing to path dependence: L4 in constrained areas can work with the old architecture, but nationwide mass-market vehicles need a more unified solution in information, models, and scalability.
7. Modular Stacks Lose Information Layer by Layer; End-to-End Keeps First-Hand Inputs in the Loop
BEV solves perception only. Traditional systems still pass perception outputs to hand-designed decision rules, then to planning and control; even excellent perception cannot tell engineers how to drive like a human in every situation.
Perception can miss objects, hallucinate objects, or discard information, meaning the decision module no longer receives first-hand data. Decision errors then feed into planning and control. 郎咸朋’s example: if each layer is off by 5 cm, 3 layers accumulate 15 cm—roughly the width of a road’s white line.
The end-to-end idea is to feed raw sensors, navigation, and other inputs into one model and output the driving route or trajectory directly. Previously the system translated the scene into “there is a person, a car, and a road boundary,” then decided whether to follow, overtake, or change lanes; now it is simply, “I see this image, so this is how I drive.”
The cost is lower explainability. 郎咸朋 admits the team usually does not know exactly what the model has learned; it can only evaluate whether the car drives well. His counterpoint is that humans do not explain every action in real time either: “I just want to drive this way” does not mean no pattern was learned.
8. Turning Every Road Scenario into a PRD Is the Software 1.0 Trap
Traditional feature development requires explicit inputs and outputs: screen size, button behavior, response time. Institutions such as SAE and NHTSA consequently defined ODD and L1 through L5, attempting to write autonomous driving as a product requirement that could be checked item by item.
Real driving has no stable set of discrete variables. Weather can be divided into sunny and rainy, and rain into heavy, moderate, and light, but 郎咸朋 jokes: “How do you distinguish light rain from moderate rain—by counting how many raindrops fall in a minute?” Add traffic, lighting, vehicle speed, and opposing road users, and the combinations quickly reach the tens of millions.
Scenarios are not orthogonal either. Changing the sunny-day right-turn rule may break right turns in rain, congestion, or the presence of pedestrians; an urban expressway may resemble a highway, while a traffic jam caused by highway construction may resemble an urban rush hour. Splitting roads into highway and urban categories creates artificial boundaries as well.
The harder problem is the long tail: a horse suddenly running onto the road, a large hole opening in the pavement, or an object that might be a plastic bag, oil, slime, or a black cat. Humans slow down, stop, or go around even when they do not know what it is. A system limited to predefined objects and scenarios, 郎咸朋 believes, “is not autonomous driving.”
9. Software 2.0 Trains Capability, Not an Enumerable Set of Features
郎咸朋 uses Karpathy’s software 2.0 framework to explain the shift: instead of humans coding line by line, models learn from human driving data and iterate. End-to-end matters not because of the label but because autonomous driving is moving for the first time from “building features” to “building capability.”
The engineer’s role changes accordingly. The team cannot learn on the model’s behalf; it can only provide the equivalent of a good school, teachers, and materials, then test whether the model is improving. “I can’t learn for him, and I don’t know how he learns.”
Models and data-driven development are “2 sides of the same coin”: using a model requires data-driven training, while turning data into capability requires a model. Parameters and baseline intelligence may gradually converge; what separates companies is the quality of their accumulated vertical data and the training compute that converts data into capability.
10. Li Auto Treated Unpurchasable Data as a Long-Term Asset as Early as 2018
郎咸朋 recalls that when he joined Li Auto in 2018, 李想 asked in the interview what would matter most in future autonomous driving. His answer was data. People can be recruited, and after a company is operating well, compute can be bought with money, but historical driving data “cannot be dug up and cannot be bought.”
That is why he agrees that “the companies left standing in the AI era will be the ones with high-quality vertical data.” Lock a genius in a dark room for 20 years and capability may still be lost; give an ordinary person continuous quality education and they may achieve a great deal. Postnatal inputs matter more than small differences in IQ.
Asked about user privacy, 郎咸朋 says Li Auto collects the external environment, not biometric information such as faces and voices inside the cabin; when data is uploaded, external faces and license plates are blurred. He traces this capability to his earlier work processing faces and license plates for Baidu Street View.
11. Baidu Street View and the BMW Project Seeded China’s Early HD-Map Team
郎咸朋 joined Baidu in April 2013, initially tasked with catching up to Tencent Street View, which was already online, and Google Street View, which came earlier. Street View built capabilities in large-scale collection, stitching, and privacy processing, and became the technical starting point for Baidu’s BMW HD-map project.
In the second half of 2013, BMW China wanted to test autonomous driving in China and evaluated Baidu, Gaode, and NavInfo as potential partners, ultimately choosing Baidu. Baidu formed a 4-person team: 倪凯 and 陶吉 from 余凯’s group, plus 郎咸朋 and 杨艳 from the map team. They served BMW while learning autonomous driving.
BMW also provided 2 test vehicles, 1 of which was later used for the 2015 Beijing Fifth Ring demonstration. By 2014, relying on Baidu’s surveying and mapping license and collection vehicles equipped with lidar, the team had essentially completed Beijing’s HD map and brought the core technology to “七七八八”—largely in place.
The paths then diverged: the research institute and ADU leaned toward Waymo-style L4 and frontier research, while 郎咸朋’s team focused more on mass production and automaker partnerships, closer to Tesla’s route. Team size did not shrink monotonically as the technology advanced, because HD maps, perception, and delivery engineering each once required large organizations.
12. The Engineering Definitions of L3, L4, and L5 Are Giving Way to Supervision Experience
In 郎咸朋’s “early memory,” L3 begins to include system responsibility: within the declared operating scope, the system assumes responsibility and must warn in advance when human takeover is required; if the human does not respond, the system should pull over or enter a safe state.
L4 is autonomous driving in a limited scenario: define the area, and the human does not need to supervise within it. L5 drives anywhere like a human. 郎咸朋 believes L5 is still “quite far away,” depending not only on the vehicle but also on society and infrastructure.
He is more aligned with Tesla’s experience-based language: supervised autonomous driving means the human must still watch the system; unsupervised means they do not have to. If unsupervised driving works only in a particular city or region, it remains L4 under the old definitions, not L5.
He cites Tesla’s plan to launch Cybercab in 2025 and initially operate it within defined urban areas as an example, arguing that its basic product boundary is the same as Waymo, Cruise, and Baidu Apollo Go: first define the area, then achieve driverless operation within it.
13. Constrained-Area L4 Can Keep the Old Technology; Nationwide Mass Production Cannot
郎咸朋 has ridden in Cruise and Pony.ai vehicles and believes that within a pre-defined area, HD maps and previous-generation modular technology have “no problem at all”: maps can be updated promptly, and common cases can be analyzed and validated relatively completely.
For passengers, whether the car uses end-to-end, lidar, or HD maps does not matter as long as it drives safely and comfortably. The technology route therefore cannot be discussed separately from the product boundary: a restricted-area Robotaxi and a mass-market vehicle sold nationwide are different engineering problems.
Li Auto needs a product that works on roads nationwide and can start from any parking space. That is why, 郎咸朋 explains, L4 companies can continue using heavy maps, rules, and lidar, while mass-market vehicles must rely more on generalizable models, data loops, and low-cost sensors.
14. Li Auto’s L3 Is Not L2 Plus Plus Plus; It Is a Lead Version of L4 Capability
郎咸朋 repeatedly emphasizes: “Our L3 is not an extension of L2; it is a lead version of L4.” Continuing to add scenarios, corner cases, and rules to L2 still cannot escape the impossibility of enumerating scenarios, interference among rules, or defining the long tail.
L3 and L4 should use the same model and R&D system; the main difference is the product’s supervision state. Li Auto is not building a separate L4 team, but training autonomous-driving capability aimed at L4 and delivering versions that have not yet reached the driverless standard in supervised form.
That is why the 2024 end-to-end system was used in an assisted-driving product even though its technical target was autonomous driving. 郎咸朋 calls it “technology lifted to a higher dimension, delivering dimensionality-reduction attacks”: if the goal were only assisted driving, HD maps, light maps, and mapless systems could all ship, with no reason to absorb the cost of a model-based transformation.
He also acknowledges that Li Auto cannot beat Huawei’s “thousands of people” in conventional engineering competition, and may not beat “genius young people” such as 楼天城. Li Auto’s differentiating resources are the data generated by mass-market vehicles and its willingness to keep investing in training compute.
15. Connect the Full Door-to-Door Experience First, Then Remove Supervision Segment by Segment
Li Auto’s first step was not to immediately declare nationwide L4, but to give the model full door-to-door capability: parks, enclosed roads, cities, highways, ramps, ETC, tidal lanes, traffic lights, and final parking all need to be connected into one chain.
“Being able to handle it” does not mean the system must have seen it before. 郎咸朋 requires the system to respond reasonably to unfamiliar road elements; otherwise it is still just memorizing scenarios. This is the generalization target shared by end-to-end and the VLM.
When capability is insufficient, the owner sits beside the system like a driving-school instructor beside a novice: let the car drive most of the time, then take over in complex situations or pay closer attention when prompted. As the model receives more high-quality inputs, reliable road segments can gradually become unsupervised.
He compares the navigation route to an all-red line: in the initial version, every segment requires supervision; over time, continuous green unsupervised segments are added. Fragmented, intermittent green sections “mean very little”—they need sufficient share and continuity before opening them up is worthwhile.
16. One Takeover Every 200 km Is the Product Threshold; Safety Has a Higher Separate Target
Li Auto’s calculated launch threshold for L3 is roughly 1 takeover every 200 km overall; broken down, highways must exceed 350 km and cities must exceed 50 km, weighted by the historical highway/city driving mix.
If a user drives 40 or 50 km a day, 1 takeover every 200 km is roughly 1 takeover per week. That takeover may not be due to danger; it could simply reflect dissatisfaction with the current speed, route, or comfort. The metric is closer to usability than to a pure accident rate.
For safety, 郎咸朋 sets a separate and stricter requirement: “10x safer than human driving.” The interview did not specify the statistical methodology, but his distinction is clear—MPI determines when the product has value, while the safety threshold cannot be replaced by takeover frequency.
17. System 1 Handles Intuitive Driving; VLM System 2 Handles Reading and Reasoning
郎咸朋 believes Tesla leads because it meets 2 conditions: it captures the essence of the product and continuously follows the AI frontier, translating general-purpose technology into an autonomous-driving architecture. Many people can imagine “driving like a human,” but do not know how to implement System 1, System 2, and the underlying technologies.
System 1 is the end-to-end behavior model. It learns from large numbers of driving cases to produce fast reactions, making following, braking to a stop, avoidance, and lane changes feel more intuitive. Common scenarios do not need lengthy reasoning; they should output driving behavior with low latency.
System 2 is handled by the VLM, which processes situations the end-to-end model may not have seen and that require knowledge and reasoning—for example, reading a sign saying “bus lane unavailable from 9:00 to 17:00,” understanding a variable lane, and passing the conclusion back to the end-to-end system for a joint decision.
Li Auto began explaining System 1 and System 2 to 李想 in weekly AI seminars in the second half of 2023. 郎咸朋 says the key is not copying a particular Tesla network, but learning its way of thinking: first identify the product problem’s essence, then find a technological opportunity that can solve it by “lifting the dimension.”
18. The World Model Is a Capability Exam, Not Another Set of Driving Rules
Models train capability, while old testing still checks whether features meet PRD specifications. 郎咸朋 rejects the mismatch in one line: “Using a Qing-dynasty sword to execute Ming-dynasty officials,” because without capability evaluation, there is no way to know whether an end-to-end release truly improved or merely memorized cases.
The world-model exam covers 5 dimensions—safety, compliance, comfort, navigation, and efficiency—with basic, intermediate, and advanced questions in each category. Safety has a hard threshold; other experience dimensions can tolerate limited deductions, but a new release’s total score cannot fall below the old release’s.
A fixed question bank lets the model “memorize the answers,” so the test must preserve real questions while continuously generating simulated ones. Reconstruction uses 3DGS Gaussian reconstruction to restore real roads as repeatable test environments; generation uses Transformer Diffusion, resembling video generation but with an emphasis on greater realism and closer compliance with requirements.
A real right-turn, traffic jam, or obstacle-avoidance scenario can be reconstructed and then lightly regenerated into a superficially different question. The testing objective stays the same while the input changes, revealing whether the model learned driving capability or a specific image.
19. Real Roads, the World Model, and Rewards Form a Reinforcement-Learning Loop
Each vehicle-side model release first takes its exam in the world model; only after clearing every threshold does it reach real roads. New problems encountered in the real world then flow back as training data and test questions. This “go out and encounter a problem—come back and be penalized—retrain” loop is the complete RL architecture 郎咸朋 describes.
The easier rewards to define include safety and comfort: a collision clearly deserves a safety penalty, while a user takeover suggests a possible safety or comfort problem. But determining which category a particular takeover belongs to requires further analysis.
Li Auto built a dedicated model to analyze takeover causes. 郎咸朋 says autonomous-driving rewards are “relatively easy to design,” while preserving the key caveat: clear reward dimensions do not mean every real-world behavior can be attributed accurately.
20. The 2024 End-to-End Project Accelerated from the March Crisis to Cars in May and June
In the second half of 2023, the team had already begun preliminary end-to-end research. 郎咸朋 distinguishes the concept from the tool: end-to-end is the idea of reducing modules while preserving first-hand information; large models provide the implementation conditions for encoding images, navigation, lidar, and “everything in the world” into a jointly trained system.
At the March 2024 strategy meeting, 李想 lost his temper over autonomous-driving performance and demanded a reversal by September. 郎咸朋 did not defend metrics close to competitors; he proposed model-based, end-to-end driving as the “only solution” he could see, then elevated the preliminary work into a delivery project.
On April 15, about 180 people entered a closed development sprint in Beijing’s Zhongguancun and did not return until July. Around the May Day holiday, the first model could already drive from near Zhongguancun to Beijing Jiaotong University. 郎咸朋 emphasizes that there was “not a single line of rules” on this route; the model drove according to however it had been trained.
His recollection of the date of 李想’s first ride varies slightly: he first said June, then placed it around May 20-something, while also mentioning a June 8 trip to Chongqing. 李想 brought 经纬张颖 in the front passenger seat, drove for more than an hour with only 1 or 2 takeovers, and became more excited the longer he drove. 张小珺 noted that 李想 called getting the end-to-end model running the “aha moment” of his AI experience over the previous year.
21. The First Model, with Fewer Than 1M KPS, Already Beat Years of Handwritten Control
The first model 李想 tested had fewer than 1M KPS; 郎咸朋 could not remember whether it was about 600,000 or 800,000 KPS, with training data of several hundred thousand clips. By the time of the interview, model scale had reached several million KPS.
The architecture change shrank the roughly 2M lines of mapless or light-map code to fewer than 200,000, a reduction of about 90%. This was not simple code deletion; large amounts of scenario judgment, behavior selection, and longitudinal control moved from explicit rules into the model.
What surprised 郎咸朋 most was longitudinal braking. The team had repeatedly tuned rules yet still struggled to balance safety distance with the human feeling of braking smoothly—fast at first, slower toward the stop. On its first drive, the initial end-to-end model outperformed the in-house and competitor systems he had driven. The engineers told him: “Dr. Lang, we really didn’t tune it.”
22. 2024 Sales Pressure Turned Autonomous-Driving Investment from Catch-Up Spending into a Major Bet
郎咸朋 believes 2024 was the year intelligent driving began to materially affect vehicle sales. Li Auto’s autonomous driving was not outstanding in 2023, but the vehicles still sold well; in the first half of 2024, customers switching to brands such as Aito for a better intelligent-driving experience became an important source of lost deals.
Some team members used metrics to show that Li Auto had nearly caught Huawei and XPeng. 郎咸朋 acknowledged that this might be true but accepted 李想’s judgment: a latecomer being “about the same” as the leader is meaningless. It must be clearly better to change consumer perceptions and competitive standing.
Team size first expanded and then contracted: the mapless and light-map phases reached about 1,000 people at one point, versus more than 800 at the time of the interview. 郎咸朋 estimates Tesla’s autonomous-driving team at roughly 200 to 300 people. Headcount is not the core efficiency metric; model architecture, data infrastructure, and training resources matter more.
He says Li Auto spent tens of billions of RMB on autonomous driving that year. The cost saved through headcount optimization was negligible; the real increase was in training cards, and it multiplied. The organization shifted from stacking engineers to write rules toward using a compute platform to amplify data.
23. A 2021 Supplier Ultimatum Forced Li Auto’s First Major In-House Leap
In 2021, Li Auto wanted to launch an upgraded Li Auto ONE. Its exterior could not be radically changed within a few months, so intelligence had to become the main upgrade. But the incumbent ADAS supplier demanded an expensive development fee and refused to deliver the code as a black box.
The harder condition was that Li Auto would have to hand all future autonomous-driving R&D to the supplier and dissolve its own team. 郎咸朋 says the supplier believed Li Auto had neither the time nor an alternative: “You have to kneel down and beg.”
What he admired most about 李想’s attitude was: “Even if I die, I’ll die standing. I will never kneel down and beg for mercy.” On the first day of the Lunar New Year, 李想 called to confirm the decision. 郎咸朋 replied, “I’ve been here 3 years waiting for this moment,” and said he was willing to resign if the team failed.
After the call, 李想 immediately added all the partners to a WeChat group and announced that autonomous driving would be developed in-house, led by 郎咸朋, with full-company support. After the holiday, the team assembled more than 100 people from across departments, held a mobilization meeting on February 26, and began racing to deliver the Li Auto ONE upgrade.
24. Slow Investment from 2018 to 2021 Reflected the Company’s Survival Priorities—and Left Work to Be Done
张小珺’s question was pointed: if in-house development was so important, why did 郎咸朋 not pursue it at full force for 3 years after joining Li Auto in 2018, while XPeng invested heavily from day 1? 郎咸朋 did not evade it. He quoted 李想’s answer at the time: “Our main task right now is to deliver the cars. We don’t have that much money.”
Early capital had to go first to vehicle development, manufacturing, stores, and the sales network. Autonomous driving received a limited annual budget, with the initial goal simply to beat traditional luxury vehicles at the same price point. 郎咸朋 gradually understood in the first half of 2019 that this was not a rejection of autonomous driving, but strategic pacing under resource constraints.
He also explains why Li Auto went through map-based, light-map, and mapless systems even after it understood end-to-end: jumping from arithmetic straight to calculus before learning the pitfalls of arithmetic is not necessarily wise. The team had to build each generation itself to understand why Tesla abandoned the previous method.
Every delivery generated practical value: light-map was better than heavy-map, and mapless was better than light-map. Even while the end-to-end team developed in parallel, Li Auto still delivered the mapless version at scale in July 2024. Without real users and real-world feedback, the company could not obtain the data or organizational consensus needed for the next generation.
25. Tesla’s Core Lead Is Its Ability to Keep Translating AI Progress into Automotive Architecture
The fact that Transformer was invented by Google does not diminish Tesla’s contribution. 郎咸朋 uses calculus and the mass-energy equation as analogies: who invented the tool is not the only point; what matters is seeing the core contradiction in autonomous driving first and applying general-purpose technology where it can change the product.
Tesla first used pure vision to escape heavy maps and lidar, then BEV to escape late fusion, and finally end-to-end to escape hand-designed decision rules. Each step was not a repair within the existing dimension, but “lifting itself into a higher dimension” before competing with the old solution through dimensionality reduction.
As of this interview in December 2024, 郎咸朋’s explicitly competitive judgment was: “There are only 2 real end-to-end products in the world today—Tesla and Li Auto.” He also attributed Li Auto’s 2024 sales inflection to “the right time, the right place, and the right people, along with luck and coincidence,” without claiming that its lead was secure.
26. Mind GPT Is Intended to Be the Common Foundation for Autonomous Driving, the Cockpit, and Industrial Intelligence
Faced with existing models such as GPT, Qwen, and ERNIE Bot, 郎咸朋 still believes Li Auto must build its own foundation model. The first reason is control: an external model may be discontinued and will optimize for all customers rather than being tailored to Li Auto’s long-term needs.
The deeper concern is value capture. He believes AI will not ultimately “let a hundred flowers bloom,” but will concentrate among a small number of companies that truly own foundation-model capabilities. If upstream models are sufficiently powerful, other companies may simply work for them and fail to reach the top tier of their industries.
Li Auto’s envisioned Mind GPT must be multimodal and cross-domain, absorbing video, sound, text, and images alongside data from autonomous driving, the cockpit, factories, and the internet. Today, each field trains separate models and knowledge cannot move between them; a unified foundation could first understand “this is a manhole cover, there is a hole underneath, and this is a dangerous area,” then transfer that understanding into driving behavior.
Autonomous driving, Li Auto Tongxue, intelligent industry, and intelligent commerce should all grow from this foundation rather than training isolated models. Preliminary VLA research has begun in autonomous driving; 郎咸朋 expects results “within 1 to 3 years,” but his answer on whether VLA is the final destination is: “I don’t know.”
27. 郎咸朋 Chose Li Auto to Turn Research into a Product on Millions of Vehicles
Looking back on 5 years at Baidu, what excited 郎咸朋 most was not the title but seeing Street View go live: he could ask his family to open Baidu Maps and see images processed by his code. When he left Baidu, he considered only automakers because he wanted the technology in millions of households rather than permanently confined to L4 research prototypes.
He believes the goal has at least “passed”: at the time of the interview, at least 1M vehicles on the market were using autonomous-driving capabilities developed by his team. Whether the team can move from passing to good or excellent depends on whether L3 and L4 are actually delivered.
Asked why he could stay so long at a team that often looked like a “poor student,” 郎咸朋 says he trusted the company’s strategic pacing and completed a transition at Li Auto from technical leader to manager: first set the strategy, then define the business playbook, and finally build the organization’s execution tactics.
On the disputes of autonomous driving’s first 10 years, he uses the Wright brothers and AC versus DC as analogies: before a technology is validated, disagreement and conflict are inevitable; in the end, only the route that genuinely changes people’s lives remains. 李想 began studying AI more seriously in 2024, driven both by industry advances from OpenAI and others and by seeing the speed of model improvement firsthand after riding in an end-to-end vehicle.