New Year Livestream 2: Tesla FSD and the Commercial War in Autonomous Driving
Summary
于振华认为V14终于把FSD从“技术可用”推到“产品足够好”,并因此收回了过去“用户不懂”的判断。 On V14.2, he saw broad improvements in right turns, parking-lot entry and exit, wide highway-lane changes, and driver monitoring. A New Year’s Eve account from a vehicle owner who drove from the US West Coast to the East Coast without taking over once was another example of the shift in user experience cited by 泓君.
旧金山停电成为两位嘉宾判断端到端路线优于rule-based路线的关键样本。 于振华 argued that Waymo might stop when faced with an undefined traffic-light failure, while Tesla could prioritize right-of-way rules; end-to-end models can also learn from large volumes of human driving behavior. He acknowledged, however, that rule-based systems can fix known problems quickly, while end-to-end systems may take longer to recollect data and retrain.
Robotaxi竞争已经从“能不能开”转向扩张速度、安全员移除与资产效率。 Waymo spent years expanding to only 5 cities, added 5 more at year-end, and is seeking roughly $15B in financing. 大卫 used a same-trip comparison—about $4 for Tesla versus about $18 for Waymo, which avoided Highway 101—to illustrate the service gap, though 泓君 clarified that he had not actually taken the latter ride. 于振华 said that, to his knowledge, Texas does not require another approval to remove safety drivers; 泓君 later summarized the bottleneck as “mainly technical for now,” leaving the room without a fully unified view. 于振华 sees the fleet ramp-up after Cybercab enters mass production in April 2026, followed by safety-driver removal, as a reasonable window.
于振华的核心判断是:Tesla的领先首先来自算法与“智能密度”,数据和算力优势还没进入同维度比较。 Since V12 in the second half of 2023, he has not seen another player come close to Tesla algorithmically. Even Chinese manufacturers using labels such as end-to-end, VLA, or world models remain, in his long-term testing, increasingly far behind. “The data and compute advantage comes later.”
Tesla的护城河被归结为技术思想、内部人才、敢于押注,以及软硬件co-design,而不只是更多GPU。 The fourth-generation chip was designed before end-to-end systems existed, while the fifth generation incorporated end-to-end and large-language-model requirements. Although Dojo has been canceled, 于振华 still believes the vehicle-side inference chip carries “decisive weight”—you cannot make something out of nothing(巧妇难为无米之炊).
按节目给出的约$100B估值和$15B融资规模,两位嘉宾都表示不会投资Waymo。 泓君 cited figures of $1.43 per kilometer for Waymo versus a possible $0.81 for Tesla, and forecast that the new Waymo vehicle could bring cost per mile down to $0.99–$1.08. 大卫 emphasized charging staff, sensor maintenance, mechanical lidar life, and depreciation, concluding that “$100B is simply too expensive.”
所谓从L2“跳到”L4,在嘉宾看来更像产品体系打通,而非跨越一个可靠的行业刻度。 A single vehicle model, consistent data, vehicle-side compute, and management priority are all prerequisites. If a company maintains 20 models at once, or places battery swapping ahead of autonomous driving, a direct upgrade is difficult to make work. 大卫 also questioned the definition of the “first half” and “second half”: “You’re not the referee, so how do you define the halves?”
Deep dive
1. V14 Turns “Users Don’t Get It” into an Obsolete Explanation
泓君 opened by disclosing the program’s bias: both guests are Tesla supporters, so the discussion clearly leans toward Tesla. A full counterargument from Waymo and other L4 manufacturers will be left to a later episode rather than voiced on their behalf here.
于振华 was using V14.2, not the later V14.2.2.2. He disliked the initial V14’s overly cautious right turns; by around V14.1.4, he had already seen a clear improvement, though he explicitly said he was unsure of the exact minor version.
What changed his mind was the complete experience loop: driver monitoring became less intrusive, parking-lot entry and exit felt natural, and wide highway-lane changes became far less concerning. He used to ask, “If FSD is this good, why doesn’t everyone use it?” His answer now is: “It’s not that users don’t get it; the product wasn’t good enough.”
泓君 also cited a vehicle owner’s New Year’s Eve account: a drive from the US West Coast to the East Coast, including charging and parking, with no takeover at any point and no dangerous incident reported. This was a customer account, not a test conducted by either guest.
2. A $4 vs. $18 Comparison Exposes the Limits of the Service
Before Christmas, 大卫 opened Tesla Robotaxi and Waymo at the same time. Tesla cost a little over $4 from his home to the gym; Waymo could not take that stretch of Highway 101, detoured, and quoted roughly $18. His product standard is simple: “A good autonomous vehicle should be boring.”
泓君 preserved an important caveat: 大卫 did not actually take that $18 Waymo ride. 于振华 also noted that parts of Waymo’s routes already include highways; the issue may have been that a particular entrance on 大卫’s route was not covered that day. A single route difference cannot be generalized into a complete absence of highway capability.
Operational maturity must also be kept separate. Waymo can operate without a driver in the Bay Area, while Tesla at the time still required a safety driver in the front seat. Austin was gradually testing safety-driver removal, but most vehicles still operated with safety drivers.
大卫 believes Tesla’s use of mass-produced Model Y vehicles makes it easier to work through existing vehicle systems with the DMV and SFO airport authorities. Waymo, by contrast, must expand its vehicle and fleet-operations infrastructure. Highway driving is relatively simple technically, but because the consequences of an accident are severe, highways are often opened last.
3. Cybercab Production Becomes the Real Clock for Removing Safety Drivers
于振华 said that, to his knowledge, Texas should not require another regulatory approval, allowing Austin to put vehicles on the road first and regulate afterward; California is the opposite, requiring a license before deployment. 泓君 later summarized the current bottleneck as “mainly technical for now,” leaving tension between the technical and regulatory assessments.
On Musk’s failure to deliver the promised year-end timing, 于振华’s explanation was that the team was still finishing V14. If the technical team had not completed that work, it was unlikely to make major simultaneous progress on safety-driver removal. He expected a delay of some kind, rather than an extension of the deadline to Lunar New Year.
He puts more weight on Musk’s confirmation that Cybercab will enter mass production in April 2026. Cybercab is the real vehicle without a steering wheel. Tesla does not have Alphabet’s financial flexibility, so cars that cannot go on the road and generate fares create inventory risk. 大卫 therefore considers it “a very reasonable timing” to ramp up the fleet over 1–2 months before or after production begins, then remove the safety drivers.
4. The San Francisco Blackout Put Two Technical Paths under Pressure
于振华 had assumed that Waymo, even with a rule-based system, would have prepared multiple backups for traffic-light failures: at an unsignaled US intersection, vehicles could still follow first-come, first-served right-of-way rules. He was surprised that Waymo had not handled this right-of-way logic, rather than believing the vehicles depended on V2X or a vehicle-road-cloud system.
泓君 summarized the guests’ view this way: Tesla’s end-to-end model behaves like a black box that learns from large volumes of human driving, while Waymo relies on rules and high-definition maps and may stop when confronted with an undefined change. 于振华 added that San Francisco is roughly the size of Shenzhen’s Nanshan District or one district of Shanghai’s Puxi, yet Waymo has tested there for years and can still encounter such corner cases—an obvious risk for commercial operations.
大卫 once watched a Waymo get stuck making a left turn at an intersection, blocking traffic through 6 or 7 signal cycles before suddenly moving through quickly. He suspected a remote takeover. 于振华 said that any genuinely commercial operation would be equipped with remote-control capabilities.
The trade-off between the two approaches is clear: Waymo can quickly patch a rule gap once it discovers one, but no one knows what the next unforeseen scenario will be. End-to-end systems may perform more “smoothly and naturally” across many situations, but they cannot guarantee that every specific problem will be fixed immediately.
5. Puddles, Wet Roads, and Multiple Parking Rows Show Why Corner Cases Never End
泓君 relayed Ashok’s “mini trolley problem”: facing a large puddle, the vehicle must weigh water depth, oncoming traffic, and the risk of briefly crossing the centerline. It is difficult to hard-code every combination with rule-based logic, while 于振华 believes both humans and AI can handle such cases through trade-offs.
于振华 cited an accident involving an autonomous taxi in a city in Hunan. A water truck left the road slippery, but the system’s parameters did not include that factor and therefore failed to account sufficiently for the longer braking distance. His inference is that human driving behavior in rain becomes part of end-to-end training data, so the system is less likely to become completely helpless merely because slippery conditions slightly extend braking distance.
于振华 relayed Andrej Karpathy’s recollection: Tesla initially wrote a rule to prevent false braking for a single row of cars parked along the roadside, then encountered double parking. In other countries, it could face triple parking or even more rows. “This is endless.”
He retained the counterargument as well: rule-based systems fix known gaps quickly, while end-to-end systems may need fresh data collection and retraining. That is why the jump from V13 to V14 took roughly a year. Its defining feature is not frequent incremental rule changes, but that “every update is revolutionary.”
6. Explainability Belongs Offline, Not inside the Driving Loop
泓君 believes a vehicle does not need to verbalize a red trash can before acting; humans often register it subconsciously as an obstacle. She is skeptical of online explainability, arguing that sufficiently large datasets can do the job—“brute force solves it.”
She used Li Auto’s first-generation “dual system” as an example: a slow-thinking layer tried to explain its judgment to people like ChatGPT, while a fast-thinking layer handled the quicker driving loop. That approach was later abandoned. 泓君 believes passengers do not need the machine to report what it is thinking in real time; online explanations may simply waste time.
于振华 added the distinction: step-by-step explanation is unnecessary online, but explanation is required offline. Regulators, model evaluation, reviews of anomalous actions, and the next round of data collection all need sufficient compute to answer why the system acted that way at that moment. That was also the key point he retained from Ashok’s presentation.
7. Algorithms Have Not Converged, So Data and Compute Advantages Are Premature
于振华 believes no player has come close to Tesla algorithmically since V12 in the second half of 2023. Li Auto, XPeng, and NIO have all claimed to use end-to-end systems, but based on his long-term observations and testing, they remain very far from genuine end-to-end performance.
His comparison therefore follows a strict sequence: only once everyone has mastered the same known algorithm do Tesla’s data and compute become the next layer of advantage. Autonomous driving is unlike LLMs, where talent moves quickly and methods are publicly discussed. Apart from one technical presentation by Ashok, Tesla has shared little specific technology.
Ashok’s simulator demonstration impressed him most: the driver turns a steering wheel as if playing a game, while new scenes are generated on the other side and the resulting data is collected. That turns part of the data-collection problem for different corner cases into a constructible training problem.
大卫 reduced the persistence of rule-based systems to two reasons: “First, they can’t build the alternative. Second, the baggage is too heavy.” He believes labels such as VLA and world models should not create “technology anxiety.” The real question is the system’s “intelligence density” in solving problems.
8. Tesla’s Talent Soil Matters More Than Star Résumés
于振华’s list includes technical vision, talent, the courage to “put everything on the line,” and compute. He believes Tesla is currently the only company with all of these conditions, while the conviction behind the direction comes primarily from Musk.
Ashok was the first person to join the FSD team. After more senior people gradually left for Nvidia, Waymo, Aurora, and other companies, he stayed with the internal core team to keep pushing forward. When the end-to-end results emerged, the main contributors were not academic stars with long publication lists, but a team that had grown inside the company over many years.
泓君 said Tesla’s acquisition of DeepScale was an acqui-hire, bringing in roughly a dozen people. She considers the acquisition a failure: the founder may have joined for 1–2 years before being fired, while the rest of the team also failed to integrate well. 于振华 added that the problem was not only a workload “more startup than a startup,” but also whether people could adapt to a problem-driven culture that uses fewer new terms and centers on solving problems.
9. Inference Chips Are the Decisive Constraint on Bringing End-to-End Onboard
Tesla has separate tracks for training and vehicle-side inference. According to the program, Dojo has been canceled and its head, Peter Bannon, has left; Nvidia still supplies most of the chips used for end-to-end training. 于振华 also cited Jensen Huang’s praise of Musk: Musk can rapidly build an xAI training cluster, including power generation and liquid cooling, while Tesla needs fewer GPUs than xAI and Musk responds directly when the team asks for “more GPUs.”
于振华 worked on Tesla’s fourth- and fifth-generation inference chips. When the fourth generation was designed, end-to-end systems did not yet exist, so requirements came mainly from rule-based systems. The fifth generation incorporated end-to-end and large-language-model workloads; sixth- and seventh-generation chips will follow.
The challenge in software-hardware co-design is not simply how much compute to pack in, but whether the software team can articulate its needs for the next 3, 5, or even 6 years. Tesla’s Silicon Valley software team and Austin hardware team are separate teams, but closely connected. Hardware can only be designed well when software requirements are clear.
Musk once spent 34 hours in meetings at the hardware department with his son. 于振华 believes this reflected Musk’s thinking about how xAI would compete and how to build cheaper inference, which is why the fifth-generation chip considered not only autonomous driving but also large-language-model applications.
In his view, Dojo’s failure does not invalidate the broader strategy. Training chips are far more complex than inference chips, and Nvidia was always available as a supplier, so Tesla lacked a genuine do-or-die incentive. If Nvidia moves too slowly or diverges from Tesla’s requirements, a training chip program could still return.
10. The Waymo Valuation Debate Comes Down to Full-Lifecycle Costs
The program opened by noting that Waymo spent years expanding to only 5 cities, added 5 more at year-end, and is seeking roughly $15B in financing at a valuation above $100B. 泓君 then cited Morgan Stanley figures putting Waymo’s current cost at $1.43 per kilometer versus a possible $0.81 for Tesla. The forecast for Waymo’s new vehicle was $0.99–$1.08 per mile, with the original figures using different units.
大卫’s refusal to invest was not based only on the headline cost of a single ride. Charging, sensor maintenance, fleet personnel, and depreciation all scale with the business. At a Bay Area charging station, he saw 3 employees charging vehicles, while Cybercab uses wireless charging. For mechanical lidar, he thinks roughly 2 years may be acceptable but 3 years may not be, which would also affect lifecycle cost.
于振华 would “bet everything” on Tesla, while still liking Alphabet and Gemini. His distinction is that Alphabet has not gone all in on Waymo either. At the roughly 2,000-vehicle scale cited on the program, Waymo must add a fleet every time it enters a city; if users grow 10x, costs could also rise sharply. This is not pure software-style economies of scale.
11. L2 to L4 Is Not a Level Jump; It Is Whether the Safety Driver Can Be Removed
In response to concerns that FSD is compute-constrained, 于振华 pointed out that its inference process does not recursively rerun token by token like an LLM. The input is large, but it does not carry the same autoregressive latency burden. V14 “has not yet used up the compute” on fourth-generation hardware, while fifth- and sixth-generation chips should address continued demand growth.
大卫 questioned the explanatory power of the SAE levels. Tesla may have some weaknesses that only qualify as L2, while other strengths already look like L4; calling the whole system L2 based on its weakest link can be misleading. Adding “L2.5” in China does not solve the definitional problem.
He agrees with the direction expressed by 何小鹏: connecting assisted driving with an end-to-end Robotaxi solution. But the prerequisites are a single vehicle model, unified data, sufficient compute, and management willing to “bet everything.” Deploying separate AI systems across 20 vehicle models with inconsistent data, or giving battery swapping priority over autonomous driving, makes a direct upgrade difficult.
泓君’s final point was that Tesla is not pursuing labels such as L3 first; it is focused directly on whether the safety driver can be removed.