Working Notes on Agent Systems/Brad Zhang

@teach_fireworks / X longform

FSD enters China: Autonomous driving begins to fight the "long-tail war". After China is included in Tesla's

FSD enters China: Autonomous driving begins to fight the "long-tail war". After China is included in Tesla's FSD Supervised country list, the focus of autonomous driving...

May 21, 2026 · 5 min read

FSD enters China: Autonomous driving begins to fight the "long-tail war". After China is included in Tesla's
Figure 1 / source image

FSD enters China: Autonomous driving begins to fight the "long-tail war". After China is included in Tesla's FSD Supervised country list, the focus of autonomous driving will shift from "whether a certain function is easy to use" to a harder level: who can continue to eat up the long-tail scenario in the real world. In the past few years, domestic smart driving has been growing very fast. Urban NOA, mapless solution, end-to-end, VLA, lidar, memory commuting, automatic parking, are piled on the car layer by layer. What ordinary users look at is whether they can change lanes on their own, whether they can bypass electric vehicles, and whether they look like experienced drivers at intersections; what the engineering side looks at is even colder: how the system sees the world, how to handle exceptions, how to reduce the takeover rate, how to update, and how to spread costs on a large scale. Tesla FSD’s route is very clear: camera first, model first, and unified software stack. It treats cars as mobile data nodes. The video stream enters the neural network, and the model learns the real driving trajectory, and then outputs steering, acceleration, deceleration, and braking. Every takeover, every hesitation, and every strange intersection is an opportunity to enter a closed loop of training. The temperament of this route is very Silicon Valley: the hardware is as convergent as possible, the complexity is pushed into the model, and the same stack is rolled out to be upgraded in more models and more regions through OTA. Its strongest point is called large-scale learning, and its most difficult point is also here. The purely visual route has extremely high requirements on model understanding, data coverage, video quality and safety boundaries. Seeing accurately is only the first step. The real difficulty is to continue to make stable movements in scenes such as occlusion, backlighting, rainy nights, mixed traffic, construction, temporary stops, and traffic jams. Huawei ADS has a completely different temperament. It is more like a city-level intelligent driving engineering system that uses multi-source information such as cameras, millimeter-wave radar, ultrasonics, and lidar for redundancy, and then superimposes the high-density experience of China's urban roads. Its advantages lie in complex urban areas and engineering reliability: ghost probes, special-shaped obstacles, narrow road intersections, mixed flow of non-motor vehicles, and construction sections are naturally suitable for multi-sensor fusion and local adjustment. The key words of Huawei's route are stability, thickness, and controllability, so it is particularly suitable for deeply binding to car companies and becoming an intelligent base for car companies. Xpeng VLA represents another ambition: pushing the car from a “rule system” to a “model driver”. The imagination of VLA is that the system not only recognizes discrete objects such as lane lines, pedestrians, and traffic lights, but also understands scene semantics and driving intentions. VLA 2.0 emphasizes vision-to-action, allowing visual input to generate driving decisions more directly and reducing intermediate translation layers. This direction is very similar to putting autonomous driving into the large framework of Physical AI. The car is just the first landing form, and can later spill over to robotaxi, robots, and flying cars. Ideal MindVLA focuses more on 3D spatial understanding. The really difficult part of driving on the road often escalates from "seeing objects" to "understanding spatial relationships." An electric vehicle approaches from the right rear, a pedestrian stands at the edge of the blind corner, and a large car blocks the intersection from crossing traffic. All these require joint judgments of spatial structure, movement trends, and driving intentions. MindVLA puts 3D ViT, verbal reasoning and behavior generation together, indicating that Ideal wants to make the "spatial brain" thicker. BYD's God's Eye's approach is more commercial and more aggressive: delegating high-end assisted driving from high-priced cars to large-market models. Its lethality lies not only in its single-bike capabilities, but also in its cost curve and loading scale. Once autonomous driving changes from a few flagship configurations to large-scale standard configurations, data, user habits, car company procurement, and supply chain prices will all change together. The key to BYD's line is not to show off its skills, but to turn smart driving into basic configurations like seat belts and airbags. So what’s really worth looking at in this picture is that several routes are on the table at the same time: Tesla is betting on a unified software stack and a real-world data flywheel. Huawei is betting on sensor redundancy and local engineering depth. Xiaopeng bets on VLA and pushes driving decision-making to large-scale models. Ideal bets on 3D spatial intelligence, allowing the car to first understand the physical world. BYD is betting on popularization of scale and reducing the cost of smart driving. In the short term, Chinese manufacturers have advantages in China's roads, hardware redundancy, vehicle model coverage and functional implementation. China's road conditions are too complicated, with food delivery trucks, pedestrians, construction, tidal lanes, temporary stops, and traffic jams all posing new questions to the model every day. In the long run, the real winner will come back to one question: Who can continuously compress the real-world long tail into trainable, verifiable, and OTA-ready model capabilities. After FSD enters China, the industry will become better. It will force domestic manufacturers to continue to improve their end-to-end capabilities, and it will also force Tesla to accept the real torture of China's road conditions. There is a high probability that the outcome of autonomous driving will not be determined by a certain press conference, but by every takeover moment in tens of millions of real-life driving. One final word: The FSD Supervised, urban NOA, Eye of the God, ADS, and VLA discussed today are still essentially assisted driving that requires driver supervision. No matter how powerful the technology is, the responsibility at the steering wheel still lies with the people. Supervised requires active supervision and mentions the boundaries of capabilities such as the 360-degree view of the on-board camera and continuous software updates; the owner's manual also emphasizes that the driver must remain focused and ready to take over at any time. XPeng’s official description of VLA 2.0 is a vision-to-action architecture that reduces the intermediate translation layer while still being a supervised intelligent driving system. Ideal MindVLA’s public information emphasizes 3D spatial understanding, VLM, behavior generation and vehicle-side real-time trajectory optimization. Reuters' review of China's smart driving market mentioned BYD's multi-speed solution, low-priced model coverage, and the differences in vision/lidar versions of Huawei models.

Visual summary

Article argument map

Generated from the post's content graph

FORMATTOPICCAPABILITYMARKETcoverscoverscoverscoverssignalssignalssignalssignalsFORMATarticle featureTOPICai workbenchesTOPICtechnical distributionTOPICmemoryTOPICretrievalCAPABILITYAI-native workbenchCAPABILITYevaluationCAPABILITYtechnical distributionMARKETAI startupMARKETopen-source builders
Mermaid outline
flowchart LR
  format-article["article feature"]
  topic-ai-workbenches["ai workbenches"]
  topic-technical-distribution["technical distribution"]
  topic-memory["memory"]
  topic-retrieval["retrieval"]
  capability-ai-native-workbench["AI-native workbench"]
  capability-evaluation["evaluation"]
  capability-technical-distribution["technical distribution"]
  market-ai-startup["AI startup"]
  market-open-source-builders["open-source builders"]
  format-article -->|covers| topic-ai-workbenches
  format-article -->|covers| topic-technical-distribution
  format-article -->|covers| topic-memory
  format-article -->|covers| topic-retrieval
  format-article -->|signals| capability-ai-native-workbench
  format-article -->|signals| capability-evaluation
  format-article -->|signals| capability-technical-distribution
  format-article -->|signals| market-ai-startup

Visual structure

Essay structure map

Built from summary and key paragraph positions

FSD enters China: Autonomous driving begins to fight the "long-tail war". After China...THESISFSD enters China:Autonomous drivingbegins to fight the"long-tail war". AfterSIGNALFSD enters China:Autonomous drivingbegins to fight the"long-tail war". AfterOPERATORFSD enters China:Autonomous drivingbegins to fight the"long-tail war". AfterIMPLICATIONFSD enters China:Autonomous drivingbegins to fight the"long-tail war". After
Mermaid outline
flowchart LR
  thesis["FSD enters China: Autonomous driving begins to fight the \"long-tail war\". After China is included in Tesla'..."]
  signal["FSD enters China: Autonomous driving begins to fight the \"long-tail war\". After China is included in Tesla'..."]
  operator["FSD enters China: Autonomous driving begins to fight the \"long-tail war\". After China is included in Tesla'..."]
  implication["FSD enters China: Autonomous driving begins to fight the \"long-tail war\". After China is included in Tesla'..."]
  thesis -->|frames| signal
  signal -->|develops| operator
  operator -->|lands in| implication

Source: View the original post