Pony.ai at WAIC 2026: How PonyWorld 2.0 Helps AI Become a Smarter Driver
As foundation models reshape industries, autonomous driving is undergoing a fundamental shift of its own. In April, Pony.ai unveiled PonyWorld 2.0—not simply an upgraded simulation environment, but the latest evolution of a complete reinforcement learning system spanning cloud and vehicle that the company has built since 2020.
Built as a virtual training environment, it allows AI to practice repeatedly, with a particular focus on one of the hardest skills to master on real roads: navigating complex interactions with other road users. At the WAIC 2026 Autonomous Driving Innovation and Development Forum in Shanghai on July 19, Bo Xiao, Vice President of Engineering and Head of AI R&D at Pony.ai, explained the technology behind PonyWorld 2.0—from data generation and training efficiency to the development of increasingly capable, self-improving AI agents.
Building Better Data, Not Just More Data
Real-world data generated by Robotaxi operations remains one of the most valuable resources for autonomous driving development. Yet different training needs require different approaches to producing and processing that data.
Pony.ai’s data pipeline spans three complementary layers.
The first is extracting greater value from existing fleet data. Through automated annotation and semantic tagging, raw driving logs can be structured into usable training inputs and indexed for efficient retrieval of specific objects, behaviors and scenarios.
The second is scenario reconstruction and editing. Once a real-world interaction has been reconstructed in virtual environment, elements such as the positions, actions and trajectories of different road users can be modified to create new variations of the same event. A single interaction at an intersection, for example, can be transformed into hundreds or even thousands of training scenarios.
The third is generative data. Some of the most important situations for an autonomous driving system to understand, including extreme weather and encounters with rare road users, occur too infrequently to be captured at scale through road testing alone. Drawing on broader internet data, generative models can create new combinations of objects and scenarios around specific training needs, allowing the system to gain experience from situations it has not directly encountered.
Together, these approaches expand both the diversity of training data and the value that can be extracted from every mile driven.
Rethinking the Efficiency of AI Training


Long-tail scenarios remain one of the defining challenges of L4 autonomous driving. Conventional reinforcement learning relies on extensive sampling and repeated trial and error. Given the near-infinite complexity of real-world traffic, this process can quickly become computationally expensive and inefficient.To address this challenge, Pony.ai has developed a Differentiable World Model to complement conventional reinforcement learning.
Traditional reinforcement learning relies heavily on trial and error. The system takes an action, evaluates the outcome and gradually learns how to make better decisions. Because the world model is differentiable, it can provide more direct guidance on how each decision should be adjusted, reducing unproductive trial and error and improving training efficiency by an order of magnitude.
The two approaches have different strengths and complement each other. Conventional reinforcement learning remains valuable for complex situations that are difficult to model explicitly, while the Differentiable World Model can significantly accelerate training across more structured and frequently encountered scenarios. Used together, they provide a more efficient way to address the breadth and complexity of real-world driving.
From a Trained Model to a Self-Improving Agent
The advances in data generation and training efficiency are important, but the more fundamental step forward in PonyWorld 2.0 is the shift from an AI model that is trained to an AI agent that can increasingly direct its own learning. The goal is not only a model that drives better, but an intelligent system that can understand how the physical world works and drive its own improvement.
Traditionally, engineers have had to identify weaknesses in a model, determine what additional data should be collected and design the next round of training. By combining world models with increasingly capable AI agents, Pony.ai is developing a new learning paradigm in which the system can diagnose gaps in its own capabilities, generate targeted training scenarios and even help guide engineering teams on what to develop or collect next.
As autonomous driving enters its next phase, mileage alone will no longer define technological leadership. The greater differentiator will be whether AI can truly understand the physical world and continuously advance beyond its previous capabilities.
PonyWorld 2.0 is designed to create a reinforcing cycle: stronger AI accelerates learning, while faster learning produces stronger AI. As data generation, model training and R&D iteration become increasingly AI-driven, the pace of improvement can move beyond what is possible under a traditional, primarily human-led development process. This is the direction Pony.ai is pursuing—and the foundation for a new phase of autonomous driving.