Skip to main content
Artificial Intelligence

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

| Source: arXiv

Preprint — not peer-reviewed. Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains. We present TerraZero, a procedural driving simulator and self-play training stack. A configurable C engine runs simulation on the CPU and policy inference on the GPU over a zero-copy path, sustaining 1.3M agent-steps per se

Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains. We present TerraZero, a procedural driving simulator and self-play training stack. A configurable C engine runs simulation on the CPU and policy inference on the GPU over a zero-copy path, sustaining 1.3M agent-steps per second on a single server-grade GPU, far faster than existing object-level simulators, while keeping fidelity lighter single-agent systems omit: heterogeneous agents, multiple dynamics models, and full traffic-rule enforcement. TerraZero treats logged data only as a source of real-world map geometry, populating each map with randomized rule-based road users and signal controllers and randomizing agent dynamics, rewards, and sizes per episode, so a map yields an unbounded set of scenarios. Every reported policy trains from scratch by reinforcement learning alone on a compute-efficient self-play recipe across GPUs, with zero human demonstrations and no fallback planner at inference. Policies generalize zero-shot across cities and datasets, including emergent left-hand-traffic driving without explicit supervision. As an ego policy, TerraZero is the first fully learned policy to top the InterPlan long-tail benchmark, ahead of larger learned planners; on routine-driving val14 it ranks among the best approaches and is the safest, posting the best collision and time-to-collision scores. On Waymo Open Sim Agents realism the same recipe outperforms other demonstration-free methods and is competitive with the strongest reference-anchored self-play method. One stack serves both roles: driving policies across dynamics for cars and trucks, and sim agents that jointly control vehicles, pedestrians, and cyclists.

Read the original source →

Related Stories

Artificial Intelligence

Plant fiber exploitation contributed to the emergence of ground-edged cutting tools in North China.

Ground-edged cutting tools are widely regarded as technological hallmarks of early agriculture in North China, commonly hypothesized to have been developed primarily for cereal harvesting. Yet the functional foundations of this innovation remain insufficiently tested. This study integrates use-wear and microfossil analyses of 140 stone tools from the Peiligang site, with experimental data, to reassess their roles within long-term trajectories of technological change from the Late Paleolithic to

Continue reading
Artificial Intelligence

Neural network-augmented Pfaffian wave-functions for scalable simulations of interacting fermions.

Developing accurate numerical methods for strongly interacting fermions is crucial for improving our understanding of various quantum many-body phenomena, especially unconventional superconductivity. Recently, neural quantum states have emerged as a promising approach for studying correlated fermions, highlighted by the hidden fermion and backflow methods, which use neural networks to model corrections to fermionic quasiparticle orbitals. In this work, we expand these ideas to the space of Pfaff

Continue reading
Artificial Intelligence

Why friends in common reveal network stars.

The Friendship Paradox states that, on average, your friends have more friends than you do. We extend this to common friends-those who appear in multiple people's friend lists. We show that the more people who share a common friend, the more connected that person tends to be, and we derive an expression quantifying this progression. In a regional Facebook network, a common friend to three randomly sampled individuals has on average more friends than 99.9% of the network. In a citation network, a

Continue reading
Artificial Intelligence

Sensory context improves language prediction in humans and LLMs.

Language is a fundamental human capacity. Large language models (LLMs) have presented the first viable model of language outside of humans, yet how these models learn and use language differs significantly from humans. Here, we compare LLMs and humans predicting language with varying levels of sensory information-from disembodied written text to audiovisual videos of speakers-to demonstrate that, in both humans and LLMs, sensory context is critical for optimal performance. We asked human partici

Continue reading
Artificial Intelligence

Treatment Decisions in Multiple Myeloma.

Revolutions in transplantation and targeted and immune therapies have transformed multiple myeloma from a disease with an associated survival of a few years into one for which functional cure is an emerging goal. This abundance of effective therapies has created clinical complexity. Here we provide a practical framework, anchored in trial evidence and informed by emerging biologic discoveries, for the navigation of treatment decisions across the disease spectrum. We outline how cytogenetic and g

Continue reading