doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving
Parthib Roy, Yash Tandon, Marcus Blennemann, Giovanni Tapia Lopez, Angel Martinez-Sanchez, Mohan M. Trivedi, Ross Greer
- Published
- Sep 29, 2026 — 17:08 UTC
Problem
The paper identifies a critical gap in existing datasets for long-horizon passenger intent in autonomous driving, particularly in the context of language-conditioned planning. Current datasets primarily focus on short prediction horizons, which limits the ability to model complex driving scenarios that require understanding of longer-term passenger instructions. This work presents doPlan, a new dataset designed to fill this void, enabling more sophisticated planning models that can interpret and act on extended language inputs.
Method
The core contribution of this work is the doPlan dataset, which consists of 5,154 human-annotated instructions derived from the nuPlan dataset. The dataset encompasses 169.1 hours of contextual driving data, with 50.9 hours specifically dedicated to driving scenarios. Annotations in doPlan are provided with varying windows, ranging from 30.0 to 508.8 seconds, allowing for a diverse range of planning horizons. The authors evaluate four language-conditioned driving models on this dataset to assess their performance in interpreting and executing passenger instructions over extended time frames.
Results
The evaluation reveals that the median time to first maneuver in the doPlan dataset is 24.6 seconds, which is significantly longer than the typical 5 seconds associated with existing prediction horizons. Additionally, only 9.8% of maneuvers occur within the 5 seconds prediction horizon, highlighting the inadequacy of current models when faced with longer-term planning tasks. The available text does not report quantitative results for the performance of the evaluated models against specific baselines.
Limitations
The authors note that the sensitivity of passenger language does not consistently translate into reliable driving behavior, indicating a potential challenge in developing models that can generalize across different language inputs. Furthermore, the dataset's reliance on human annotations may introduce biases or inconsistencies that could affect model training and evaluation. The paper does not address potential limitations related to the diversity of driving scenarios or the representativeness of the passenger instructions.
Why it matters
The introduction of the doPlan dataset has significant implications for the field of autonomous driving and language-conditioned planning. By providing a resource that focuses on long-horizon planning, this work encourages the development of more advanced models capable of understanding and executing complex passenger instructions. This could lead to improved safety and user experience in autonomous vehicles, as well as inspire further research into the integration of natural language processing with real-time decision-making in dynamic environments.
By Turing Wire Research Desk · Sep 29, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
