CoRL 2026 Half-Day Workshop

Grounded 4D Multimodal World Models for Autonomous Driving Decision Making

A problem-driven workshop on 4D occupancy, multi-view video, LiDAR, and scenario generation that are useful for planning, control, safety validation, and real autonomous systems.

November 12, 2026 JW Marriott Austin In person

Motivation

From realistic generation to grounded 4D multimodal simulation.

Autonomous driving is a real-world robot learning domain where perception, prediction, decision making, control, safety validation, and human interaction meet. 4D multimodal driving world models now generate semantic occupancy, multi-view video, LiDAR, and interactive scenarios, but the field still lacks shared criteria for when a generated world is useful for action.

This workshop asks participants to define and stress-test grounded 4D multimodal world models: models whose scene structure, dynamics, sensor outputs, and evaluations are tied to real autonomous driving decisions.

01

Representations for action

Occupancy, BEV, object states, video, LiDAR, trajectories, and hybrid state abstractions.

02

Grounded 4D multimodal worlds

Geometry, semantics, temporal coherence, sensor rigs, and physically plausible interactions.

03

Decision-aware evaluation

Planning utility, no at-fault collisions, drivable area compliance, robustness, and failure transfer.

Technical anchors

Two concrete entry points for a broader workshop conversation.

The workshop is not limited to these systems, but they provide shared examples for participants: occupancy-centric 4D multimodal world modeling and photorealistic adversarial scenario generation.

UniSceneV2 occupancy-centric pipeline for dynamic 4D occupancy, multi-view video, and LiDAR generation
UniSceneV2 connects 4D semantic occupancy to multi-view video, LiDAR, downstream perception, and planning evaluation.
Challenger adversarial driving video and trajectory stress-testing visualization
Challenger turns physically plausible adversarial maneuvers into photorealistic driving videos for safety testing.

Core challenges

Focused questions for the half-day workshop.

Decision-grounded 4D representations

What should a 4D multimodal driving world model predict or generate to support planning and control?

Physical and sensor grounding

How do generated scenes remain consistent across cameras, LiDAR rigs, maps, and ego motion?

Evaluation beyond visual realism

Which metrics should complement FID and FVD when the goal is robust decision making?

Long-tail but solvable interactions

How can models generate rare and safety-critical scenarios that are challenging but not impossible?

From offline generation to robot learning

How should generated data support policy training, audits, monitoring, and sim-to-real validation?

Format

A workshop that actually workshops.

The half-day session centers on collaborative outputs: evaluation checklists, scenario taxonomies, and an open-problems whitepaper.

Opening and problem framing

Define decision-centric 4D multimodal world models and the working goals for the day.

Two focused talks

Structured world modeling and decision-aware evaluation, each followed by guided questions.

Evaluation clinic

Breakout groups design protocols for synthetic-data training, planning evaluation, and stress testing.

Paper spotlights and posters

Short spotlights emphasize open problems, negative results, benchmark proposals, and new ideas.

Scenario-design breakout

Participants specify challenging but solvable driving interactions and the metrics needed to evaluate them.

Fishbowl discussion

Converge on the post-workshop artifact: open problems in grounded 4D multimodal driving world models.

Starter resources

Route each modality and topic to a runnable codebase, dataset, or checkpoint source.

These resources are intended as optional starting points for workshop contributors. Participants are encouraged to build on them, compare against them, or challenge their assumptions.

Nuplan-Occ dataset comparison graphic
Dataset and checkpoints

Nuplan-Occ and UniSceneV2 pretrained models

Large-scale semantic occupancy data with released checkpoints for 4D occupancy, video, and LiDAR generation.

UniSceneV2 occupancy generation module
World model representation

4D occupancy generation

Use this path for spatial expansion, temporal forecasting, and occupancy-centric 4D scene representations.

UniSceneV2 multi-view video generation module
Sensor-realistic simulation

Multi-view video and LiDAR generation

Use occupancy-derived conditions for temporally coherent surround-view video and sensor-specific LiDAR outputs.

Challenger adversarial driving scene visualization
Safety-critical scenarios

Challenger and Adv-nuSc

Generate physically plausible adversarial maneuvers and evaluate vision-based end-to-end driving models under stress.

Challenger downstream evaluation overview
Downstream evaluation

Perception and planning baselines

Connect generated data to semantic occupancy prediction and autonomous driving evaluation pipelines.

Workshop submissions

Suggested starting points

Submit methods, datasets, benchmarks, negative results, position papers, or resource reports that tie 4D multimodal world models to decisions.

Case study media

Grounded 4D multimodal generation spans structure, sensors, and safety-critical interaction.

Nuplan-Occ dataset comparison showing scale and sensor modalities

Large-scale occupancy data

Nuplan-Occ supports 4D scene generation, semantic occupancy prediction, and multimodal simulation.

Cross-dataset sensor-rig example
NuPlan-Occ example

4D multimodal simulation across datasets

UniSceneV2 treats occupancy as the shared world representation, then renders sensor-level outputs for different datasets and sensor rigs.

  • 4D occupancy generation Spatial expansion and temporal forecasting of semantic driving scenes.
  • Multi-view video generation Surround-view visual synthesis conditioned on generated scene structure.
  • Sensor-specific LiDAR generation LiDAR point clouds aligned with occupancy and camera-space observations.

Adversarial interaction

Scenario generation can expose rare failure modes in vision-based end-to-end driving systems.

Call for papers

Short papers, benchmarks, datasets, positions, and early ideas.

We plan to solicit 2-4 page short papers with optional appendices. Submissions may include new methods, benchmark proposals, dataset reports, position papers, negative results, or early-stage ideas. We welcome work that makes a concrete connection between world modeling and autonomous driving decisions.

Accepted papers will be presented as posters, with selected short spotlights focused on open questions and community value rather than polished final results.

Topics of interest

  • Occupancy-centric 4D, object-centric, BEV, video, LiDAR, or hybrid driving world models.
  • Multimodal and sensor-realistic simulation for robot learning and autonomous driving.
  • Decision-aware generative modeling for planning, policy training, and closed-loop evaluation.
  • Long-tail, adversarial, counterfactual, or safety-critical scenario generation.
  • Benchmarks and metrics for sim-to-real validity, planning utility, and failure transfer.
  • Dataset curation, annotation, and scaling laws for grounded 4D multimodal driving world models.

Important dates

Timeline.

Workshop paper submission TBD
Acceptance notification TBD
Camera-ready deadline TBD
Workshop date November 12, 2026

People

Speakers and organizers.

Confirmed speakers and the current organizing team spanning autonomous driving, robot learning, world modeling, simulation, and safety evaluation.

Confirmed speakers

Invited perspectives on world models, autonomous driving, simulation, and safety-critical evaluation.

Organizers

Workshop organization, submissions, resources, and post-workshop artifact coordination.

Post-workshop artifact

Open Problems in Grounded 4D Multimodal World Models for Autonomous Driving Decision Making

The workshop will produce a community whitepaper summarizing evaluation protocols, scenario taxonomies, resource gaps, and research directions proposed by attendees.

Suggested citations

Please cite the source projects when using their code, data, checkpoints, or generated scenarios.

Workshop participants building on the starter resources are encouraged to cite UniScene, UniSceneV2 / Nuplan-Occ, OccScene, OmniNWM, and Challenger as applicable.

UniScene

Unified Occupancy-centric Driving Scene Generation

Recommended when using the original UniScene framework or occupancy-centric generation baseline.

@inproceedings{li2025uniscene,
  title={UniScene: Unified Occupancy-centric Driving Scene Generation},
  author={Li, Bohan and Guo, Jiazhe and Liu, Hongsi and Zou, Yingshuang and Ding, Yikang and Chen, Xiwu and Zhu, Hu and Tan, Feiyang and Zhang, Chi and Wang, Tiancai and others},
  booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},
  pages={11971--11981},
  year={2025}
}
UniSceneV2 / Nuplan-Occ

Scaling Up Occupancy-centric Driving Scene Generation

Recommended when using UniSceneV2 code, Nuplan-Occ data, checkpoints, or generated multimodal assets.

@article{li2026scaling,
  title={Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method},
  author={Li, Bohan and Jin, Xin and Zhu, Hu and Liu, Hongsi and Li, Ruikai and Guo, Jiazhe and Cai, Kaiwen and Ma, Chao and Jin, Yueming and Zhao, Hao and others},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
  year={2026}
}
OccScene

Semantic Occupancy-based Cross-task Mutual Learning

Recommended for semantic occupancy-guided 3D scene generation and cross-task mutual learning.

@article{li2025occscene,
  title={OccScene: Semantic occupancy-based cross-task mutual learning for 3D scene generation},
  author={Li, Bohan and Jin, Xin and Wang, Jianan and Shi, Yukai and Sun, Yasheng and Wang, Xiaofeng and Ma, Zhuang and Xie, Baao and Ma, Chao and Yang, Xiaokang and others},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
  year={2025},
  publisher={IEEE}
}
OmniNWM

Omniscient Driving Navigation World Models

Recommended for navigation-conditioned driving world models and decision-making research.

@article{li2025omninwm,
  title={OmniNWM: Omniscient Driving Navigation World Models},
  author={Li, Bohan and Ma, Zhuang and Du, Dalong and Peng, Baorui and Liang, Zhujin and Liu, Zhenqiang and Ma, Chao and Jin, Yueming and Zhao, Hao and Zeng, Wenjun and others},
  journal={arXiv preprint arXiv:2510.18313},
  year={2025}
}
Challenger

Affordable Adversarial Driving Video Generation

Recommended when using Challenger code, Adv-nuSc, or adversarial driving video generation resources.

@article{xu2025challenger,
  title={Challenger: Affordable Adversarial Driving Video Generation},
  author={Xu, Zhiyuan and Li, Bohan and Gao, Huan-ang and Gao, Mingju and Chen, Yong and Liu, Ming and Yan, Chenxu and Zhao, Hang and Feng, Shuo and Zhao, Hao},
  journal={arXiv preprint arXiv:2505.15880},
  year={2025}
}