CoRL 2026 Half-Day Workshop

Grounded 4D Multimodal World Models for Autonomous Driving Decision Making

A problem-driven workshop on 4D occupancy, multi-view video, LiDAR, and scenario generation that are useful for planning, control, safety validation, and real autonomous systems.

November 12, 2026 JW Marriott Austin In person

Submissions open August 15, 2026 at 12:00 UTC · Deadline October 12, 2026 at 11:59 UTC

Call for contributions

Contributed work will anchor posters, spotlights, and open discussions.

We invite anonymous contributions that connect world modeling to autonomous driving decisions: short papers, extended abstracts, benchmark or dataset reports, resource or challenge proposals, position papers, negative results, and early-stage ideas. Contributions may present new results or sharpen the questions, evidence, and shared resources the community needs next.

This is a lightweight, double-blind, non-archival workshop. All accepted contributions will be presented as posters; selected submissions will be invited for short spotlight talks designed to seed panel questions, breakout discussion, and shared community takeaways.

Submission deadline October 12, 2026, 11:59 UTC Equivalent to October 11, 23:59 Anywhere on Earth.
Open the submission venue
Format and length

Anonymous 2-4 page contributions

The limit applies to the main text; references and an optional appendix are excluded. Please use the current CoRL 2026 LaTeX template available through the official CoRL 2026 site.

Review process

Double-blind and private

Submissions remain private during review and are assigned manually. Anonymous supporting code, data, video, or project pages are welcome when neither the URL nor its contents reveal author identities.

Presentation and release

Posters, selected spotlights

Every accepted contribution is presented as a poster. Selected papers receive a short spotlight. Authors choose whether the accepted PDF is public on OpenReview or remains private.

Non-archival

Accepted work will not appear in CoRL or PMLR proceedings.

In-person participation

At least one author should attend and present in Austin, or promptly contact the organizers if circumstances change.

Author-controlled PDF release

Authors explicitly choose public OpenReview release or a private PDF with title, authors, and an author-provided public link listed by the workshop.

Create profiles early

New OpenReview profiles without an institutional email may require moderation for up to two weeks; institutional-email profiles are activated automatically.

Topics of interest

  • Occupancy-centric 4D, object-centric, BEV, video, LiDAR, or hybrid driving world models.
  • Multimodal and sensor-realistic simulation for robot learning and autonomous driving.
  • Decision-aware generative modeling for planning, policy training, and closed-loop evaluation.
  • Long-tail, adversarial, counterfactual, or safety-critical scenario generation.
  • Benchmarks and metrics for sim-to-real validity, planning utility, and failure transfer.
  • Dataset curation, annotation, and scaling laws for grounded 4D multimodal driving world models.
  • Resource and challenge tracks based on open code, datasets, pretrained weights, or reproducible evaluation suites.

Motivation

From realistic generation to grounded 4D multimodal simulation.

Autonomous driving is a real-world robot learning domain where perception, prediction, decision making, control, safety validation, and human interaction meet. 4D multimodal driving world models now generate semantic occupancy, multi-view video, LiDAR, and interactive scenarios, but the field still lacks shared criteria for when a generated world is useful for action.

This workshop asks participants to define and stress-test grounded 4D multimodal world models: models whose scene structure, dynamics, sensor outputs, and evaluations are tied to real autonomous driving decisions.

01

Representations for action

Occupancy, BEV, object states, video, LiDAR, trajectories, and hybrid state abstractions.

02

Grounded 4D multimodal worlds

Geometry, semantics, temporal coherence, sensor rigs, and physically plausible interactions.

03

Decision-aware evaluation

Planning utility, no at-fault collisions, drivable area compliance, robustness, and failure transfer.

Technical anchors

Two concrete entry points for a broader workshop conversation.

The workshop is not limited to these systems, but they provide shared examples for participants: occupancy-centric 4D multimodal world modeling and photorealistic adversarial scenario generation.

UniSceneV2 occupancy-centric pipeline for dynamic 4D occupancy, multi-view video, and LiDAR generation
UniSceneV2 connects 4D semantic occupancy to multi-view video, LiDAR, downstream perception, and planning evaluation.
Challenger adversarial driving video and trajectory stress-testing visualization
Challenger turns physically plausible adversarial maneuvers into photorealistic driving videos for safety testing.

Core challenges

Focused questions for the half-day workshop.

Decision-grounded 4D representations

What should a 4D multimodal driving world model predict or generate to support planning and control?

Physical and sensor grounding

How do generated scenes remain consistent across cameras, LiDAR rigs, maps, and ego motion?

Evaluation beyond visual realism

Which metrics should complement FID and FVD when the goal is robust decision making?

Long-tail but solvable interactions

How can models generate rare and safety-critical scenarios that are challenging but not impossible?

From offline generation to robot learning

How should generated data support policy training, audits, monitoring, and sim-to-real validation?

Format

A workshop that actually workshops.

The half-day session centers on collaborative outputs: evaluation checklists, scenario taxonomies, and an open-problems whitepaper.

Opening and problem framing

Define decision-centric 4D multimodal world models and the working goals for the day.

Two focused talks

Structured world modeling and decision-aware evaluation, each followed by guided questions.

Evaluation clinic

Breakout groups design protocols for synthetic-data training, planning evaluation, and stress testing.

Paper spotlights and posters

Short spotlights emphasize open problems, negative results, benchmark proposals, and new ideas.

Scenario-design breakout

Participants specify challenging but solvable driving interactions and the metrics needed to evaluate them.

Fishbowl discussion

Converge on the post-workshop artifact: open problems in grounded 4D multimodal driving world models.

Starter resources

Route each modality and topic to a runnable codebase, dataset, or checkpoint source.

These resources are intended as optional starting points for workshop contributors. Participants are encouraged to build on them, compare against them, or challenge their assumptions.

Nuplan-Occ dataset comparison graphic
Dataset and checkpoints

Nuplan-Occ and UniSceneV2 pretrained models

Large-scale semantic occupancy data with released checkpoints for 4D occupancy, video, and LiDAR generation.

UniSceneV2 occupancy generation module
World model representation

4D occupancy generation

Use this path for spatial expansion, temporal forecasting, and occupancy-centric 4D scene representations.

UniSceneV2 multi-view video generation module
Sensor-realistic simulation

Multi-view video and LiDAR generation

Use occupancy-derived conditions for temporally coherent surround-view video and sensor-specific LiDAR outputs.

Challenger adversarial driving scene visualization
Safety-critical scenarios

Challenger and Adv-nuSc

Generate physically plausible adversarial maneuvers and evaluate vision-based end-to-end driving models under stress.

Challenger downstream evaluation overview
Downstream evaluation

Perception and planning baselines

Connect generated data to semantic occupancy prediction and autonomous driving evaluation pipelines.

Workshop submissions

Suggested starting points

Submit methods, datasets, benchmarks, negative results, position papers, or resource reports that tie 4D multimodal world models to decisions.

Case study media

Grounded 4D multimodal generation spans structure, sensors, and safety-critical interaction.

Nuplan-Occ dataset comparison showing scale and sensor modalities

Large-scale occupancy data

Nuplan-Occ supports 4D scene generation, semantic occupancy prediction, and multimodal simulation.

Cross-dataset sensor-rig example
NuPlan-Occ example

4D multimodal simulation across datasets

UniSceneV2 treats occupancy as the shared world representation, then renders sensor-level outputs for different datasets and sensor rigs.

  • 4D occupancy generation Spatial expansion and temporal forecasting of semantic driving scenes.
  • Multi-view video generation Surround-view visual synthesis conditioned on generated scene structure.
  • Sensor-specific LiDAR generation LiDAR point clouds aligned with occupancy and camera-space observations.

Adversarial interaction

Scenario generation can expose rare failure modes in vision-based end-to-end driving systems.

Important dates

Submission and workshop timeline.

All submission-system times below are stated in UTC.

Submission opens August 15, 2026 12:00 UTC
Submission deadline October 12, 2026 11:59 UTC · October 11, 23:59 AoE
Author notification October 26, 2026
Final version and poster confirmation November 2, 2026
Spotlight material due November 6, 2026
Workshop date November 12, 2026 JW Marriott Austin · In person

Submission FAQ

What authors should know before submitting.

Is the workshop archival?

No. Accepted contributions will not appear in the CoRL or PMLR proceedings. The workshop is designed for timely exchange, constructive feedback, and community discussion.

Are preliminary work, negative results, position papers, resources, and challenge proposals welcome?

Yes. We explicitly welcome early-stage ideas, negative results, benchmark or dataset reports, resource or challenge proposals, and position papers, provided they make a clear contribution to the workshop conversation.

Can the work be submitted to another venue?

Because the workshop is non-archival, concurrent work may be considered, but authors are responsible for complying with every other venue's policies. A submission must not already be accepted to or published at the CoRL 2026 main conference.

Will an accepted PDF automatically become public?

No. The submission form asks authors to choose between public release of the accepted PDF on OpenReview and keeping the PDF private while allowing the workshop to list the title, authors, and an author-provided public link.

Is in-person attendance required?

If accepted, at least one author should attend the workshop in Austin and present the contribution. If circumstances change, authors should contact the organizers promptly; the workshop does not currently promise remote presentation.

How should anonymous code, data, video, or project pages be shared?

Use the optional anonymous-resource field in the submission form. The URL, repository, page content, account names, commit history, and file metadata must not reveal author identities during double-blind review.

When should authors create an OpenReview profile?

Create or update the profile well before the deadline. New profiles without an institutional email may require moderation for up to two weeks, while profiles created with an institutional email are activated automatically.

People

Speakers and organizers.

Confirmed speakers and the current organizing team spanning autonomous driving, robot learning, world modeling, simulation, and safety evaluation.

Confirmed speakers

Invited perspectives on world models, autonomous driving, simulation, and safety-critical evaluation.

Organizers

Workshop organization, submissions, resources, and post-workshop artifact coordination.

Post-workshop artifact

Open Problems in Grounded 4D Multimodal World Models for Autonomous Driving Decision Making

The workshop will produce a community whitepaper summarizing evaluation protocols, scenario taxonomies, resource gaps, and research directions proposed by attendees.

Suggested citations

Please cite the source projects when using their code, data, checkpoints, or generated scenarios.

Workshop participants building on the starter resources are encouraged to cite UniScene, UniSceneV2 / Nuplan-Occ, OccScene, OmniNWM, and Challenger as applicable.

UniScene

Unified Occupancy-centric Driving Scene Generation

Recommended when using the original UniScene framework or occupancy-centric generation baseline.

@inproceedings{li2025uniscene,
  title={UniScene: Unified Occupancy-centric Driving Scene Generation},
  author={Li, Bohan and Guo, Jiazhe and Liu, Hongsi and Zou, Yingshuang and Ding, Yikang and Chen, Xiwu and Zhu, Hu and Tan, Feiyang and Zhang, Chi and Wang, Tiancai and others},
  booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},
  pages={11971--11981},
  year={2025}
}
UniSceneV2 / Nuplan-Occ

Scaling Up Occupancy-centric Driving Scene Generation

Recommended when using UniSceneV2 code, Nuplan-Occ data, checkpoints, or generated multimodal assets.

@article{li2026scaling,
  title={Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method},
  author={Li, Bohan and Jin, Xin and Zhu, Hu and Liu, Hongsi and Li, Ruikai and Guo, Jiazhe and Cai, Kaiwen and Ma, Chao and Jin, Yueming and Zhao, Hao and others},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
  year={2026}
}
OccScene

Semantic Occupancy-based Cross-task Mutual Learning

Recommended for semantic occupancy-guided 3D scene generation and cross-task mutual learning.

@article{li2025occscene,
  title={OccScene: Semantic occupancy-based cross-task mutual learning for 3D scene generation},
  author={Li, Bohan and Jin, Xin and Wang, Jianan and Shi, Yukai and Sun, Yasheng and Wang, Xiaofeng and Ma, Zhuang and Xie, Baao and Ma, Chao and Yang, Xiaokang and others},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
  year={2025},
  publisher={IEEE}
}
OmniNWM

Omniscient Driving Navigation World Models

Recommended for navigation-conditioned driving world models and decision-making research.

@article{li2025omninwm,
  title={OmniNWM: Omniscient Driving Navigation World Models},
  author={Li, Bohan and Ma, Zhuang and Du, Dalong and Peng, Baorui and Liang, Zhujin and Liu, Zhenqiang and Ma, Chao and Jin, Yueming and Zhao, Hao and Zeng, Wenjun and others},
  journal={arXiv preprint arXiv:2510.18313},
  year={2025}
}
Challenger

Affordable Adversarial Driving Video Generation

Recommended when using Challenger code, Adv-nuSc, or adversarial driving video generation resources.

@article{xu2025challenger,
  title={Challenger: Affordable Adversarial Driving Video Generation},
  author={Xu, Zhiyuan and Li, Bohan and Gao, Huan-ang and Gao, Mingju and Chen, Yong and Liu, Ming and Yan, Chenxu and Zhao, Hang and Feng, Shuo and Zhao, Hao},
  journal={arXiv preprint arXiv:2505.15880},
  year={2025}
}