English
Related papers

Related papers: WorldAct: Activating Monolithic 3D Worlds into Int…

200 papers

Unbounded 3D world generation is emerging as a foundational task for scene modeling in computer vision, graphics, and robotics. In this work, we present WorldFlow3D, a novel method capable of generating unbounded 3D worlds. Building upon a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Amogh Joshi , Julian Ost , Felix Heide

We propose a spatial-constraint approach for modeling spatial-based interactions and enabling interactive visualizations, which involves the manipulation of visualizations through selection, filtering, navigation, arrangement, and…

Human-Computer Interaction · Computer Science 2024-03-21 Can Liu , Yu Zhang , Cong Wu , Chen Li , Xiaoru Yuan

In this work, we propose a framework that creates a lively virtual dynamic scene with contextual motions of multiple humans. Generating multi-human contextual motion requires holistic reasoning over dynamic relationships among human-human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Donggeun Lim , Jinseok Bae , Inwoo Hwang , Seungmin Lee , Hwanhee Lee , Young Min Kim

3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Jaeyoung Chung , Suyoung Lee , Jianfeng Xiang , Jiaolong Yang , Kyoung Mu Lee

Multimodal large language models are evolving toward multimodal agents capable of proactively executing tasks. Most agent research focuses on GUI or embodied scenarios, which correspond to agents interacting with 2D virtual worlds or 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Longrong Yang , Zhixiong Zeng , Yufeng Zhong , Jing Huang , Liming Zheng , Lei Chen , Haibo Qiu , Zequn Qin , Lin Ma , Xi Li

LangDriveCTRL is a natural-language-controllable framework for editing real-world driving videos to synthesize diverse traffic scenarios. It represents each video as an explicit 3D scene graph, decomposing the scene into a static background…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Yun He , Francesco Pittaluga , Ziyu Jiang , Matthias Zwicker , Manmohan Chandraker , Zaid Tasneem

Generative world models are reshaping embodied AI, enabling agents to synthesize realistic 4D driving environments that look convincing but often fail physically or behaviorally. Despite rapid progress, the field still lacks a unified way…

This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works that lack support for unified shared world construction with multi-agent interaction. ShareVerse…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Jiayi Zhu , Jianing Zhang , Yiying Yang , Wei Cheng , Xiaoyun Yuan

3D simulated environments play a critical role in Embodied AI, but their creation requires expertise and extensive manual effort, restricting their diversity and scope. To mitigate this limitation, we present Holodeck, a system that…

Generating realistic 3D indoor scenes from user inputs remains a challenging problem in computer vision and graphics, requiring careful balance of geometric consistency, spatial relationships, and visual realism. While neural generation…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Mengqi Zhou , Xipeng Wang , Yuxi Wang , Zhaoxiang Zhang

A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action Models (WAMs) largely miss this reciprocity, treating world prediction and action…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Hongbo Lu , Liang Yao , Chenghao He , Haoyu Wang , Xiang Gu , Xianfei Li , Wenlong Liao , Tao He , Pai Peng

Recent advances in generative world models have enabled remarkable progress in creating open-ended game environments, evolving from static scene synthesis toward dynamic, interactive simulation. However, current approaches remain limited by…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Junshu Tang , Jiacheng Liu , Jiaqi Li , Longhuang Wu , Haoyu Yang , Penghao Zhao , Siruis Gong , Xiang Yuan , Shuai Shao , Linfeng Zhang , Qinglin Lu

Despite recent advancements in text-to-image generation, most existing methods struggle to create images with multiple objects and complex spatial relationships in the 3D world. To tackle this limitation, we introduce a generic AI system,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yanbo Ding , Shaobin Zhuang , Kunchang Li , Zhengrong Yue , Yu Qiao , Yali Wang

World models are progressively being employed across diverse fields, extending from basic environment simulation to complex scenario construction. However, existing models are mainly trained on domain-specific states and actions, and…

Artificial Intelligence · Computer Science 2024-10-01 Zhiqi Ge , Hongzhe Huang , Mingze Zhou , Juncheng Li , Guoming Wang , Siliang Tang , Yueting Zhuang

Generating articulated objects, such as laptops and microwaves, is a crucial yet challenging task with extensive applications in Embodied AI and AR/VR. Current image-to-3D methods primarily focus on surface geometry and texture, neglecting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Ruijie Lu , Yu Liu , Jiaxiang Tang , Junfeng Ni , Yuxiang Wang , Diwen Wan , Gang Zeng , Yixin Chen , Siyuan Huang

We tackle the problem of generating long-term 3D human motion from multiple action labels. Two main previous approaches, such as action- and motion-conditioned methods, have limitations to solve this problem. The action-conditioned methods…

Computer Vision and Pattern Recognition · Computer Science 2023-02-20 Taeryung Lee , Gyeongsik Moon , Kyoung Mu Lee

Embodied AI depends on interactive 3D environments that support meaningful activities for diverse users, yet assessing their functional affordances remains a core challenge. We introduce SceneTeract, a framework that verifies 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Léopold Maillard , Francis Engelmann , Tom Durand , Boxiao Pan , Yang You , Or Litany , Leonidas Guibas , Maks Ovsjanikov

Video--based world models have emerged along two dominant paradigms: video generation and 3D reconstruction. However, existing evaluation benchmarks either focus narrowly on visual fidelity and text--video alignment for generative models,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Meiqi Wu , Zhixin Cai , Fufangchen Zhao , Xiaokun Feng , Rujing Dang , Bingze Song , Ruitian Tian , Jiashu Zhu , Jiachen Lei , Hao Dou , Jing Tang , Lei Sun , Jiahong Wu , Xiangxiang Chu , Zeming Liu , Kaiqi Huang

Procedurally generating cohesive and interesting game environments is challenging and time-consuming. In order for the relationships between the game elements to be natural, common-sense has to be encoded into arrangement of the elements.…

Learning latent actions from action-free video has emerged as a powerful paradigm for scaling up controllable world model learning. Latent actions provide a natural interface for users to iteratively generate and manipulate videos. However,…

Machine Learning · Computer Science 2026-05-26 Zizhao Wang , Chang Shi , Jiaheng Hu , Kevin Rohling , Roberto Martín-Martín , Amy Zhang , Peter Stone