English
Related papers

Related papers: Inferring Compositional 4D Scenes without Ever See…

200 papers

Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D compositional objects from single images--particularly under occlusions--remains challenging. Existing methods often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Hui Shan , Keyang Luo , Ming Li , Sizhe Zheng , Yanwei Fu , Zhen Chen , Xiangru Huang

Accurate capture of human-object interaction from ubiquitous sensors like RGB cameras is important for applications in human understanding, gaming, and robot learning. However, inferring 4D interactions from a single RGB view is highly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xianghui Xie , Bowen Wen , Yan Chang , Hesam Rabeti , Jiefeng Li , Ye Yuan , Gerard Pons-Moll , Stan Birchfield

Generative models have demonstrated remarkable abilities in generating high-fidelity visual content. In this work, we explore how generative models can further be used not only to synthesize visual content but also to understand the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Yanbo Wang , Justin Dauwels , Yilun Du

Recent advancements in foundation models for 2D vision have substantially improved the analysis of dynamic scenes from monocular videos. However, despite their strong generalization capabilities, these models often lack 3D consistency, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Haoran Zhou , Gim Hee Lee

This paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth prediction, motion tracking, and segmentation models, offer promising…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 David Yifan Yao , Albert J. Zhai , Shenlong Wang

Recovering 4D from monocular video, which jointly estimates dynamic geometry and camera poses, is an inevitably challenging problem. While recent pointmap-based 3D reconstruction methods (e.g., DUSt3R) have made great progress in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Shizun Wang , Zhenxiang Jiang , Xingyi Yang , Xinchao Wang

Text-to-3D form plays a crucial role in creating editable 3D scenes for AR/VR. Recent advances have shown promise in merging neural radiance fields (NeRFs) with pre-trained diffusion models for text-to-3D object generation. However, one…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Haotian Bai , Yuanhuiyi Lyu , Lutao Jiang , Sijia Li , Haonan Lu , Xiaodong Lin , Lin Wang

Large-scale diffusion generative models are greatly simplifying image, video and 3D asset creation from user-provided text prompts and images. However, the challenging problem of text-to-4D dynamic 3D scene generation with diffusion…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Yufeng Zheng , Xueting Li , Koki Nagano , Sifei Liu , Karsten Kreis , Otmar Hilliges , Shalini De Mello

Tracking non-rigidly deforming scenes using range sensors has numerous applications including computer vision, AR/VR, and robotics. However, due to occlusions and physical limitations of range sensors, existing methods only handle the…

Computer Vision and Pattern Recognition · Computer Science 2021-05-06 Yang Li , Hikari Takehara , Takafumi Taketomi , Bo Zheng , Matthias Nießner

We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a…

Recent advances in diffusion-based generative models have established a new paradigm for image and video relighting. However, extending these capabilities to 4D relighting remains challenging, due primarily to the scarcity of paired 4D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Zhenghuang Wu , Kang Chen , Zeyu Zhang , Hao Tang

This study presents a framework for capturing human attention in the spatio-temporal domain using eye-tracking glasses. Attention mapping is a key technology for human perceptual activity analysis or Human-Robot Interaction (HRI) to support…

Robotics · Computer Science 2021-07-09 Shuji Oishi , Kenji Koide , Masashi Yokozuka , Atsuhiko Banno

3D object detection with surround-view images is an essential task for autonomous driving. In this work, we propose DETR4D, a Transformer-based framework that explores sparse attention and direct feature query for 3D object detection in…

Computer Vision and Pattern Recognition · Computer Science 2022-12-16 Zhipeng Luo , Changqing Zhou , Gongjie Zhang , Shijian Lu

We present a new approach to instill 4D dynamic object priors into learned 3D representations by unsupervised pre-training. We observe that dynamic movement of an object through an environment provides important cues about its objectness,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Yujin Chen , Matthias Nießner , Angela Dai

Understanding and reasoning about objects' physical properties in the natural world is a fundamental challenge in artificial intelligence. While some properties like colors and shapes can be directly observed, others, such as mass and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Zhenfang Chen , Shilong Dong , Kexin Yi , Yunzhu Li , Mingyu Ding , Antonio Torralba , Joshua B. Tenenbaum , Chuang Gan

We study the problem of synthesizing a long-term dynamic video from only a single image. This is challenging since it requires consistent visual content movements given large camera motions. Existing methods either hallucinate inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Liao Shen , Xingyi Li , Huiqiang Sun , Juewen Peng , Ke Xian , Zhiguo Cao , Guosheng Lin

The advancement of diffusion models has pushed the boundary of text-to-3D object generation. While it is straightforward to composite objects into a scene with reasonable geometry, it is nontrivial to texture such a scene perfectly due to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Qi Wang , Ruijie Lu , Xudong Xu , Jingbo Wang , Michael Yu Wang , Bo Dai , Gang Zeng , Dan Xu

Visual scenes are composed of visual concepts and have the property of combinatorial explosion. An important reason for humans to efficiently learn from diverse visual scenes is the ability of compositional perception, and it is desirable…

Machine Learning · Computer Science 2023-06-16 Jinyang Yuan , Tonglin Chen , Bin Li , Xiangyang Xue

We introduce Drag4D, an interactive framework that integrates object motion control within text-driven 3D scene generation. This framework enables users to define 3D trajectories for the 3D objects generated from a single image, seamlessly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Minjun Kang , Inkyu Shin , Taeyeop Lee , In So Kweon , Kuk-Jin Yoon

Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D scenes. Replicating such capability, termed compositional 3D reconstruction, is pivotal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Mingyu Dong , Chong Xia , Mingyuan Jia , Weichen Lyu , Long Xu , Zheng Zhu , Yueqi Duan