English
Related papers

Related papers: PanopticQuery: Unified Query-Time Reasoning for 4D…

200 papers

Performing single image holistic understanding and 3D reconstruction is a central task in computer vision. This paper presents an integrated system that performs dense scene labeling, object detection, instance segmentation, depth…

Computer Vision and Pattern Recognition · Computer Science 2021-12-01 Sainan Liu , Vincent Nguyen , Yuan Gao , Subarna Tripathi , Zhuowen Tu

Despite advances in Multimodal LLMs (MLLMs), their ability to reason over 3D structures and temporal dynamics remains limited, constrained by weak 4D perception and temporal understanding. Existing 3D and 4D Video Question Answering (VQA)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Chiao-An Yang , Ryo Hachiuma , Sifei Liu , Subhashree Radhakrishnan , Raymond A. Yeh , Yu-Chiang Frank Wang , Min-Hung Chen

The blooming of virtual reality and augmented reality (VR/AR) technologies has driven an increasing demand for the creation of high-quality, immersive, and dynamic environments. However, existing generative techniques either focus solely on…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Renjie Li , Panwang Pan , Bangbang Yang , Dejia Xu , Shijie Zhou , Xuanyang Zhang , Zeming Li , Achuta Kadambi , Zhangyang Wang , Zhengzhong Tu , Zhiwen Fan

Panoptic scene understanding and tracking of dynamic agents are essential for robots and automated vehicles to navigate in urban environments. As LiDARs provide accurate illumination-independent geometric depictions of the scene, performing…

Computer Vision and Pattern Recognition · Computer Science 2021-12-28 Whye Kit Fong , Rohit Mohan , Juana Valeria Hurtado , Lubing Zhou , Holger Caesar , Oscar Beijbom , Abhinav Valada

Recent advancements in 2D and 3D generative models have expanded the capabilities of computer vision. However, generating high-quality 4D dynamic content from a single static image remains a significant challenge. Traditional methods have…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Jing Yang , Yufeng Yang

Novel view synthesis has shown rapid progress recently, with methods capable of producing increasingly photorealistic results. 3D Gaussian Splatting has emerged as a promising method, producing high-quality renderings of scenes and enabling…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Richard Shaw , Michal Nazarczuk , Jifei Song , Arthur Moreau , Sibi Catley-Chandar , Helisa Dhamo , Eduardo Perez-Pellitero

Rendering novel view images in dynamic scenes is a crucial yet challenging task. Current methods mainly utilize NeRF-based methods to represent the static scene and an additional time-variant MLP to model scene deformations, resulting in…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Diwen Wan , Ruijie Lu , Gang Zeng

We present Free4D, a novel tuning-free framework for 4D scene generation from a single image. Existing methods either focus on object-level generation, making scene-level generation infeasible, or rely on large-scale multi-view video…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Tianqi Liu , Zihao Huang , Zhaoxi Chen , Guangcong Wang , Shoukang Hu , Liao Shen , Huiqiang Sun , Zhiguo Cao , Wei Li , Ziwei Liu

Recent advancements in 3D Gaussian Splatting (3D-GS) enable high-quality 3D scene reconstruction from RGB images. Many studies extend this paradigm for language-driven open-vocabulary scene understanding. However, most of them simply…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Jiazhong Cen , Xudong Zhou , Jiemin Fang , Changsong Wen , Lingxi Xie , Xiaopeng Zhang , Wei Shen , Qi Tian

Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spatiotemporal scale. Typically, existing 4D generative models directly embed macro scale…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Haonan Wang , Hanyu Zhou , Tao Gu , Luxin Yan

Instruction-guided generative models, especially those using text-to-image (T2I) and text-to-video (T2V) diffusion frameworks, have advanced the field of content editing in recent years. To extend these capabilities to 4D scene, we…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Hasan Iqbal , Nazmul Karim , Umar Khalid , Azib Farooq , Zichun Zhong , Chen Chen , Jing Hua

The reconstruction of immersive and realistic 3D scenes holds significant practical importance in various fields of computer vision and computer graphics. Typically, immersive and realistic scenes should be free from obstructions by dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Zilong Huang , Jun He , Junyan Ye , Lihan Jiang , Weijia Li , Yiping Chen , Ting Han

Reconstructing dynamic 3D scenes from sparse multi-view videos is highly ill-posed, often leading to geometric collapse, trajectory drift, and floating artifacts. Recent attempts introduce generative priors to hallucinate missing content,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Zhenlong Wu , Zihan Zheng , Xuanxuan Wang , Qianhe Wang , Hua Yang , Xiaoyun Zhang , Qiang Hu , Wenjun Zhang

Feed-forward 3D reconstruction for autonomous driving has advanced rapidly, yet existing methods struggle with the joint challenges of sparse, non-overlapping camera views and complex scene dynamics. We present UniSplat, a general…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Chen Shi , Shaoshuai Shi , Xiaoyang Lyu , Chunyang Liu , Kehua Sheng , Bo Zhang , Li Jiang

Reconstructing large-scale dynamic driving scenes remains challenging due to the coexistence of static environments with extreme depth variation and diverse dynamic actors exhibiting complex motions. Existing Gaussian Splatting based…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Cong Wang , Ruiqi Song , Wei Tian , Chenming Zhang , Lingxi Li , Long Chen

This paper proposes Panoptic Narrative Grounding, a spatially fine and general formulation of the natural language visual grounding problem. We establish an experimental framework for the study of this new task, including new ground truth…

Computer Vision and Pattern Recognition · Computer Science 2021-09-13 C. González , N. Ayobi , I. Hernández , J. Hernández , J. Pont-Tuset , P. Arbeláez

Recent advancements in camera-based occupancy prediction have focused on the simultaneous prediction of 3D semantics and scene flow, a task that presents significant challenges due to specific difficulties, e.g., occlusions and unbalanced…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Ziyue Zhu , Shenlong Wang , Jin Xie , Jiang-jiang Liu , Jingdong Wang , Jian Yang

Learning 3D scene geometry and semantics from images is a core challenge in computer vision and a key capability for autonomous driving. Since large-scale 3D annotation is prohibitively expensive, recent work explores self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Adam Lilja , Ji Lan , Junsheng Fu , Lars Hammarstrand

3D semantic scene understanding is a fundamental challenge in computer vision. It enables mobile agents to autonomously plan and navigate arbitrary environments. SSC formalizes this challenge as jointly estimating dense geometry and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Adrian Hayler , Felix Wimbauer , Dominik Muhle , Christian Rupprecht , Daniel Cremers

We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing approaches that typically decouple motion from geometry or produce limited 4D attributes such as sparse trajectories or two-view…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Yihang Luo , Shangchen Zhou , Yushi Lan , Xingang Pan , Chen Change Loy
‹ Prev 1 4 5 6 7 8 10 Next ›