English
Related papers

Related papers: Canonical Space Representation for 4D Panoptic Seg…

200 papers

We propose an end-to-end trainable, cross-category method for reconstructing multiple man-made articulated objects from a single RGBD image, focusing on part-level shape reconstruction and pose and kinematics estimation. We depart from…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuki Kawana , Tatsuya Harada

Perception is a key building block of autonomously acting vision systems such as autonomous vehicles. It is crucial that these systems are able to understand their surroundings in order to operate safely and robustly. Additionally,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Matteo Sodano , Federico Magistri , Jens Behley , Cyrill Stachniss

Modeling, understanding, and reconstructing the real world are crucial in XR/VR. Recently, 3D Gaussian Splatting (3D-GS) methods have shown remarkable success in modeling and understanding 3D scenes. Similarly, various 4D representations…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Shengxiang Ji , Guanjun Wu , Jiemin Fang , Jiazhong Cen , Taoran Yi , Wenyu Liu , Qi Tian , Xinggang Wang

4D LiDAR semantic segmentation, also referred to as multi-scan semantic segmentation, plays a crucial role in enhancing the environmental understanding capabilities of autonomous vehicles or robots. It classifies the semantic category of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Neng Wang , Ruibin Guo , Chenghao Shi , Ziyue Wang , Hui Zhang , Huimin Lu , Zhiqiang Zheng , Xieyuanli Chen

Capturing 4D spatiotemporal surroundings is crucial for the safe and reliable operation of robots in dynamic environments. However, most existing methods address only one side of the problem: they either provide coarse geometric tracking…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Maximilian Luz , Rohit Mohan , Thomas Nürnberg , Yakov Miron , Daniele Cattaneo , Abhinav Valada

Articulated objects are ubiquitous in daily life. Our goal is to achieve a high-quality reconstruction, segmentation of independent moving parts, and analysis of articulation. Recent methods analyse two different articulation states and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Hao Ai , Wenjie Chang , Jianbo Jiao , Ales Leonardis , Ofek Eyal

Recovering 4D from monocular video, which jointly estimates dynamic geometry and camera poses, is an inevitably challenging problem. While recent pointmap-based 3D reconstruction methods (e.g., DUSt3R) have made great progress in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Shizun Wang , Zhenxiang Jiang , Xingyi Yang , Xinchao Wang

In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making. However, object motion and ego-motion often induce cross-frame spatiotemporal inconsistencies in BEV-based detectors, leading to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Wenxuan Li , Qin Zou , Shoubing Chen , Chi Chen , Yingyi Yang , Shoubing Chen , Qingxiang Meng

Learning 4D language fields to enable time-sensitive, open-ended language queries in dynamic scenes is essential for many real-world applications. While LangSplat successfully grounds CLIP features into 3D Gaussian representations,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Wanhua Li , Renping Zhou , Jiawei Zhou , Yingwei Song , Johannes Herter , Minghan Qin , Gao Huang , Hanspeter Pfister

We present SAM4D, a multi-modal and temporal foundation model designed for promptable segmentation across camera and LiDAR streams. Unified Multi-modal Positional Encoding (UMPE) is introduced to align camera and LiDAR features in a shared…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Jianyun Xu , Song Wang , Ziqian Ni , Chunyong Hu , Sheng Yang , Jianke Zhu , Qiang Li

We present a new approach to instill 4D dynamic object priors into learned 3D representations by unsupervised pre-training. We observe that dynamic movement of an object through an environment provides important cues about its objectness,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Yujin Chen , Matthias Nießner , Angela Dai

Recent advances in unsupervised learning for object detection, segmentation, and tracking hold significant promise for applications in robotics. A common approach is to frame these tasks as inference in probabilistic latent-variable models.…

Robotics · Computer Science 2021-09-14 Yizhe Wu , Oiwi Parker Jones , Martin Engelcke , Ingmar Posner

Dynamic Gaussian Splatting approaches have achieved remarkable performance for 4D scene reconstruction. However, these approaches rely on dense-frame video sequences for photorealistic reconstruction. In real-world scenarios, due to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Changyue Shi , Chuxiao Yang , Xinyuan Hu , Minghao Chen , Wenwen Pan , Yan Yang , Jiajun Ding , Zhou Yu , Jun Yu

To determine the 3D orientation and 3D location of objects in the surroundings of a camera mounted on a robot or mobile device, we developed two powerful algorithms in object detection and temporal tracking that are combined seamlessly for…

Computer Vision and Pattern Recognition · Computer Science 2017-09-06 David Joseph Tan , Nassir Navab , Federico Tombari

Persistent dynamic scene modeling for tracking and novel-view synthesis remains challenging due to the difficulty of capturing accurate deformations while maintaining computational efficiency. We propose SCas4D, a cascaded optimization…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Jipeng Lyu , Jiahua Dong , Yu-Xiong Wang

Perception is an essential part of robotic manipulation in a semi-structured environment. Traditional approaches produce a narrow task-specific prediction (e.g., object's 6D pose), that cannot be adapted to other tasks and is ill-suited for…

Robotics · Computer Science 2023-03-03 Benjamin Joffe , Konrad Ahlin

Instance segmentation and panoptic segmentation is being paid more and more attention in recent years. In comparison with bounding box based object detection and semantic segmentation, instance segmentation can provide more analytical…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Xiaolong Liu , Yuqing Hou , Anbang Yao , Yurong Chen , Keqiang Li

We present Stable Part Diffusion 4D (SP4D), a framework for generating paired RGB and kinematic part videos from monocular inputs. Unlike conventional part segmentation methods that rely on appearance-based semantic cues, SP4D learns to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-06 Hao Zhang , Chun-Han Yao , Simon Donné , Narendra Ahuja , Varun Jampani

Recent studies have shown that video-level representation learning is crucial to the capture and understanding of the long-range temporal structure for video action recognition. Most existing 3D convolutional neural network (CNN)-based…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Mohammad Al-Saad , Lakshmish Ramaswamy , Suchendra Bhandarkar

Directly learning to model 4D content, including shape, color, and motion, is challenging. Existing methods rely on pose priors for motion control, resulting in limited motion diversity and continuity in details. To address this, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Qitong Yang , Mingtao Feng , Zijie Wu , Shijie Sun , Weisheng Dong , Yaonan Wang , Ajmal Mian