English
Related papers

Related papers: Zero-Shot Multi-Object Scene Completion

200 papers

We study the challenging problem of unsupervised multi-object segmentation on single images. Existing methods, which rely on image reconstruction objectives to learn objectness or leverage pretrained image features to group similar pixels,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Yafei Yang , Zihui Zhang , Bo Yang

Mapping and understanding complex 3D environments is fundamental to how autonomous systems perceive and interact with the physical world, requiring both precise geometric reconstruction and rich semantic comprehension. While existing 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Naman Patel , Prashanth Krishnamurthy , Farshad Khorrami

In this paper, we propose a novel iterative multi-task framework to complete the segmentation mask of an occluded vehicle and recover the appearance of its invisible parts. In particular, to improve the quality of the segmentation…

Computer Vision and Pattern Recognition · Computer Science 2019-07-23 Xiaosheng Yan , Yuanlong Yu , Feigege Wang , Wenxi Liu , Shengfeng He , Jia Pan

Recent advances in text-to-image (T2I) diffusion models have significantly improved semantic image editing, yet most methods fall short in performing 3D-aware object manipulation. In this work, we present FFSE, a 3D-aware autoregressive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xincheng Shuai , Zhenyuan Qin , Henghui Ding , Dacheng Tao

The perception of transparent objects for grasp and manipulation remains a major challenge, because existing robotic grasp methods which heavily rely on depth maps are not suitable for transparent objects due to their unique visual…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Yifan Zhou , Wanli Peng , Zhongyu Yang , He Liu , Yi Sun

Image matching is a fundamental and critical task in various visual applications, such as Simultaneous Localization and Mapping (SLAM) and image retrieval, which require accurate pose estimation. However, most existing methods ignore the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Miao Fan , Mingrui Chen , Chen Hu , Shuchang Zhou

Accurate depth information is crucial for enhancing the performance of multi-view 3D object detection. Despite the success of some existing multi-view 3D detectors utilizing pixel-wise depth supervision, they overlook two significant…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Jinghua Hou , Tong Wang , Xiaoqing Ye , Zhe Liu , Shi Gong , Xiao Tan , Errui Ding , Jingdong Wang , Xiang Bai

3D scene understanding for robotic applications exhibits a unique set of requirements including real-time inference, object-centric latent representation learning, accurate 6D pose estimation and 3D reconstruction of objects. Current…

Robotics · Computer Science 2024-02-27 Yizhe Wu , Haitz Sáez de Ocáriz Borde , Jack Collins , Oiwi Parker Jones , Ingmar Posner

Realizing unified 3D object detection, including both indoor and outdoor scenes, holds great importance in applications like robot navigation. However, involving various scenarios of data to train models poses challenges due to their…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Zhuoling Li , Xiaogang Xu , SerNam Lim , Hengshuang Zhao

Capturing general deforming scenes from monocular RGB video is crucial for many computer graphics and vision applications. However, current approaches suffer from drawbacks such as struggling with large scene deformations, inaccurate shape…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Erik C. M. Johnson , Marc Habermann , Soshi Shimada , Vladislav Golyanik , Christian Theobalt

Despite monocular 3D object detection having recently made a significant leap forward thanks to the use of pre-trained depth estimators for pseudo-LiDAR recovery, such two-stage methods typically suffer from overfitting and are incapable of…

Computer Vision and Pattern Recognition · Computer Science 2022-11-03 Yongzhi Su , Yan Di , Fabian Manhardt , Guangyao Zhai , Jason Rambach , Benjamin Busam , Didier Stricker , Federico Tombari

In this paper, we address the inverse problem of reconstructing a scene as well as the camera motion from the image sequence taken by an omni-directional camera. Our structure from motion results give sharp conditions under which the…

Computer Vision and Pattern Recognition · Computer Science 2007-08-21 Oliver Knill , Jose Ramirez-Herran

3D geometry is a very informative cue when interacting with and navigating an environment. This writing proposes a new approach to 3D reconstruction and scene understanding, which implicitly learns 3D geometry from depth maps pairing a deep…

Computer Vision and Pattern Recognition · Computer Science 2018-08-22 Dario Rethage , Federico Tombari , Felix Achilles , Nassir Navab

Three-dimensional scene inpainting is crucial for applications from virtual reality to architectural visualization, yet existing methods struggle with view consistency and geometric accuracy in 360{\deg} unbounded scenes. We present…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Chung-Ho Wu , Yang-Jung Chen , Ying-Huan Chen , Jie-Ying Lee , Bo-Hsu Ke , Chun-Wei Tuan Mu , Yi-Chuan Huang , Chin-Yang Lin , Min-Hung Chen , Yen-Yu Lin , Yu-Lun Liu

Multi-camera 3D perception has emerged as a prominent research field in autonomous driving, offering a viable and cost-effective alternative to LiDAR-based solutions. The existing multi-camera algorithms primarily rely on monocular 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Chen Min , Liang Xiao , Dawei Zhao , Yiming Nie , Bin Dai

Modern multi-object tracking (MOT) systems usually model the trajectories by associating per-frame detections. However, when camera motion, fast motion, and occlusion challenges occur, it is difficult to ensure long-range tracking or even…

Computer Vision and Pattern Recognition · Computer Science 2020-09-21 Shoudong Han , Piao Huang , Hongwei Wang , En Yu , Donghaisheng Liu , Xiaofeng Pan , Jun Zhao

Recent works in hand-object reconstruction mainly focus on the single-view and dense multi-view settings. On the one hand, single-view methods can leverage learned shape priors to generalise to unseen objects but are prone to inaccuracies…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Yik Lung Pang , Changjae Oh , Andrea Cavallaro

Multi-object tracking from RGB-D video sequences is a challenging problem due to the combination of changing viewpoints, motion, and occlusions over time. We observe that having the complete geometry of objects aids in their tracking, and…

Computer Vision and Pattern Recognition · Computer Science 2020-12-17 Norman Müller , Yu-Shiang Wong , Niloy J. Mitra , Angela Dai , Matthias Nießner

This paper proposes an online multi-camera multi-object tracker that only requires monocular detector training, independent of the multi-camera configurations, allowing seamless extension/deletion of cameras without retraining effort. The…

Computer Vision and Pattern Recognition · Computer Science 2020-10-28 Jonah Ong , Ba Tuong Vo , Ba Ngu Vo , Du Yong Kim , Sven Nordholm

We propose EscherNet++, a masked fine-tuned diffusion model that can synthesize novel views of objects in a zero-shot manner with amodal completion ability. Existing approaches utilize multiple stages and complex pipelines to first…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Xinan Zhang , Muhammad Zubair Irshad , Anthony Yezzi , Yi-Chang Tsai , Zsolt Kira