English
Related papers

Related papers: CM-EVS: Sparse Panoramic RGB-D-Pose Data for Compl…

200 papers

The limited dynamic range of the detector can impede coherent diffractive imaging (CDI) schemes from achieving diffraction-limited resolution. To overcome this limitation, a straightforward approach is to utilize high dynamic range (HDR)…

Image and Video Processing · Electrical Eng. & Systems 2024-06-11 Shantanu Kodgirwar , Lars Loetgering , Chang Liu , Aleena Joseph , Leona Licht , Daniel S. Penagos Molina , Wilhelm Eschen , Jan Rothhardt , Michael Habeck

Multimodal Large Language Models (MLLMs) require comprehensive visual inputs to achieve dense understanding of the physical world. While existing MLLMs demonstrate impressive world understanding capabilities through limited field-of-view…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Yikang Zhou , Tao Zhang , Dizhe Zhang , Shunping Ji , Xiangtai Li , Lu Qi

Spherical cameras capture scenes in a holistic manner and have been used for room layout estimation. Recently, with the availability of appropriate datasets, there has also been progress in depth estimation from a single omnidirectional…

Computer Vision and Pattern Recognition · Computer Science 2022-06-24 Nikolaos Zioulis , Federico Alvarez , Dimitrios Zarpalas , Petros Daras

To advance the state of the art in the creation of 3D foundation models, this paper introduces the ConDense framework for 3D pre-training utilizing existing pre-trained 2D networks and large-scale multi-view datasets. We propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-09-02 Xiaoshuai Zhang , Zhicheng Wang , Howard Zhou , Soham Ghosh , Danushen Gnanapragasam , Varun Jampani , Hao Su , Leonidas Guibas

We aim to tackle sparse-view reconstruction of a 360 3D scene using priors from latent diffusion models (LDM). The sparse-view setting is ill-posed and underconstrained, especially for scenes where the camera rotates 360 degrees around a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Soumava Paul , Christopher Wewer , Bernt Schiele , Jan Eric Lenssen

Imitation learning is promising for robotic manipulation, but \emph{precise insertion} in the real world remains difficult due to contact-rich dynamics, tight clearances, and limited demonstrations. Many existing visuomotor policies depend…

Robotics · Computer Science 2026-03-25 Han Sun , Sheng Liu , Yizhao Wang , Zhenning Zhou , Shuai Wang , Haibo Yang , Jingyuan Sun , Qixin Cao

In computer vision, estimating the six-degree-of-freedom pose from an RGB image is a fundamental task. However, this task becomes highly challenging in multi-object scenes. Currently, the best methods typically employ an indirect strategy,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Xin Liu , Hao Wang , Shibei Xue , Dezong Zhao

Current video-based computer vision (CV) applications typically suffer from high energy consumption due to reading and processing all pixels in a frame, regardless of their significance. While previous works have attempted to reduce this…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Md Abdullah-Al Kaiser , Sreetama Sarkar , Peter A. Beerel , Akhilesh R. Jaiswal , Gourav Datta

We propose a novel technique to register sparse 3D scans in the absence of texture. While existing methods such as KinectFusion or Iterative Closest Points (ICP) heavily rely on dense point clouds, this task is particularly challenging…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Siddhant Ranade , Xin Yu , Shantnu Kakkar , Pedro Miraldo , Srikumar Ramalingam

End-to-End Autonomous Driving (E2EAD) methods typically rely on supervised perception tasks to extract explicit scene information (e.g., objects, maps). This reliance necessitates expensive annotations and constrains deployment and data…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Peidong Li , Dixiao Cui

Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Sayan Deb Sarkar , Ondrej Miksik , Marc Pollefeys , Daniel Barath , Iro Armeni

We study the problem of learning to estimate the 3D object pose from a few labelled examples and a collection of unlabelled data. Our main contribution is a learning framework, neural view synthesis and matching, that can transfer the 3D…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Angtian Wang , Shenxiao Mei , Alan Yuille , Adam Kortylewski

Given a single scene image, this paper proposes a method of Category-level 6D Object Pose and Size Estimation (COPSE) from the point cloud of the target object, without external real pose-annotated training data. Specifically, beyond the…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Haitao Lin , Zichang Liu , Chilam Cheang , Yanwei Fu , Guodong Guo , Xiangyang Xue

6D Object Pose Estimation is a crucial yet challenging task in computer vision, suffering from a significant lack of large-scale datasets. This scarcity impedes comprehensive evaluation of model performance, limiting research advancements.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Jiyao Zhang , Weiyao Huang , Bo Peng , Mingdong Wu , Fei Hu , Zijian Chen , Bo Zhao , Hao Dong

3D single object tracking remains a challenging problem due to the sparsity and incompleteness of the point clouds. Existing algorithms attempt to address the challenges in two strategies. The first strategy is to learn dense geometric…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Jingwen Zhang , Zikun Zhou , Guangming Lu , Jiandong Tian , Wenjie Pei

Event cameras continue to attract interest due to desirable characteristics such as high dynamic range, low latency, virtually no motion blur, and high energy efficiency. One of the potential applications that would benefit from these…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Tobias Fischer , Michael Milford

Sparse-view 3D reconstruction is essential for applications in which dense image acquisition is impractical, such as robotics, augmented/virtual reality (AR/VR), and autonomous systems. In these settings, minimal image overlap prevents…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Tanveer Younis , Zhanglin Cheng

Compressive Sensing (CS) stipulates that a sparse signal can be recovered from a small number of linear measurements, and that this recovery can be performed efficiently in polynomial time. The framework of model-based compressive sensing…

Information Theory · Computer Science 2015-04-22 Chinmay Hegde , Piotr Indyk , Ludwig Schmidt

Diffusion policies generate robot motions by learning to denoise action-space trajectories conditioned on observations. These observations are commonly streams of RGB images, whose high dimensionality includes substantial task-irrelevant…

Robotics · Computer Science 2025-09-18 Xiatao Sun , Yinxing Chen , Daniel Rakita

We present a deep model that can accurately produce dense depth maps given an RGB image with known depth at a very sparse set of pixels. The model works simultaneously for both indoor/outdoor scenes and produces state-of-the-art dense depth…

Computer Vision and Pattern Recognition · Computer Science 2018-12-11 Zhao Chen , Vijay Badrinarayanan , Gilad Drozdov , Andrew Rabinovich