English
Related papers

Related papers: Constructing Category-Specific Models for Monocula…

200 papers

3D object detection with surrounding cameras has been a promising direction for autonomous driving. In this paper, we present SimMOD, a Simple baseline for Multi-camera Object Detection, to solve the problem. To incorporate multi-view…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Yunpeng Zhang , Wenzhao Zheng , Zheng Zhu , Guan Huang , Jie Zhou , Jiwen Lu

Most SLAM algorithms are based on the assumption that the scene is static. However, in practice, most scenes are dynamic which usually contains moving objects, these methods are not suitable. In this paper, we introduce DymSLAM, a dynamic…

Computer Vision and Pattern Recognition · Computer Science 2020-03-11 Chenjie Wang , Bin Luo , Yun Zhang , Qing Zhao , Lu Yin , Wei Wang , Xin Su , Yajun Wang , Chengyuan Li

Due to its cost-effectiveness and widespread availability, monocular 3D object detection, which relies solely on a single camera during inference, holds significant importance across various applications, including autonomous driving and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Bonan Ding , Jin Xie , Jing Nie , Jiale Cao , Xuelong Li , Yanwei Pang

We propose a framework that allows a mobile robot to build a map of an indoor scenario, identifying and highlighting objects that may be considered a hindrance to people with limited mobility. The map is built by combining recent…

Robotics · Computer Science 2021-11-25 V. Ayala-Alfaro , J. A. Vilchis-Mar , F. E. Correa-Tome , J. P. Ramirez-Paredes

Object detection traditionally relies on fixed category sets, requiring costly re-training to handle novel objects. While Open-World and Open-Vocabulary Object Detection (OWOD and OVOD) improve flexibility, OWOD lacks semantic labels for…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Furkan Mumcu , Michael J. Jones , Anoop Cherian , Yasin Yilmaz

Accurate perception of objects in the environment is important for improving the scene understanding capability of SLAM systems. In robotic and augmented reality applications, object maps with semantic and metric information show attractive…

Robotics · Computer Science 2023-11-21 Xiao Han , Houxuan Liu , Yunchao Ding , Lu Yang

Motion segmentation from a single moving camera presents a significant challenge in the field of computer vision. This challenge is compounded by the unknown camera movements and the lack of depth information of the scene. While deep…

Computer Vision and Pattern Recognition · Computer Science 2024-06-28 Yuxiang Huang , Yuhao Chen , John Zelek

An accurate and computationally efficient SLAM algorithm is vital for modern autonomous vehicles. To make a lightweight the algorithm, most SLAM systems rely on feature detection from images for vision SLAM or point cloud for laser-based…

Robotics · Computer Science 2021-03-22 Waqas Ali , Peilin Liu , Rendong Ying , Zheng Gong

Traditional Visual Simultaneous Localization and Mapping (VSLAM) systems assume a static environment, which makes them ineffective in highly dynamic settings. To overcome this, many approaches integrate semantic information from deep…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Sanghyoup Gu , Ratnesh Kumar

Cooperative Simultaneous Localization and Mapping (C-SLAM) enables multiple agents to work together in mapping unknown environments while simultaneously estimating their own positions. This approach enhances robustness, scalability, and…

Robotics · Computer Science 2025-08-28 Joshua Bird , Jan Blumenkamp , Amanda Prorok

Accurate and robust 3D scene reconstruction from casual, in-the-wild videos can significantly simplify robot deployment to new environments. However, reliable camera pose estimation and scene reconstruction from such unconstrained videos…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Shuo Sun , Torsten Sattler , Malcolm Mielle , Achim J. Lilienthal , Martin Magnusson

In this work, we tackle the challenging problem of category-level object pose and size estimation from a single depth image. Although previous fully-supervised works have demonstrated promising performance, collecting ground-truth pose…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Yisheng He , Haoqiang Fan , Haibin Huang , Qifeng Chen , Jian Sun

Many monocular visual SLAM algorithms are derived from incremental structure-from-motion (SfM) methods. This work proposes a novel monocular SLAM method which integrates recent advances made in global SfM. In particular, we present two main…

Computer Vision and Pattern Recognition · Computer Science 2017-10-20 Chengzhou Tang , Oliver Wang , Ping Tan

We present DropD-SLAM, a real-time monocular SLAM system that achieves RGB-D-level accuracy without relying on depth sensors. The system replaces active depth input with three pretrained vision modules: a monocular metric depth estimator, a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Mert Kiray , Alican Karaomer , Benjamin Busam

Category-level articulated object pose estimation focuses on the pose estimation of unknown articulated objects within known categories. Despite its significance, this task remains challenging due to the varying shapes and poses of objects,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yuchen Che , Ryo Furukawa , Asako Kanezaki

A natural approach to generative modeling of videos is to represent them as a composition of moving objects. Recent works model a set of 2D sprites over a slowly-varying background, but without considering the underlying 3D scene that gives…

Computer Vision and Pattern Recognition · Computer Science 2021-03-26 Paul Henderson , Christoph H. Lampert

A spatial AI that can perform complex tasks through visual signals and cooperate with humans is highly anticipated. To achieve this, we need a visual SLAM that easily adapts to new scenes without pre-training and generates dense maps for…

Monocular 3D object detection is a crucial and challenging task for autonomous driving vehicle, while it uses only a single camera image to infer 3D objects in the scene. To address the difficulty of predicting depth using only pictorial…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Jia-Quan Yu , Soo-Chang Pei

The application of monocular dense Simultaneous Localization and Mapping (SLAM) is often hindered by high latency, large GPU memory consumption, and reliance on camera calibration. To relax this constraint, we propose EC3R-SLAM, a novel…

Robotics · Computer Science 2025-10-03 Lingxiang Hu , Naima Ait Oufroukh , Fabien Bonardi , Raymond Ghandour

The precise localization of 3D objects from a single image without depth information is a highly challenging problem. Most existing methods adopt the same approach for all objects regardless of their diverse distributions, leading to…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Yunpeng Zhang , Jiwen Lu , Jie Zhou