English
Related papers

Related papers: Speak the Same Language: Global LiDAR Registration…

200 papers

Point cloud registration aligns multiple unposed point clouds into a common reference frame and is a core step for 3D reconstruction and robot localization without initial guess. In this work, we cast registration as conditional generation:…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yue Pan , Tao Sun , Liyuan Zhu , Lucas Nunes , Iro Armeni , Jens Behley , Cyrill Stachniss

Joint image-text embedding is the bedrock for most Vision-and-Language (V+L) tasks, where multimodality inputs are simultaneously processed for joint visual and textual understanding. In this paper, we introduce UNITER, a UNiversal…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Yen-Chun Chen , Linjie Li , Licheng Yu , Ahmed El Kholy , Faisal Ahmed , Zhe Gan , Yu Cheng , Jingjing Liu

Global visual localization in LiDAR-maps, crucial for autonomous driving applications, remains largely unexplored due to the challenging issue of bridging the cross-modal heterogeneity gap. Popular multi-modal learning approach Contrastive…

Robotics · Computer Science 2023-12-29 Sai Shubodh Puligilla , Mohammad Omama , Husain Zaidi , Udit Singh Parihar , Madhava Krishna

Place recognition is critical for both offline mapping and online localization. However, current single-sensor based place recognition still remains challenging in adverse conditions. In this paper, a heterogeneous measurements based…

Computer Vision and Pattern Recognition · Computer Science 2021-06-21 Huan Yin , Xuecheng Xu , Yue Wang , Rong Xiong

Human-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predict HOI triplets. Despite the challenges posed by…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Yichao Cao , Qingfei Tang , Feng Yang , Xiu Su , Shan You , Xiaobo Lu , Chang Xu

Recent advances have demonstrated that Language Vision Models (LVMs) surpass the existing State-of-the-Art (SOTA) in two-dimensional (2D) computer vision tasks, motivating attempts to apply LVMs to three-dimensional (3D) data. While LVMs…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 June Moh Goo , Zichao Zeng , Jan Boehm

Emotion recognition in conversations is challenging due to the multi-modal nature of the emotion expression. We propose a hierarchical cross-attention model (HCAM) approach to multi-modal emotion recognition using a combination of recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-10 Soumya Dutta , Sriram Ganapathy

Cross-modal data registration has long been a critical task in computer vision, with extensive applications in autonomous driving and robotics. Accurate and robust registration methods are essential for aligning data from different…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Yuanchao Yue , Hui Yuan , Qinglong Miao , Xiaolong Mao , Raouf Hamzaoui , Peter Eisert

Most models tasked to ground referential utterances in 2D and 3D scenes learn to select the referred object from a pool of object proposals provided by a pre-trained detector. This is limiting because an utterance may refer to visual…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Ayush Jain , Nikolaos Gkanatsios , Ishita Mediratta , Katerina Fragkiadaki

Autonomous robots that assist humans in day to day living tasks are becoming increasingly popular. Autonomous mobile robots operate by sensing and perceiving their surrounding environment to make accurate driving decisions. A combination of…

Computer Vision and Pattern Recognition · Computer Science 2018-08-24 Varuna De Silva , Jamie Roche , Ahmet Kondoz

In this paper, we unleash the potential of the powerful monodepth model in camera-LiDAR calibration and propose CLAIM, a novel method of aligning data from the camera and LiDAR. Given the initial guess and pairs of images and LiDAR point…

Robotics · Computer Science 2026-03-18 Zhuo Zhang , Yonghui Liu , Meijie Zhang , Feiyang Tan , Yikang Ding

This work reports a novel multi-frame Bundle Adjustment (BA) framework called RKHS-BA. It uses continuous landmark representations that encode RGB-D/LiDAR and semantic observations in a Reproducing Kernel Hilbert Space (RKHS). With a…

Robotics · Computer Science 2024-12-03 Ray Zhang , Jingwei Song , Xiang Gao , Junzhe Wu , Tianyi Liu , Jinyuan Zhang , Ryan Eustice , Maani Ghaffari

Currently, GPS is by far the most popular global localization method. However, it is not always reliable or accurate in all environments. SLAM methods enable local state estimation but provide no means of registering the local map to a…

Lidar-based simultaneous localization and mapping (SLAM) approaches have obtained considerable success in autonomous robotic systems. This is in part owing to the high-accuracy of robust SLAM algorithms and the emergence of new and…

Robotics · Computer Science 2022-10-04 Ha Sier , Li Qingqing , Yu Xianjia , Jorge Peña Queralta , Zhuo Zou , Tomi Westerlund

One-shot LiDAR localization refers to the ability to estimate the robot pose from one single point cloud, which yields significant advantages in initialization and relocalization processes. In the point cloud domain, the topic has been…

Robotics · Computer Science 2023-09-19 Pengyu Yin , Haozhi Cao , Thien-Minh Nguyen , Shenghai Yuan , Shuyang Zhang , Kangcheng Liu , Lihua Xie

To navigate through urban roads, an automated vehicle must be able to perceive and recognize objects in a three-dimensional environment. A high-level contextual understanding of the surroundings is necessary to plan and execute accurate…

Robotics · Computer Science 2020-03-05 Julie Stephany Berrio , Mao Shan , Stewart Worrall , James Ward , Eduardo Nebot

The ability to build maps is a key functionality for the majority of mobile robots. A central ingredient to most mapping systems is the registration or alignment of the recorded sensor data. In this paper, we present a general methodology…

Computer Vision and Pattern Recognition · Computer Science 2017-09-19 Bartolomeo Della Corte , Igor Bogoslavskyi , Cyrill Stachniss , Giorgio Grisetti

Existing learning methods for LiDAR-based applications use 3D points scanned under a pre-determined beam configuration, e.g., the elevation angles of beams are often evenly distributed. Those fixed configurations are task-agnostic, so…

Robotics · Computer Science 2023-03-29 Niclas Vödisch , Ozan Unal , Ke Li , Luc Van Gool , Dengxin Dai

3D object detection is an important task that has been widely applied in autonomous driving. To perform this task, a new trend is to fuse multi-modal inputs, i.e., LiDAR and camera. Under such a trend, recent methods fuse these two…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yang Song , Lin Wang

3D occupancy prediction aims to infer dense, voxel-wise scene semantics from sensor observations, where the 2D-to-3D view transformation serves as a crucial step in bridging image features and volumetric representations. Most previous…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Yuan Wu , Zhiqiang Yan , Jiawei Lian , Zhengxue Wang , Jian Yang
‹ Prev 1 3 4 5 6 7 10 Next ›