English
Related papers

Related papers: CoNav: Collaborative Cross-Modal Reasoning for Emb…

200 papers

Text-image cross-modal retrieval is a challenging task in the field of language and vision. Most previous approaches independently embed images and sentences into a joint embedding space and compare their similarities. However, previous…

Computer Vision and Pattern Recognition · Computer Science 2019-09-13 Zihao Wang , Xihui Liu , Hongsheng Li , Lu Sheng , Junjie Yan , Xiaogang Wang , Jing Shao

Vision-based bird's-eye-view (BEV) 3D object detection has advanced significantly in autonomous driving by offering cost-effectiveness and rich contextual information. However, existing methods often construct BEV representations by…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Jicheng Yuan , Manh Nguyen Duc , Qian Liu , Manfred Hauswirth , Danh Le Phuoc

Modeling the cognitive and experiential factors of human navigation is central to deepening our understanding of human-environment interaction and to enabling safe social navigation and effective assistive wayfinding. Most existing methods…

Machine Learning · Computer Science 2026-03-09 Zhiwen Qiu , Ziang Liu , Wenqian Niu , Tapomayukh Bhattacharjee , Saleh Kalantari

Audiovisual embodied navigation enables robots to locate audio sources by dynamically integrating visual observations from onboard sensors with the auditory signals emitted by the target. The core challenge lies in effectively leveraging…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Yinfeng Yu , Hailong Zhang , Meiling Zhu

Semantic communication has been introduced into collaborative perception systems for autonomous driving, offering a promising approach to enhancing data transmission efficiency and robustness. Despite its potential, existing semantic…

Signal Processing · Electrical Eng. & Systems 2025-12-30 Jipeng Gan , Le Liang , Hua Zhang , Chongtao Guo , Shi Jin

Large unimodal foundation models for vision and language encode rich semantic structures, yet aligning them typically requires computationally intensive multimodal fine-tuning. Such approaches depend on large-scale parameter updates, are…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Abhishek Dalvi , Vasant Honavar

We propose 3D Congealing, a novel problem of 3D-aware alignment for 2D images capturing semantically similar objects. Given a collection of unlabeled Internet images, our goal is to associate the shared semantic parts from the inputs and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Yunzhi Zhang , Zizhang Li , Amit Raj , Andreas Engelhardt , Yuanzhen Li , Tingbo Hou , Jiajun Wu , Varun Jampani

Socially compliant navigation requires robots to move safely and appropriately in human-centered environments by respecting social norms. However, social norms are often ambiguous, and in a single scenario, multiple actions may be equally…

Robotics · Computer Science 2025-12-29 Zishuo Wang , Xinyu Zhang , Zhuonan Liu , Tomohito Kawabata , Daeun Song , Xuesu Xiao , Ling Xiao

Understanding the geometric and semantic properties of the scene is crucial in autonomous navigation and particularly challenging in the case of Unmanned Aerial Vehicle (UAV) navigation. Such information may be by obtained by estimating…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yara AlaaEldin , Francesca Odone

Embodied AI has been recently gaining attention as it aims to foster the development of autonomous and intelligent agents. In this paper, we devise a novel embodied setting in which an agent needs to explore a previously unknown environment…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Roberto Bigazzi , Federico Landi , Marcella Cornia , Silvia Cascianelli , Lorenzo Baraldi , Rita Cucchiara

Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-language models are…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Minghui Hou , Wei-Hsing Huang , Shaofeng Liang , Daizong Liu , Tai-Hao Wen , Gang Wang , Runwei Guan , Weiping Ding

Ever more robust, accurate and detailed mapping using visual sensing has proven to be an enabling factor for mobile robots across a wide variety of applications. For the next level of robot intelligence and intuitive user interaction, maps…

Computer Vision and Pattern Recognition · Computer Science 2016-09-29 John McCormac , Ankur Handa , Andrew Davison , Stefan Leutenegger

Efficient data utilization is crucial for advancing 3D scene understanding in autonomous driving, where reliance on heavily human-annotated LiDAR point clouds challenges fully supervised methods. Addressing this, our study extends into…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Lingdong Kong , Xiang Xu , Jiawei Ren , Wenwei Zhang , Liang Pan , Kai Chen , Wei Tsang Ooi , Ziwei Liu

Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross-modal information discrepancies and data scarcity, leading to suboptimal alignment spaces that overlook…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Zeyu Chen , Jie Li , Kai Han

Unsupervised contrastive learning for indoor-scene point clouds has achieved great successes. However, unsupervised learning point clouds in outdoor scenes remains challenging because previous methods need to reconstruct the whole scene and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Runjian Chen , Yao Mu , Runsen Xu , Wenqi Shao , Chenhan Jiang , Hang Xu , Zhenguo Li , Ping Luo

Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasoning requires fine-grained evidence from both modalities. A central limitation is that…

Building cross-modal applications is challenging due to limited paired multi-modal data. Recent works have shown that leveraging a pre-trained multi-modal contrastive representation space enables cross-modal tasks to be learned from…

Machine Learning · Computer Science 2024-01-17 Yuhui Zhang , Elaine Sui , Serena Yeung-Levy

Small object detection in unmanned aerial vehicle (UAV) imagery is challenging, mainly due to scale variation, structural detail degradation, and limited computational resources. In high-altitude scenarios, fine-grained features are further…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Xuecheng Bai , Yuxiang Wang , Chuanzhi Xu , Boyu Hu , Kang Han , Ruijie Pan , Xiaowei Niu , Xiaotian Guan , Liqiang Fu , Pengfei Ye

Image-goal navigation steers an agent to a target location specified by an image in unseen environments. Existing methods primarily handle this task by learning an end-to-end navigation policy, which compares the similarities of target and…

Robotics · Computer Science 2026-04-21 Pengna Li , Kangyi Wu , Shaoqing Xu , Fang Li , Lin Zhao , Long Chen , Zhi-Xin Yang , Nanning Zheng

With the novel and fast advances in the area of deep neural networks, several challenging image-based tasks have been recently approached by researchers in pattern recognition and computer vision. In this paper, we address one of these…

Computer Vision and Pattern Recognition · Computer Science 2022-11-11 Jônatas Wehrmann , Anderson Mattjie , Rodrigo C. Barros