中文
相关论文

相关论文: UniMPR: A Unified Framework for Multimodal Place R…

200 篇论文

Predicting future trajectories of traffic agents in highly interactive environments is an essential and challenging problem for the safe operation of autonomous driving systems. On the basis of the fact that self-driving vehicles are…

计算机视觉与模式识别 · 计算机科学 2021-06-15 Chiho Choi , Joon Hee Choi , Srikanth Malla , Jiachen Li

Industrial anomaly detection based on RGB-3D multimodal data has emerged as a mainstream paradigm for intelligent quality inspection. However, existing unsupervised methods suffer from two critical limitations: ambiguous cross-modal…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Zewen Li , Shuo Ye , Zitong Yu , Weicheng Xie , Linlin Shen

Multimodal image registration is a fundamental task and a prerequisite for downstream cross-modal analysis. Despite recent progress in shared feature extraction and multi-scale architectures, two key limitations remain. First, some methods…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Chunlei Zhang , Jiahao Xia , Yun Xiao , Bo Jiang , Jian Zhang

Visual Place Recognition (VPR) is a crucial component of 6-DoF localization, visual SLAM and structure-from-motion pipelines, tasked to generate an initial list of place match hypotheses by matching global place descriptors. However,…

计算机视觉与模式识别 · 计算机科学 2022-02-21 Ahmad Khaliq , Michael Milford , Sourav Garg

This paper presents Universal Vision-Language Dense Retrieval (UniVL-DR), which builds a unified model for multi-modal retrieval. UniVL-DR encodes queries and multi-modality resources in an embedding space for searching candidates from…

信息检索 · 计算机科学 2023-02-07 Zhenghao Liu , Chenyan Xiong , Yuanhuiyi Lv , Zhiyuan Liu , Ge Yu

Mesh-based scene representation offers a promising direction for simplifying large-scale hierarchical visual localization pipelines, combining a visual place recognition step based on global features (retrieval) and a visual localization…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Gabriele Berton , Lorenz Junglas , Riccardo Zaccone , Thomas Pollok , Barbara Caputo , Carlo Masone

Universal Multimodal Retrieval (UMR) aims to map different modalities (e.g., visual and textual) into a shared embedding space for multi-modal retrieval. Existing UMR methods can be broadly divided into two categories: early-fusion…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Juan Li , Chuanghao Ding , Xujie Zhang , Cam-Tu Nguyen

Vision Language Place Recognition (VLVPR) enhances robot localization performance by incorporating natural language descriptions from images. By utilizing language information, VLVPR directs robot place matching, overcoming the constraint…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Tianyi Shang , Zhenyu Li , Pengjie Xu , Jinwei Qiao

Image-to-point cloud cross-modal Visual Place Recognition (VPR) is a challenging task where the query is an RGB image, and the database samples are LiDAR point clouds. Compared to single-modal VPR, this approach benefits from the widespread…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Jianyi Peng , Fan Lu , Bin Li , Yuan Huang , Sanqing Qu , Guang Chen

Safe manipulation-oriented navigation for humanoid robots requires scene memory that remains reliable under locomotion-induced perceptual distortion, environmental changes, and interaction-level geometric safety constraints. Existing…

机器人学 · 计算机科学 2026-05-22 Peifeng Jiang , Hong Liu , Jin Jin , Wenshuai Wang , Xia Li

Grid maps are widely established for the representation of static objects in robotics and automotive applications. Though, incorporating velocity information is still widely examined because of the increased complexity of dynamic grids…

Visual place recognition (VPR) is critical in not only localization and mapping for autonomous driving vehicles, but also in assistive navigation for the visually impaired population. To enable a long-term VPR system on a large scale,…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Diwei Sheng , Yuxiang Chai , Xinru Li , Chen Feng , Jianzhe Lin , Claudio Silva , John-Ross Rizzo

Universal multimodal retrieval (UMR), which aims to address complex retrieval tasks where both queries and candidates span diverse modalities, has been significantly advanced by the emergence of MLLMs. While state-of-the-art MLLM-based…

信息检索 · 计算机科学 2026-02-17 Xiaojie Li , Chu Li , Shi-Zhe Chen , Xi Chen

The ability to accurately recognize, localize and separate sound sources is fundamental to any audio-visual perception task. Historically, these abilities were tackled separately, with several methods developed independently for each task.…

声音 · 计算机科学 2023-06-01 Shentong Mo , Pedro Morgado

Multimodal representation learning has demonstrated remarkable potential in enabling models to process and integrate diverse data modalities, such as text and images, for improved understanding and performance. While the medical domain can…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Shuvendu Roy , Franklin Ogidi , Ali Etemad , Elham Dolatabadi , Arash Afkanpour

Image feature matching, a foundational task in computer vision, remains challenging for multimodal image applications, often necessitating intricate training on specific datasets. In this paper, we introduce a Unified Feature Matching…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Yide Di , Yun Liao , Hao Zhou , Kaijun Zhu , Qing Duan , Junhui Liu , Mingyu Lu

In the context of autonomous driving, the significance of effective feature learning is widely acknowledged. While conventional 3D self-supervised pre-training methods have shown widespread success, most methods follow the ideas originally…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Honghui Yang , Sha Zhang , Di Huang , Xiaoyang Wu , Haoyi Zhu , Tong He , Shixiang Tang , Hengshuang Zhao , Qibo Qiu , Binbin Lin , Xiaofei He , Wanli Ouyang

Medical multi-modal pre-training has revealed promise in computer-aided diagnosis by leveraging large-scale unlabeled datasets. However, existing methods based on masked autoencoders mainly rely on data-level reconstruction tasks, but lack…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Yupei Zhang , Li Pan , Qiushi Yang , Tan Li , Zhen Chen

Conventional person re-identification (ReID) research is often limited to single-modality sensor data from static cameras, which fails to address the complexities of real-world scenarios where multi-modal signals are increasingly prevalent.…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Ruiyang Ha , Songyi Jiang , Bin Li , Bikang Pan , Yihang Zhu , Junjie Zhang , Xiatian Zhu , Shaogang Gong , Jingya Wang

Sparse annotations fundamentally constrain multimodal remote sensing: even recent state-of-the-art supervised methods such as MSFMamba are limited by the availability of labeled data, restricting their practical deployment despite…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Yuzhen Hu , Saurabh Prasad
‹ 上一页 1 8 9 10 下一页 ›