中文
相关论文

相关论文: VCR: Learning Valid Contextual Representation for …

200 篇论文

Visual Place Recognition (VPR) enables robust localization through image retrieval based on learned descriptors. However, drastic appearance variations of images at the same place caused by viewpoint changes can lead to inconsistent…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Qiwen Gu , Xufei Wang , Junqiao Zhao , Siyue Tao , Tiantian Feng , Ziqiao Wang , Guang Chen

Concept-based models aim to explain model decisions with human-understandable concepts. However, most existing approaches treat concepts as numerical attributes, without providing complementary visual explanations that could localize the…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Cristiano Patrício , Luís F. Teixeira , João C. Neves

High-precision medical diagnosis relies not only on static imaging features but also on the implicit diagnostic memory experts instantly invoke during image interpretation. We pinpoint a fundamental cognitive misalignment in medical VLMs…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Chunzheng Zhu , Jiaqi Zeng , Junyu Jiang , Jianxin Lin , Yijun Wang

Learning common subspace is prevalent way in cross-modal retrieval to solve the problem of data from different modalities having inconsistent distributions and representations that cannot be directly compared. Previous cross-modal retrieval…

多媒体 · 计算机科学 2021-10-27 Donghuo Zeng , Jianming Wu , Gen Hattori , Yi Yu , Rong Xu

Audio-visual learning has been a major pillar of multi-modal machine learning, where the community mostly focused on its modality-aligned setting, i.e., the audio and visual modality are both assumed to signal the prediction target. With…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Yung-Hsuan Lai , Yen-Chun Chen , Yu-Chiang Frank Wang

Visible-infrared cross-modality person re-identification is a challenging ReID task, which aims to retrieve and match the same identity's images between the heterogeneous visible and infrared modalities. Thus, the core of this task is to…

计算机视觉与模式识别 · 计算机科学 2023-05-24 Tengfei Liang , Yi Jin , Yajun Gao , Wu Liu , Songhe Feng , Tao Wang , Yidong Li

Cross-modal contrastive distillation has recently been explored for learning effective 3D representations. However, existing methods focus primarily on modality-shared features, neglecting the modality-specific features during the…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Yifan Zhang , Junhui Hou

Sport-related concussion (SRC) depends on sensory information from visual, vestibular, and somatosensory systems. At the same time, the current clinical administration of Vestibular/Ocular Motor Screening (VOMS) is subjective and deviates…

图像与视频处理 · 电气工程与系统科学 2022-10-18 Khondker Fariha Hossain , Sharif Amit Kamran , Prithul Sarker , Philip Pavilionis , Isayas Adhanom , Nicholas Murray , Alireza Tavakkoli

Multimodal large language models (MLLMs) trained with visual instruction tuning have achieved strong performance across diverse tasks, yet they remain limited in vision-centric tasks such as object counting or spatial reasoning. We…

Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using only coarse-grained labels. Faced with the challenging task settings, existing research…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Huilai Li , Xiaomeng Di , Ying Xing , Yonghao Dang , Yiming Wang , Jianqin Yin

Unsupervised deformable image registration requires aligning complex anatomical structures without reference labels, making interpretability and reliability critical. Existing deep learning methods achieve considerable accuracy but often…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Zafar Iqbal , Anwar Ul Haq , Srimannarayana Grandhi

Radio astronomy is an indispensable discipline for observing distant celestial objects. Measurements of wave signals from radio telescopes, called visibility, need to be transformed into images for astronomical observations. These dirty…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Kai Cheng , Ruoqi Wang , Qiong Luo

Visual Place Recognition (VPR) systems often have imperfect performance, affecting the `integrity' of position estimates and subsequent robot navigation decisions. Previously, SVM classifiers have been used to monitor VPR integrity. This…

计算机视觉与模式识别 · 计算机科学 2024-11-20 Owen Claxton , Connor Malone , Helen Carson , Jason Ford , Gabe Bolton , Iman Shames , Michael Milford

Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Suyang Xi , Songtao Hu , Yuxiang Lai , Wangyun Dan , Yaqi Liu , Shansong Wang , Xiaofeng Yang

Visible Light Communication~(VLC) systems provide not only illumination and data communication, but also indoor monitoring services if the effect that different events create on the received optical signal is properly tracked. For this…

信号处理 · 电气工程与系统科学 2021-01-27 Mehmet C. Ilter , Alexis A. Dowhuszko , Jyri Hämäläinen , Risto Wichman

Representing a signal as a continuous function parameterized by neural network (a.k.a. Implicit Neural Representations, INRs) has attracted increasing attention in recent years. Neural Processes (NPs), which model the distributions over…

机器学习 · 计算机科学 2023-02-22 Zongyu Guo , Cuiling Lan , Zhizheng Zhang , Yan Lu , Zhibo Chen

Cross-modal retrieval (CMR) typically involves learning common representations to directly measure similarities between multimodal samples. Most existing CMR methods commonly assume multimodal samples in pairs and employ joint training to…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Ruitao Pu , Yang Qin , Dezhong Peng , Xiaomin Song , Huiming Zheng

Vision-language models (VLMs) allow to embed texts and images in a shared representation space. However, it has been shown that these models are subject to a modality gap phenomenon meaning there exists a clear separation between the…

计算机视觉与模式识别 · 计算机科学 2025-05-07 François Role , Sébastien Meyer , Victor Amblard

One of the most critical aspects of multimodal Reinforcement Learning (RL) is the effective integration of different observation modalities. Having robust and accurate representations derived from these modalities is key to enhancing the…

机器人学 · 计算机科学 2024-06-21 Fotios Lygerakis , Vedant Dave , Elmar Rueckert

Monocular 3D object detection poses a significant challenge in 3D scene understanding due to its inherently ill-posed nature in monocular depth estimation. Existing methods heavily rely on supervised learning using abundant 3D labels,…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Zihua Liu , Hiroki Sakuma , Masatoshi Okutomi