中文
相关论文

相关论文: Modality-missing RGBT Tracking: Invertible Prompt …

200 篇论文

We explore a new language model inversion problem under strict black-box, zero-shot, and limited data conditions. We propose a novel training-free framework that reconstructs prompts using only a limited number of text outputs from a…

计算与语言 · 计算机科学 2025-02-18 Hanqing Li , Diego Klabjan

This paper introduces a generalized federated prompt-tuning framework for practical scenarios where local datasets are multi-modal and exhibit different distributional patterns of missing features at the input level. The proposed framework…

Although data-free incremental learning methods are memory-friendly, accurately estimating and counteracting representation shifts is challenging in the absence of historical data. This paper addresses this thorny problem by proposing a…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Zhiheng Ma , Xiaopeng Hong , Beinan Liu , Yabin Wang , Pinyue Guo , Huiyun Li

Addressing missing modalities and limited labeled data is crucial for advancing robust multimodal learning. We propose Robult, a scalable framework designed to mitigate these challenges by preserving modality-specific information and…

机器学习 · 计算机科学 2025-09-26 Duy A. Nguyen , Abhi Kamboj , Minh N. Do

Multimodal perception systems for robotics and embodied AI often assume reliable RGB-D sensing, but in practice, depth is frequently missing, noisy, or corrupted. We thus present GeomPrompt, a lightweight cross-modal adaptation module that…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Krishna Jaganathan , Patricio Vela

In autonomous driving, environment perception has significantly advanced with the utilization of deep learning techniques for diverse sensors such as cameras, depth sensors, or infrared sensors. The diversity in the sensor stack increases…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Niharika Hegde , Shishir Muralidhara , René Schuster , Didier Stricker

Multimodal learning has recently gained significant popularity, demonstrating impressive performance across various zero-shot classification tasks and a range of perceptive and generative applications. Models such as Contrastive…

机器学习 · 计算机科学 2026-02-16 Can Yaras , Siyi Chen , Peng Wang , Qing Qu

Multi-object tracking and segmentation (MOTS) is a critical task for autonomous driving applications. The existing MOTS studies face two critical challenges: 1) the published datasets inadequately capture the real-world complexity for…

计算机视觉与模式识别 · 计算机科学 2021-10-18 Yiming Cui , Zhiwen Cao , Yixin Xie , Xingyu Jiang , Feng Tao , Yingjie Chen , Lin Li , Dongfang Liu

Visible-infrared cross-modality person re-identification is a challenging ReID task, which aims to retrieve and match the same identity's images between the heterogeneous visible and infrared modalities. Thus, the core of this task is to…

计算机视觉与模式识别 · 计算机科学 2023-05-24 Tengfei Liang , Yi Jin , Yajun Gao , Wu Liu , Songhe Feng , Tao Wang , Yidong Li

We propose an end-to-end tracking framework for fusing the RGB and TIR modalities in RGB-T tracking. Our baseline tracker is DiMP (Discriminative Model Prediction), which employs a carefully designed target prediction network trained…

计算机视觉与模式识别 · 计算机科学 2019-09-02 Lichao Zhang , Martin Danelljan , Abel Gonzalez-Garcia , Joost van de Weijer , Fahad Shahbaz Khan

Learning based on multimodal data has attracted increasing interest recently. While a variety of sensory modalities can be collected for training, not all of them are always available in development scenarios, which raises the challenge to…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Shicai Wei , Yang Luo , Chunbo Luo

Visual prompting, an efficient method for transfer learning, has shown its potential in vision tasks. However, previous works focus exclusively on VP from standard source models, it is still unknown how it performs under the scenario of a…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Qi Li , Liangzhi Li , Zhouqiang Jiang , Bowen Wang

Multimodal learning, which integrates diverse data sources such as images, text, and structured data, has proven superior to unimodal counterparts in high-stakes decision-making. However, while performance gains remain the gold standard for…

人工智能 · 计算机科学 2025-05-07 Kishore Sampath , Pratheesh , Ayaazuddin Mohammad , Resmi Ramachandranpillai

The effectiveness of Multimodal Chain-of-Thought (MCoT) prompting is often limited by the use of randomly or manually selected examples. These examples fail to account for both model-specific knowledge distributions and the intrinsic…

计算与语言 · 计算机科学 2025-10-14 Xinglong Yang , Quan Feng , Zhongying Pan , Xiang Chen , Yu Tian , Wentong Li , Shuofei Qiao , Yuxia Geng , Xingyu Zhao , Sheng-Jun Huang

Multimedia-based recommendation provides personalized item suggestions by learning the content preferences of users. With the proliferation of digital devices and APPs, a huge number of new items are created rapidly over time. How to…

信息检索 · 计算机科学 2024-05-28 Haoyue Bai , Le Wu , Min Hou , Miaomiao Cai , Zhuangzhuang He , Yuyang Zhou , Richang Hong , Meng Wang

Continual Visual Question Answering (CVQA) based on pre-trained models(PTMs) has achieved promising progress by leveraging prompt tuning to enable continual multi-modal learning. However, most existing methods adopt cross-modal prompt…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Xu Li , Fan Lyu

Multi-modal learning has made significant advances across diverse pattern recognition applications. However, handling missing modalities, especially under imbalanced missing rates, remains a major challenge. This imbalance triggers a…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Binyu Zhao , Wei Zhang , Zhaonian Zou

Semantic segmentation across arbitrary sensor modalities faces significant challenges due to diverse sensor characteristics, and the traditional configurations for this task result in redundant development efforts. We address these…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Jiong Liu , Yingjie Xu , Xingcheng Zhou , Rui Song , Walter Zimmer , Alois Knoll , Hu Cao

In recent years, the deployment of large-scale pre-trained models in audio-visual downstream tasks has yielded remarkable outcomes. However, these models, primarily trained on single-modality unconstrained datasets, still encounter…

机器学习 · 计算机科学 2023-12-22 Haoyi Duan , Yan Xia , Mingze Zhou , Li Tang , Jieming Zhu , Zhou Zhao

Standard multi-modal models assume the use of the same modalities in training and inference stages. However, in practice, the environment in which multi-modal models operate may not satisfy such assumption. As such, their performances…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Sangmin Woo , Sumin Lee , Yeonju Park , Muhammad Adi Nugroho , Changick Kim