中文
相关论文

相关论文: Dynamic Weighted Combiner for Mixed-Modal Image Re…

200 篇论文

Advances in computer vision and deep learning have blurred the line between deepfakes and authentic media, undermining multimedia credibility through audio-visual forgery. Current multimodal detection methods remain limited by unbalanced…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Zihan Xiong , Xiaohua Wu , Lei Chen , Fangqi Lou

Universal Multimodal Retrieval (UMR) aims to enable search across various modalities using a unified model, where queries and candidates can consist of pure text, images, or a combination of both. Previous work has attempted to adopt…

计算与语言 · 计算机科学 2025-04-02 Xin Zhang , Yanzhao Zhang , Wen Xie , Mingxin Li , Ziqi Dai , Dingkun Long , Pengjun Xie , Meishan Zhang , Wenjie Li , Min Zhang

The success of speech-image retrieval relies on establishing an effective alignment between speech and image. Existing methods often model cross-modal interaction through simple cosine similarity of the global feature of each modality,…

计算与语言 · 计算机科学 2024-09-12 Lifeng Zhou , Yuke Li , Rui Deng , Yuting Yang , Haoqi Zhu

Multi-modal object tracking integrates auxiliary modalities such as depth, thermal infrared, event flow, and language to provide additional information beyond RGB images, showing great potential in improving tracking stabilization in…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Shiyu Xuan , Zechao Li , Jinhui Tang

Multi-modal multi-label emotion recognition (MMER) aims to identify relevant emotions from multiple modalities. The challenge of MMER is how to effectively capture discriminative features for multiple labels from heterogeneous data. Recent…

多媒体 · 计算机科学 2024-01-17 Cheng Peng , Ke Chen , Lidan Shou , Gang Chen

To enhance the controllability of text-to-image diffusion models, current ControlNet-like models have explored various control signals to dictate image attributes. However, existing methods either handle conditions inefficiently or use a…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Qingdong He , Jinlong Peng , Pengcheng Xu , Boyuan Jiang , Xiaobin Hu , Donghao Luo , Yong Liu , Yabiao Wang , Chengjie Wang , Xiangtai Li , Jiangning Zhang

Low dose computed tomography (LDCT) is desirable for both diagnostic imaging and image guided interventions. Denoisers are openly used to improve the quality of LDCT. Deep learning (DL)-based denoisers have shown state-of-the-art…

图像与视频处理 · 电气工程与系统科学 2020-12-08 Ti Bai , Biling Wang , Dan Nguyen , Bao Wang , Bin Dong , Wenxiang Cong , Mannudeep K. Kalra , Steve Jiang

Multi-modality image fusion aims to combine different modalities to produce fused images that retain the complementary features of each modality, such as functional highlights and texture details. To leverage strong generative priors and…

计算机视觉与模式识别 · 计算机科学 2023-08-24 Zixiang Zhao , Haowen Bai , Yuanzhi Zhu , Jiangshe Zhang , Shuang Xu , Yulun Zhang , Kai Zhang , Deyu Meng , Radu Timofte , Luc Van Gool

With the increasing multimedia information, multimodal recommendation has received extensive attention. It utilizes multimodal information to alleviate the data sparsity problem in recommendation systems, thus improving recommendation…

信息检索 · 计算机科学 2024-03-01 Jinfeng Xu , Zheyu Chen , Shuo Yang , Jinze Li , Hewei Wang , Edith C. -H. Ngai

Learning multi-view data is an emerging problem in machine learning research, and nonnegative matrix factorization (NMF) is a popular dimensionality-reduction method for integrating information from multiple views. These views often provide…

机器学习 · 统计学 2023-04-26 Shuo Shuo Liu , Lin Lin

Multimodal Sentiment Analysis (MSA) aims to predict sentiment from language, acoustic, and visual data in videos. However, imbalanced unimodal performance often leads to suboptimal fused representations. Existing approaches typically adopt…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Dingkang Yang , Mingcheng Li , Xuecheng Wu , Zhaoyu Chen , Kaixun Jiang , Keliang Liu , Peng Zhai , Lihua Zhang

The task of generating natural language descriptions from images has received a lot of attention in recent years. Consequently, it is becoming increasingly important to evaluate such image captioning approaches in an automatic manner. In…

计算与语言 · 计算机科学 2016-12-23 Mert Kilickaya , Aykut Erdem , Nazli Ikizler-Cinbis , Erkut Erdem

As a fundamental and challenging task in bridging language and vision domains, Image-Text Retrieval (ITR) aims at searching for the target instances that are semantically relevant to the given query from the other modality, and its key…

计算机视觉与模式识别 · 计算机科学 2023-01-18 Yan Zhang , Zhong Ji , Di Wang , Yanwei Pang , Xuelong Li

Multimodal pathological images are usually in clinical diagnosis, but computer vision-based multimodal image-assisted diagnosis faces challenges with modality fusion, especially in the absence of expert-annotated data. To achieve the…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Qinghua Lin , Guang-Hai Liu , Zuoyong Li , Yang Li , Yuting Jiang , Xiang Wu

Multi-modal creative writing (MMCW) aims to produce illustrated articles. Unlike common multi-modal generative (MMG) tasks such as storytelling or caption generation, MMCW is an entirely new and more abstract challenge where textual and…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Jiahao Chen , Zhiyong Ma , Wenbiao Du , Qingyuan Chuai

We present a multiway fusion algorithm capable of directly processing uncertain pairwise affinities. In contrast to existing works that require initial pairwise associations, our MIXER algorithm improves accuracy by leveraging the…

计算机视觉与模式识别 · 计算机科学 2023-05-04 Parker C. Lusk , Kaveh Fathian , Jonathan P. How

The rapid advancement of unsupervised representation learning and large-scale pre-trained vision-language models has significantly improved cross-modal retrieval tasks. However, existing multi-modal information retrieval (MMIR) studies lack…

信息检索 · 计算机科学 2025-10-20 Zirui Li , Siwei Wu , Yizhi Li , Xingyu Wang , Yi Zhou , Chenghua Lin

In recent years, Multimodal Emotion Recognition (MER) has made substantial progress. Nevertheless, most existing approaches neglect the semantic inconsistencies that may arise across modalities, such as conflicting emotional cues between…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Guowei Zhong , Junjie Li , Huaiyu Zhu , Ruohong Huan , Yun Pan

Recently, prompt learning has demonstrated remarkable success in adapting pre-trained Vision-Language Models (VLMs) to various downstream tasks such as image classification. However, its application to the downstream Image-Text Retrieval…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Yifan Wang , Tao Wang , Chenwei Tang , Caiyang Yu , Zhengqing Zang , Mengmi Zhang , Shudong Huang , Jiancheng Lv

Audio-text retrieval enables semantic alignment between audio content and natural language queries, supporting applications in multimedia search, accessibility, and surveillance. However, current state-of-the-art approaches struggle with…