中文
相关论文

相关论文: Adaptive Context Matters: Towards Provable Multi-M…

200 篇论文

Face Anti-Spoofing (FAS) is crucial for securing face recognition systems against presentation attacks. With advancements in sensor manufacture and multi-modal learning techniques, many multi-modal FAS approaches have emerged. However, they…

计算机视觉与模式识别 · 计算机科学 2024-03-06 Xun Lin , Shuai Wang , Rizhao Cai , Yizhong Liu , Ying Fu , Zitong Yu , Wenzhong Tang , Alex Kot

The recent Segment Anything Model (SAM) represents a significant breakthrough in scaling segmentation models, delivering strong performance across various downstream applications in the RGB modality. However, directly applying SAM to…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Chenyang Zhu , Bin Xiao , Lin Shi , Shoukun Xu , Xu Zheng

Multimodal learning leverages complementary information derived from different modalities, thereby enhancing performance in medical image segmentation. However, prevailing multimodal learning methods heavily rely on extensive well-annotated…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Xiaogen Zhou , Yiyou Sun , Min Deng , Winnie Chiu Wing Chu , Qi Dou

Multimodal emotion recognition (MER) benefits from combining text, audio, and vision, yet standard fusion often fails when modalities conflict. Crucially, conflicts differ in resolvability: benign conflicts stem from missing, weak, or…

多媒体 · 计算机科学 2026-05-07 Yangchen Yu , Qian Chen , Jia Li , Zhenzhen Hu , Jinpeng Hu , Lizi Liao , Erik Cambria , Richang Hong

In this paper we propose a vision system that performs image Super Resolution (SR) with selectivity. Conventional SR techniques, either by multi-image fusion or example-based construction, have failed to capitalize on the intrinsic…

计算机视觉与模式识别 · 计算机科学 2010-10-28 Ju Sun , Qiang Chen , Shuicheng Yan , Loong-Fah Cheong

Continuous dimensional speech emotion recognition captures affective variation along valence, arousal, and dominance, providing finer-grained representations than categorical approaches. Yet most multimodal methods rely solely on global…

声音 · 计算机科学 2026-01-27 Haoxun Li , Yuqing Sun , Hanlei Shi , Yu Liu , Leyuan Qu , Taihao Li

Video captioning is a challenging task that necessitates a thorough comprehension of visual scenes. Existing methods follow a typical one-to-one mapping, which concentrates on a limited sample space while ignoring the intrinsic semantic…

计算机视觉与模式识别 · 计算机科学 2022-05-20 Xiaoya Chen , Jingkuan Song , Pengpeng Zeng , Lianli Gao , Heng Tao Shen

Modern recommender systems face critical challenges in handling information overload while addressing the inherent limitations of multimodal representation learning. Existing methods suffer from three fundamental limitations: (1) restricted…

信息检索 · 计算机科学 2025-08-15 Zheyu Chen , Jinfeng Xu , Hewei Wang , Shuo Yang , Zitong Wan , Haibo Hu

Emotion Recognition in Conversations (ERC) presents unique challenges, requiring models to capture the temporal flow of multi-turn dialogues and to effectively integrate cues from multiple modalities. We propose Mixture of Speech-Text…

计算与语言 · 计算机科学 2026-02-27 Soumya Dutta , Smruthi Balaji , Sriram Ganapathy

Recent advancements in multimodal large language models (MLLMs) have demonstrated considerable potential for comprehensive 3D scene understanding. However, existing approaches typically utilize only one or a limited subset of 3D modalities,…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Yue Zhang , Yingzhao Jian , Hehe Fan , Yi Yang , Roger Zimmermann

Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant representations.…

声音 · 计算机科学 2023-02-03 Chen Chen , Yuchen Hu , Qiang Zhang , Heqing Zou , Beier Zhu , Eng Siong Chng

Multimodal Large Language Models demonstrate strong performance on multimodal benchmarks, yet often exhibit poor robustness when exposed to spurious modality interference, such as irrelevant text in vision understanding, or irrelevant…

机器学习 · 计算机科学 2026-01-30 Rui Cai , Bangzheng Li , Xiaofei Wen , Muhao Chen , Zhe Zhao

Foundational feed-forward visual geometry models enable accurate and efficient camera pose estimation and scene reconstruction by learning strong scene priors from massive RGB datasets. However, their effectiveness drops when applied to…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Vsevolod Skorokhodov , Chenghao Xu , Shuo Sun , Olga Fink , Malcolm Mielle

High-order tensor methods that employ Taylor-based local models (of degree $p\ge 3$) within adaptive regularization frameworks have been recently proposed for both convex and nonconvex optimization problems. They have been shown to have…

最优化与控制 · 数学 2024-04-19 Wenqi Zhu , Coralia Cartis

A non-linear system governed by multi-spatial and multi-temporal physics scales cannot be fully understood with a single diagnostic, as each provides only a partial view, leading to information loss. Combining multiple diagnostics may also…

Multimodal fusion is often treated as an optimization-balancing problem, where training signals are adjusted to prevent one modality from dominating the others. However, balanced optimization does not fully determine the geometry of…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Zixuan Xia , Hao Wang , Pengcheng Weng , Yanyu Qian , Yangxin Xu , William Dan , Fei Wang

Recently, we introduced Relative Resolution as a hybrid formalism for fluid mixtures [1]. The essence of this approach is that it switches molecular resolution in terms or relative separation: While nearest neighbors are characterized by a…

统计力学 · 物理学 2019-10-09 Aviel Chaimovich , Kurt Kremer , Christine Peter

Assuming a known degradation model, the performance of a learned image super-resolution (SR) model depends on how well the variety of image characteristics within the training set matches those in the test set. As a result, the performance…

计算机视觉与模式识别 · 计算机科学 2022-09-20 Cansu Korkmaz , A. Murat Tekalp , Zafer Dogan

Visual place recognition (VPR) remains challenging due to significant viewpoint changes and appearance variations. Mainstream works tackle these challenges by developing various feature aggregation methods to transform deep features into…

计算机视觉与模式识别 · 计算机科学 2024-07-10 Teng Wang , Lingquan Meng , Lei Cheng , Changyin Sun

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual inputs, yet modality alignment remains one of the most challenging aspects. Current MLLMs typically rely on simple adapter…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Yuanyang Yin , Yaqi Zhao , Yajie Zhang , Yuanxing Zhang , Ke Lin , Jiahao Wang , Xin Tao , Pengfei Wan , Wentao Zhang , Feng Zhao