English
Related papers

Related papers: Conditional Information Bottleneck for Multimodal …

200 papers

Existing multimodal methods typically assume that different modalities share the same category set. However, in real-world applications, the category distributions in multimodal data exhibit inconsistencies, which can hinder the model's…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yangrui Zhu , Junhua Bao , Yipan Wei , Yapeng Li , Bo Du

Link prediction aims to identify potential missing triples in knowledge graphs. To get better results, some recent studies have introduced multimodal information to link prediction. However, these methods utilize multimodal information…

Artificial Intelligence · Computer Science 2023-03-21 Xinhang Li , Xiangyu Zhao , Jiaxing Xu , Yong Zhang , Chunxiao Xing

The literature in automated sarcasm detection has mainly focused on lexical, syntactic and semantic-level analysis of text. However, a sarcastic sentence can be expressed with contextual presumptions, background and commonsense knowledge.…

Computation and Language · Computer Science 2018-05-17 Devamanyu Hazarika , Soujanya Poria , Sruthi Gorantla , Erik Cambria , Roger Zimmermann , Rada Mihalcea

In recent years, speech emotion recognition technology is of great significance in industrial applications such as call centers, social robots and health care. The combination of speech recognition and speech emotion recognition can improve…

Artificial Intelligence · Computer Science 2023-02-20 Ying Zhou , Xuefeng Liang , Yu Gu , Yifei Yin , Longshan Yao

Existing rumor detection methods often neglect the content within images as well as the inherent relationships between contexts and images across different visual scales, thereby resulting in the loss of critical information pertinent to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Bin Ma , Yifei Zhang , Yongjin Xian , Qi Li , Linna Zhou , Gongxun Miao

Multi-modal stance detection (MSD) aims to determine an author's stance toward a given target using both textual and visual content. While recent methods leverage multi-modal fusion and prompt-based learning, most fail to distinguish…

Multimedia · Computer Science 2026-01-30 Zhiyu Xie , Fuqiang Niu , Genan Dai , Qianlong Wang , Li Dong , Bowen Zhang , Hu Huang

Automatically recognising apparent emotions from face and voice is hard, in part because of various sources of uncertainty, including in the input data and the labels used in a machine learning framework. This paper introduces an…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Mani Kumar Tellamekala , Shahin Amiriparian , Björn W. Schuller , Elisabeth André , Timo Giesbrecht , Michel Valstar

Learning invariant (causal) features for out-of-distribution (OOD) generalization has attracted extensive attention recently, and among the proposals invariant risk minimization (IRM) is a notable solution. In spite of its theoretical…

Machine Learning · Computer Science 2023-02-01 Bin Deng , Kui Jia

Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based multimodal fusion methods ignore the inter-modality…

Computer Vision and Pattern Recognition · Computer Science 2023-05-17 Heqing Zou , Meng Shen , Chen Chen , Yuchen Hu , Deepu Rajan , Eng Siong Chng

We address the challenge of detecting questionable content in online media, specifically the subcategory of comic mischief. This type of content combines elements such as violence, adult content, or sarcasm with humor, making it difficult…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Elaheh Baharlouei , Mahsa Shafaei , Yigeng Zhang , Hugo Jair Escalante , Thamar Solorio

Multimodal fusion is considered a key step in multimodal tasks such as sentiment analysis, emotion detection, question answering, and others. Most of the recent work on multimodal fusion does not guarantee the fidelity of the multimodal…

Machine Learning · Computer Science 2019-08-19 Navonil Majumder , Soujanya Poria , Gangeshwar Krishnamurthy , Niyati Chhaya , Rada Mihalcea , Alexander Gelbukh

Concept Bottleneck Models (CBMs) assume that training examples (e.g., x-ray images) are annotated with high-level concepts (e.g., types of abnormalities), and perform classification by first predicting the concepts, followed by predicting…

Computation and Language · Computer Science 2023-12-19 Danis Alukaev , Semen Kiselev , Ilya Pershin , Bulat Ibragimov , Vladimir Ivanov , Alexey Kornaev , Ivan Titov

The presence of sarcasm in conversational systems and social media like chatbots, Facebook, Twitter, etc. poses several challenges for downstream NLP tasks. This is attributed to the fact that the intended meaning of a sarcastic text is…

Computation and Language · Computer Science 2022-02-08 Aditya Shah , Chandresh Kumar Maurya

Multimodal deception detection aims to identify deceptive behavior by analyzing audiovisual cues for forensics and security. In these high-stakes settings, investigators need verifiable evidence connecting audiovisual cues to final…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Jiajian Huang , Dongliang Zhu , Zitong YU , Hui Ma , Jiayu Zhang , Chunmei Zhu , Xiaochun Cao

The Information Bottleneck (IB) principle has emerged as a promising approach for enhancing the generalization, robustness, and interpretability of deep neural networks, demonstrating efficacy across image segmentation, document clustering,…

Information Theory · Computer Science 2025-04-18 Hanzhe Yang , Youlong Wu , Dingzhu Wen , Yong Zhou , Yuanming Shi

As multimodal misinformation becomes more sophisticated, its detection and grounding are crucial. However, current multimodal verification methods, relying on passive holistic fusion, struggle with sophisticated misinformation. Due to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Zizhao Chen , Ping Wei , Ziyang Ren , Huan Li , Xiangru Yin

Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to obtain unimodal features, and risk being too costly or…

Sound · Computer Science 2025-07-11 Sidong Zhang , Shiv Shankar , Trang Nguyen , Andrea Fanelli , Madalina Fiterau

Multimodal video understanding plays a crucial role in tasks such as action recognition and emotion classification by combining information from different modalities. However, multimodal models are prone to overfitting strong modalities,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Xiaoyu Ma , Ding Ding , Hao Chen

The contextual multi-armed bandit (MAB) problem is crucial in sequential decision-making. A line of research, known as online clustering of bandits, extends contextual MAB by grouping similar users into clusters, utilizing shared features…

Machine Learning · Computer Science 2025-01-03 Zhuohua Li , Maoli Liu , Xiangxiang Dai , John C. S. Lui

Multimodal learning is an essential paradigm for addressing complex real-world problems, where individual data modalities are typically insufficient to accurately solve a given modelling task. While various deep learning approaches have…

Machine Learning · Computer Science 2023-07-04 Gabriele Dominici , Pietro Barbiero , Lucie Charlotte Magister , Pietro Liò , Nikola Simidjievski