English
Related papers

Related papers: DLF: Disentangled-Language-Focused Multimodal Sent…

200 papers

Ensuring fairness across demographic groups in medical diagnosis is essential for equitable healthcare, particularly under distribution shifts caused by variations in imaging equipment and clinical practice. Vision-language models (VLMs)…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Yuexuan Xia , Benteng Ma , Jiang He , Zhiyong Wang , Qi Dou , Yong Xia

Federated learning (FL) is severely challenged by non-independent and identically distributed (non-IID) client data, a problem that degrades global model performance, especially in multimodal perception settings. Conventional methods often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Jing Liu , Zhengliang Guo , Yan Wang , Xiaoguang Zhu , Yao Du , Zehua Wang , Victor C. M. Leung

Disentangled Representation Learning aims to improve the explainability of deep learning methods by training a data encoder that identifies semantically meaningful latent variables in the data generation process. Nevertheless, there is no…

Machine Learning · Computer Science 2024-10-08 Ruoyu Wang , Lina Yao

Multi-modal emotion recognition in conversations is a challenging problem due to the complex and complementary interactions between different modalities. Audio and textual cues are particularly important for understanding emotions from a…

Sound · Computer Science 2025-04-02 Jiachen Luo , Huy Phan , Lin Wang , Joshua Reiss

People usually have different intents for choosing items, while their preferences under the same intent may also different. In traditional collaborative filtering approaches, both intent and preference factors are usually entangled in the…

Information Retrieval · Computer Science 2023-05-19 Chao Wang , Hengshu Zhu , Dazhong Shen , Wei wu , Hui Xiong

Existing deepfake analysis methods are primarily based on discriminative models, which significantly limit their application scenarios. This paper aims to explore interactive deepfake analysis by performing instruction tuning on multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Lixiong Qin , Ning Jiang , Yang Zhang , Yuhan Qiu , Dingheng Zeng , Jiani Hu , Weihong Deng

Cross-modal retrieval is crucial in understanding latent correspondences across modalities. However, existing methods implicitly assume well-matched training data, which is impractical as real-world data inevitably involves imperfect…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Zhuohang Dang , Minnan Luo , Jihong Wang , Chengyou Jia , Haochen Han , Herun Wan , Guang Dai , Xiaojun Chang , Jingdong Wang

Disentangled representation learning in speech processing has lagged behind other domains, largely due to the lack of datasets with annotated generative factors for robust evaluation. To address this, we propose SynSpeech, a novel…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Yusuf Brima , Ulf Krumnack , Simone Pika , Gunther Heidemann

Multimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However, in real-world applications, some practical factors cause…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Mingcheng Li , Dingkang Yang , Xiao Zhao , Shuaibing Wang , Yan Wang , Kun Yang , Mingyang Sun , Dongliang Kou , Ziyun Qian , Lihua Zhang

Sequential Recommender Systems (SRS) aim to predict users' next interaction based on their historical behaviors, while still facing the challenge of data sparsity. With the rapid advancement of Multimodal Large Language Models (MLLMs),…

Information Retrieval · Computer Science 2026-02-17 Mingyao Huang , Qidong Liu , Wenxuan Yang , Moranxin Wang , Yuqi Sun , Haiping Zhu , Feng Tian , Yan Chen

Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating information from multiple modalities, such as text, audio, and visual data. However, existing methods often suffer from spurious correlations both within…

Machine Learning · Computer Science 2026-05-21 Menghua Jiang , Yuxia Lin , Baoliang Chen , Haifeng Hu , Yuncheng Jiang , Sijie Mai

Disentangled representation learning aims to capture the underlying explanatory factors of observed data, enabling a principled understanding of the data-generating process. Recent advances in generative modeling have introduced new…

Machine Learning · Computer Science 2026-05-12 Jinjin Chi , Taoping Liu , Mengtao Yin , Ximing Li , Yongcheng Jing , Jialie Shen , Leszek Rutkowski , Dacheng Tao

Multimodal Affective Computing (MAC) aims to recognize and interpret human emotions by integrating information from diverse modalities such as text, video, and audio. Recent advancements in Multimodal Large Language Models (MLLMs) have…

Artificial Intelligence · Computer Science 2025-08-05 Miaosen Luo , Jiesen Long , Zequn Li , Yunying Yang , Yuncheng Jiang , Sijie Mai

Due to the rapid advancements of sensory and computing technology, multi-modal data sources that represent the same pattern or phenomenon have attracted growing attention. As a result, finding means to explore useful information from these…

Machine Learning · Computer Science 2021-03-10 Lei Gao , Ling Guan

Multi-modal MRIs are widely used in neuroimaging applications since different MR sequences provide complementary information about brain structures. Recent works have suggested that multi-modal deep learning analysis can benefit from…

Computer Vision and Pattern Recognition · Computer Science 2021-06-14 Jiahong Ouyang , Ehsan Adeli , Kilian M. Pohl , Qingyu Zhao , Greg Zaharchuk

This paper introduces a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of LLMs. By decomposing the evaluation along the dimensions of linguistic medium and cultural context, this framework enables a…

Computation and Language · Computer Science 2025-06-02 Jiahao Ying , Wei Tang , Yiran Zhao , Yixin Cao , Yu Rong , Wenxuan Zhang

Scene text images contain not only style information (font, background) but also content information (character, texture). Different scene text tasks need different information, but previous representation learning methods use tightly…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Boqiang Zhang , Hongtao Xie , Zuan Gao , Yuxin Wang

Individual neurons participate in the representation of multiple high-level concepts. To what extent can different interpretability methods successfully disentangle these roles? To help address this question, we introduce RAVEL (Resolving…

Computation and Language · Computer Science 2024-08-28 Jing Huang , Zhengxuan Wu , Christopher Potts , Mor Geva , Atticus Geiger

As multi-agent AI systems evolve from simple chatbots to autonomous swarms, debugging semantic failures requires reasoning about knowledge, belief, causality, and obligation, precisely what modal logic was designed to formalize. However,…

Artificial Intelligence · Computer Science 2026-02-13 Antonin Sulc

Multimodal large language models (MLLMs) have gained significant attention due to their impressive ability to integrate vision and language modalities. Recent advancements in MLLMs have primarily focused on improving performance through…

Computation and Language · Computer Science 2025-09-19 Chenkun Tan , Pengyu Wang , Shaojun Zhou , Botian Jiang , Zhaowei Li , Dong Zhang , Xinghao Wang , Yaqian Zhou , Xipeng Qiu
‹ Prev 1 4 5 6 7 8 10 Next ›