中文
相关论文

相关论文: Conditional Information Bottleneck for Multimodal …

200 篇论文

Sarcasm is often expressed through several verbal and non-verbal cues, e.g., a change of tone, overemphasis in a word, a drawn-out syllable, or a straight looking face. Most of the recent work in sarcasm detection has been carried out on…

Learning effective joint embedding for cross-modal data has always been a focus in the field of multimodal machine learning. We argue that during multimodal fusion, the generated multimodal embedding may be redundant, and the discriminative…

机器学习 · 计算机科学 2022-12-06 Sijie Mai , Ying Zeng , Haifeng Hu

Interpreting figurative language such as sarcasm across multi-modal inputs presents unique challenges, often requiring task-specific fine-tuning and extensive reasoning steps. However, current Chain-of-Thought approaches do not efficiently…

计算与语言 · 计算机科学 2025-08-26 Aashish Anantha Ramakrishnan , Aadarsh Anantha Ramakrishnan , Dongwon Lee

Multimodal sarcasm detection (MSD) aims to identify sarcastic intent from semantic incongruity between text and image. Although recent methods have improved MSD through cross-modal interaction and incongruity reasoning, most still treat…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Zhenyu Wang , Weichen Cheng , Weijia Li , Junjie Mou , Zongyou Zhao , Guoying Zhang

Multi-modal sarcasm detection has attracted much recent attention. Nevertheless, the existing benchmark (MMSD) has some shortcomings that hinder the development of reliable multi-modal sarcasm detection system: (1) There are some spurious…

计算与语言 · 计算机科学 2023-07-17 Libo Qin , Shijue Huang , Qiguang Chen , Chenran Cai , Yudi Zhang , Bin Liang , Wanxiang Che , Ruifeng Xu

Multimodal sarcasm detection, which aims to precisely identify pragmatic incongruities between literal text and nonverbal cues, has gained substantial attention in multimodal understanding. Recent advancements have predominantly relied on…

计算与语言 · 计算机科学 2026-05-05 Maoheng Li , Ling Zhou , Xiaohua Huang , Rubing Huang , Wenming Zheng , Guoying Zhao

Multimodal sarcasm detection (MSD) is essential for various downstream tasks. Existing MSD methods tend to rely on spurious correlations. These methods often mistakenly prioritize non-essential features yet still make correct predictions,…

计算与语言 · 计算机科学 2024-12-10 Diandian Guo , Cong Cao , Fangfang Yuan , Yanbing Liu , Guangjie Zeng , Xiaoyan Yu , Hao Peng , Philip S. Yu

Multi-modal semantic understanding requires integrating information from different modalities to extract users' real intention behind words. Most previous work applies a dual-encoder structure to separately encode image and text, but fails…

计算与语言 · 计算机科学 2024-03-12 Ming Zhang , Ke Chang , Yunfang Wu

In the past decade, sarcasm detection has been intensively conducted in a textual scenario. With the popularization of video communication, the analysis in multi-modal scenarios has received much attention in recent years. Therefore,…

计算与语言 · 计算机科学 2021-10-01 Xiaoqiang Zhang , Ying Chen , Guangyuan Li

Multimodal data has significantly advanced recommendation systems by integrating diverse information sources to model user preferences and item characteristics. However, these systems often struggle with redundant and irrelevant…

信息检索 · 计算机科学 2025-09-25 Hui Wang , Jinghui Qin , Wushao Wen , Qingling Li , Shanshan Zhong , Zhongzhan Huang

Effectively leveraging multimodal data such as various images, laboratory tests and clinical information is gaining traction in a variety of AI-based medical diagnosis and prognosis tasks. Most existing multi-modal techniques only focus on…

图像与视频处理 · 电气工程与系统科学 2023-11-28 Yingying Fang , Shuang Wu , Sheng Zhang , Chaoyan Huang , Tieyong Zeng , Xiaodan Xing , Simon Walsh , Guang Yang

Human Multimodal Language Understanding (MLU) aims to infer human intentions by integrating related cues from heterogeneous modalities. Existing works predominantly follow a ``learning to attend" paradigm, which maximizes mutual information…

计算与语言 · 计算机科学 2025-09-29 Menghua Jiang , Yuncheng Jiang , Haifeng Hu , Sijie Mai

Deep multimodal semantic understanding that goes beyond the mere superficial content relation mining has received increasing attention in the realm of artificial intelligence. The challenges of collecting and annotating high-quality…

计算与语言 · 计算机科学 2024-03-26 Zichen Wu , Hsiu-Yuan Huang , Fanyi Qu , Yunfang Wu

Despite progress in multimodal sarcasm detection, existing datasets and methods predominantly focus on single-image scenarios, overlooking potential semantic and affective relations across multiple images. This leaves a gap in modeling…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Haochen Zhao , Yuyao Kong , Yongxiu Xu , Gaopeng Gou , Hongbo Xu , Yubin Wang , Haoliang Zhang

Sarcasm in social media, frequently conveyed through the interplay of text and images, presents significant challenges for sentiment analysis and intention mining. Existing multi-modal sarcasm detection approaches have been shown to…

计算与语言 · 计算机科学 2025-11-14 Junjie Chen , Hang Yu , Subin Huang , Sanmin Liu , Linfeng Zhang

The prevalence of sarcasm in multimodal dialogues on the social platforms presents a crucial yet challenging task for understanding the true intent behind online content. Comprehensive sarcasm analysis requires two key aspects: Multimodal…

计算与语言 · 计算机科学 2026-03-31 Diandian Guo , Fangfang Yuan , Cong Cao , Xixun Lin , Chuan Zhou , Hao Peng , Yanan Cao , Yanbing Liu

While sentiment and emotion analysis have been studied extensively, the relationship between sarcasm and emotion has largely remained unexplored. A sarcastic expression may have a variety of underlying emotions. For example, "I love being…

计算与语言 · 计算机科学 2022-06-07 Anupama Ray , Shubham Mishra , Apoorva Nunna , Pushpak Bhattacharyya

Benefiting from large-scale pretrained vision language models (VLMs), the performance of visual question answering (VQA) has approached human oracles. However, finetuning such models on limited data often suffers from overfitting and poor…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Jingjing Jiang , Ziyi Liu , Nanning Zheng

This paper studies the multimodal named entity recognition (MNER) and multimodal relation extraction (MRE), which are important for multimedia social platform analysis. The core of MNER and MRE lies in incorporating evident visual…

多媒体 · 计算机科学 2024-02-12 Shiyao Cui , Jiangxia Cao , Xin Cong , Jiawei Sheng , Quangang Li , Tingwen Liu , Jinqiao Shi

Multimodal learning is an emerging yet challenging research area. In this paper, we deal with multimodal sarcasm and humor detection from conversational videos and image-text pairs. Being a fleeting action, which is reflected across the…

计算机视觉与模式识别 · 计算机科学 2021-10-22 Shraman Pramanick , Aniket Roy , Vishal M. Patel
‹ 上一页 1 2 3 10 下一页 ›