中文
相关论文

相关论文: Multi-Modal Semantic Inconsistency Detection in So…

200 篇论文

We propose detection of deepfake videos based on the dissimilarity between the audio and visual modalities, termed as the Modality Dissonance Score (MDS). We hypothesize that manipulation of either modality will lead to dis-harmony between…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Komal Chugh , Parul Gupta , Abhinav Dhall , Ramanathan Subramanian

In the past few years, there has been a surge of interest in multi-modal problems, from image captioning to visual question answering and beyond. In this paper, we focus on hate speech detection in multi-modal memes wherein memes pose an…

计算机视觉与模式识别 · 计算机科学 2021-01-01 Abhishek Das , Japsimar Singh Wahi , Siyao Li

News image captioning aims to produce journalistically informative descriptions by combining visual content with contextual cues from associated articles. Despite recent advances, existing methods struggle with three key challenges: (1)…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Xiaoxing You , Qiang Huang , Lingyu Li , Chi Zhang , Xiaopeng Liu , Min Zhang , Jun Yu

Large-scale dissemination of disinformation online intended to mislead or deceive the general population is a major societal problem. Rapid progression in image, video, and natural language generative models has only exacerbated this…

人工智能 · 计算机科学 2022-05-27 Reuben Tan , Bryan A. Plummer , Kate Saenko

Indoor scene recognition is a growing field with great potential for behaviour understanding, robot localization, and elderly monitoring, among others. In this study, we approach the task of scene recognition from a novel standpoint, using…

计算机视觉与模式识别 · 计算机科学 2021-12-24 Andreea Glavan , Estefania Talavera

Sarcasm is a rhetorical device that is used to convey the opposite of the literal meaning of an utterance. Sarcasm is widely used on social media and other forms of computer-mediated communication motivating the use of computational models…

计算与语言 · 计算机科学 2024-10-25 Shafkat Farabi , Tharindu Ranasinghe , Diptesh Kanojia , Yu Kong , Marcos Zampieri

The classification of indoor scenes is a critical component in various applications, such as intelligent robotics for assistive living. While deep learning has significantly advanced this field, models often suffer from reduced performance…

计算机视觉与模式识别 · 计算机科学 2024-08-26 Willams de Lima Costa , Raul Ismayilov , Nicola Strisciuglio , Estefania Talavera Martinez

Multimodal sarcasm detection (MSD) aims to identify sarcasm within image-text pairs by modeling semantic incongruities across modalities. Existing methods often exploit cross-modal embedding misalignment to detect inconsistency but struggle…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Shuguang Zhang , Junhong Lian , Guoxin Yu , Baoxun Xu , Xiang Ao

Multi-modal approaches employ data from multiple input streams such as textual and visual domains. Deep neural networks have been successfully employed for these approaches. In this paper, we present a novel multi-modal approach that fuses…

计算机视觉与模式识别 · 计算机科学 2018-10-05 Ignazio Gallo , Alessandro Calefati , Shah Nawaz , Muhammad Kamran Janjua

Prevalent multimodal fake news detection relies on consistency-based fusion, yet this paradigm fundamentally misinterprets critical cross-modal discrepancies as noise, leading to over-smoothing, which dilutes critical evidence of…

Semantic location prediction aims to derive meaningful location insights from multimodal social media posts, offering a more contextual understanding of daily activities than using GPS coordinates. This task faces significant challenges due…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Zhizhen Zhang , Ning Wang , Haojie Li , Zhihui Wang

Multi-modal semantic understanding requires integrating information from different modalities to extract users' real intention behind words. Most previous work applies a dual-encoder structure to separately encode image and text, but fails…

计算与语言 · 计算机科学 2024-03-12 Ming Zhang , Ke Chang , Yunfang Wu

Over the last years, there has been an unprecedented proliferation of fake news. As a consequence, we are more susceptible to the pernicious impact that misinformation and disinformation spreading can have in different segments of our…

计算与语言 · 计算机科学 2021-12-10 Santiago Alonso-Bartolome , Isabel Segura-Bedmar

Considerable advancements have been made to tackle the misrepresentation of information derived from reference articles in the domains of fact-checking and faithful summarization. However, an unaddressed aspect remains - the identification…

计算与语言 · 计算机科学 2024-12-03 Kung-Hsiang Huang , Hou Pong Chan , Kathleen McKeown , Heng Ji

Automatically generating a natural language sentence to describe the content of an input video is a very challenging problem. It is an essential multimodal task in which auditory and visual contents are equally important. Although audio…

计算机视觉与模式识别 · 计算机科学 2018-12-10 Yapeng Tian , Chenxiao Guan , Justin Goodman , Marc Moore , Chenliang Xu

The rapid advancement of deepfake technology poses a significant threat to digital media integrity. Deepfakes, synthetic media created using AI, can convincingly alter videos and audio to misrepresent reality. This creates risks of…

计算机视觉与模式识别 · 计算机科学 2024-10-21 Kashish Gandhi , Prutha Kulkarni , Taran Shah , Piyush Chaudhari , Meera Narvekar , Kranti Ghag

The rapid growth of social media has led to the widespread dissemination of fake news across multiple content forms, including text, images, audio, and video. Compared to unimodal fake news detection, multimodal fake news detection benefits…

多媒体 · 计算机科学 2025-04-15 Moyang Liu , Kaiying Yan , Yukun Liu , Ruibo Fu , Zhengqi Wen , Xuefei Liu , Chenxing Li

This study investigates how fake news uses a thumbnail for a news article with a focus on whether a news article's thumbnail represents the news content correctly. A news article shared with an irrelevant thumbnail can mislead readers into…

计算与语言 · 计算机科学 2022-04-28 Hyewon Choi , Yejun Yoon , Seunghyun Yoon , Kunwoo Park

Multimodal Misinformation Recognition has become an urgent task with the emergence of huge multimodal fake content on social media platforms. Previous studies mainly focus on complex feature extraction and fusion to learn discriminative…

多媒体 · 计算机科学 2025-10-15 Hengyang Zhou , Yiwei Wei , Jian Yang , Zhenyu Zhang

Multiple modalities represent different aspects by which information is conveyed by a data source. Modern day social media platforms are one of the primary sources of multimodal data, where users use different modes of expression by posting…

计算机视觉与模式识别 · 计算机科学 2018-07-17 Mayank Meghawat , Satyendra Yadav , Debanjan Mahata , Yifang Yin , Rajiv Ratn Shah , Roger Zimmermann