中文
相关论文

相关论文: QuTI! Quantifying Text-Image Consistency in Multim…

200 篇论文

The increasing realism of multimodal content has made misinformation more subtle and harder to detect, especially in news media where images are frequently paired with bilingual (e.g., Chinese-English) subtitles. Such content often includes…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Yiwei He , Zhenglin Huang , Haiquan Wen , Tianxiao Li , Yi Dong , Hao Fei , Baoyuan Wu , Guangliang Cheng

Charts go hand in hand with text to communicate complex data and are widely adopted in news articles, online blogs, and academic papers. They provide graphical summaries of the data, while text explains the message and context. However,…

人机交互 · 计算机科学 2021-08-10 Shahid Latif , Zheng Zhou , Yoon Kim , Fabian Beck , Nam Wook Kim

In recent years, the problem of misinformation on the web has become widespread across languages, countries, and various social media platforms. Although there has been much work on automated fake news detection, the role of images and…

计算与语言 · 计算机科学 2022-05-05 Gullal S. Cheema , Sherzod Hakimov , Abdul Sittar , Eric Müller-Budack , Christian Otto , Ralph Ewerth

In recent years, the rampant spread of misinformation on social media has made accurate detection of multimodal fake news a critical research focus. However, previous research has not adequately understood the semantics of images, and…

多媒体 · 计算机科学 2025-07-18 Peican Zhu , Yubo Jing , Le Cheng , Keke Tang , Yangming Guo

Feature modeling of different modalities is a basic problem in current research of cross-modal information retrieval. Existing models typically project texts and images into one embedding space, in which semantically similar information…

多媒体 · 计算机科学 2019-06-13 Jing Yu , Chenghao Yang , Zengchang Qin , Zhuoqian Yang , Yue Hu , Weifeng Zhang

Multimodal misinformation increasingly leverages visual persuasion, where repurposed or manipulated images strengthen misleading text. We introduce \textbf{RW-Post}, a post-aligned \textbf{text--image benchmark} for real-world multimodal…

多媒体 · 计算机科学 2026-05-13 Danni Xu , Shaojing Fan , Harry Cheng , Mohan Kankanhalli

Interacting and understanding with text heavy visual content with multiple images is a major challenge for traditional vision models. This paper is on enhancing vision models' capability to comprehend or understand and learn from images…

计算机视觉与模式识别 · 计算机科学 2024-08-31 Adithya TG , Adithya SK , Abhinav R Bharadwaj , Abhiram HA , Surabhi Narayan

Web-scale visual entity recognition, the task of associating images with their corresponding entities within vast knowledge bases like Wikipedia, presents significant challenges due to the lack of clean, large-scale training data. In this…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Mathilde Caron , Alireza Fathi , Cordelia Schmid , Ahmet Iscen

In the battle against widespread online misinformation, a growing problem is text-image inconsistency, where images are misleadingly paired with texts with different intent or meaning. Existing classification-based methods for text-image…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Mingzhen Huang , Shan Jia , Zhou Zhou , Yan Ju , Jialing Cai , Siwei Lyu

Point-of-interest (POI) type prediction is the task of inferring the type of a place from where a social media post was shared. Inferring a POI's type is useful for studies in computational social science including sociolinguistics,…

计算与语言 · 计算机科学 2021-09-03 Danae Sánchez Villegas , Nikolaos Aletras

Multimodal image-text memes are prevalent on the internet, serving as a unique form of communication that combines visual and textual elements to convey humor, ideas, or emotions. However, some memes take a malicious turn, promoting hateful…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Giovanni Burbi , Alberto Baldrati , Lorenzo Agnolucci , Marco Bertini , Alberto Del Bimbo

Text-to-image generation has attracted significant interest from researchers and practitioners in recent years due to its widespread and diverse applications across various industries. Despite the progress made in the domain of vision and…

计算机视觉与模式识别 · 计算机科学 2023-05-25 Yutong Zhou , Nobutaka Shimada

Multimodal misinformation increasingly leverages visual persuasion, where repurposed or manipulated images strengthen misleading text. We introduce RW-Post, a post-aligned text--image benchmark for real-world multimodal fact-checking with…

人工智能 · 计算机科学 2026-05-13 Danni Xu , Shaojing Fan , Harry Cheng , Mohan Kankanhalli

In mixed-initiative conversational search systems, clarifying questions are used to help users who struggle to express their intentions in a single query. These questions aim to uncover user's information needs and resolve query…

计算与语言 · 计算机科学 2024-02-13 Yifei Yuan , Clemencia Siro , Mohammad Aliannejadi , Maarten de Rijke , Wai Lam

This paper addresses the critical challenge of assessing the representativeness of news thumbnail images, which often serve as the first visual engagement for readers when an article is disseminated on social media. We focus on whether a…

计算与语言 · 计算机科学 2024-06-10 Yejun Yoon , Seunghyun Yoon , Kunwoo Park

This paper presents a novel crowd-sourced resource for multimodal discourse: our resource characterizes inferences in image-text contexts in the domain of cooking recipes in the form of coherence relations. Like previous corpora annotating…

计算与语言 · 计算机科学 2019-04-17 Malihe Alikhani , Sreyasi Nag Chowdhury , Gerard de Melo , Matthew Stone

Emojis are small images that are commonly included in social media text messages. The combination of visual and textual content in the same message builds up a modern way of communication, that automatic systems are not used to deal with.…

计算与语言 · 计算机科学 2018-04-18 Francesco Barbieri , Miguel Ballesteros , Francesco Ronzano , Horacio Saggion

Interacting with the legal system and the government requires the assembly and analysis of various pieces of information that can be spread across different (paper) documents, such as forms, certificates and contracts (e.g. leases). This…

计算与语言 · 计算机科学 2024-12-23 Hannes Westermann , Jaromir Savelka

Given the massive market of advertising and the sharply increasing online multimedia content (such as videos), it is now fashionable to promote advertisements (ads) together with the multimedia content. It is exhausted to find relevant ads…

多媒体 · 计算机科学 2020-01-06 Huaizheng Zhang , Yong Luo , Qiming Ai , Yonggang Wen

The proliferation of disinformation, particularly in multimodal contexts combining text and images, presents a significant challenge across digital platforms. This study investigates the potential of large multimodal models (LMMs) in…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Yasmina Kheddache , Marc Lalonde