中文
相关论文

相关论文: QuTI! Quantifying Text-Image Consistency in Multim…

200 篇论文

Recently advancements in large multimodal models have led to significant strides in image comprehension capabilities. Despite these advancements, there is a lack of the robust benchmark specifically for assessing the Image-to-Web conversion…

The content of today's social media is becoming more and more rich, increasingly mixing text, images, videos, and audio. It is an intriguing research question to model the interplay between these different modes in attracting user attention…

社会与信息网络 · 计算机科学 2017-03-07 Jack Hessel , Lillian Lee , David Mimno

The pervasiveness of the dissemination of fake news through social media platforms poses critical risks to the trust of the general public, societal stability, and democratic institutions. This challenge calls for novel methodologies in…

计算与语言 · 计算机科学 2025-02-04 Jingyuan Yi , Zeqiu Xu , Tianyi Huang , Peiyang Yu

Social Media Popularity Prediction is a complex multimodal task that requires effective integration of images, text, and structured information. However, current approaches suffer from inadequate visual-textual alignment and fail to capture…

信息检索 · 计算机科学 2025-08-25 Ao Zhou , Mingsheng Tu , Luping Wang , Tenghao Sun , Zifeng Cheng , Yafeng Yin , Zhiwei Jiang , Qing Gu

Humans have an incredible ability to process and understand information from multiple sources such as images, video, text, and speech. Recent success of deep neural networks has enabled us to develop algorithms which give machines the…

计算机视觉与模式识别 · 计算机科学 2019-03-18 Dheeraj Peri , Shagan Sah , Raymond Ptucha

Automated frame analysis of political communication is a popular task in computational social science that is used to study how authors select aspects of a topic to frame its reception. So far, such studies have been narrow, in that they…

计算与语言 · 计算机科学 2025-05-30 Arnav Arora , Srishti Yadav , Maria Antoniak , Serge Belongie , Isabelle Augenstein

Interpreting uncertain data can be difficult, particularly if the data presentation is complex. We investigate the efficacy of different modalities for representing data and how to combine the strengths of each modality to facilitate the…

人机交互 · 计算机科学 2024-04-15 Chase Stokes , Chelsea Sanker , Bridget Cogley , Vidya Setlur

Social media platforms have grown into an important medium to spread information about an event published by the traditional media, such as news articles. Grouping such diverse sources of information that discuss the same topic in varied…

计算与语言 · 计算机科学 2017-10-26 Aditya Mogadala , Dominik Jung , Achim Rettinger

There has been an explosion of multimodal content generated on social media networks in the last few years, which has necessitated a deeper understanding of social media content and user behavior. We present a novel content-independent…

信息检索 · 计算机科学 2019-06-12 Karan Sikka , Lucas Van Bramer , Ajay Divakaran

Developers are increasingly sharing images in social coding environments alongside the growth in visual interactions within social networks. The analysis of the ratio between the textual and visual content of Mozilla's change requests and…

软件工程 · 计算机科学 2020-01-20 Maleknaz Nayebi

Textual claims are often accompanied by images to enhance their credibility and spread on social media, but this also raises concerns about the spread of misinformation. Existing datasets for automated verification of image-text claims…

计算与语言 · 计算机科学 2025-10-08 Rui Cao , Zifeng Ding , Zhijiang Guo , Michael Schlichtkrull , Andreas Vlachos

Large language models (LLMs) have recently demonstrated remarkable advancements in embodying diverse personas, enhancing their effectiveness as conversational agents and virtual assistants. Consequently, LLMs have made significant strides…

计算与语言 · 计算机科学 2025-03-03 Julius Broomfield , Kartik Sharma , Srijan Kumar

Online misinformation is often multimodal in nature, i.e., it is caused by misleading associations between texts and accompanying images. To support the fact-checking process, researchers have been recently developing automatic multimodal…

Making image retrieval methods practical for real-world search applications requires significant progress in dataset scales, entity comprehension, and multimodal information fusion. In this work, we introduce \textbf{E}ntity-\textbf{D}riven…

计算与语言 · 计算机科学 2023-10-24 Siqi Liu , Weixi Feng , Tsu-jui Fu , Wenhu Chen , William Yang Wang

With the growing adoption of short-form video by social media platforms, reducing the spread of misinformation through video posts has become a critical challenge for social media providers. In this paper, we develop methods to detect…

计算机视觉与模式识别 · 计算机科学 2022-08-02 Kehan Wang , David Chan , Seth Z. Zhao , John Canny , Avideh Zakhor

Multimodal interfaces, combining the use of speech, graphics, gestures, and facial expressions in input and output, promise to provide new possibilities to deal with information in more effective and efficient ways, supporting for instance:…

计算与语言 · 计算机科学 2009-09-24 Harry Bunt , Laurent Romary

We present a study on predicting the factuality of reporting and bias of news media. While previous work has focused on studying the veracity of claims or documents, here we are interested in characterizing entire news media. These are…

信息检索 · 计算机科学 2018-10-04 Ramy Baly , Georgi Karadzhov , Dimitar Alexandrov , James Glass , Preslav Nakov

Multimodal target/aspect sentiment classification combines multimodal sentiment analysis and aspect/target sentiment classification. The goal of the task is to combine vision and language to understand the sentiment towards a target entity…

计算与语言 · 计算机科学 2021-08-09 Zaid Khan , Yun Fu

In this work we present an approach for generating alternative text (or alt-text) descriptions for images shared on social media, specifically Twitter. More than just a special case of image captioning, alt-text is both more literally…

计算机视觉与模式识别 · 计算机科学 2024-03-04 Nikita Srivatsan , Sofia Samaniego , Omar Florez , Taylor Berg-Kirkpatrick

Neural topic models can successfully find coherent and diverse topics in textual data. However, they are limited in dealing with multimodal datasets (e.g., images and text). This paper presents the first systematic and comprehensive…

计算与语言 · 计算机科学 2024-03-27 Felipe González-Pizarro , Giuseppe Carenini