English
Related papers

Related papers: Validating Multimedia Content Moderation Software …

200 papers

Multimodal sarcasm detection (MSD) aims to identify sarcastic intent from semantic incongruity between text and image. Although recent methods have improved MSD through cross-modal interaction and incongruity reasoning, most still treat…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Zhenyu Wang , Weichen Cheng , Weijia Li , Junjie Mou , Zongyou Zhao , Guoying Zhang

We present a holistic approach to building a robust and useful natural language classification system for real-world content moderation. The success of such a system relies on a chain of carefully designed and executed steps, including the…

Computation and Language · Computer Science 2023-02-16 Todor Markov , Chong Zhang , Sandhini Agarwal , Tyna Eloundou , Teddy Lee , Steven Adler , Angela Jiang , Lilian Weng

Multimodal sentiment analysis (MSA) integrates various modalities, such as text, image, and audio, to provide a more comprehensive understanding of sentiment. However, effective MSA is challenged by alignment and fusion issues. Alignment…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yuhua Wen , Qifei Li , Yingying Zhou , Yingming Gao , Zhengqi Wen , Jianhua Tao , Ya Li

Textual interaction networks (TINs) are an omnipresent data structure used to model the interplay between users and items on e-commerce websites, social networks, etc., where each interaction is associated with a text description.…

Computation and Language · Computer Science 2025-04-08 Hongtao Wang , Renchi Yang , Hewen Wang , Haoran Zheng , Jianliang Xu

Multimedia learning using text and images has been shown to improve learning outcomes compared to text-only instruction. But conversational AI systems in education predominantly rely on text-based interactions while multimodal conversations…

Human-Computer Interaction · Computer Science 2025-04-22 Karan Taneja , Anjali Singh , Ashok K. Goel

Text-to-image (T2I) diffusion models have the ability to build high-quality pictures from text prompts, but they pose safety concerns because they can generate offensive or disturbing imagery when provided with harmful inputs. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Chi Zhang , Changjia Zhu , Xiaowen Li , Yao Liu , Zhuo Lu

With the rapid development of Internet and multimedia services in the past decade, a huge amount of user-generated and service provider-generated multimedia data become available. These data are heterogeneous and multi-modal in nature,…

Multimedia · Computer Science 2020-01-07 Wenwu Zhu , Xin Wang , Hongzhi Li

Social media platforms face escalating challenges in detecting harmful content that promotes muscle dysmorphic behaviors and cognitions (bigorexia). This content can evade moderation by camouflaging as legitimate fitness advice and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Minh Duc Chu , Kshitij Pawar , Zihao He , Roxanna Sharifi , Ross Sonnenblick , Magdalayna Curry , Laura D'Adamo , Lindsay Young , Stuart B Murray , Kristina Lerman

With the development of Artificial Intelligence (AI) and Internet of Things (IoT) technologies, network communications based on the Shannon-Nyquist theorem gradually reveal their limitations due to the neglect of semantic information in the…

Information Theory · Computer Science 2025-04-10 Linhan Xia , Jiaxin Cai , Ricky Yuen-Tan Hou , Seon-Phil Jeong

With the proliferation of social media posts in recent years, the need to detect sentiments in multimodal (image-text) content has grown rapidly. Since posts are user-generated, the image and text from the same post can express different or…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Daiqing Wu , Dongbao Yang , Huawen Shen , Can Ma , Yu Zhou

Classifying products into categories precisely and efficiently is a major challenge in modern e-commerce. The high traffic of new products uploaded daily and the dynamic nature of the categories raise the need for machine learning models…

Computer Vision and Pattern Recognition · Computer Science 2016-11-30 Tom Zahavy , Alessandro Magnani , Abhinandan Krishnan , Shie Mannor

Traditional online content moderation systems struggle to classify modern multimodal means of communication, such as memes, a highly nuanced and information-dense medium. This task is especially hard in a culturally diverse society like…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Cao Yuxuan , Wu Jiayang , Alistair Cheong Liang Chuen , Bryan Shan Guanrong , Theodore Lee Chong Jen , Sherman Chann Zhi Shen

Text-guided semantic manipulation refers to semantically editing an image generated from a source prompt to match a target prompt, enabling the desired semantic changes (e.g., addition, removal, and style transfer) while preserving…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Yu Hong , Xiao Cai , Pengpeng Zeng , Shuai Zhang , Jingkuan Song , Lianli Gao , Heng Tao Shen

With the continuous emergence of various social media platforms frequently used in daily life, the multimodal meme understanding (MMU) task has been garnering increasing attention. MMU aims to explore and comprehend the meanings of memes…

Computation and Language · Computer Science 2025-03-18 Li Zheng , Hao Fei , Ting Dai , Zuquan Peng , Fei Li , Huisheng Ma , Chong Teng , Donghong Ji

We present a scalable and agile approach for ads image content moderation at Google, addressing the challenges of moderating massive volumes of ads with diverse content and evolving policies. The proposed method utilizes human-curated…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Enming Luo , Wei Qiao , Katie Warren , Jingxiang Li , Eric Xiao , Krishna Viswanathan , Yuan Wang , Yintao Liu , Jimin Li , Ariel Fuxman

Visually rich documents (e.g. leaflets, banners, magazine articles) are physical or digital documents that utilize visual cues to augment their semantics. Information contained in these documents are ad-hoc and often incomplete. Existing…

Machine Learning · Computer Science 2024-04-02 Ritesh Sarkhel , Arnab Nandi

The Fediverse, a group of interconnected servers providing a variety of interoperable services (e.g. micro-blogging in Mastodon) has gained rapid popularity. This sudden growth, partly driven by Elon Musk's acquisition of Twitter, has…

Social and Information Networks · Computer Science 2025-01-13 Haris Bin Zia , Aravindh Raman , Ignacio Castro , Gareth Tyson

Social media platforms are increasingly dominated by long-form multimodal content, where harmful narratives are constructed through a complex interplay of audio, visual, and textual cues. While automated systems can flag hate speech with…

Artificial Intelligence · Computer Science 2026-05-29 Girish A. Koushik , Helen Treharne , Diptesh Kanojia

With the rapid progression of deep learning technologies, multi-modality image fusion has become increasingly prevalent in object detection tasks. Despite its popularity, the inherent disparities in how different sources depict scene…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Xingyuan Li , Yang Zou , Jinyuan Liu , Zhiying Jiang , Long Ma , Xin Fan , Risheng Liu

Diffusion models have been widely used for conditional data cross-modal generation tasks such as text-to-image and text-to-video. However, state-of-the-art models still fail to align the generated visual concepts with high-level semantics…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Zizhao Hu , Shaochong Jia , Mohammad Rostami
‹ Prev 1 4 5 6 7 8 10 Next ›