中文
相关论文

相关论文: Reasoning-Aware Multimodal Fusion for Hateful Vide…

200 篇论文

Video captioning models convert frames into visual tokens and generate descriptions with large language models (LLMs). Since encoding all frames is prohibitively expensive, uniform sampling is the default choice, but it enforces equal…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Lianying Chao , Linfeng Yin , Peiyu Ren , Yifan Jiang , Qiaoyu Ren , Dingcheng Shan , Jing-cheng Pang , Sijie Wu , Xubin Li , Kai Zhang , Xin Chen

In this work, we examine hateful memes from three complementary angles - how to detect them, how to explain their content and how to intervene them prior to being posted - by applying a range of strategies built on top of generative AI…

计算与语言 · 计算机科学 2026-01-09 Naquee Rizwan , Subhankar Swain , Paramananda Bhaskar , Gagan Aryan , Shehryaar Shah Khan , Animesh Mukherjee

Contrastive decoding strategies are widely used to mitigate object hallucinations in multimodal large language models (MLLMs). By reducing over-reliance on language priors, these strategies ensure that generated content remains closely…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Hao Yin , Guangzong Si , Zilei Wang

Both functional and structural magnetic resonance imaging (fMRI and sMRI) are widely used for the diagnosis of mental disorder. However, combining complementary information from these two modalities is challenging due to their…

图像与视频处理 · 电气工程与系统科学 2024-04-02 Ziyu Zhou , Anton Orlichenko , Gang Qu , Zening Fu , Vince D Calhoun , Zhengming Ding , Yu-Ping Wang

In the pathway toward Artificial General Intelligence (AGI), understanding human's affection is essential to enhance machine's cognition abilities. For achieving more sensual human-AI interaction, Multimodal Affective Computing (MAC) in…

计算机视觉与模式识别 · 计算机科学 2024-08-15 Ronghao Lin , Ying Zeng , Sijie Mai , Haifeng Hu

Video-based person re-identification (ReID) is challenging due to the presence of various interferences in video frames. Recent approaches handle this problem using temporal aggregation strategies. In this work, we propose a novel Context…

计算机视觉与模式识别 · 计算机科学 2022-07-07 Kan Wang , Changxing Ding , Jianxin Pang , Xiangmin Xu

Detecting hate speech in online content is essential to ensuring safer digital spaces. While significant progress has been made in text and meme modalities, video-based hate speech detection remains under-explored, hindered by a lack of…

计算机视觉与模式识别 · 计算机科学 2025-01-28 Han Wang , Rui Yang Tan , Roy Ka-Wei Lee

Technology videos contain rich multi-modal information. In cross-modal information search, the data features of different modalities cannot be compared directly, so the semantic gap between different modalities is a key problem that needs…

信息检索 · 计算机科学 2022-10-12 Xiangbin Liu , Junping Du , Meiyu Liang , Ang Li

There is a growing trend in placing video advertisements on social platforms for online marketing, which demands automatic approaches to understand the contents of advertisements effectively. Taking the 2021 TAAC competition as an…

计算机视觉与模式识别 · 计算机科学 2021-08-31 Zejia Weng , Lingchen Meng , Rui Wang , Zuxuan Wu , Yu-Gang Jiang

Internet memes have become a dominant method of communication; at the same time, however, they are also increasingly being used to advocate extremism and foster derogatory beliefs. Nonetheless, we do not have a firm understanding as to…

Traditional sentiment analysis has long been a unimodal task, relying solely on text. This approach overlooks non-verbal cues such as vocal tone and prosody that are essential for capturing true emotional intent. We introduce Dynamic…

计算与语言 · 计算机科学 2025-09-30 Sadia Abdulhalim , Muaz Albaghdadi , Moshiur Farazi

Comprehending long videos remains a significant challenge for Large Multi-modal Models (LMMs). Current LMMs struggle to process even minutes to hours videos due to their lack of explicit memory and retrieval mechanisms. To address this…

计算机视觉与模式识别 · 计算机科学 2025-05-07 Sameer Malik , Moyuru Yamada , Ayush Singh , Dishank Aggarwal

Text-embedded images can serve as a means of spreading hate speech, propaganda, and extremist beliefs. Throughout the Russia-Ukraine war, both opposing factions heavily relied on text-embedded images as a vehicle for spreading propaganda…

计算与语言 · 计算机科学 2023-07-27 Umitcan Sahin , Izzet Emre Kucukkaya , Oguzhan Ozcelik , Cagri Toraman

We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in the image or video. In…

计算机视觉与模式识别 · 计算机科学 2021-02-10 Linwei Ye , Mrigank Rochan , Zhi Liu , Xiaoqin Zhang , Yang Wang

Visual reasoning abilities play a crucial role in understanding complex multimodal data, advancing both domain-specific applications and artificial general intelligence (AGI). Existing methods enhance Vision-Language Models (VLMs) through…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Huajie Tan , Yuheng Ji , Xiaoshuai Hao , Xiansheng Chen , Pengwei Wang , Zhongyuan Wang , Shanghang Zhang

Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple videos into a single…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yue Zhang , Liqiang Jing , Jia Li , Yapeng Tian , Xinya Du , Yunhui Guo , Vibhav Gogate

Emotion recognition is a core research area at the intersection of artificial intelligence and human communication analysis. It is a significant technical challenge since humans display their emotions through complex idiosyncratic…

人机交互 · 计算机科学 2018-09-14 Paul Pu Liang , Amir Zadeh , Louis-Philippe Morency

Cross-modality fusing complementary information of multispectral remote sensing image pairs can improve the perception ability of detection algorithms, making them more robust and reliable for a wider range of applications, such as…

计算机视觉与模式识别 · 计算机科学 2021-12-07 Qingyun Fang , Zhaokui Wang

Recently, weakly supervised video anomaly detection (WS-VAD) has emerged as a contemporary research direction to identify anomaly events like violence and nudity in videos using only video-level labels. However, this task has substantial…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Ayush Ghadiya , Purbayan Kar , Vishal Chudasama , Pankaj Wasnik

Infrared and visible image fusion is a powerful technique that combines complementary information from different modalities for downstream semantic perception tasks. Existing learning-based methods show remarkable performance, but are…

计算机视觉与模式识别 · 计算机科学 2023-08-09 Zhu Liu , Jinyuan Liu , Benzhuang Zhang , Long Ma , Xin Fan , Risheng Liu