English
Related papers

Related papers: Differential Attention for Multimodal Crisis Event…

200 papers

This paper focuses on the challenging crowd counting task. As large-scale variations often exist within crowd images, neither fixed-size convolution kernel of CNN nor fixed-size attention of recent vision transformers can well handle this…

Computer Vision and Pattern Recognition · Computer Science 2022-03-08 Hui Lin , Zhiheng Ma , Rongrong Ji , Yaowei Wang , Xiaopeng Hong

In recent years, multimodal large language models (MLLMs) have shown remarkable capabilities in tasks like visual question answering and common sense reasoning, while visual perception models have made significant strides in perception…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Guanqun Wang , Xinyu Wei , Jiaming Liu , Ray Zhang , Yichi Zhang , Kevin Zhang , Maurice Chong , Shanghang Zhang

Vision-Language Models (VLMs) achieve strong cross-modal performance, yet recent evidence suggests they over-rely on textual descriptions while under-utilizing visual evidence -- a phenomenon termed ``text shortcut learning.'' We propose an…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Lijie Zhou

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

Multimedia · Computer Science 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan

Weakly supervised video anomaly detection (WS-VAD) is a crucial area in computer vision for developing intelligent surveillance systems. This system uses three feature streams: RGB video, optical flow, and audio signals, where each stream…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuta Kaneko , Abu Saleh Musa Miah , Najmul Hassan , Hyoun-Sup Lee , Si-Woong Jang , Jungpil Shin

Detecting bias in multimodal news requires models that reason over text--image pairs, not just classify text. In response, we present ViLBias, a VQA-style benchmark and framework for detecting and reasoning about bias in multimodal news.…

Large multimodal models (LMMs) excel in scene understanding but struggle with fine-grained spatiotemporal reasoning due to weak alignment between linguistic and visual representations. Existing methods map textual positions and durations…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Hanyu Zhou , Gim Hee Lee

In this new digital era, social media has created a severe impact on the lives of people. In recent times, fake news content on social media has become one of the major challenging problems for society. The dissemination of fabricated and…

Multimedia · Computer Science 2023-03-14 Ashima Yadav , Shivani Gaba , Haneef Khan , Ishan Budhiraja , Akansha Singh , Krishan Kant Singh

Recent advancements in Large Language Models (LLMs) have significantly enhanced their capacity to process long contexts. However, effectively utilizing this long context remains a challenge due to the issue of distraction, where irrelevant…

Computation and Language · Computer Science 2024-11-12 Zijun Wu , Bingyuan Liu , Ran Yan , Lei Chen , Thomas Delteil

Multimodal sentiment analysis has attracted increasing attention with broad application prospects. The existing methods focuses on single modality, which fails to capture the social media content for multiple modalities. Moreover, in…

Multimedia · Computer Science 2022-05-11 Ashima Yadav , Dinesh Kumar Vishwakarma

Crash detection from video feeds is a critical problem in intelligent transportation systems. Recent developments in large language models (LLMs) and vision-language models (VLMs) have transformed how we process, reason about, and summarize…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Sanjeda Akter , Ibne Farabi Shihab , Anuj Sharma

Learning discriminative representations for subtle localized details plays a significant role in Fine-grained Visual Categorization (FGVC). Compared to previous attention-based works, our work does not explicitly define or localize the part…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Ranran Huang , Yu Wang , Huazhong Yang

Large Vision-Language Models (LVLMs) have achieved impressive performance in multimodal tasks, but they still suffer from hallucinations, i.e., generating content that is grammatically accurate but inconsistent with visual inputs. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Chenxi Li , Yichen Guo , Benfang Qian , Jinhao You , Kai Tang , Yaosong Du , Zonghao Zhang , Xiande Huang

Soft attention is a critical mechanism powering LLMs to locate relevant parts within a given context. However, individual attention weights are determined by the similarity of only a single query and key token vector. This "single token…

Computation and Language · Computer Science 2025-07-14 Olga Golovneva , Tianlu Wang , Jason Weston , Sainbayar Sukhbaatar

Timely interpretation of satellite imagery is critical for disaster response, yet existing vision-language benchmarks for remote sensing largely focus on coarse labels and image-level recognition, overlooking the functional understanding…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Sara Tehrani , Yonghao Xu , Leif Haglund , Amanda Berg , Michael Felsberg

Social media has a significant impact on people's lives. Hate speech on social media has emerged as one of society's most serious issues in recent years. Text and pictures are two forms of multimodal data that are distributed within…

Computation and Language · Computer Science 2024-09-18 Anusha Chhabra , Dinesh Kumar Vishwakarma

Recently, while large language models (LLMs) have demonstrated impressive results, they still suffer from hallucination, i.e., the generation of false information. Model editing is the task of fixing factual mistakes in LLMs; yet, most…

Computation and Language · Computer Science 2024-06-03 Taolin Zhang , Qizhou Chen , Dongyang Li , Chengyu Wang , Xiaofeng He , Longtao Huang , Hui Xue , Jun Huang

Weakly supervised violence detection refers to the technique of training models to identify violent segments in videos using only video-level labels. Among these approaches, multimodal violence detection, which integrates modalities such as…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Wenping Jin , Li Zhu , Jing Sun

Large vision-language models (LVLMs) are markedly proficient in deriving visual representations guided by natural language. Recent explorations have utilized LVLMs to tackle zero-shot visual anomaly detection (VAD) challenges by pairing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Jiaqi Zhu , Shaofeng Cai , Fang Deng , Beng Chin Ooi , Junran Wu

Emotion represents an essential aspect of human speech that is manifested in speech prosody. Speech, visual, and textual cues are complementary in human communication. In this paper, we study a hybrid fusion method, referred to as…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Zexu Pan , Zhaojie Luo , Jichen Yang , Haizhou Li