中文
相关论文

相关论文: CoLoRSMamba: Conditional LoRA-Steered Mamba for Su…

200 篇论文

Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, which simulates an ``Audio-Visual Confusion'' scene by modifying…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Qilang Ye , Wei Zeng , Meng Liu , Jie Zhang , Yupeng Hu , Zitong Yu , Yu Zhou

Existing Video Anomaly Detection (VAD) methods typically rely on task-specific training, leading to strong domain dependency and high training costs. Moreover, most existing methods output only scalar anomaly scores, providing limited…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Hyeongmuk Lim , Youngbum Hur

Heterogeneous multi-modal remote sensing object detection aims to accurately detect objects from diverse sensors (e.g., RGB, SAR, Infrared). Existing approaches largely adopt a late alignment paradigm, in which modality alignment and…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Yuxuan Li , Yuming Chen , Yunheng Li , Ming-Ming Cheng , Xiang Li , Jian Yang

Video anomaly detection (VAD) is an essential task in the image processing community with prospects in video surveillance, which faces fundamental challenges in balancing detection accuracy with computational efficiency. As video content…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Yang Liu , Boan Chen , Xiaoguang Zhu , Jing Liu , Peng Sun , Wei Zhou

Although change detection using MODIS time series is critical for environmental monitoring, it is a highly challenging task due to key MODIS difficulties, e.g., mixed pixels, spatial-spectral-temporal information coupling effect, and…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Zhengsen Xu , Yimin Zhu , Zack Dewis , Mabel Heffring , Motasem Alkayid , Saeid Taleghanidoozdoozan , Lincoln Linlin Xu

While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the visual domain, one…

声音 · 计算机科学 2025-09-22 Qiaolin Wang , Xilin Jiang , Linyang He , Junkai Wu , Nima Mesgarani

In multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the temporal dynamics of…

音频与语音处理 · 电气工程与系统科学 2025-01-15 Wenze Ren , Haibin Wu , Yi-Cheng Lin , Xuanjun Chen , Rong Chao , Kuo-Hsuan Hung , You-Jin Li , Wen-Yuan Ting , Hsin-Min Wang , Yu Tsao

How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this problem, we develop a two-stage audiovisual learning framework…

计算机视觉与模式识别 · 计算机科学 2020-07-15 Rui Qian , Di Hu , Heinrich Dinkel , Mengyue Wu , Ning Xu , Weiyao Lin

Audio-Visual Segmentation (AVS) aims to extract the sounding object from a video frame, which is represented by a pixel-wise segmentation mask for application scenarios such as multi-modal video editing, augmented reality, and intelligent…

图像与视频处理 · 电气工程与系统科学 2024-12-25 Zhaofeng Shi , Qingbo Wu , Fanman Meng , Linfeng Xu , Hongliang Li

We present StyleMamba, an efficient image style transfer framework that translates text prompts into corresponding visual styles while preserving the content integrity of the original images. Existing text-guided stylization requires…

计算机视觉与模式识别 · 计算机科学 2024-05-09 Zijia Wang , Zhi-Song Liu

State-space models (SSMs) have recently shown promise in capturing long-range dependencies with subquadratic computational complexity, making them attractive for various applications. However, purely SSM-based models face critical…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Abdelrahman Shaker , Syed Talal Wasim , Salman Khan , Juergen Gall , Fahad Shahbaz Khan

The ability to capture and segment sounding objects in dynamic visual scenes is crucial for the development of Audio-Visual Segmentation (AVS) tasks. While significant progress has been made in this area, the interaction between audio and…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Kai Peng , Yunzhe Shen , Miao Zhang , Leiye Liu , Yidong Han , Wei Ji , Jingjing Li , Yongri Piao , Huchuan Lu

In this article, we describe Conditioned Localizer and Classifier (CoLoC) which is a novel solution for Sound Event Localization and Detection (SELD). The solution constitutes of two stages: the localization is done first and is followed by…

声音 · 计算机科学 2022-10-26 Sławomir Kapka , Jakub Tkaczuk

Existing RGB-Event detection methods process the low-information regions of both modalities (background in images and non-event regions in event data) uniformly during feature extraction and fusion, resulting in high computational costs and…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Nan Yang , Yang Wang , Zhanwen Liu , Yuchao Dai , Yang Liu , Xiangmo Zhao

Existing works have made strides in video generation, but the lack of sound effects (SFX) and background music (BGM) hinders a complete and immersive viewer experience. We introduce a novel semantically consistent v ideo-to-audio generation…

多媒体 · 计算机科学 2024-04-29 Gehui Chen , Guan'an Wang , Xiaowen Huang , Jitao Sang

Vision-language models (VLMs) have shown strong performance in video anomaly detection (VAD) while providing interpretable predictions. However, existing VLM-based VAD methods suffer from a fundamental mismatch between training and…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Darryl Cherian Jacob , Xinyu Liu , Kai Wang , Pan He

Recent Mamba-based models have shown promise in speech enhancement by efficiently modeling long-range temporal dependencies. However, models like Speech Enhancement Mamba (SEMamba) remain limited to single-speaker scenarios and struggle in…

声音 · 计算机科学 2025-10-01 Rong Chao , Wenze Ren , You-Jin Li , Kuo-Hsuan Hung , Sung-Feng Huang , Szu-Wei Fu , Wen-Huang Cheng , Yu Tsao

We introduce a new music source separation model tailored for accurate vocal isolation. Unlike Transformer-based approaches, which often fail to capture intermittently occurring vocals, our model leverages Mamba2, a recent state space…

声音 · 计算机科学 2026-01-01 Euiyeon Kim , Yong-Hoon Choi

Audio-visual segmentation (AVS) is an emerging task that aims to accurately segment sounding objects based on audio-visual cues. The success of AVS learning systems depends on the effectiveness of cross-modal interaction. Such a requirement…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Yuanhong Chen , Chong Wang , Yuyuan Liu , Hu Wang , Gustavo Carneiro

While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored. In this work, we…

声音 · 计算机科学 2026-01-05 Akanksha Chuchra , Shukesh Reddy , Sudeepta Mishra , Abhijit Das , Abhinav Dhall