中文
相关论文

相关论文: BAVS: Bootstrapping Audio-Visual Segmentation by I…

200 篇论文

Audio-Visual Target Speaker Extraction (AVTSE) aims to isolate a target speaker's voice in a multi-speaker environment with visual cues as auxiliary. Most of the existing AVTSE methods encode visual and audio features simultaneously,…

声音 · 计算机科学 2025-11-13 Zixuan Li , Xueliang Zhang , Lei Miao , Zhipeng Yan , Ying Sun , Chong Zhu

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Ruohan Gao , Kristen Grauman

Audio-visual segmentation (AVS) plays a critical role in multimodal machine learning by effectively integrating audio and visual cues to precisely segment objects or regions within visual scenes. Recent AVS methods have demonstrated…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Yunzhe Shen , Kai Peng , Leiye Liu , Wei Ji , Jingjing Li , Miao Zhang , Yongri Piao , Huchuan Lu

Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from ambiguous…

声音 · 计算机科学 2025-03-18 Chen Liu , Peike Li , Liying Yang , Dadong Wang , Lincheng Li , Xin Yu

Sound localization aims to find the source of the audio signal in the visual scene. However, it is labor-intensive to annotate the correlations between the signals sampled from the audio and visual modalities, thus making it difficult to…

计算机视觉与模式识别 · 计算机科学 2021-04-02 Yan-Bo Lin , Hung-Yu Tseng , Hsin-Ying Lee , Yen-Yu Lin , Ming-Hsuan Yang

Never having seen an object and heard its sound simultaneously, can the model still accurately localize its visual position from the input audio? In this work, we concentrate on the Audio-Visual Localization and Segmentation tasks but under…

计算机视觉与模式识别 · 计算机科学 2024-02-05 Yaoting Wang , Weisong Liu , Guangyao Li , Jian Ding , Di Hu , Xi Li

Adding visual cues to audio-based speech separation can improve separation performance. This paper introduces AV-CrossNet, an audiovisual (AV) system for speech enhancement, target speaker extraction, and multi-talker speaker separation.…

音频与语音处理 · 电气工程与系统科学 2024-06-18 Vahid Ahmadi Kalkhorani , Cheng Yu , Anurag Kumar , Ke Tan , Buye Xu , DeLiang Wang

Open-vocabulary semantic segmentation (OVS) aims to segment images of arbitrary categories specified by class labels or captions. However, most previous best-performing methods, whether pixel grouping methods or region recognition methods,…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Yuan Wang , Rui Sun , Naisong Luo , Yuwen Pan , Tianzhu Zhang

In this paper, we consider the problem of open-vocabulary semantic segmentation (OVS), which aims to segment objects of arbitrary classes instead of pre-defined, closed-set categories. The main contributions are as follows: First, we…

计算机视觉与模式识别 · 计算机科学 2023-03-06 Jilan Xu , Junlin Hou , Yuejie Zhang , Rui Feng , Yi Wang , Yu Qiao , Weidi Xie

Constructing an embedding space for musical instrument sounds that can meaningfully represent new and unseen instruments is important for downstream music generation tasks such as multi-instrument synthesis and timbre transfer. The…

音频与语音处理 · 电气工程与系统科学 2021-12-28 Xuan Shi , Erica Cooper , Junichi Yamagishi

We study the merit of transfer learning for two sound recognition problems, i.e., audio tagging and sound event detection. Employing feature fusion, we adapt a baseline system utilizing only spectral acoustic inputs to also make use of…

音频与语音处理 · 电气工程与系统科学 2022-09-27 Wim Boes , Hugo Van hamme

Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Yang Liu , Mengyuan Liu , Shudong Huang , Jiancheng Lv

The rapid development of deep learning has driven significant progress in image semantic segmentation - a fundamental task in computer vision. Semantic segmentation algorithms often depend on the availability of pixel-level labels (i.e.,…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Zhaozheng Chen , Qianru Sun

Semantic segmentation plays a crucial role in enabling comprehensive scene understanding for robotic systems. However, generating annotations is challenging, requiring labels for every pixel in an image. In scenarios like autonomous…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Mostafa ElAraby , Ali Harakeh , Liam Paull

With the rapid advancement of autonomous driving, vehicle perception, particularly detection and segmentation, has placed increasingly higher demands on algorithmic performance. Pre-trained large segmentation models, especially Segment…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Xiao Wang , Ziwen Wang , Wentao Wu , Anjie Wang , Jiashu Wu , Yantao Pan , Chenglong Li

Referring Video Object Segmentation (RVOS) aims to segment target objects in videos based on natural language descriptions. However, fixed keyframe-based approaches that couple a vision language model with a separate propagation module…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Jihwan Hong , Jaeyoung Do

Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introducing adverse…

音频与语音处理 · 电气工程与系统科学 2026-03-09 Linzhi Wu , Xingyu Zhang , Hao Yuan , Yakun Zhang , Changyan Zheng , Liang Xie , Tiejun Liu , Erwei Yin

Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audio-visual data…

计算机视觉与模式识别 · 计算机科学 2019-10-28 Chuang Gan , Hang Zhao , Peihao Chen , David Cox , Antonio Torralba

Open-vocabulary remote sensing image segmentation (OVRSIS) remains underexplored due to fragmented datasets, limited training diversity, and the lack of evaluation benchmarks that reflect realistic geospatial application demands. Our…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Bingyu Li , Tao Huo , Haocheng Dong , Da Zhang , Zhiyuan Zhao , Junyu Gao , Xuelong Li

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source and target…

多媒体 · 计算机科学 2023-11-27 Arjun Somayazulu , Changan Chen , Kristen Grauman