中文
相关论文

相关论文: A Unified Audio-Visual Learning Framework for Loca…

200 篇论文

Audio-visual recognition (AVR) has been considered as a solution for speech recognition tasks when the audio is corrupted, as well as a visual recognition method used for speaker verification in multi-speaker scenarios. The approach of AVR…

计算机视觉与模式识别 · 计算机科学 2017-11-01 Amirsina Torfi , Seyed Mehdi Iranmanesh , Nasser M. Nasrabadi , Jeremy Dawson

In this paper we propose a multi-modal multi-correlation learning framework targeting at the task of audio-visual speech separation. Although previous efforts have been extensively put on combining audio and visual modalities, most of them…

声音 · 计算机科学 2022-07-05 Xiaoyu Wang , Xiangyu Kong , Xiulian Peng , Yan Lu

We address prevailing challenges of the brain-powered research, departing from the observation that the literature hardly recover accurate spatial information and require subject-specific models. To address these challenges, we propose…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Weihao Xia , Raoul de Charette , Cengiz Öztireli , Jing-Hao Xue

Visual and audio events simultaneously occur and both attract attention. However, most existing saliency prediction works ignore the influence of audio and only consider vision modality. In this paper, we propose a multitask learning method…

计算机视觉与模式识别 · 计算机科学 2021-11-17 Minglang Qiao , Yufan Liu , Mai Xu , Xin Deng , Bing Li , Weiming Hu , Ali Borji

The audio-visual sound separation field assumes visible sources in videos, but this excludes invisible sounds beyond the camera's view. Current methods struggle with such sounds lacking visible cues. This paper introduces a novel…

计算机视觉与模式识别 · 计算机科学 2023-10-19 Yiyang Su , Ali Vosoughi , Shijian Deng , Yapeng Tian , Chenliang Xu

One-shot voice conversion (VC), which performs conversion across arbitrary speakers with only a single target-speaker utterance for reference, can be effectively achieved by speech representation disentanglement. Existing work generally…

音频与语音处理 · 电气工程与系统科学 2021-07-22 Disong Wang , Liqun Deng , Yu Ting Yeung , Xiao Chen , Xunying Liu , Helen Meng

Developing Audio-Visual Large Language Models (AV-LLMs) for unified scene understanding is pivotal in multimodal intelligence. While instruction tuning enables pre-trained models with multi-task abilities, we observe that conventional…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Dongnuan Cai , Henghui Du , Chang Zhou , Xi Chen , Dan Guo , Hongyuan Zhang , Xuelong Li , Di Hu

Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of…

计算机视觉与模式识别 · 计算机科学 2026-05-06 You Qin , Kai Liu , Shengqiong Wu , Kai Wang , Shijian Deng , Yapeng Tian , Junbin Xiao , Yazhou Xing , Yinghao Ma , Bobo Li , Roger Zimmermann , Lei Cui , Furu Wei , Jiebo Luo , Hao Fei

Autonomous driving systems require a comprehensive understanding of the environment, achieved by extracting visual features essential for perception, planning, and control. However, models trained solely on single-task objectives or generic…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Huy-Dung Nguyen , Anass Bairouk , Mirjana Maras , Wei Xiao , Tsun-Hsuan Wang , Patrick Chareyre , Ramin Hasani , Marc Blanchon , Daniela Rus

Unsupervised learning from visual data is one of the most difficult challenges in computer vision, being a fundamental task for understanding how visual recognition works. From a practical point of view, learning from unsupervised visual…

计算机视觉与模式识别 · 计算机科学 2017-04-03 Ioana Croitoru , Simion-Vlad Bogolin , Marius Leordeanu

During the performance of sound source localization which uses both visual and aural information, it presently remains unclear how much either image or sound modalities contribute to the result, i.e. do we need both image and sound for…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Takashi Oya , Shohei Iwase , Ryota Natsume , Takahiro Itazuri , Shugo Yamaguchi , Shigeo Morishima

Separating an audio scene into isolated sources is a fundamental problem in computer audition, analogous to image segmentation in visual scene analysis. Source separation systems based on deep learning are currently the most successful…

声音 · 计算机科学 2018-11-07 Prem Seetharaman , Gordon Wichern , Jonathan Le Roux , Bryan Pardo

Medical Visual Question Answering (Medical-VQA) aims to to answer clinical questions regarding radiology images, assisting doctors with decision-making options. Nevertheless, current Medical-VQA models learn cross-modal representations…

计算机视觉与模式识别 · 计算机科学 2023-09-28 Chenlu Zhan , Peng Peng , Hongsen Wang , Tao Chen , Hongwei Wang

Audio-visual segmentation is a challenging task that aims to predict pixel-level masks for sound sources in a video. Previous work applied a comprehensive manually designed architecture with countless pixel-wise accurate masks as…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Shentong Mo , Bhiksha Raj

Recent years have witnessed the success of deep learning on the visual sound separation task. However, existing works follow similar settings where the training and testing datasets share the same musical instrument categories, which to…

多媒体 · 计算机科学 2022-03-28 Xinchi Zhou , Dongzhan Zhou , Wanli Ouyang , Hang Zhou , Ziwei Liu , Di Hu

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events with their mixture and…

Audio-Visual Learning (AVL) is one fundamental task of multi-modality learning and embodied intelligence, displaying the vital role in scene understanding and interaction. However, previous researchers mostly focus on exploring downstream…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Muyi Sun , Yixuan Wang , Hong Wang , Chen Su , Man Zhang , Xingqun Qi , Qi Li , Zhenan Sun

Recent audio-visual generative models have made substantial progress in generating images from audio. However, existing approaches focus on generating images from single-class audio and fail to generate images from mixed audio. To address…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Minjae Kang , Martim Brandão

Audio visual segmentation (AVS) aims to segment the sounding objects for each frame of a given video. To distinguish the sounding objects from silent ones, both audio-visual semantic correspondence and temporal interaction are required. The…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Shaofei Huang , Han Li , Yuqing Wang , Hongji Zhu , Jiao Dai , Jizhong Han , Wenge Rong , Si Liu

Pre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predominant pre-training…

音频与语音处理 · 电气工程与系统科学 2025-03-04 Alexander H. Liu , Sang-gil Lee , Chao-Han Huck Yang , Yuan Gong , Yu-Chiang Frank Wang , James R. Glass , Rafael Valle , Bryan Catanzaro