中文
相关论文

相关论文: Exploiting Audio-Visual Consistency with Partial S…

200 篇论文

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

Masked modeling has emerged as a powerful self-supervised learning framework, but existing methods largely rely on random masking, disregarding the structural properties of different modalities. In this work, we introduce structured…

机器学习 · 计算机科学 2025-03-21 Aritra Bhowmik , Fida Mohammad Thoker , Carlos Hinojosa , Bernard Ghanem , Cees G. M. Snoek

Generating combined visual and auditory sensory experiences is critical for the consumption of immersive content. Recent advances in neural generative models have enabled the creation of high-resolution content across multiple modalities…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Rishit Dagli , Shivesh Prakash , Robert Wu , Houman Khosravani

We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (limited) video data…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Ho Kei Cheng , Masato Ishii , Akio Hayakawa , Takashi Shibuya , Alexander Schwing , Yuki Mitsufuji

Monaural speech enhancement has achieved remarkable progress recently. However, its performance has been constrained by the limited spatial cues available at a single microphone. To overcome this limitation, we introduce a strategy to map…

音频与语音处理 · 电气工程与系统科学 2024-03-05 Xinmeng Xu , Yuhong Yang , Weiping Tu

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very…

音频与语音处理 · 电气工程与系统科学 2020-02-21 Abhinav Shukla , Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localization without…

计算机视觉与模式识别 · 计算机科学 2021-12-23 Di Hu , Yake Wei , Rui Qian , Weiyao Lin , Ruihua Song , Ji-Rong Wen

We present a multimodal framework to learn general audio representations from videos. Existing contrastive audio representation learning methods mainly focus on using the audio modality alone during training. In this work, we show that…

声音 · 计算机科学 2021-04-29 Luyu Wang , Pauline Luc , Adria Recasens , Jean-Baptiste Alayrac , Aaron van den Oord

Given an input sound signal and a target virtual sound source, sound spatialisation algorithms manipulate the signal so that a listener perceives it as though it were emitted from the target source. There exist several established…

声音 · 计算机科学 2017-11-28 Ali Tarzan , Marco Alunno , Paolo Bientinesi

The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings. In this work, we show that ambient sounds can be used as a supervisory signal for learning visual…

计算机视觉与模式识别 · 计算机科学 2017-12-21 Andrew Owens , Jiajun Wu , Josh H. McDermott , William T. Freeman , Antonio Torralba

When watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlying correlation between audio and visual events, which can be…

多媒体 · 计算机科学 2020-08-19 Ying Cheng , Ruize Wang , Zhihao Pan , Rui Feng , Yuejie Zhang

Auditory and visual signals usually present together and correlate with each other, not only in natural environments but also in clinical settings. However, the audio-visual modelling in the latter case can be more challenging, due to the…

计算机视觉与模式识别 · 计算机科学 2023-10-26 Jianbo Jiao , Mohammad Alsharid , Lior Drukker , Aris T. Papageorghiou , Andrew Zisserman , J. Alison Noble

Spatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized equipment and skills, posing a high barrier for amateur video creators. We…

人机交互 · 计算机科学 2024-04-24 Zheng Ning , Zheng Zhang , Jerrick Ban , Kaiwen Jiang , Ruohong Gan , Yapeng Tian , Toby Jia-Jun Li

Spatial audio quality is a highly multifaceted concept, with many interactions between environmental, geometrical, anatomical, psychological, and contextual considerations. Methods for characterization or evaluation of the geometrical…

音频与语音处理 · 电气工程与系统科学 2024-08-27 Karn N. Watcharasupat , Alexander Lerch

Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality was approached separately, which poses limitations. The…

We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce the target video, then…

多媒体 · 计算机科学 2026-03-18 Masato Ishii , Akio Hayakawa , Takashi Shibuya , Yuki Mitsufuji

Speaker diarization, the process of segmenting an audio stream or transcribed speech content into homogenous partitions based on speaker identity, plays a crucial role in the interpretation and analysis of human speech. Most existing…

机器学习 · 计算机科学 2024-08-23 Luyao Cheng , Hui Wang , Siqi Zheng , Yafeng Chen , Rongjie Huang , Qinglin Zhang , Qian Chen , Xihao Li

The capture and reproduction of spatial audio is becoming increasingly popular, with the mushrooming of applications in teleconferencing, entertainment and virtual reality. Many binaural reproduction methods have been developed and studied…

音频与语音处理 · 电气工程与系统科学 2023-11-27 Ami Berger , Vladimir Tourbabin , Jacob Donley , Zamir Ben-Hur , Boaz Rafaely

In this paper, we propose a quality-aware end-to-end audio-visual neural speaker diarization framework, which comprises three key techniques. First, our audio-visual model takes both audio and visual features as inputs, utilizing a series…

多媒体 · 计算机科学 2024-10-31 Mao-Kui He , Jun Du , Shu-Tong Niu , Qing-Feng Liu , Chin-Hui Lee

Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listening effort. This survey reviews recent advances in deep…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Xikun Lu , Yunda Chen , Zehua Chen , Jie Wang , Mingxing Liu , Hongmei Hu , Chengshi Zheng , Stefan Bleeck , Jinqiu Sang