中文
相关论文

相关论文: CCStereo: Audio-Visual Contextual and Contrastive …

200 篇论文

Recent advances in audio-text cross-modal contrastive learning have shown its potential towards zero-shot learning. One possibility for this is by projecting item embeddings from pre-trained backbone neural networks into a cross-modal space…

声音 · 计算机科学 2025-09-29 Tiago Tavares , Fabio Ayres , Zhepei Wang , Paris Smaragdis

Sounding Video Generation (SVG) is an audio-video joint generation task challenged by high-dimensional signal spaces, distinct data formats, and different patterns of content information. To address these issues, we introduce a novel…

计算机视觉与模式识别 · 计算机科学 2024-10-03 Mingzhen Sun , Weining Wang , Yanyuan Qiao , Jiahui Sun , Zihan Qin , Longteng Guo , Xinxin Zhu , Jing Liu

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

声音 · 计算机科学 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement…

音频与语音处理 · 电气工程与系统科学 2024-03-11 Vikas Tokala , Eric Grinstein , Mike Brookes , Simon Doclo , Jesper Jensen , Patrick A. Naylor

In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a…

声音 · 计算机科学 2018-05-04 Bin Liu , Shuai Nie , Yaping Zhang , Dengfeng Ke , Shan Liang , Wenju Liu1

This paper introduces a novel approach to speech restoration by integrating a context-related conditioning strategy. Specifically, we employ the diffusion-based generative restoration model, UNIVERSE++, as a backbone to evaluate the…

音频与语音处理 · 电气工程与系统科学 2025-08-13 Soo-Whan Chung , Min-Seok Choi

Adversarial perturbations are noise-like patterns that can subtly change the data, while failing an otherwise accurate classifier. In this paper, we propose to use such perturbations within a novel contrastive learning setup to build…

计算机视觉与模式识别 · 计算机科学 2020-04-17 Jue Wang , Anoop Cherian

The objective optimization of medical imaging systems requires full characterization of all sources of randomness in the measured data, which includes the variability within the ensemble of objects to-be-imaged. This can be accomplished by…

图像与视频处理 · 电气工程与系统科学 2020-01-28 Weimin Zhou , Sayantan Bhadra , Frank J. Brooks , Hua Li , Mark A. Anastasio

Bootstrap-based Self-Supervised Learning (SSL) has achieved remarkable progress in audio understanding. However, existing methods typically operate at a single level of granularity, limiting their ability to model the diverse temporal and…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Bing Han , Chushu Zhou , Yifan Yang , Wei Wang , Chenda Li , Wangyou Zhang , Yanmin Qian

A binaural rendering framework for personal sound zones (PSZs) is proposed to enable multiple head-tracked listeners to receive fully independent stereo audio programs. Current PSZ systems typically rely on monophonic rendering and…

音频与语音处理 · 电气工程与系统科学 2026-01-13 Hao Jiang , Edgar Choueiri

Understanding how infants perceive speech sounds and language structures is still an open problem. Previous research in artificial neural networks has mainly focused on large dataset-dependent generative models, aiming to replicate…

人工智能 · 计算机科学 2024-12-24 Xiaodan Chen , Alexandre Pitti , Mathias Quoy , Nancy F Chen

Deep learning has the potential to enhance speech signals and increase their intelligibility for users of hearing aids. Deep models suited for real-world application should feature a low computational complexity and low processing delay of…

音频与语音处理 · 电气工程与系统科学 2024-10-31 Nils L. Westhausen , Hendrik Kayser , Theresa Jansen , Bernd T. Meyer

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

声音 · 计算机科学 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li

While interest in models that generalize at test time to new compositions has risen in recent years, benchmarks in the visually-grounded domain have thus far been restricted to synthetic images. In this work, we propose COVR, a new test-bed…

计算与语言 · 计算机科学 2021-09-23 Ben Bogin , Shivanshu Gupta , Matt Gardner , Jonathan Berant

For fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate representations extracted from speech should serve as a…

音频与语音处理 · 电气工程与系统科学 2023-12-19 Chunyu Qiang , Hao Li , Yixin Tian , Ruibo Fu , Tao Wang , Longbiao Wang , Jianwu Dang

In a recent paper, we have presented a generative adversarial network (GAN)-based model for unconditional generation of the mel-spectrograms of singing voices. As the generator of the model is designed to take a variable-length sequence of…

音频与语音处理 · 电气工程与系统科学 2021-05-13 Jen-Yu Liu , Yu-Hua Chen , Yin-Cheng Yeh , Yi-Hsuan Yang

Sound event detection (SED) and Acoustic scene classification (ASC) are two widely researched audio tasks that constitute an important part of research on acoustic scene analysis. Considering shared information between sound events and…

声音 · 计算机科学 2022-09-14 Daniel Aleksander Krause , Annamaria Mesaros

We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-to-audio models achieve strong semantic…

Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous speakers across…

声音 · 计算机科学 2023-02-17 Sang-gil Lee , Wei Ping , Boris Ginsburg , Bryan Catanzaro , Sungroh Yoon

Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech signals, which is then…

声音 · 计算机科学 2020-12-18 Mostafa Sadeghi , Simon Leglaive , Xavier Alameda-PIneda , Laurent Girin , Radu Horaud