中文
相关论文

相关论文: Sound Source Localization for a Source inside a St…

200 篇论文

Recent work has shown that recurrent neural networks can be trained to separate individual speakers in a sound mixture with high fidelity. Here we explore convolutional neural network models as an alternative and show that they achieve…

声音 · 计算机科学 2018-05-29 Shariq Mobin , Brian Cheung , Bruno Olshausen

A key challenge in segmentation in digital histopathology is inter- and intra-stain variations as it reduces model performance. Labelling each stain is expensive and time-consuming so methods using stain transfer via CycleGAN, have been…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Zeeshan Nisar , Friedrich Feuerhake , Thomas Lampert

The separation of single-channel underwater acoustic signals is a challenging problem with practical significance. Few existing studies focus on the source separation problem with unknown numbers of signals, and how to evaluate the…

声音 · 计算机科学 2024-05-29 Qinggang Sun , Kejun Wang

Conventional approaches to sound source localization require at least two microphones. It is known, however, that people with unilateral hearing loss can also localize sounds. Monaural localization is possible thanks to the scattering by…

音频与语音处理 · 电气工程与系统科学 2018-08-29 Dalia El Badawy , Ivan Dokmanić

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provides a natural and…

音频与语音处理 · 电气工程与系统科学 2024-12-03 Xubo Liu , Qiuqiang Kong , Yan Zhao , Haohe Liu , Yi Yuan , Yuzhuo Liu , Rui Xia , Yuxuan Wang , Mark D. Plumbley , Wenwu Wang

Environmental Sound Classification (ESC) is an active research area in the audio domain and has seen a lot of progress in the past years. However, many of the existing approaches achieve high accuracy by relying on domain-specific features…

计算机视觉与模式识别 · 计算机科学 2020-04-17 Andrey Guzhov , Federico Raue , Jörn Hees , Andreas Dengel

Sound source localization in visual scenes aims to localize objects emitting the sound in a given image. Recent works showing impressive localization performance typically rely on the contrastive learning framework. However, the random…

计算机视觉与模式识别 · 计算机科学 2022-03-28 Zengjie Song , Yuxi Wang , Junsong Fan , Tieniu Tan , Zhaoxiang Zhang

Learning how to localize and separate individual object sounds in the audio channel of the video is a difficult task. Current state-of-the-art methods predict audio masks from artificially mixed spectrograms, known as Mix-and-Separate…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Tanzila Rahman , Leonid Sigal

This paper proposes Scyclone, a high-quality voice conversion (VC) technique without parallel data training. Scyclone improves speech naturalness and speaker similarity of the converted speech by introducing CycleGAN-based spectrogram…

音频与语音处理 · 电气工程与系统科学 2020-05-08 Masaya Tanaka , Takashi Nose , Aoi Kanagaki , Ryohei Shimizu , Akira Ito

Our previous work, the unified source-filter GAN (uSFGAN) vocoder, introduced a novel architecture based on the source-filter theory into the parallel waveform generative adversarial network to achieve high voice quality and pitch…

声音 · 计算机科学 2023-02-28 Reo Yoneyama , Yi-Chiao Wu , Tomoki Toda

The problem of source localization with ad hoc microphone networks in noisy and reverberant enclosures, given a training set of prerecorded measurements, is addressed in this paper. The training set is assumed to consist of a limited number…

声音 · 计算机科学 2016-10-18 Bracha Laufer-Goldshtein , Ronen Talmon , Sharon Gannot

Sound source tracking is commonly performed using classical array-processing algorithms, while machine-learning approaches typically rely on precise source position labels that are expensive or impractical to obtain. This paper introduces a…

音频与语音处理 · 电气工程与系统科学 2026-02-12 Luan Vinícius Fiorio , Ivana Nikoloska , Bruno Defraene , Alex Young , Johan David , Ronald M. Aarts

We showcase an unsupervised method that repurposes deep models trained for music generation and music tagging for audio source separation, without any retraining. An audio generation model is conditioned on an input mixture, producing a…

声音 · 计算机科学 2021-10-26 Ethan Manilow , Patrick O'Reilly , Prem Seetharaman , Bryan Pardo

This paper introduces a new framework for supervised sound source localization referred to as virtually-supervised learning. An acoustic shoe-box room simulator is used to generate a large number of binaural single-source audio scenes.…

声音 · 计算机科学 2017-03-21 Saurabh Kataria , Clément Gaultier , Antoine Deleforge

This paper introduces a novel neural network-based speech coding system that can process noisy speech effectively. The proposed source-aware neural audio coding (SANAC) system harmonizes a deep autoencoder-based source separation model and…

音频与语音处理 · 电气工程与系统科学 2020-11-11 Haici Yang , Kai Zhen , Seungkwon Beack , Minje Kim

Accurate sound source localization (SSL), such as direction-of-arrival (DoA) estimation, relies on consistent multichannel data. However, batteryless systems often suffer from missing data due to the stochastic nature of energy harvesting,…

机器学习 · 计算机科学 2025-07-21 Subrata Biswas , Mohammad Nur Hossain Khan , Violet Colwell , Jack Adiletta , Bashima Islam

Separating vocal elements from musical tracks is a longstanding challenge in audio signal processing. This study tackles the distinct separation of vocal components from musical spectrograms. We employ the Short Time Fourier Transform…

声音 · 计算机科学 2024-05-31 Adam Sorrenti

Despite progress in audio classification, a generalization gap remains between speech and other sound domains, such as environmental sounds and music. Models trained for speech tasks often fail to perform well on environmental or musical…

声音 · 计算机科学 2024-06-14 Heinrich Dinkel , Zhiyong Yan , Yongqing Wang , Junbo Zhang , Yujun Wang , Bin Wang

Transfer learning of StyleGAN has recently shown great potential to solve diverse tasks, especially in domain translation. Previous methods utilized a source model by swapping or freezing weights during transfer learning, however, they have…

计算机视觉与模式识别 · 计算机科学 2023-03-23 Dongyeun Lee , Jae Young Lee , Doyeon Kim , Jaehyun Choi , Junmo Kim

The image source method (ISM) is often used to simulate room acoustics due to its ease of use and computational efficiency. The standard ISM is limited to simulations of room impulse responses between point sources and omnidirectional…

音频与语音处理 · 电气工程与系统科学 2023-09-08 Zeyu Xu , Adrian Herzog , Alexander Lodermeyer , Emanuël A. P. Habets , Albert G. Prinn