中文
相关论文

相关论文: Microphone Conversion: Mitigating Device Variabili…

200 篇论文

CycleGAN provides a framework to train image-to-image translation with unpaired datasets using cycle consistency loss [4]. While results are great in many applications, the pixel level cycle consistency can potentially be problematic and…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Tongzhou Wang , Yihan Lin

Singing technique conversion (STC) refers to the task of converting from one voice technique to another while leaving the original singer identity, melody, and linguistic components intact. Previous STC studies, as well as singing voice…

音频与语音处理 · 电气工程与系统科学 2023-08-22 Tung-Cheng Su , Yung-Chuan Chang , Yi-Wen Liu

This technical report describes the systems submitted to the DCASE2022 challenge task 3: sound event localization and detection (SELD). The task aims to detect occurrences of sound events and specify their class, furthermore estimate their…

声音 · 计算机科学 2025-12-30 Jin Sob Kim , Hyun Joon Park , Wooseok Shin , Sung Won Han

Continuously learning new classes without catastrophic forgetting is a challenging problem for on-device environmental sound classification given the restrictions on computation resources (e.g., model size, running memory). To address this…

声音 · 计算机科学 2022-07-19 Yang Xiao , Xubo Liu , James King , Arshdeep Singh , Eng Siong Chng , Mark D. Plumbley , Wenwu Wang

Current research in synthesized speech detection primarily focuses on the generalization of detection systems to unknown spoofing methods of noise-free speech. However, the performance of anti-spoofing countermeasures (CM) system is often…

声音 · 计算机科学 2024-07-30 Yikang Wang , Xingming Wang , Hiromitsu Nishizaki , Ming Li

Deep learning-based sound event localization and classification is an emerging research area within wireless acoustic sensor networks. However, current methods for sound event localization and classification typically rely on a single…

This paper presents a configurable version of Extreme Bandwidth Extension Network (EBEN), a Generative Adversarial Network (GAN) designed to improve audio captured with body-conduction microphones. We show that although these microphones…

音频与语音处理 · 电气工程与系统科学 2024-07-18 Julien Hauret , Thomas Joubaud , Véronique Zimpfer , Éric Bavu

Auscultation plays a pivotal role in early respiratory and pulmonary disease diagnosis. Despite the emergence of deep learning-based methods for automatic respiratory sound classification post-Covid-19, limited datasets impede performance…

声音 · 计算机科学 2025-03-04 Yun Chu , Qiuhao Wang , Enze Zhou , Ling Fu , Qian Liu , Gang Zheng

This work provides a solution to the challenge of small amounts of training data in Non-Destructive Ultrasonic Testing for composite components. It was demonstrated that direct simulation alone is ineffective at producing training data that…

图像与视频处理 · 电气工程与系统科学 2023-11-06 Shaun McKnight , S. Gareth Pierce , Ehsan Mohseni , Christopher MacKinnon , Charles MacLeod , Tom OHare , Charalampos Loukas

Consistency regularization (CR), which enforces agreement between model predictions on augmented views, has found recent benefits in automatic speech recognition [1]. In this paper, we propose the use of consistency regularization for audio…

声音 · 计算机科学 2025-09-15 Shanmuka Sadhu , Weiran Wang

Generative adversarial networks (GAN) have recently been shown to be efficient for speech enhancement. However, most, if not all, existing speech enhancement GANs (SEGAN) make use of a single generator to perform one-stage enhancement…

机器学习 · 计算机科学 2020-10-28 Huy Phan , Ian V. McLoughlin , Lam Pham , Oliver Y. Chén , Philipp Koch , Maarten De Vos , Alfred Mertins

Purpose: The objective of this work is to introduce an advanced framework designed to enhance ultrasound images, especially those captured by portable hand-held devices, which often produce lower quality images due to hardware constraints.…

图像与视频处理 · 电气工程与系统科学 2024-12-19 Shreeram Athreya , Ashwath Radhachandran , Vedrana Ivezić , Vivek Sant , Corey W. Arnold , William Speier

Speech enhancement aims to obtain speech signals with high intelligibility and quality from noisy speech. Recent work has demonstrated the excellent performance of time-domain deep learning methods, such as Conv-TasNet. However, these…

声音 · 计算机科学 2021-09-21 Feiyang Xiao , Jian Guan , Qiuqiang Kong , Wenwu Wang

Training of speech enhancement systems often does not incorporate knowledge of human perception and thus can lead to unnatural sounding results. Incorporating psychoacoustically motivated speech perception metrics as part of model training…

声音 · 计算机科学 2022-06-16 George Close , Thomas Hain , Stefan Goetze

Non-parallel training is a difficult but essential task for DNN-based speech enhancement methods, for the lack of adequate noisy and paired clean speech corpus in many real scenarios. In this paper, we propose a novel adaptive…

声音 · 计算机科学 2021-09-15 Guochen Yu , Yutian Wang , Chengshi Zheng , Hui Wang , Qin Zhang

The expression of emotion is highly individualistic. However, contemporary speech emotion recognition (SER) systems typically rely on population-level models that adopt a `one-size-fits-all' approach for predicting emotion. Moreover,…

计算与语言 · 计算机科学 2025-04-11 Andreas Triantafyllopoulos , Björn Schuller

In recent years, Generative Adversarial Networks (GANs) have produced significantly improved results in speech enhancement (SE) tasks. They are difficult to train, however. In this work, we introduce several improvements to the GAN training…

声音 · 计算机科学 2022-10-27 Vasily Zadorozhnyy , Qiang Ye , Kazuhito Koishida

In Sound Event Detection (SED) systems, the lengths of median filters for post-processing have never been optimized during training due to several problems. No gradient is received by the lengths so they cannot be learned during…

音频与语音处理 · 电气工程与系统科学 2021-03-23 Fengnian Zhao , Ruwei Li , Xin Liu , Liwen Xu

Joint sound event localization and detection (SELD) is an integral part of developing context awareness into communication interfaces of mobile robots, smartphones, and home assistants. For example, an automatic audio focus for video…

音频与语音处理 · 电气工程与系统科学 2021-06-29 Pasi Pertilä , Emre Cakir , Aapo Hakala , Eemi Fagerlund , Tuomas Virtanen , Archontis Politis , Antti Eronen

Recent research has delved into speech enhancement (SE) approaches that leverage audio embeddings from pre-trained models, diverging from time-frequency masking or signal prediction techniques. This paper introduces an efficient and…

音频与语音处理 · 电气工程与系统科学 2025-06-16 Xingwei Sun , Heinrich Dinkel , Yadong Niu , Linzhang Wang , Junbo Zhang , Jian Luan