中文
相关论文

相关论文: Perceptual audio loss function for deep learning

200 篇论文

Automatic subjective speech quality assessment (SSQA) traditionally estimates speech quality on an utterance or system level. While this resolution was adequate for older transmission or synthesis systems that produced speech signals of…

音频与语音处理 · 电气工程与系统科学 2026-05-21 Michael Kuhlmann , Tobias Cord-Landwehr , Reinhold Haeb-Umbach

In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. Its aim is to improve speaker recognition performance when speech signals are corrupted by noise. Instead of…

计算与语言 · 计算机科学 2020-05-25 Yanpei Shi , Qiang Huang , Thomas Hain

This paper investigates several aspects of training a RNN (recurrent neural network) that impact the objective and subjective quality of enhanced speech for real-time single-channel speech enhancement. Specifically, we focus on a RNN that…

音频与语音处理 · 电气工程与系统科学 2020-02-14 Yangyang Xia , Sebastian Braun , Chandan K. A. Reddy , Harishchandra Dubey , Ross Cutler , Ivan Tashev

Despite the growing popularity of metric learning approaches, very little work has attempted to perform a fair comparison of these techniques for speaker verification. We try to fill this gap and compare several metric learning loss…

机器学习 · 计算机科学 2020-04-02 Juan M. Coria , Hervé Bredin , Sahar Ghannay , Sophie Rosset

Content and style representations have been widely studied in the field of style transfer. In this paper, we propose a new loss function using speaker content representation for audio source separation, and we call it speaker representation…

声音 · 计算机科学 2020-02-28 Seongkyu Mun , Soyeon Choe , Jaesung Huh , Joon Son Chung

Objective speech quality measures are widely used to assess the performance of video conferencing platforms and telecommunication systems. They predict human-rated speech quality and are crucial for assessing the systems quality of…

音频与语音处理 · 电气工程与系统科学 2025-05-23 Javier Perez , Dimme de Groot , Jorge Martinez

Constructing an embedding space for musical instrument sounds that can meaningfully represent new and unseen instruments is important for downstream music generation tasks such as multi-instrument synthesis and timbre transfer. The…

音频与语音处理 · 电气工程与系统科学 2021-12-28 Xuan Shi , Erica Cooper , Junichi Yamagishi

In this paper, we propose StableQuant, a novel adaptive post-training quantization (PTQ) algorithm for widely used speech foundation models (SFMs). While PTQ has been successfully employed for compressing large language models (LLMs) due to…

音频与语音处理 · 电气工程与系统科学 2025-04-22 Yeona Hong , Hyewon Han , Woo-jin Chung , Hong-Goo Kang

Recent years have witnessed the success of the deep learning-based technique in research of no-reference point cloud quality assessment (NR-PCQA). For a more accurate quality prediction, many previous studies have attempted to capture…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Yujie Zhang , Qi Yang , Ziyu Shan , Yiling Xu

After constructing a deep neural network for urban sound classification, this work focuses on the sensitive application of assisting drivers suffering from hearing loss. As such, clear etiology justifying and interpreting model predictions…

声音 · 计算机科学 2021-11-22 Marco Colussi , Stavros Ntalampiras

Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been…

音频与语音处理 · 电气工程与系统科学 2021-03-16 Daniel Michelsanti , Zheng-Hua Tan , Shi-Xiong Zhang , Yong Xu , Meng Yu , Dong Yu , Jesper Jensen

The objective of this paper is 'open-set' speaker recognition of unseen speakers, where ideal embeddings should be able to condense information into a compact utterance-level representation that has small intra-speaker and large…

音频与语音处理 · 电气工程与系统科学 2020-11-05 Joon Son Chung , Jaesung Huh , Seongkyu Mun , Minjae Lee , Hee Soo Heo , Soyeon Choe , Chiheon Ham , Sunghwan Jung , Bong-Jin Lee , Icksang Han

Recent advances in deep learning have facilitated the design of speaker verification systems that directly input raw waveforms. For example, RawNet extracts speaker embeddings from raw waveforms, which simplifies the process pipeline and…

音频与语音处理 · 电气工程与系统科学 2020-05-08 Jee-weon Jung , Seung-bin Kim , Hye-jin Shim , Ju-ho Kim , Ha-Jin Yu

Supervised learning based on a deep neural network recently has achieved substantial improvement on speech enhancement. Denoising networks learn mapping from noisy speech to clean one directly, or to a spectrum mask which is the ratio…

声音 · 计算机科学 2023-03-10 Jaeyoung Kim , Mostafa El-Khamy , Jungwon Lee

Traditionally, the vision community has devised algorithms to estimate the distance between an original image and images that have been subject to perturbations. Inspiration was usually taken from the human visual perceptual system and how…

机器学习 · 计算机科学 2020-11-18 Alexander Hepburn , Valero Laparra , Jesús Malo , Ryan McConville , Raul Santos-Rodriguez

The learning of interpretable representations from raw data presents significant challenges for time series data like speech. In this work, we propose a relevance weighting scheme that allows the interpretation of the speech representations…

音频与语音处理 · 电气工程与系统科学 2020-11-05 Purvi Agrawal , Sriram Ganapathy

Detecting machine malfunctions at an early stage is crucial for reducing interruptions in operational processes within industrial settings. Recently, the deep learning approach has started to be preferred for the detection of failures in…

声音 · 计算机科学 2023-12-05 Mustafa Yurdakul , Sakir Tasdemir

Many audio processing tasks require perceptual assessment. The ``gold standard`` of obtaining human judgments is time-consuming, expensive, and cannot be used as an optimization criterion. On the other hand, automated metrics are efficient…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Pranay Manocha , Adam Finkelstein , Richard Zhang , Nicholas J. Bryan , Gautham J. Mysore , Zeyu Jin

Neural audio compression has emerged as a promising technology for efficiently representing speech, music, and general audio. However, existing methods suffer from significant performance degradation at limited bitrates, where the available…

声音 · 计算机科学 2026-05-08 Jin Wang , Wenbin Jiang , Xiangbo Wang , Yubo You , Sheng Fang

Detection of common events and scenes from audio is useful for extracting and understanding human contexts in daily life. Prior studies have shown that leveraging knowledge from a relevant domain is beneficial for a target acoustic event…

‹ 上一页 1 8 9 10 下一页 ›