English
Related papers

Related papers: Enhanced Generative Machine Listener

200 papers

This paper presents the Deep learning-based Perceptual Audio Quality metric (DeePAQ) for evaluating general audio quality. Our approach leverages metric learning together with the music foundation model MERT, guided by surrogate labels, to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-15 Guanxin Jiang , Andreas Brendel , Pablo M. Delgado , Jürgen Herre

The real-world capabilities of objective speech quality measures are limited since current measures (1) are developed from simulated data that does not adequately model real environments; or they (2) predict objective scores that are not…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-03 Xuan Dong , Donald S. Williamson

Learning music representations that are general-purpose offers the flexibility to finetune several downstream tasks using smaller datasets. The wav2vec 2.0 speech representation model showed promising results in many downstream speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-28 Alessandro Ragano , Emmanouil Benetos , Andrew Hines

This paper presents a self-supervised learning framework, named MGF, for general-purpose speech representation learning. In the design of MGF, speech hierarchy is taken into consideration. Specifically, we propose to use generative learning…

Sound · Computer Science 2021-02-04 Yucheng Zhao , Dacheng Yin , Chong Luo , Zhiyuan Zhao , Chuanxin Tang , Wenjun Zeng , Zheng-Jun Zha

Remixing separated audio sources trades off interferer attenuation against the amount of audible deteriorations. This paper proposes a non-intrusive audio quality estimation method for controlling this trade-off in a signal-adaptive manner.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-24 Matteo Torcoli , Jouni Paulus , Thorsten Kastner , Christian Uhle

Speech quality assessment is a problem for every researcher working on models that produce or process speech. Human subjective ratings, the gold standard in speech quality assessment, are expensive and time-consuming to acquire in a…

Sound · Computer Science 2023-05-25 Lorenz Diener , Marju Purin , Sten Sootla , Ando Saabas , Robert Aichner , Ross Cutler

For enhancing noisy signals, machine-learning based single-channel speech enhancement schemes exploit prior knowledge about typical speech spectral structures. To ensure a good generalization and to meet requirements in terms of…

Sound · Computer Science 2018-01-17 Robert Rehr , Timo Gerkmann

Safe deployment of large language models (LLMs) may benefit from a reliable method for assessing their generated content to determine when to abstain or to selectively generate. While likelihood-based metrics such as perplexity are widely…

Computation and Language · Computer Science 2023-12-18 Jie Ren , Yao Zhao , Tu Vu , Peter J. Liu , Balaji Lakshminarayanan

Multimodal Generative Models (MGMs) have rapidly evolved beyond text generation, now spanning diverse output modalities including images, music, video, human motion, and 3D objects, by integrating language with other sensory modalities…

Multimedia · Computer Science 2025-11-25 Longzhen Han , Awes Mubarak , Almas Baimagambetov , Nikolaos Polatidis , Thar Baker

Existing objective evaluation metrics for voice conversion (VC) are not always correlated with human perception. Therefore, training VC models with such criteria may not effectively improve naturalness and similarity of converted speech. In…

Sound · Computer Science 2022-03-01 Chen-Chou Lo , Szu-Wei Fu , Wen-Chin Huang , Xin Wang , Junichi Yamagishi , Yu Tsao , Hsin-Min Wang

PESQ and POLQA , are standards are standards for automated assessment of voice quality of speech as experienced by human beings. The predictions of those objective measures should come as close as possible to subjective quality scores as…

Sound · Computer Science 2017-08-22 Dan Elbaz , Michael Zibulevsky

We introduce a new cross-modal fusion technique designed for generative error correction in automatic speech recognition (ASR). Our methodology leverages both acoustic information and external linguistic representations to generate accurate…

Speech quality assessment (SQA) aims to predict the perceived quality of speech signals under a wide range of distortions. It is inherently connected to speech enhancement (SE), which seeks to improve speech quality by removing unwanted…

Sound · Computer Science 2025-08-25 Wei Wang , Wangyou Zhang , Chenda Li , Jiatong Shi , Shinji Watanabe , Yanmin Qian

Generative Spoken Language Modeling research focuses on optimizing speech Language Models (LMs) using raw audio recordings without accessing any textual supervision. Such speech LMs usually operate over discrete units obtained from…

Computation and Language · Computer Science 2023-05-30 Itai Gat , Felix Kreuk , Tu Anh Nguyen , Ann Lee , Jade Copet , Gabriel Synnaeve , Emmanuel Dupoux , Yossi Adi

Audio super-resolution is the task of constructing a high-resolution (HR) audio from a low-resolution (LR) audio by adding the missing band. Previous methods based on convolutional neural networks and mean squared error training objective…

Sound · Computer Science 2021-06-17 Kexun Zhang , Yi Ren , Changliang Xu , Zhou Zhao

The rapid development of generative audio raises ethical and security concerns stemming from forged data, making deepfake sound detection an important safeguard against the malicious use of such technologies. Although prior studies have…

Sound · Computer Science 2025-09-29 Zeyu Xie , Yaoyun Zhang , Xuenan Xu , Yongkang Yin , Chenxing Li , Mengyue Wu , Yuexian Zou

Natural Language Generation (NLG) refers to the operation of expressing the calculation results of a system in human language. Since the quality of generated sentences from an NLG model cannot be fully represented using only quantitative…

Computation and Language · Computer Science 2022-08-04 Dojun Park , Youngjin Jang , Harksoo Kim

Large language models (LLMs) have revolutionized lots of fields of research. Although it is well-known that fine-tuning is essential for enhancing the capabilities of LLMs, existing research suggests that there is potential redundancy in…

Artificial Intelligence · Computer Science 2025-02-14 Haoling Li , Xin Zhang , Xiao Liu , Yeyun Gong , Yifan Wang , Qi Chen , Peng Cheng

Analysing music in the field of machine learning is a very difficult problem with numerous constraints to consider. The nature of audio data, with its very high dimensionality and widely varying scales of structure, is one of the primary…

Sound · Computer Science 2022-05-17 Tracy Qian , Jackson Kaunismaa , Tony Chung

One of the most promising areas of research to obtain practical advantage is Quantum Machine Learning which was born as a result of cross-fertilisation of ideas between Quantum Computing and Classical Machine Learning. In this paper, we…

Quantum Physics · Physics 2021-11-08 N. Schetakis , D. Aghamalyan , M. Boguslavsky , P. Griffin
‹ Prev 1 8 9 10 Next ›