English
Related papers

Related papers: Adapting Frechet Audio Distance for Generative Mus…

200 papers

In audio-related creative tasks, sound designers often seek to extend and morph different sounds from their libraries. Generative audio models, capable of creating audio using examples as references, offer promising solutions. By masking…

Sound · Computer Science 2026-02-20 Prem Seetharaman , Oriol Nieto , Justin Salamon

Fr\'echet Inception Distance (FID) is a widely used metric for assessing synthetic image quality. It relies on an ImageNet-based feature extractor, making its applicability to medical imaging unclear. A recent trend is to adapt FID to…

Conditional Generative Adversarial Networks (cGANs) are finding increasingly widespread use in many application domains. Despite outstanding progress, quantitative evaluation of such models often involves multiple distinct metrics to assess…

Computer Vision and Pattern Recognition · Computer Science 2019-12-25 Terrance DeVries , Adriana Romero , Luis Pineda , Graham W. Taylor , Michal Drozdzal

Automatic Music Transcription (AMT) -- the task of converting music audio into note representations -- has seen rapid progress, driven largely by deep learning systems. Due to the limited availability of richly annotated music datasets,…

Sound · Computer Science 2026-01-27 Lukáš Samuel Marták , Patricia Hu , Gerhard Widmer

In this paper, we propose and investigate the use of neural audio codec language models for the automatic generation of sample-based musical instruments based on text or reference audio prompts. Our approach extends a generative audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-23 Shahan Nercessian , Johannes Imort , Ninon Devis , Frederik Blang

Generative artificial intelligence (AI) models in smart grids have advanced significantly in recent years due to their ability to generate large amounts of synthetic data, which would otherwise be difficult to obtain in the real world due…

Machine Learning · Computer Science 2025-10-27 Yuting Cai , Shaohuai Liu , Chao Tian , Le Xie

The rapid advancement of Generative Adversarial Networks (GANs) necessitates the need to robustly evaluate these models. Among the established evaluation criteria, the Fr\'{e}chetInception Distance (FID) has been widely adopted due to its…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Lorenzo Luzi , Helen Jenne , Ryan Murray , Carlos Ortiz Marrero

Speech synthesis is an important practical generative modeling problem that has seen great progress over the last few years, with likelihood-based autoregressive neural models now outperforming traditional concatenative systems. A downside…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-26 Alexey A. Gritsenko , Tim Salimans , Rianne van den Berg , Jasper Snoek , Nal Kalchbrenner

This work evaluates the robustness of quality measures of generative models such as Inception Score (IS) and Fr\'echet Inception Distance (FID). Analogous to the vulnerability of deep models against a variety of adversarial attacks, we show…

Machine Learning · Computer Science 2022-07-21 Motasem Alfarra , Juan C. Pérez , Anna Frühstück , Philip H. S. Torr , Peter Wonka , Bernard Ghanem

The advent of Music-Language Models has greatly enhanced the automatic music generation capability of AI systems, but they are also limited in their coverage of the musical genres and cultures of the world. We present a study of the…

Recent years have seen considerable advances in audio synthesis with deep generative models. However, the state-of-the-art is very difficult to quantify; different studies often use different evaluation methodologies and different metrics…

Sound · Computer Science 2022-09-02 Ashvala Vinay , Alexander Lerch

Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio Detection (FAD), as we expect that MOS can be used to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-26 Wangjin Zhou , Zhengdong Yang , Chenhui Chu , Sheng Li , Raj Dabre , Yi Zhao , Tatsuya Kawahara

Audio applications involving environmental sound analysis increasingly use general-purpose audio representations, also known as embeddings, for transfer learning. Recently, Holistic Evaluation of Audio Representations (HEAR) evaluated…

The rapid advancement of natural language processing, information retrieval (IR), computer vision, and other technologies has presented significant challenges in evaluating the performance of these systems. One of the main challenges is the…

Information Retrieval · Computer Science 2024-02-20 Negar Arabzadeh , Charles L. A. Clarke

Modern metrics for generative learning like Fr\'echet Inception Distance (FID) and DINOv2-Fr\'echet Distance (FD-DINOv2) demonstrate impressive performance. However, they suffer from various shortcomings, like a bias towards specific…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Lokesh Veeramacheneni , Moritz Wolter , Hildegard Kuehne , Juergen Gall

Evaluating text-to-image and text-to-video models is challenging due to a fundamental disconnect: established metrics fail to jointly measure visual quality and semantic alignment with text, leading to a poor correlation with human…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Jaywon Koo , Jefferson Hernandez , Moayed Haji-Ali , Ziyan Yang , Vicente Ordonez

Generative adversarial networks (GANs) and diffusion models have recently achieved state-of-the-art performance in audio super-resolution (ADSR), producing perceptually convincing wideband audio from narrowband inputs. However, existing…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-08 Mikhail Silaev , Konstantinos Drossos , Tuomas Virtanen

Diffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited…

Sound · Computer Science 2023-08-04 Ke Chen , Yusong Wu , Haohe Liu , Marianna Nezhurina , Taylor Berg-Kirkpatrick , Shlomo Dubnov

With the rapid advancement of generative audio models, distinguishing between human-composed and generated music is becoming increasingly challenging. As a response, models for detecting fake music have been proposed. In this work, we…

Sound · Computer Science 2025-07-15 Tomasz Sroka , Tomasz Wężowicz , Dominik Sidorczuk , Mateusz Modrzejewski

In this article, we highlight what appears to be major issue of Variational Autoencoders, evinced from an extensive experimentation with different network architectures and datasets: the variance of generated data is significantly lower…

Machine Learning · Computer Science 2020-05-26 Andrea Asperti