English
Related papers

Related papers: Adapting Frechet Audio Distance for Generative Mus…

200 papers

Audio embeddings are crucial tools in understanding large catalogs of music. Typically embeddings are evaluated on the basis of the performance they provide in a wide range of downstream tasks, however few studies have investigated the…

Several methods have been developed to assess the perceptual quality of audio under transforms like lossy compression. However, they require paired reference signals of the unaltered content, limiting their use in applications where…

Sound · Computer Science 2021-04-06 Agrin Hilmkil , Carl Thomé , Anders Arpteg

This work is an update of a previous paper on the same topic published a few years ago. With the dramatic progress in generative modeling, a suite of new quantitative and qualitative techniques to evaluate models has emerged. Although some…

Machine Learning · Computer Science 2021-10-05 Ali Borji

Automatic sample identification (ASID), the detection and identification of portions of audio recordings that have been reused in new musical works, is an essential but challenging task in the field of audio query-based retrieval. While a…

Sound · Computer Science 2025-06-23 Aditya Bhattacharjee , Ivan Meresman Higgs , Mark Sandler , Emmanouil Benetos

Cross-domain few-shot learning (CD-FSL) requires models to generalize from limited labeled samples under significant distribution shifts. While recent methods enhance adaptability through lightweight task-specific modules, they operate…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Ruixiao Shi , Fu Feng , Yucheng Xie , Jing Wang , Xin Geng

We show that Fr\'echet Distance (FD), long considered impractical as a training objective, can in fact be effectively optimized in the representation space. Our idea is simple: decouple the population size for FD estimation (e.g., 50k) from…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Jiawei Yang , Zhengyang Geng , Xuan Ju , Yonglong Tian , Yue Wang

We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-04 Dyah A. M. G. Wisnu , Ryandhimas E. Zezario , Stefano Rini , Hsin-Min Wang , Yu Tsao

Evaluating the performance of generative models in image synthesis is a challenging task. Although the Fr\'echet Inception Distance is a widely accepted evaluation metric, it integrates different aspects (e.g., fidelity and diversity) of…

Computer Vision and Pattern Recognition · Computer Science 2021-06-07 Ryoungwoo Jang , Minjee Kim , Da-in Eun , Kyungjin Cho , Jiyeon Seo , Namkug Kim

Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene Synthesis challenge held as part of the Detection and…

Image reconstruction and synthesis have witnessed remarkable progress thanks to the development of generative models. Nonetheless, gaps could still exist between the real and generated images, especially in the frequency domain. In this…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Liming Jiang , Bo Dai , Wayne Wu , Chen Change Loy

Collecting large, aligned cross-modal datasets for music-flavor research is difficult because perceptual experiments are costly and small by design. We address this bottleneck through two complementary experiments. The first tests whether…

Sound · Computer Science 2026-04-14 Matteo Spanio , Valentina Frezzato , Antonio Rodà

Generative audio models are rapidly advancing in both capabilities and public utilization -- several powerful generative audio models have readily available open weights, and some tech companies have released high quality generative audio…

The development of models for learning music similarity and feature extraction from audio media files is an increasingly important task for the entertainment industry. This work proposes a novel music classification model based on metric…

Sound · Computer Science 2019-09-19 Angelo C. Mendes da Silva , Mauricio A. Nunes , Raul Fonseca Neto

Modern audio systems universally employ mel-scale representations derived from 1940s Western psychoacoustic studies, potentially encoding cultural biases that create systematic performance disparities. We present a comprehensive evaluation…

Sound · Computer Science 2026-04-14 Shivam Chauhan , Ajay Pundhir

Embedding acoustic information into fixed length representations is of interest for a whole range of applications in speech and audio technology. Two novel unsupervised approaches to generate acoustic embeddings by modelling of acoustic…

Computation and Language · Computer Science 2021-02-08 Yanpei Shi , Thomas Hain

We introduce a new metric to assess the quality of generated images that is more reliable, data-efficient, compute-efficient, and adaptable to new domains than the previous metrics, such as Fr\'echet Inception Distance (FID). The proposed…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Pranav Jeevan , Neeraj Nixon , Amit Sethi

Every artist has a creative process that draws inspiration from previous artists and their works. Today, "inspiration" has been automated by generative music models. The black box nature of these models obscures the identity of the works…

Sound · Computer Science 2024-01-29 Julia Barnett , Hugo Flores Garcia , Bryan Pardo

Audio diffusion models can synthesize a wide variety of sounds. Existing models often operate on the latent domain with cascaded phase recovery modules to reconstruct waveform. This poses challenges when generating high-fidelity audio. In…

Sound · Computer Science 2023-11-21 Ge Zhu , Yutong Wen , Marc-André Carbonneau , Zhiyao Duan

Foley sound generation, the art of creating audio for multimedia, has recently seen notable advancements through text-conditioned latent diffusion models. These systems use multimodal text-audio representation models, such as Contrastive…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-15 Tornike Karchkhadze , Hassan Salami Kavaki , Mohammad Rasool Izadi , Bryce Irvin , Mikolaj Kegler , Ari Hertz , Shuo Zhang , Marko Stamenovic

Audio embeddings enable large scale comparisons of the similarity of audio files for applications such as search and recommendation. Due to the subjectivity of audio similarity, it can be desirable to design systems that answer not only…