English
Related papers

Related papers: Adapting Frechet Audio Distance for Generative Mus…

200 papers

Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative models may introduce…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-19 Paul A. Bereuter , Benjamin Stahl , Mark D. Plumbley , Alois Sontacchi

Recent advancements have brought generated music closer to human-created compositions, yet evaluating these models remains challenging. While human preference is the gold standard for assessing quality, translating these subjective…

Machine Learning · Computer Science 2025-06-25 Florian Grötschla , Ahmet Solak , Luca A. Lanzendörfer , Roger Wattenhofer

The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding…

Sound · Computer Science 2021-10-12 Cheng Gong , Longbiao Wang , Zhenhua Ling , Ju Zhang , Jianwu Dang

We present two new metrics for evaluating generative models in the class-conditional image generation setting. These metrics are obtained by generalizing the two most popular unconditional metrics: the Inception Score (IS) and the Fre'chet…

Computer Vision and Pattern Recognition · Computer Science 2021-02-09 Yaniv Benny , Tomer Galanti , Sagie Benaim , Lior Wolf

Generative models are designed to address the data scarcity problem. Even with the exploding amount of data, due to computational advancements, some applications (e.g., health care, weather forecast, fault detection) still suffer from data…

Machine Learning · Computer Science 2024-05-07 Alireza Koochali , Maria Walch , Sankrutyayan Thota , Peter Schichtel , Andreas Dengel , Sheraz Ahmed

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-17 Dimitrios Bralios , Gordon Wichern , François G. Germain , Zexu Pan , Sameer Khurana , Chiori Hori , Jonathan Le Roux

Recent audio generation models typically rely on Variational Autoencoders (VAEs) and perform generation within the VAE latent space. Although VAEs excel at compression and reconstruction, their latents inherently encode low-level acoustic…

Sound · Computer Science 2026-02-27 Zeyu Xie , Chenxing Li , Qiao Jin , Xuenan Xu , Guanrou Yang , Wenfu Wang , Mengyue Wu , Dong Yu , Yuexian Zou

The criteria for measuring music similarity are important for developing a flexible music recommendation system. Some data-driven methods have been proposed to calculate music similarity from only music signals, such as metric learning…

Sound · Computer Science 2022-11-16 Yuka Hashizume , Li Li , Tomoki Toda

This paper explores a specific sub-task of cross-modal music retrieval. We consider the delicate task of retrieving a performance or rendition of a musical piece based on a description of its style, expressive character, or emotion from a…

Sound · Computer Science 2024-01-29 Shreyan Chowdhury , Gerhard Widmer

Although perceptual (dis)similarity between sensory stimuli seems akin to distance, measuring the Euclidean distance between vector representations of auditory stimuli is a poor estimator of subjective dissimilarity. In hearing, nonlinear…

Neurons and Cognition · Quantitative Biology 2020-11-03 Sarah Oh , Elijah FW Bowen , Antonio Rodriguez , Damian Sowinski , Eva Childers , Annemarie Brown , Laura Ray , Richard Granger

Fake audio detection is a growing concern and some relevant datasets have been designed for research. However, there is no standard public Chinese dataset under complex conditions.In this paper, we aim to fill in the gap and design a…

Sound · Computer Science 2023-07-19 Haoxin Ma , Jiangyan Yi , Chenglong Wang , Xinrui Yan , Jianhua Tao , Tao Wang , Shiming Wang , Ruibo Fu

Timbre spaces have been used in music perception to study the perceptual relationships between instruments based on dissimilarity ratings. However, these spaces do not generalize to novel examples and do not provide an invertible mapping,…

Sound · Computer Science 2018-10-02 Philippe Esling , Axel Chemla--Romeu-Santos , Adrien Bitton

While both the data volume and heterogeneity of the digital music content is huge, it has become increasingly important and convenient to build a recommendation or search system to facilitate surfacing these content to the user or consumer…

The growth of generative adversarial network (GAN) models has increased the ability of image processing and provides numerous industries with the technology to produce realistic image transformations. However, with the field being recently…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Ricardo de Deijn , Aishwarya Batra , Brandon Koch , Naseef Mansoor , Hema Makkena

This paper presents a boundary element method (BEM) for computing the energy transmittance of a singly-periodic grating in 2D for a wide frequency band, which is of engineering interest in various fields with possible applications to…

Numerical Analysis · Mathematics 2023-05-16 Yuta Honshuku , Hiroshi Isakari

Speech embeddings are fixed-size acoustic representations of variable-length speech sequences. They are increasingly used for a variety of tasks ranging from information retrieval to unsupervised term discovery and speech segmentation.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-09 Robin Algayres , Mohamed Salah Zaiem , Benoit Sagot , Emmanuel Dupoux

Training GANs under limited data often leads to discriminator overfitting and memorization issues, causing divergent training. Existing approaches mitigate the overfitting by employing data augmentations, model regularization, or attention…

Computer Vision and Pattern Recognition · Computer Science 2022-10-12 Mengping Yang , Zhe Wang , Ziqiu Chi , Yanbing Zhang

Large language models reveal deep comprehension and fluent generation in the field of multi-modality. Although significant advancements have been achieved in audio multi-modality, existing methods are rarely leverage language model for…

Sound · Computer Science 2024-08-06 Hualei Wang , Jianguo Mao , Zhifang Guo , Jiarui Wan , Hong Liu , Xiangdong Wang

High-fidelity general audio compression at ultra-low bitrates is crucial for applications ranging from low-bandwidth communication to generative audio-language modeling. Traditional audio compression methods and contemporary neural codecs…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-03 Hao Ma , Ruihao Jing , Shansong Liu , Cheng Gong , Chi Zhang , Xiao-Lei Zhang , Xuelong Li

In late 2011, Fado was elevated to the oral and intangible heritage of humanity by UNESCO. This study aims to develop a tool for automatic detection of Fado music based on the audio signal. To do this, frequency spectrum-related…

Sound · Computer Science 2014-06-18 Pedro Girão Antunes , David Martins de Matos , Ricardo Ribeiro , Isabel Trancoso