中文
相关论文

相关论文: Adapting Frechet Audio Distance for Generative Mus…

200 篇论文

Perceptual metrics are traditionally used to evaluate the quality of natural signals, such as images and audio. They are designed to mimic the perceptual behaviour of human observers and usually reflect structures found in natural signals.…

声音 · 计算机科学 2023-12-07 Tashi Namgyal , Alexander Hepburn , Raul Santos-Rodriguez , Valero Laparra , Jesus Malo

Synthetic images are an option for augmenting limited medical imaging datasets to improve the performance of various machine learning models. A common metric for evaluating synthetic image quality is the Fr\'echet Inception Distance (FID)…

图像与视频处理 · 电气工程与系统科学 2025-07-30 Thomas Wallace , Ik Siong Heng , Senad Subasic , Chris Messenger

The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused…

This study focuses on the perception of music performances when contextual factors, such as room acoustics and instrument, change. We propose to distinguish the concept of "performance" from the one of "interpretation", which expresses the…

声音 · 计算机科学 2022-03-08 Federico Simonetta , Federico Avanzini , Stavros Ntalampiras

Significant advancements have been made in video generative models recently. Unlike image generation, video generation presents greater challenges, requiring not only generating high-quality frames but also ensuring temporal consistency…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Jiahe Liu , Youran Qu , Qi Yan , Xiaohui Zeng , Lele Wang , Renjie Liao

The Fr\'echet Inception Distance (FID) has been used to evaluate hundreds of generative models. We introduce FastFID, which can efficiently train generative models with FID as a loss function. Using FID as an additional loss for Generative…

机器学习 · 计算机科学 2021-04-15 Alexander Mathiasen , Frederik Hvilshøj

The evaluation of audio fingerprinting at a realistic scale is limited by the scarcity of large public music databases. We present an audio-free approach that synthesises latent fingerprints which approximate the distribution of real…

声音 · 计算机科学 2025-09-24 Aditya Bhattacharjee , Marco Pasini , Emmanouil Benetos

Generative Adversarial Networks (GANs) are by far the most successful generative models. Learning the transformation which maps a low dimensional input noise to the data distribution forms the foundation for GANs. Although they have been…

机器学习 · 计算机科学 2020-04-16 Manisha Padala , Debojit Das , Sujit Gujar

Chord progression generation is practically important but understudied. Most large-scale symbolic music systems target melody, multi-track arrangement, or audio synthesis, and chord-only models tend to be relegated to conditioning…

声音 · 计算机科学 2026-05-07 Jinju Lee

With success on controlled tasks, generative models are being increasingly applied to humanitarian applications [1,2]. In this paper, we focus on the evaluation of a conditional generative model that illustrates the consequences of climate…

机器学习 · 计算机科学 2019-10-23 Sharon Zhou , Alexandra Luccioni , Gautier Cosne , Michael S. Bernstein , Yoshua Bengio

Perceptual similarity representations enable music retrieval systems to determine which songs sound most similar to listeners. State-of-the-art approaches based on task-specific training via self-supervised metric learning show promising…

声音 · 计算机科学 2026-01-28 Arhan Vohra , Taketo Akama

Detecting AI-generated music is crucial for preserving artistic authenticity and preventing the misuse of generative music technologies. However, existing discriminative detectors typically rely on generated samples during training and…

声音 · 计算机科学 2026-05-19 Chaolei Han , Hongsong Wang , Jie Gui

The escalating challenges of managing vast sensor-generated data, particularly in audio applications, necessitate innovative solutions. Current systems face significant computational and storage demands, especially in real-time applications…

Recent approaches in music generation rely on disentangled representations, often labeled as structure and timbre or local and global, to enable controllable synthesis. Yet the underlying properties of these embeddings remain underexplored.…

Generative systems of musical accompaniments are rapidly growing, yet there are no standardized metrics to evaluate how well generations align with the conditional audio prompt. We introduce a distribution-based measure called…

声音 · 计算机科学 2025-04-09 Maarten Grachten , Javier Nistal

Over the past few decades, computational methods have been developed to estimate perceptual audio quality. These methods, also referred to as objective quality measures, are usually developed and intended for a specific application domain.…

音频与语音处理 · 电气工程与系统科学 2021-10-25 Matteo Torcoli , Thorsten Kastner , Jürgen Herre

Extracting individual elements from music mixtures is a valuable tool for music production and practice. While neural networks optimized to mask or transform mixture spectrograms into the individual source(s) have been the leading approach,…

声音 · 计算机科学 2025-11-26 Genís Plaja-Roglans , Yun-Ning Hung , Xavier Serra , Igor Pereira

The evaluation of deep generative models has been extensively studied in the centralized setting, where the reference data are drawn from a single probability distribution. On the other hand, several applications of generative models…

机器学习 · 计算机科学 2024-06-12 Zixiao Wang , Farzan Farnia , Zhenghao Lin , Yunheng Shen , Bei Yu

Singing voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authenticity. Unlike Audio…

声音 · 计算机科学 2023-09-07 Yuankun Xie , Jingjing Zhou , Xiaolin Lu , Zhenghao Jiang , Yuxin Yang , Haonan Cheng , Long Ye

Foley sound presents the background sound for multimedia content and the generation of Foley sound involves computationally modelling sound effects with specialized techniques. In this work, we proposed a system for DCASE 2023 challenge…

声音 · 计算机科学 2023-09-18 Yi Yuan , Haohe Liu , Xubo Liu , Xiyuan Kang , Mark D. Plumbley , Wenwu Wang