中文
相关论文

相关论文: A Benchmarking Initiative for Audio-Domain Music G…

200 篇论文

Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous speakers across…

声音 · 计算机科学 2023-02-17 Sang-gil Lee , Wei Ping , Boris Ginsburg , Bryan Catanzaro , Sungroh Yoon

Audio-language models (ALMs) generate linguistic descriptions of sound-producing events and scenes. Advances in dataset creation and computational power have led to significant progress in this domain. This paper surveys 69 datasets used to…

声音 · 计算机科学 2025-02-10 Gijs Wijngaard , Elia Formisano , Michele Esposito , Michel Dumontier

The goal of this paper to generate a visually appealing video that responds to music with a neural network so that each frame of the video reflects the musical characteristics of the corresponding audio clip. To achieve the goal, we propose…

声音 · 计算机科学 2021-02-10 Dasaem Jeong , Seungheon Doh , Taegyun Kwon

Music generation has emerged as a significant topic in artificial intelligence and machine learning. While recurrent neural networks (RNNs) have been widely employed for sequence generation, generative adversarial networks (GANs) remain…

声音 · 计算机科学 2025-12-30 Pratik Nag

In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a…

声音 · 计算机科学 2018-05-04 Bin Liu , Shuai Nie , Yaping Zhang , Dengfeng Ke , Shan Liang , Wenju Liu1

Current methods for creating drum loop audio in digital music production, such as using one-shot samples or resampling, often demand non-trivial efforts of creators. While recent generative models achieve high fidelity and adhere to text,…

Evaluating generative models remains a fundamental challenge, particularly when the goal is to reflect human preferences. In this paper, we use music generation as a case study to investigate the gap between automatic evaluation metrics and…

音频与语音处理 · 电气工程与系统科学 2025-05-01 Huan Zhang , Jinhua Liang , Huy Phan , Wenwu Wang , Emmanouil Benetos

A conversational music retrieval system can help users discover music that matches their preferences through dialogue. To achieve this, a conversational music retrieval system should seamlessly engage in multi-turn conversation by 1)…

声音 · 计算机科学 2024-11-13 SeungHeon Doh , Keunwoo Choi , Daeyong Kwon , Taesu Kim , Juhan Nam

In this paper, we propose a novel score-base generative model for unconditional raw audio synthesis. Our proposal builds upon the latest developments on diffusion process modeling with stochastic differential equations, which already…

声音 · 计算机科学 2021-06-15 Simon Rouard , Gaëtan Hadjeres

The intelligibility of speech severely degrades in the presence of environmental noise and reverberation. In this paper, we propose a novel deep learning based system for modifying the speech signal to increase its intelligibility under the…

音频与语音处理 · 电气工程与系统科学 2021-09-17 Haoyu Li , Junichi Yamagishi

Audio codecs are typically transform-domain based and efficiently code stationary audio signals, but they struggle with speech and signals containing dense transient events such as applause. Specifically, with these two classes of signals…

音频与语音处理 · 电气工程与系统科学 2020-01-28 Arijit Biswas , Dai Jia

Music auto-tagging is crucial for enhancing music discovery and recommendation. Existing models in Music Information Retrieval (MIR) struggle with real-world noise such as environmental and speech sounds in multimedia content. This study…

声音 · 计算机科学 2024-01-30 Haesun Joung , Kyogu Lee

Separating two sources from an audio mixture is an important task with many applications. It is a challenging problem since only one signal channel is available for analysis. In this paper, we propose a novel framework for singing voice…

声音 · 计算机科学 2017-11-15 Zhe-Cheng Fan , Yen-Lin Lai , Jyh-Shing Roger Jang

In this paper, we compare different audio signal representations, including the raw audio waveform and a variety of time-frequency representations, for the task of audio synthesis with Generative Adversarial Networks (GANs). We conduct the…

音频与语音处理 · 电气工程与系统科学 2020-06-18 Javier Nistal , Stefan Lattner , Gaël Richard

Controllable generation using StyleGANs is usually achieved by training the model using labeled data. For audio textures, however, there is currently a lack of large semantically labeled datasets. Therefore, to control generation, we…

音频与语音处理 · 电气工程与系统科学 2024-10-08 Purnima Kamath , Chitralekha Gupta , Lonce Wyse , Suranga Nanayakkara

Loops, seamlessly repeatable musical segments, are a cornerstone of modern music production. Contemporary artists often mix and match various sampled or pre-recorded loops based on musical criteria such as rhythm, harmony and timbral…

声音 · 计算机科学 2021-05-24 Pritish Chandna , António Ramires , Xavier Serra , Emilia Gómez

Diffusion models have shown promising results for a wide range of generative tasks with continuous data, such as image and audio synthesis. However, little progress has been made on using diffusion models to generate discrete symbolic music…

声音 · 计算机科学 2023-10-24 Jincheng Zhang , György Fazekas , Charalampos Saitis

Deep generative models have emerged as a promising approach in the medical image domain to address data scarcity. However, their use for sequential data like respiratory sounds is less explored. In this work, we propose a straightforward…

声音 · 计算机科学 2023-11-14 June-Woo Kim , Chihyeon Yoon , Miika Toikkanen , Sangmin Bae , Ho-Young Jung

Deep generative models have recently achieved impressive performance in speech and music synthesis. However, compared to the generation of those domain-specific sounds, generating general sounds (such as siren, gunshots) has received less…

音频与语音处理 · 电气工程与系统科学 2021-10-07 Xubo Liu , Turab Iqbal , Jinzheng Zhao , Qiushi Huang , Mark D. Plumbley , Wenwu Wang

Deep learning requires large datasets for training (convolutional) networks with millions of parameters. In neuroimaging, there are few open datasets with more than 100 subjects, which makes it difficult to, for example, train a classifier…

图像与视频处理 · 电气工程与系统科学 2020-01-27 Anders Eklund