English
Related papers

Related papers: AudioGen: Textually Guided Audio Generation

200 papers

We introduce Audio-Agent, a multimodal framework for audio generation, editing and composition based on text or video inputs. Conventional approaches for text-to-audio (TTA) tasks often make single-pass inferences from text descriptions.…

Sound · Computer Science 2025-01-15 Zixuan Wang , Chi-Keung Tang , Yu-Wing Tai

Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input audio as a unimodal retrieval query. In contrast, we propose…

Sound · Computer Science 2025-06-11 Choi Changin , Lim Sungjun , Rhee Wonjong

The ability to automatically generate music that appropriately matches an arbitrary input track is a challenging task. We present a novel controllable system for generating single stems to accompany musical mixes of arbitrary length. At the…

Sound · Computer Science 2024-02-05 Marco Pasini , Maarten Grachten , Stefan Lattner

In recent years, with the realistic generation results and a wide range of personalized applications, diffusion-based generative models gain huge attention in both visual and audio generation areas. Compared to the considerable advancements…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Shiqi Yang , Zhi Zhong , Mengjie Zhao , Shusuke Takahashi , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

We introduce MAGNeT, a masked generative sequence modeling method that operates directly over several streams of audio tokens. Unlike prior work, MAGNeT is comprised of a single-stage, non-autoregressive transformer. During training, we…

Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene Synthesis challenge held as part of the Detection and…

Text-to-speech models trained on large-scale datasets have demonstrated impressive in-context learning capabilities and naturalness. However, control of speaker identity and style in these models typically requires conditioning on reference…

Sound · Computer Science 2024-02-08 Dan Lyth , Simon King

With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite producing high-quality outputs, existing text-to-audio models…

Sound · Computer Science 2026-04-28 Yi Yuan , Xubo Liu , Haohe Liu , Xiyuan Kang , Zhuo Chen , Yuxuan Wang , Mark D. Plumbley , Wenwu Wang

Generative AI has demonstrated impressive performance in various fields, among which speech synthesis is an interesting direction. With the diffusion model as the most popular generative model, numerous works have attempted two active…

The field of text-to-audio generation has seen significant advancements, and yet the ability to finely control the acoustic characteristics of generated audio remains under-explored. In this paper, we introduce a novel yet simple approach…

Sound · Computer Science 2024-12-16 Sonal Kumar , Prem Seetharaman , Justin Salamon , Dinesh Manocha , Oriol Nieto

The design of diffusion-based audio generation systems has been investigated from diverse perspectives, such as data space, network architecture, and conditioning techniques, while most of these innovations require model re-training. In…

Sound · Computer Science 2026-04-10 Junyou Wang , Zehua Chen , Binjie Yuan , Kaiwen Zheng , Chang Li , Yuxuan Jiang , Jun Zhu

This paper studies constrained text generation, which is to generate sentences under certain pre-conditions. We focus on CommonGen, the task of generating text based on a set of concepts, as a representative task of constrained text…

Computation and Language · Computer Science 2021-03-15 Yixian Liu , Liwen Zhang , Wenjuan Han , Yue Zhang , Kewei Tu

Autoregressive (AR) models with diffusion heads have recently achieved strong text-to-audio performance, yet their iterative decoding and multi-step sampling process introduce high-latency issues. To address this bottleneck, we propose a…

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-17 Dimitrios Bralios , Gordon Wichern , François G. Germain , Zexu Pan , Sameer Khurana , Chiori Hori , Jonathan Le Roux

End-to-end generation of musical audio using deep learning techniques has seen an explosion of activity recently. However, most models concentrate on generating fully mixed music in response to abstract conditioning information. In this…

We introduce EgoSonics, a method to generate semantically meaningful and synchronized audio tracks conditioned on silent egocentric videos. Generating audio for silent egocentric videos could open new applications in virtual reality,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Aashish Rai , Srinath Sridhar

It is an open challenge to obtain high quality training data, especially captions, for text-to-audio models. Although prior methods have leveraged \textit{text-only language models} to augment and improve captions, such methods have…

Computation and Language · Computer Science 2024-07-10 Zhifeng Kong , Sang-gil Lee , Deepanway Ghosal , Navonil Majumder , Ambuj Mehrish , Rafael Valle , Soujanya Poria , Bryan Catanzaro

Separating the individual elements in a musical mixture is an essential process for music analysis and practice. While this is generally addressed using neural networks optimized to mask or transform the time-frequency representation of a…

Sound · Computer Science 2025-11-27 Genís Plaja-Roglans , Yun-Ning Hung , Xavier Serra , Igor Pereira

Recently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events. Compared with classification labels, the advantages of such…

Sound · Computer Science 2024-03-08 Xuenan Xu , Xiaohang Xu , Zeyu Xie , Pingyue Zhang , Mengyue Wu , Kai Yu

We present AGGGEN (pronounced 'again'), a data-to-text model which re-introduces two explicit sentence planning stages into neural data-to-text systems: input ordering and input aggregation. In contrast to previous work using sentence…

Computation and Language · Computer Science 2021-06-18 Xinnuo Xu , Ondřej Dušek , Verena Rieser , Ioannis Konstas