English
Related papers

Related papers: Neural Music Synthesis for Flexible Timbre Control

200 papers

This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded in neural waveform…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-22 Takenori Yoshimura , Shinji Takaki , Kazuhiro Nakamura , Keiichiro Oura , Yukiya Hono , Kei Hashimoto , Yoshihiko Nankaku , Keiichi Tokuda

While Large Language Models (LLMs) make symbolic music generation increasingly accessible, producing music with distinctive composition and rich expressiveness remains a significant challenge. Many studies have introduced emotion models to…

Sound · Computer Science 2025-11-19 Dengyun Huang , Yonghua Zhu

Music classification has been one of the most popular tasks in the field of music information retrieval. With the development of deep learning models, the last decade has seen impressive improvements in a wide range of classification tasks.…

Sound · Computer Science 2023-07-03 Yiwei Ding , Alexander Lerch

Customizing voice and speaking style in a speech synthesis system with intuitive and fine-grained controls is challenging, given that little data with appropriate labels is available. Furthermore, editing an existing human's voice also…

Sound · Computer Science 2023-10-27 Florian Lux , Pascal Tilli , Sarina Meyer , Ngoc Thang Vu

We present in this paper PerformacnceNet, a neural network model we proposed recently to achieve score-to-audio music generation. The model learns to convert a music piece from the symbolic domain to the audio domain, assigning…

Sound · Computer Science 2019-05-29 Yu-Hua Chen , Bryan Wang , Yi-Hsuan Yang

Controlling the variations of sound effects using neural audio synthesis models has been a difficult task. Differentiable digital signal processing (DDSP) provides a lightweight solution that achieves high-quality sound synthesis while…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Yunyi Liu , Craig Jin , David Gunawan

Neural Audio Synthesis (NAS) models offer interactive musical control over high-quality, expressive audio generators. While these models can operate in real-time, they often suffer from high latency, making them unsuitable for intimate…

Sound · Computer Science 2025-04-15 Franco Caspe , Jordie Shier , Mark Sandler , Charalampos Saitis , Andrew McPherson

In audio processing applications, the generation of expressive sounds based on high-level representations demonstrates a high demand. These representations can be used to manipulate the timbre and influence the synthesis of creative…

Sound · Computer Science 2023-01-19 Anastasia Natsiou , Luca Longo , Sean O'Leary

This early example of neural synthesis is a proof-of-concept for how machine learning can drive new types of music software. Creating music can be as simple as specifying a set of music influences on which a model trains. We demonstrate a…

Sound · Computer Science 2018-11-19 CJ Carr , Zack Zukowski

While most music generation models generate a mixture of stems (in mono or stereo), we propose to train a multi-stem generative model with 3 stems (bass, drums and other) that learn the musical dependencies between them. To do so, we train…

Sound · Computer Science 2025-01-08 Simon Rouard , Robin San Roman , Yossi Adi , Axel Roebel

Our goal is to be able to build a generative model from a deep neural network architecture to try to create music that has both harmony and melody and is passable as music composed by humans. Previous work in music generation has mainly…

Machine Learning · Computer Science 2016-06-16 Allen Huang , Raymond Wu

Audio textures are a subset of environmental sounds, often defined as having stable statistical characteristics within an adequately large window of time but may be unstructured locally. They include common everyday sounds such as from…

Sound · Computer Science 2020-11-26 M. Huzaifah , L. Wyse

Physical models of rigid bodies are used for sound synthesis in applications from virtual environments to music production. Traditional methods such as modal synthesis often rely on computationally expensive numerical solvers, while recent…

Sound · Computer Science 2022-10-31 Rodrigo Diaz , Ben Hayes , Charalampos Saitis , György Fazekas , Mark Sandler

The ''pretraining-and-finetuning'' paradigm has become a norm for training domain-specific models in natural language processing and computer vision. In this work, we aim to examine this paradigm for symbolic music generation through…

Sound · Computer Science 2023-11-22 Weihan Xu , Julian McAuley , Shlomo Dubnov , Hao-Wen Dong

Emotional and controllable speech synthesis is a topic that has received much attention. However, most studies focused on improving the expressiveness and controllability in the context of linguistic content, even though natural verbal…

Sound · Computer Science 2022-01-27 Hieu-Thi Luong , Junichi Yamagishi

In this paper, we introduce a simple method that can separate arbitrary musical instruments from an audio mixture. Given an unaligned MIDI transcription for a target instrument from an input mixture, we synthesize new mixtures from the midi…

Sound · Computer Science 2020-09-30 Ethan Manilow , Bryan Pardo

Synthesizer is a type of electronic musical instrument that is now widely used in modern music production and sound design. Each parameters configuration of a synthesizer produces a unique timbre and can be viewed as a unique instrument.…

Sound · Computer Science 2022-07-29 Zui Chen , Yansen Jing , Shengcheng Yuan , Yifei Xu , Jian Wu , Hang Zhao

In this study, we define the identity of the singer with two independent concepts - timbre and singing style - and propose a multi-singer singing synthesis system that can model them separately. To this end, we extend our single-singer…

Sound · Computer Science 2019-10-30 Juheon Lee , Hyeong-Seok Choi , Junghyun Koo , Kyogu Lee

This paper proposes a WaveNet-based neural excitation model (ExcitNet) for statistical parametric speech synthesis systems. Conventional WaveNet-based neural vocoding systems significantly improve the perceptual quality of synthesized…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-23 Eunwoo Song , Kyungguen Byun , Hong-Goo Kang

A sound synthesis model for woodwind instruments is developed using modal decomposition of the input impedance, accounting for viscothermal losses as well as localized nonlinear losses at the end of the resonator. To extend the definition…

Classical Physics · Physics 2024-01-12 N Szwarcberg , T Colinot , C Vergez , M Jousserand