English
Related papers

Related papers: Research on Piano Timbre Transformation System Bas…

200 papers

Recent advances in text-to-music generation models have opened new avenues in musical creativity. However, music generation usually involves iterative refinements, and how to edit the generated music remains a significant challenge. This…

A data set of recorded single played tones of a concert grand piano is investigated using Machine Learning (ML) on psychoacoustic timbre features. The examined instrument has been recorded at two stages: firstly right after manufacture and…

Neurons and Cognition · Quantitative Biology 2021-12-17 Niko Plath , Rolf Bader

Advances in neural network design and the availability of large-scale labeled datasets have driven major improvements in piano transcription. Existing approaches target either offline applications, with no restrictions on computational…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-10 Patricia Hu , Silvan David Peter , Jan Schlüter , Gerhard Widmer

Taking long-term spectral and temporal dependencies into account is essential for automatic piano transcription. This is especially helpful when determining the precise onset and offset for each note in the polyphonic piano content. In this…

Sound · Computer Science 2023-07-11 Keisuke Toyama , Taketo Akama , Yukara Ikemiya , Yuhta Takida , Wei-Hsiang Liao , Yuki Mitsufuji

Neural audio synthesis is an actively researched topic, having yielded a wide range of techniques that leverages machine learning architectures. Google Magenta elaborated a novel approach called Differential Digital Signal Processing (DDSP)…

Existing methods for expressive music performance rendering rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and…

Sound · Computer Science 2025-12-03 Hong-Jie You , Jie-Jing Shao , Xiao-Wen Yang , Lin-Han Jia , Lan-Zhe Guo , Yu-Feng Li

Neural style transfer, allowing to apply the artistic style of one image to another, has become one of the most widely showcased computer vision applications shortly after its introduction. In contrast, related tasks in the music audio…

Sound · Computer Science 2021-06-11 Ondřej Cífka , Alexey Ozerov , Umut Şimşekli , Gaël Richard

Deep learning researches on the transformation problems for image and text have raised great attention. However, present methods for music feature transfer using neural networks are far from practical application. In this paper, we initiate…

Sound · Computer Science 2021-08-05 Xutan Peng , Chen Li , Zhi Cai , Faqiang Shi , Yidan Liu , Jianxin Li

Voice conversion is a common speech synthesis task which can be solved in different ways depending on a particular real-world scenario. The most challenging one often referred to as one-shot many-to-many voice conversion consists in copying…

Sound · Computer Science 2022-08-05 Vadim Popov , Ivan Vovk , Vladimir Gogoryan , Tasnima Sadekova , Mikhail Kudinov , Jiansheng Wei

Piano cover generation aims to automatically transform a pop song into a piano arrangement. While numerous deep learning approaches have been proposed, existing models often fail to maintain structural consistency with the original song,…

Sound · Computer Science 2026-01-26 Tse-Yang Chen , Yuh-Jzer Joung

Although recent speech processing technologies have achieved significant improvements in objective metrics, there still remains a gap in human perceptual quality. This paper proposes Diffiner, a novel solution that utilizes the powerful…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-11 Masato Hirano , Ryosuke Sawata , Naoki Murata , Shusuke Takahashi , Yuki Mitsufuji

In recent years, the burgeoning interest in diffusion models has led to significant advances in image and speech generation. Nevertheless, the direct synthesis of music waveforms from unrestricted textual prompts remains a relatively…

Sound · Computer Science 2023-09-22 Pengfei Zhu , Chao Pang , Yekun Chai , Lei Li , Shuohuan Wang , Yu Sun , Hao Tian , Hua Wu

Diffusion based Text-To-Music (TTM) models generate music corresponding to text descriptions. Typically UNet based diffusion models condition on text embeddings generated from a pre-trained large language model or from a cross-modality…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-28 Jisi Zhang , Pablo Peso Parada , Md Asif Jalal , Karthikeyan Saravanan

Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text prompts. We show that by training a generative model on long…

Sound · Computer Science 2024-07-30 Zach Evans , Julian D. Parker , CJ Carr , Zack Zukowski , Josiah Taylor , Jordi Pons

Existing music generation models are mostly language-based, neglecting the frequency continuity property of notes, resulting in inadequate fitting of rare or never-used notes and thus reducing the diversity of generated samples. We argue…

Sound · Computer Science 2024-08-06 Shipei Liu , Xiaoya Fan , Guowei Wu

Recent advancements in text-to-image diffusion models have significantly transformed visual content generation, yet their application in specialized fields such as interior design remains underexplored. In this paper, we present…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Zhaowei Wang , Ying Hao , Hao Wei , Qing Xiao , Lulu Chen , Yulong Li , Yue Yang , Tianyi Li

This paper aims to apply a new deep learning approach to the task of generating raw audio files. It is based on diffusion models, a recent type of deep generative model. This new type of method has recently shown outstanding results with…

Sound · Computer Science 2023-07-21 Svetlana Pavlova

Diffusion probabilistic models have demonstrated an outstanding capability to model natural images and raw audio waveforms through a paired diffusion and reverse processes. The unique property of the reverse process (namely, eliminating…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-23 Yen-Ju Lu , Yu Tsao , Shinji Watanabe

Text-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this process demands extensive, high-quality data, which are often…

Sound · Computer Science 2025-06-18 Chang Li , Ruoyu Wang , Lijuan Liu , Jun Du , Yixuan Sun , Zilu Guo , Zhenrong Zhang , Yuan Jiang , Jianqing Gao , Feng Ma

Recent advances in latent diffusion models have demonstrated state-of-the-art performance in high-dimensional time-series data synthesis while providing flexible control through conditioning and guidance. However, existing methodologies…

Machine Learning · Computer Science 2025-11-11 Matteo Pettenó , Alessandro Ilic Mezza , Alberto Bernardini
‹ Prev 1 4 5 6 7 8 10 Next ›