English
Related papers

Related papers: MIDI-VALLE: Improving Expressive Piano Performance…

200 papers

A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instantaneous velocity,…

Sound · Computer Science 2025-09-09 Xiaoran Yang , Jianxuan Yang , Xinyue Guo , Haoyu Wang , Ningning Pan , Gongping Huang

Automated piano performance evaluation traditionally relies on symbolic (MIDI) representations, which capture note-level information but miss the acoustic nuances that characterize expressive playing. I propose using pre-trained audio…

Sound · Computer Science 2026-01-28 Jai Dhiman

This paper explores a specific sub-task of cross-modal music retrieval. We consider the delicate task of retrieving a performance or rendition of a musical piece based on a description of its style, expressive character, or emotion from a…

Sound · Computer Science 2024-01-29 Shreyan Chowdhury , Gerhard Widmer

Emotions are fundamental to the creation and perception of music performances. However, achieving human-like expression and emotion through machine learning models for performance rendering remains a challenging task. In this work, we…

Sound · Computer Science 2025-11-06 Ilya Borovik , Dmitrii Gavrilev , Vladimir Viro

We present a controllable neural audio synthesizer based on Gaussian Mixture Variational Autoencoders (GM-VAE), which can generate realistic piano performances in the audio domain that closely follows temporal conditions of two essential…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-14 Hao Hao Tan , Yin-Jyun Luo , Dorien Herremans

Existing methods for expressive music performance rendering rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and…

Sound · Computer Science 2025-12-03 Hong-Jie You , Jie-Jing Shao , Xiao-Wen Yang , Lin-Han Jia , Lan-Zhe Guo , Yu-Feng Li

With the help of discrete neural audio codecs, large language models (LLM) have increasingly been recognized as a promising methodology for zero-shot Text-to-Speech (TTS) synthesis. However, sampling based decoding strategies bring…

Computation and Language · Computer Science 2024-06-13 Bing Han , Long Zhou , Shujie Liu , Sanyuan Chen , Lingwei Meng , Yanming Qian , Yanqing Liu , Sheng Zhao , Jinyu Li , Furu Wei

Robot musicians require precise control to obtain proper note accuracy, sound quality, and musical expression. Performance of string instruments, such as violin and cello, presents a significant challenge due to the precise control required…

This paper proposes a controllable singing voice synthesis system capable of generating expressive singing voice with two novel methodologies. First, a local style token module, which predicts frame-level style tokens from an input pitch…

Sound · Computer Science 2022-04-08 Juheon Lee , Hyeong-Seok Choi , Kyogu Lee

Generating musical audio directly with neural networks is notoriously difficult because it requires coherently modeling structure at many different timescales. Fortunately, most music is also highly structured and can be represented as…

Recent advances in discrete audio codecs have significantly improved speech representation modeling, while codec language models have enabled in-context learning for zero-shot speech synthesis. Inspired by this, we propose a voice…

Sound · Computer Science 2025-09-30 Junchuan Zhao , Xintong Wang , Ye Wang

This comprehensive paper delves into the forefront of personalized voice synthesis within artificial intelligence (AI), spotlighting the Dynamic Individual Voice Synthesis Engine (DIVSE). DIVSE represents a groundbreaking leap in…

Sound · Computer Science 2024-01-01 Fan Shi

We propose a cross-lingual neural codec language model, VALL-E X, for cross-lingual speech synthesis. Specifically, we extend VALL-E and train a multi-lingual conditional codec language model to predict the acoustic token sequences of the…

Computation and Language · Computer Science 2023-03-08 Ziqiang Zhang , Long Zhou , Chengyi Wang , Sanyuan Chen , Yu Wu , Shujie Liu , Zhuo Chen , Yanqing Liu , Huaming Wang , Jinyu Li , Lei He , Sheng Zhao , Furu Wei

Generating realistic drum audio directly from symbolic representations is a challenging task at the intersection of music perception and machine learning. We propose a system that transforms an expressive drum grid, a time-aligned MIDI…

This paper describes a computational model of loudness variations in expressive ensemble performance. The model predicts and explains the continuous variation of loudness as a function of information extracted automatically from the written…

Sound · Computer Science 2016-12-19 Thassilo Gadermaier , Maarten Grachten , Carlos Eduardo Cancino Chacón

Controllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with robustness and generalization due to data quality constraints,…

Sound · Computer Science 2025-06-27 Rui Niu , Weihao Wu , Jie Chen , Long Ma , Zhiyong Wu

Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recordings. However, not all speaking styles are easy to model:…

Autoregressive models are now capable of generating high-quality minute-long expressive MIDI piano performances. Even though this progress suggests new tools to assist music composition, we observe that generative algorithms are still not…

Sound · Computer Science 2021-07-14 Gaëtan Hadjeres , Léopold Crestel

Audio-visual synchronization aims to determine whether the mouth movements and speech in the video are synchronized. VocaLiST reaches state-of-the-art performance by incorporating multimodal Transformers to model audio-visual interact…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Xuanjun Chen , Haibin Wu , Chung-Che Wang , Hung-yi Lee , Jyh-Shing Roger Jang

This study focuses on the perception of music performances when contextual factors, such as room acoustics and instrument, change. We propose to distinguish the concept of "performance" from the one of "interpretation", which expresses the…

Sound · Computer Science 2022-03-08 Federico Simonetta , Federico Avanzini , Stavros Ntalampiras