中文
相关论文

相关论文: MIDI-VALLE: Improving Expressive Piano Performance…

200 篇论文

A text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model the acoustic conditions explicitly, in order to disentangle…

Music creation is typically composed of two parts: composing the musical score, and then performing the score with instruments to make sounds. While recent work has made much progress in automatic music generation in the symbolic domain,…

声音 · 计算机科学 2018-11-13 Bryan Wang , Yi-Hsuan Yang

Automatic Music Transcription (AMT) is a vital technology in the field of music information processing. Despite recent enhancements in performance due to machine learning techniques, current methods typically attain high accuracy in domains…

声音 · 计算机科学 2024-07-04 Gakusei Sato , Taketo Akama

The recent success of raw audio waveform synthesis models like WaveNet motivates a new approach for music synthesis, in which the entire process --- creating audio samples from a score and instrument information --- is modeled using…

声音 · 计算机科学 2018-11-02 Jong Wook Kim , Rachel Bittner , Aparna Kumar , Juan Pablo Bello

Neural audio synthesis methods now allow specifying ideas in natural language. However, these methods produce results that cannot be easily tweaked, as they are based on large latent spaces and up to billions of uninterpretable parameters.…

声音 · 计算机科学 2024-06-04 Manuel Cherep , Nikhil Singh , Jessica Shand

Performance-score synchronization is an integral task in signal processing, which entails generating an accurate mapping between an audio recording of a performance and the corresponding musical score. Traditional synchronization methods…

声音 · 计算机科学 2022-04-20 Ruchit Agrawal , Daniel Wolff , Simon Dixon

The purpose of this study is to investigate how humans interpret musical scores expressively, and then design machines that sing like humans. We consider six factors that have a strong influence on the expression of human singing. The…

声音 · 计算机科学 2015-02-17 Ju-Chiang Wang , Hung-Yan Gu , Hsin-Min Wang

Emotional voice conversion (EVC) traditionally targets the transformation of spoken utterances from one emotional state to another, with previous research mainly focusing on discrete emotion categories. This paper departs from the norm by…

音频与语音处理 · 电气工程与系统科学 2023-09-19 Kun Zhou , Berrak Sisman , Carlos Busso , Bin Ma , Haizhou Li

Expressive text-to-speech (TTS) aims to synthesize different speaking style speech according to human's demands. Nowadays, there are two common ways to control speaking styles: (1) Pre-defining a group of speaking style and using…

声音 · 计算机科学 2023-06-27 Dongchao Yang , Songxiang Liu , Rongjie Huang , Chao Weng , Helen Meng

Music can be represented in multiple forms, such as in the audio form as a recording of a performance, in the symbolic form as a computer readable score, or in the image form as a scan of the sheet music. Music synchronisation provides a…

声音 · 计算机科学 2022-06-02 Ruchit Agrawal

Research on style transfer and domain translation has clearly demonstrated the ability of deep learning-based algorithms to manipulate images in terms of artistic style. More recently, several attempts have been made to extend such…

声音 · 计算机科学 2021-06-11 Ondřej Cífka , Umut Şimşekli , Gaël Richard

Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this paper, we present a novel solution to learn prosody…

音频与语音处理 · 电气工程与系统科学 2023-06-21 Jakub Swiatkowski , Duo Wang , Mikolaj Babianski , Patrick Lumban Tobing , Ravichander Vipperla , Vincent Pollet

Musicians mostly have to rely on their ears when they want to analyze what they play, for example to detect errors. Since hearing is sequential, it is not possible to quickly grasp an overview over one or multiple recordings of a whole…

人机交互 · 计算机科学 2026-03-26 Frank Heyen , Michael Sedlmair

We present an end-to-end text-to-speech (TTS) synthesis system that generates audio and synchronized tongue motion directly from text. This is achieved by adapting a 3D model of the tongue surface to an articulatory dataset and training a…

人机交互 · 计算机科学 2018-04-17 Ingmar Steiner , Sébastien Le Maguer , Alexander Hewer

Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has been overlooked in…

声音 · 计算机科学 2025-07-03 Hiroki Kanagawa , Kenichi Fujita , Aya Watanabe , Yusuke Ijima

Automatic music transcription converts audio recordings into symbolic representations, facilitating music analysis, retrieval, and generation. A musical note is characterized by pitch, onset, and offset in an audio domain, whereas it is…

声音 · 计算机科学 2025-02-19 Leekyung Kim , Sungwook Jeon , Wan Heo , Jonghun Park

Recent advances in text-to-speech (TTS) have yielded remarkable improvements in naturalness and intelligibility. Building on these achievements, research has increasingly shifted toward enhancing the expressiveness of generated speech, such…

声音 · 计算机科学 2025-12-23 Pengchao Feng , Yao Xiao , Ziyang Ma , Zhikang Niu , Shuai Fan , Yao Li , Sheng Wang , Xie Chen

Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of discrete speech…

声音 · 计算机科学 2023-12-19 Chunyu Qiang , Hao Li , Yixin Tian , Yi Zhao , Ying Zhang , Longbiao Wang , Jianwu Dang

Expressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in recent years. Previous methods performed the expressive…

声音 · 计算机科学 2022-01-19 Yi Lei , Shan Yang , Xinsheng Wang , Lei Xie

We present an audio-to-analysis pipeline that produces composer-level information-theoretic profiles : reflecting compositional vocabulary as it emerges from aggregated performances : from raw recordings, built on a transcription layer…

声音 · 计算机科学 2026-05-11 Fred Jalbert-Desforges