中文
相关论文

相关论文: Peking Opera Synthesis via Duration Informed Atten…

200 篇论文

Recently, deep learning-based generative models have been introduced to generate singing voices. One approach is to predict the parametric vocoder features consisting of explicit speech parameters. This approach has the advantage that the…

音频与语音处理 · 电气工程与系统科学 2024-06-14 Tae-Woo Kim , Min-Su Kang , Gyeong-Hoon Lee

Representing speech as discretized units has numerous benefits in supporting downstream spoken language processing tasks. However, the approach has been less explored in speech synthesis of tonal languages like Mandarin Chinese. Our…

音频与语音处理 · 电气工程与系统科学 2024-09-04 Dehua Tao , Daxin Tan , Yu Ting Yeung , Xiao Chen , Tan Lee

This paper presents Sinsy, a deep neural network (DNN)-based singing voice synthesis (SVS) system. In recent years, DNNs have been utilized in statistical parametric SVS systems, and DNN-based SVS systems have demonstrated better…

音频与语音处理 · 电气工程与系统科学 2021-09-28 Yukiya Hono , Kei Hashimoto , Keiichiro Oura , Yoshihiko Nankaku , Keiichi Tokuda

Most singer identification methods are processed in the frequency domain, which potentially leads to information loss during the spectral transformation. In this paper, instead of the frequency domain, we propose an end-to-end architecture…

音频与语音处理 · 电气工程与系统科学 2022-05-24 Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Unison singing is the name given to an ensemble of singers simultaneously singing the same melody and lyrics. While each individual singer in a unison sings the same principle melody, there are slight timing and pitch deviations between the…

音频与语音处理 · 电气工程与系统科学 2020-09-22 Pritish Chandna , Helena Cuesta , Emilia Gómez

Chorus detection is a challenging problem in musical signal processing as the chorus often repeats more than once in popular songs, usually with rich instruments and complex rhythm forms. Most of the existing works focus on the…

音频与语音处理 · 电气工程与系统科学 2023-08-21 Qiqi He , Xiaoheng Sun , Yi Yu , Wei Li

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

声音 · 计算机科学 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

Advancing quantum information processors and building fault-tolerant architectures rely on the ability to accurately characterize the noise sources and suppress their impact on quantum devices. In practice, noise often drifts over time,…

量子物理 · 物理学 2025-11-13 Devansh Bhardwaj , Evangelia Takou , Yingjia Lin , Kenneth R. Brown

Sequence expansion between encoder and decoder is a critical challenge in sequence-to-sequence TTS. Attention-based methods achieve great naturalness but suffer from unstable issues like missing and repeating phonemes, not to mention…

声音 · 计算机科学 2022-03-21 Yunchao He , Jian Luan , Yujun Wang

Dynamic Range Compression (DRC) is a widely used audio effect that adjusts signal dynamics for applications in music production, broadcasting, and speech processing. Inverting DRC is of broad importance for restoring the original dynamics,…

声音 · 计算机科学 2025-09-11 Haoran Sun , Dominique Fourer , Hichem Maaref

We present a number of contributions to bridging the gap between supervisory control theory and coordination of services in order to explore the frontiers between coordination and control systems. Firstly, we modify the classical synthesis…

系统与控制 · 电气工程与系统科学 2023-06-22 Davide Basile , Maurice H. ter Beek , Rosario Pugliese

In this paper, we propose a fully supervised speaker diarization approach, named unbounded interleaved-state recurrent neural networks (UIS-RNN). Given extracted speaker-discriminative embeddings (a.k.a. d-vectors) from input utterances,…

音频与语音处理 · 电气工程与系统科学 2019-02-20 Aonan Zhang , Quan Wang , Zhenyao Zhu , John Paisley , Chong Wang

Music comprises of a set of complex simultaneous events organized in time. In this paper we introduce a novel framework that we call Deep Musical Information Dynamics, which combines two parallel streams - a low rate latent representation…

声音 · 计算机科学 2021-02-03 Shlomo Dubnov

This paper presents an end-to-end high-quality singing voice synthesis (SVS) system that uses bidirectional encoder representation from Transformers (BERT) derived semantic embeddings to improve the expressiveness of the synthesized singing…

声音 · 计算机科学 2023-09-01 Shaohuan Zhou , Shun Lei , Weiya You , Deyi Tuo , Yuren You , Zhiyong Wu , Shiyin Kang , Helen Meng

It is still an interesting and challenging problem to synthesize a vivid and realistic singing face driven by music signal. In this paper, we present a method for this task with natural motions of the lip, facial expression, head pose, and…

图形学 · 计算机科学 2023-03-27 Pengfei Liu , Wenjin Deng , Hengda Li , Jintai Wang , Yinglin Zheng , Yiwei Ding , Xiaohu Guo , Ming Zeng

Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like…

计算与语言 · 计算机科学 2026-01-15 Fadhil Muhammad , Alwin Djuliansah , Adrian Aryaputra Hamzah , Kurniawati Azizah

We introduce a sophisticated multi-speaker speech data simulator, specifically engineered to generate multi-speaker speech recordings. A notable feature of this simulator is its capacity to modulate the distribution of silence and overlap…

音频与语音处理 · 电气工程与系统科学 2023-10-20 Tae Jin Park , He Huang , Coleman Hooper , Nithin Koluguri , Kunal Dhawan , Ante Jukic , Jagadeesh Balam , Boris Ginsburg

We study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single…

声音 · 计算机科学 2019-05-07 Michael Michelashvili , Sagie Benaim , Lior Wolf

This paper presents a method of using autoregressive neural networks for the acoustic modeling of singing voice synthesis (SVS). Singing voice differs from speech and it contains more local dynamic movements of acoustic features, e.g.,…

声音 · 计算机科学 2019-06-24 Yuan-Hao Yi , Yang Ai , Zhen-Hua Ling , Li-Rong Dai

Collective temporal organization in complex systems is commonly attributed to synchronization, resonance, or proximity to dynamical instabilities. Here we identify a distinct mechanism by which coherent, synchronization-like behavior can…

适应与自组织系统 · 物理学 2026-03-10 V. Troude , D. Sornette