English
Related papers

Related papers: Rich Prosody Diversity Modelling with Phone-level …

200 papers

Modeling the rich prosodic variations inherent in human speech is essential for generating natural-sounding speech. While speaker embeddings are commonly used as conditioning inputs in personalized speech generation, they are typically…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-22 Ismail Rasim Ulgen , John H. L. Hansen , Carlos Busso , Berrak Sisman

Prosody modeling is an essential component in modern text-to-speech (TTS) frameworks. By explicitly providing prosody features to the TTS model, the style of synthesized utterances can thus be controlled. However, predicting natural and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-04 Chung-Ming Chien , Hung-yi Lee

The Bidirectional LSTM (BLSTM) RNN based speech synthesis system is among the best parametric Text-to-Speech (TTS) systems in terms of the naturalness of generated speech, especially the naturalness in prosody. However, the model complexity…

Computation and Language · Computer Science 2018-02-27 Mengxiao Bi , Heng Lu , Shiliang Zhang , Ming Lei , Zhijie Yan

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, existing CSS systems are…

Sound · Computer Science 2025-02-28 Weihao wu , Zhiwei Lin , Yixuan Zhou , Jingbei Li , Rui Niu , Qinghua Wu , Songjun Cao , Long Ma , Zhiyong Wu

This paper proposes a new approach to duration modelling for statistical parametric speech synthesis in which a recurrent statistical model is trained to output a phone transition probability at each timestep (acoustic frame). Unlike…

Computation and Language · Computer Science 2020-07-28 Srikanth Ronanki , Oliver Watts , Simon King , Gustav Eje Henter

Speech synthesis has significantly advanced from statistical methods to deep neural network architectures, leading to various text-to-speech (TTS) models that closely mimic human speech patterns. However, capturing nuances such as emotion…

Sound · Computer Science 2025-01-14 Shaozuo Zhang , Ambuj Mehrish , Yingting Li , Soujanya Poria

Massive multiple-input multiple-output (MIMO) offers significant advantages in spectral and energy efficiencies, positioning it as a cornerstone technology of fifth-generation (5G) wireless communication systems and a promising solution for…

Information Theory · Computer Science 2026-03-10 Zhenzhou Jin , Li You , Huibin Zhou , Yuanshuo Wang , Xiaofeng Liu , Xinrui Gong , Xiqi Gao , Derrick Wing Kwan Ng , Xiang-Gen Xia

A large part of the expressive speech synthesis literature focuses on learning prosodic representations of the speech signal which are then modeled by a prior distribution during inference. In this paper, we compare different prior…

In this paper we investigate the GMM-derived (GMMD) features for adaptation of deep neural network (DNN) acoustic models. The adaptation of the DNN trained on GMMD features is done through the maximum a posteriori (MAP) adaptation of the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-17 Natalia Tomashenko , Yuri Khokhlov , Yannick Esteve

Neural networks offer a versatile, flexible and accurate approach to loss reserving. However, such applications have focused primarily on the (important) problem of fitting accurate central estimates of the outstanding claims. In practice,…

Methodology · Statistics 2022-08-09 Muhammed Taher Al-Mudafer , Benjamin Avanzi , Greg Taylor , Bernard Wong

We propose a novel Multi-Scale Spectrogram (MSS) modelling approach to synthesise speech with an improved coarse and fine-grained prosody. We present a generic multi-scale spectrogram prediction mechanism where the system first predicts…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-01 Ammar Abbas , Bajibabu Bollepalli , Alexis Moinet , Arnaud Joly , Penny Karanasou , Peter Makarov , Simon Slangens , Sri Karlapati , Thomas Drugman

Speech language models refer to language models with speech processing and understanding capabilities. One key desirable capability for speech language models is the ability to capture the intricate interdependency between content and…

Computation and Language · Computer Science 2025-08-11 Kaizhi Qian , Xulin Fan , Junrui Ni , Slava Shechtman , Mark Hasegawa-Johnson , Chuang Gan , Yang Zhang

This paper presents sampling-based speech parameter generation using moment-matching networks for Deep Neural Network (DNN)-based speech synthesis. Although people never produce exactly the same speech even if we try to express the same…

Sound · Computer Science 2017-04-13 Shinnosuke Takamichi , Tomoki Koriyama , Hiroshi Saruwatari

Prosody contains rich information beyond the literal meaning of words, which is crucial for the intelligibility of speech. Current models still fall short in phrasing and intonation; they not only miss or misplace breaks when synthesizing…

Computation and Language · Computer Science 2024-12-20 Xiangheng He , Junjie Chen , Zixing Zhang , Björn W. Schuller

Due to the superior modeling ability of deep neural network (DNN), it is widely used in voice activity detection (VAD). However, the performance may degrade if no sufficient data especially for practical data could be used for training,…

Sound · Computer Science 2020-05-19 Lu Ma , Xiaomeng Zhang , Pei Zhao , Tengrong Su

Large language models (LLMs) effectively generate fluent text when the target output follows natural language patterns. However, structured prediction tasks confine the output format to a limited ontology, causing even very large models to…

Computation and Language · Computer Science 2023-10-19 Derek Chen , Celine Lee , Yunan Lu , Domenic Rosati , Zhou Yu

Generating a synthetic population that is both feasible and diverse is crucial for ensuring the validity of downstream activity schedule simulation in activity-based models (ABMs). While deep generative models (DGMs), such as variational…

Machine Learning · Computer Science 2025-05-08 Sung Yoo Lim , Hyunsoo Yun , Prateek Bansal , Dong-Kyu Kim , Eui-Jin Kim

This paper proposes a single-channel speech enhancement method to reduce the noise and enhance speech at low signal-to-noise ratio (SNR) levels and non-stationary noise conditions. Specifically, we focus on modeling the noise using a…

This paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice synthesis. The natural pitch variation of a human singing…

Sound · Computer Science 2019-02-12 Hiroki Tamaru , Yuki Saito , Shinnosuke Takamichi , Tomoki Koriyama , Hiroshi Saruwatari

Speech synthesis is widely used in many practical applications. In recent years, speech synthesis technology has developed rapidly. However, one of the reasons why synthetic speech is unnatural is that it often has over-smoothness. In order…

Sound · Computer Science 2018-12-18 Leyuan Sheng , Evgeniy N. Pavlovskiy