English
Related papers

Related papers: Going Retro: Astonishingly Simple Yet Effective Ru…

200 papers

Recent advancements in end-to-end speech synthesis have made it possible to generate highly natural speech. However, training these models typically requires a large amount of high-fidelity speech data, and for unseen texts, the prosody of…

Computation and Language · Computer Science 2021-11-16 Zhu Li , Yuqing Zhang , Mengxi Nie , Ming Yan , Mengnan He , Ruixiong Zhang , Caixia Gong

Recent research in zero-shot speech synthesis has made significant progress in speaker similarity. However, current efforts focus on timbre generalization rather than prosody modeling, which results in limited naturalness and…

Sound · Computer Science 2024-06-12 Yuepeng Jiang , Tao Li , Fengyu Yang , Lei Xie , Meng Meng , Yujun Wang

This paper explores predicting suitable prosodic features for fine-grained emotion analysis from the discourse-level text. To obtain fine-grained emotional prosodic features as predictive values for our model, we extract a phoneme-level…

Sound · Computer Science 2023-09-22 Xianhao Wei , Jia Jia , Xiang Li , Zhiyong Wu , Ziyi Wang

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

Sound · Computer Science 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee

Annotating and recognizing speech emotion using prompt engineering has recently emerged with the advancement of Large Language Models (LLMs), yet its efficacy and reliability remain questionable. In this paper, we conduct a systematic study…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-01 Yuanchao Li , Yuan Gong , Chao-Han Huck Yang , Peter Bell , Catherine Lai

The aim of this research is development of rule based decision model for emotion recognition. This research also proposes using the rules for augmenting inter-corporal recognition accuracy in multimodal systems that use supervised learning…

Human-Computer Interaction · Computer Science 2016-07-12 Amol Patwardhan , Gerald Knapp

Pre-trained model representations have demonstrated state-of-the-art performance in speech recognition, natural language processing, and other applications. Speech models, such as Bidirectional Encoder Representations from Transformers…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-07 Vikramjit Mitra , Vasudha Kowtha , Hsiang-Yun Sherry Chien , Erdrin Azemi , Carlos Avendano

Previous work on emotion recognition demonstrated a synergistic effect of combining several modalities such as auditory, visual, and transcribed text to estimate the affective state of a speaker. Among these, the linguistic modality is…

Computation and Language · Computer Science 2019-03-01 Egor Lakomkin , Mohammad Ali Zamani , Cornelius Weber , Sven Magg , Stefan Wermter

Code generation models are widely used in software development, yet their sensitivity to prompt phrasing remains under-examined. Identical requirements expressed with different emotions or communication styles can yield divergent outputs,…

Software Engineering · Computer Science 2025-09-18 Wei Ma , Yixiao Yang , Jingquan Ge , Xiaofei Xie , Lingxiao Jiang

Automatic synthesis of realistic co-speech gestures is an increasingly important yet challenging task in artificial embodied agent creation. Previous systems mainly focus on generating gestures in an end-to-end manner, which leads to…

Sound · Computer Science 2023-05-05 Tenglong Ao , Qingzhe Gao , Yuke Lou , Baoquan Chen , Libin Liu

Expressive speech synthesis aims to generate speech that captures a wide range of para-linguistic features, including emotion and articulation, though current research primarily emphasizes emotional aspects over the nuanced articulatory…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-24 Zehua Kcriss Li , Meiying Melissa Chen , Yi Zhong , Pinxin Liu , Zhiyao Duan

With the development of speech synthesis, recent research has focused on challenging tasks, such as speaker generation and emotion intensity control. Attribute interpolation is a common approach to these tasks. However, most previous…

Sound · Computer Science 2024-07-02 Masato Murata , Koichi Miyazaki , Tomoki Koriyama

In this paper, we present a database of emotional speech intended to be open-sourced and used for synthesis and generation purpose. It contains data for male and female actors in English and a male actor in French. The database covers 5…

Computation and Language · Computer Science 2018-06-26 Adaeze Adigwe , Noé Tits , Kevin El Haddad , Sarah Ostadabbas , Thierry Dutoit

The field of prosody transfer in speech synthesis systems is rapidly advancing. This research is focused on evaluating learning methods for adapting pre-trained monolingual text-to-speech (TTS) models to multilingual conditions, i.e.,…

Computation and Language · Computer Science 2024-06-19 Arnav Goel , Medha Hira , Anubha Gupta

Adding an emotions using prosody manipulation method for Indonesian text to speech system. Text To Speech (TTS) is a system that can convert text in one language into speech, accordance with the reading of the text in the language used. The…

Sound · Computer Science 2016-06-30 Salita Ulitia Prini , Ary Setijadi Prihatmanto

A large part of the expressive speech synthesis literature focuses on learning prosodic representations of the speech signal which are then modeled by a prior distribution during inference. In this paper, we compare different prior…

Various parametric representations have been proposed to model the speech signal. While the performance of such vocoders is well-known in the context of speech processing, their extrapolation to singing voice synthesis might not be…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-09 Onur Babacan , Thomas Drugman , Tuomo Raitio , Daniel Erro , Thierry Dutoit

In a typical voice conversion system, prior works utilize various acoustic features (e.g., the pitch, voiced/unvoiced flag, aperiodicity) of the source speech to control the prosody of generated waveform. However, the prosody is related…

Sound · Computer Science 2020-06-01 Zheng Lian , Zhengqi Wen

How important are different temporal speech modulations for speech recognition? We answer this question from two complementary perspectives. Firstly, we quantify the amount of phonetic \textit{information} in the modulation spectrum of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-24 Samik Sadhu , Hynek Hermansky

Emotion embedding space learned from references is a straightforward approach for emotion transfer in encoder-decoder structured emotional text to speech (TTS) systems. However, the transferred emotion in the synthetic speech is not…

Sound · Computer Science 2020-11-18 Tao Li , Shan Yang , Liumeng Xue , Lei Xie
‹ Prev 1 4 5 6 7 8 10 Next ›