English
Related papers

Related papers: CAMP: a Two-Stage Approach to Modelling Prosody in…

200 papers

In long-text speech synthesis, current approaches typically convert text to speech at the sentence-level and concatenate the results to form pseudo-paragraph-level speech. These methods overlook the contextual coherence of paragraphs,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-21 Zhipeng Li , Xiaofen Xing , Jingyuan Xing , Hangrui Hu , Heng Lu , Xiangmin Xu

Turn-taking is a fundamental aspect of human communication and can be described as the ability to take turns, project upcoming turn shifts, and supply backchannels at appropriate locations throughout a conversation. In this work, we…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-13 Erik Ekstedt , Gabriel Skantze

In many domains, it is difficult to obtain the race data that is required to estimate racial disparity. To address this problem, practitioners have adopted the use of proxy methods which predict race using non-protected covariates. However,…

Computers and Society · Computer Science 2024-09-04 Kweku Kwegyir-Aggrey , Naveen Durvasula , Jennifer Wang , Suresh Venkatasubramanian

While Transformer has become the de-facto standard for speech, modeling upon the fine-grained frame-level features remains an open challenge of capturing long-distance dependencies and distributing the attention weights. We propose…

Computation and Language · Computer Science 2023-05-30 Chen Xu , Yuhao Zhang , Chengbo Jiao , Xiaoqian Liu , Chi Hu , Xin Zeng , Tong Xiao , Anxiang Ma , Huizhen Wang , JingBo Zhu

Expressive speech synthesis requires vibrant prosody and well-timed pauses. We propose an effective strategy to augment a small dataset to train an expressive end-to-end Text-to-Speech model. We merge audios of emotionally congruent text…

Sound · Computer Science 2026-02-12 Raymond Chung

Recent research in zero-shot speech synthesis has made significant progress in speaker similarity. However, current efforts focus on timbre generalization rather than prosody modeling, which results in limited naturalness and…

Sound · Computer Science 2024-06-12 Yuepeng Jiang , Tao Li , Fengyu Yang , Lei Xie , Meng Meng , Yujun Wang

In spoken conversations, spontaneous behaviors like filled pause and prolongations always happen. Conversational partner tends to align features of their speech with their interlocutor which is known as entrainment. To produce human-like…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-22 Jian Cong , Shan Yang , Na Hu , Guangzhi Li , Lei Xie , Dan Su

Human speech can be characterized by different components, including semantic content, speaker identity and prosodic information. Significant progress has been made in disentangling representations for semantic content and speaker identity…

Sound · Computer Science 2023-09-27 Leyuan Qu , Taihao Li , Cornelius Weber , Theresa Pekarek-Rosin , Fuji Ren , Stefan Wermter

This paper proposes a hierarchical, fine-grained and interpretable latent variable model for prosody based on the Tacotron 2 text-to-speech model. It achieves multi-resolution modeling of prosody by conditioning finer level representations…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-11 Guangzhi Sun , Yu Zhang , Ron J. Weiss , Yuan Cao , Heiga Zen , Yonghui Wu

We introduce two rule-based models to modify the prosody of speech synthesis in order to modulate the emotion to be expressed. The prosody modulation is based on speech synthesis markup language (SSML) and can be used with any commercial…

Sound · Computer Science 2023-07-06 Felix Burkhardt , Uwe Reichel , Florian Eyben , Björn Schuller

Recently semantic parsing in context has received considerable attention, which is challenging since there are complex contextual phenomena. Previous works verified their proposed methods in limited scenarios, which motivates us to conduct…

Computation and Language · Computer Science 2020-06-16 Qian Liu , Bei Chen , Jiaqi Guo , Jian-Guang Lou , Bin Zhou , Dongmei Zhang

Vocal entrainment is a social adaptation mechanism in human interaction, knowledge of which can offer useful insights to an individual's cognitive-behavioral characteristics. We propose a context-aware approach for measuring vocal…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-08 Rimita Lahiri , Md Nasir , Catherine Lord , So Hyun Kim , Shrikanth Narayanan

Large Language Models (LLMs) can be adapted to extend their text capabilities to speech inputs. However, these speech-adapted LLMs consistently underperform their text-based counterparts--and even cascaded pipelines--on language…

Computation and Language · Computer Science 2026-02-24 Santiago Cuervo , Skyler Seto , Maureen de Seyssel , Richard He Bai , Zijin Gu , Tatiana Likhomanenko , Navdeep Jaitly , Zakaria Aldeneh

Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability. However, the acquisition of prosodic boundary labels relies on manual annotation, which is costly and time-consuming. In…

Sound · Computer Science 2022-06-17 Ziqian Dai , Jianwei Yu , Yan Wang , Nuo Chen , Yanyao Bian , Guangzhi Li , Deng Cai , Dong Yu

Conditional diffusion models rely on language-to-image alignment methods to steer the generation towards semantically accurate outputs. Despite the success of this architecture, misalignment and hallucinations remain common issues and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Vasco Ramos , Regev Cohen , Idan Szpektor , Joao Magalhaes

Automatic speech recognition (ASR) has benefited from advances in pretrained speech and language models, yet most systems remain constrained to monolingual settings and short, isolated utterances. While recent efforts in context-aware ASR…

Computation and Language · Computer Science 2026-03-09 Yuchen Zhang , Haralambos Mouratidis , Ravi Shekhar

We present a neural text-to-speech system for fine-grained prosody transfer from one speaker to another. Conventional approaches for end-to-end prosody transfer typically use either fixed-dimensional or variable-length prosody embedding via…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-05 Viacheslav Klimkov , Srikanth Ronanki , Jonas Rohnke , Thomas Drugman

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issues: 1) Multiscale…

Multimedia · Computer Science 2025-01-03 Yuan Zhao , Rui Liu , Gaoxiang Cong

Change captioning generates descriptions that explicitly describe the differences between two visually similar images. Existing methods operate on static image pairs, thus ignoring the rich temporal dynamics of the change procedure, which…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Jiayang Sun , Zixin Guo , Min Cao , Guibo Zhu , Jorma Laaksonen

Identifying whether an utterance is a statement, question, greeting, and so forth is integral to effective automatic understanding of natural dialog. Little is known, however, about how such dialog acts (DAs) can be automatically classified…

Computation and Language · Computer Science 2007-05-23 E. Shriberg , R. Bates , A. Stolcke , P. Taylor , D. Jurafsky , K. Ries , N. Coccaro , R. Martin , M. Meteer , C. Van Ess-Dykema
‹ Prev 1 4 5 6 7 8 10 Next ›