English
Related papers

Related papers: A Discourse-level Multi-scale Prosodic Model for F…

200 papers

We propose spoken sentence embeddings which capture both acoustic and linguistic content. While existing works operate at the character, phoneme, or word level, our method learns long-term dependencies by modeling speech at the sentence…

Sound · Computer Science 2019-02-22 Albert Haque , Michelle Guo , Prateek Verma , Li Fei-Fei

Speech Emotion Recognition (SER) has emerged as a critical component of the next generation human-machine interfacing technologies. In this work, we propose a new dual-level model that predicts emotions based on both MFCC features and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-24 Jianyou Wang , Michael Xue , Ryan Culhane , Enmao Diao , Jie Ding , Vahid Tarokh

We often verbally express emotions in a multifaceted manner, they may vary in their intensities and may be expressed not just as a single but as a mixture of emotions. This wide spectrum of emotions is well-studied in the structural model…

Computation and Language · Computer Science 2024-06-28 Rendi Chevi , Alham Fikri Aji

Text-to-speech is now able to achieve near-human naturalness and research focus has shifted to increasing expressivity. One popular method is to transfer the prosody from a reference speech sample. There have been considerable advances in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-22 Alexandra Torresquintero , Tian Huey Teh , Christopher G. R. Wallis , Marlene Staib , Devang S Ram Mohan , Vivian Hu , Lorenzo Foglianti , Jiameng Gao , Simon King

While the performance of cross-lingual TTS based on monolingual corpora has been significantly improved recently, generating cross-lingual speech still suffers from the foreign accent problem, leading to limited naturalness. Besides,…

Sound · Computer Science 2023-09-06 Tao Li , Chenxu Hu , Jian Cong , Xinfa Zhu , Jingbei Li , Qiao Tian , Yuping Wang , Lei Xie

Emotion shapes all aspects of our interpersonal and intellectual experiences. Its automatic analysis has there-fore many applications, e.g., human-machine interface. In this paper, we propose an emotional tonal speech dataset, namely…

Sound · Computer Science 2018-10-17 Zhongzhe Xiao , Ying Chen , Weibei Dou , Zhi Tao , Liming Chen

A text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model the acoustic conditions explicitly, in order to disentangle…

Natural Language Processing has recently made understanding human interaction easier, leading to improved sentimental analysis and behaviour prediction. However, the choice of words and vocal cues in conversations presents an underexplored…

Computers and Society · Computer Science 2022-06-24 Amna Anwar , Eiman Kanjo , Dario Ortega Anderez

Expressive speech synthesis is crucial for many human-computer interaction scenarios, such as audiobooks, podcasts, and voice assistants. Previous works focus on predicting the style embeddings at one single scale from the information…

Sound · Computer Science 2023-08-01 Shun Lei , Yixuan Zhou , Liyang Chen , Zhiyong Wu , Xixin Wu , Shiyin Kang , Helen Meng

Speech emotion recognition is a challenging task, and extensive reliance has been placed on models that use audio features in building well-performing classifiers. In this paper, we propose a novel deep dual recurrent encoder model that…

Computation and Language · Computer Science 2018-10-11 Seunghyun Yoon , Seokhyun Byun , Kyomin Jung

Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesized speech of a target speaker's timbre. In most previous methods, the synthesized fine-grained prosody features often represent…

Sound · Computer Science 2023-03-15 Chunyu Qiang , Peng Yang , Hao Che , Ying Zhang , Xiaorui Wang , Zhongyuan Wang

Metaphors play a pivotal role in expressing emotions, making them crucial for emotional intelligence. The advent of multimodal data and widespread communication has led to a proliferation of multimodal metaphors, amplifying the complexity…

Computation and Language · Computer Science 2025-05-21 Xingyuan Lu , Yuxi Liu , Dongyu Zhang , Zhiyao Wu , Jing Ren , Feng Xia

In expressive speech synthesis it is widely adopted to use latent prosody representations to deal with variability of the data during training. Same text may correspond to various acoustic realizations, which is known as a one-to-many…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-13 Mikolaj Babianski , Kamil Pokora , Raahil Shah , Rafal Sienkiewicz , Daniel Korzekwa , Viacheslav Klimkov

Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Soumya Dutta

Prosody conveys rich emotional and semantic information of the speech signal as well as individual idiosyncrasies. We propose a stand-alone model that maps text-to-prosodic features such as F0 and energy and can be used in downstream tasks…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-14 Eray Eren , Qingju Liu , Hyeongwoo Kim , Pablo Garrido , Abeer Alwan

In this paper, we propose a method of speaker adaption with intuitive prosodic features for statistical parametric speech synthesis. The intuitive prosodic features employed in this method include pitch, pitch range, speech rate and energy…

Sound · Computer Science 2022-03-03 Pengyu Cheng , Zhenhua Ling

Modeling discourse -- the linguistic phenomena that go beyond individual sentences, is a fundamental yet challenging aspect of natural language processing (NLP). However, existing evaluation benchmarks primarily focus on the evaluation of…

Computation and Language · Computer Science 2023-07-25 Longyue Wang , Zefeng Du , Donghuai Liu , Deng Cai , Dian Yu , Haiyun Jiang , Yan Wang , Leyang Cui , Shuming Shi , Zhaopeng Tu

This paper addresses the problem of modeling textual conversations and detecting emotions. Our proposed model makes use of 1) deep transfer learning rather than the classical shallow methods of word embedding; 2) self-attention mechanisms…

Computation and Language · Computer Science 2019-06-18 Waleed Ragheb , Jérôme Azé , Sandra Bringay , Maximilien Servajean

Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal and largely overlook…

Sound · Computer Science 2026-05-06 Vamshi Nallaguntla , Shruti Kshirsagar , Anderson R. Avila

Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms directly from the text. Single-stage text-to-speech often faces…

Sound · Computer Science 2022-07-14 Zhengxi Liu , Qiao Tian , Chenxu Hu , Xudong Liu , Menglin Wu , Yuping Wang , Hang Zhao , Yuxuan Wang