English
Related papers

Related papers: Efficient Emotional Adaptation for Audio-Driven Ta…

200 papers

Speech-driven facial video generation has been a complex problem due to its multi-modal aspects namely audio and video domain. The audio comprises lots of underlying features such as expression, pitch, loudness, prosody(speaking style) and…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall

Text-to-speech (TTS) has shown great progress in recent years. However, most existing TTS systems offer only coarse and rigid emotion control, typically via discrete emotion labels or a carefully crafted and detailed emotional text prompt,…

Sound · Computer Science 2025-10-28 Tianxin Xie , Shan Yang , Chenxing Li , Dong Yu , Li Liu

We propose a method for test-time adaptation of pretrained depth completion models. Depth completion models, trained on some ``source'' data, often predict erroneous outputs when transferred to ``target'' data captured in novel…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Younjoon Chung , Hyoungseob Park , Patrick Rim , Xiaoran Zhang , Jihe He , Ziyao Zeng , Safa Cicek , Byung-Woo Hong , James S. Duncan , Alex Wong

Recently, self-supervised pre-training has shown significant improvements in many areas of machine learning, including speech and NLP. We propose using large self-supervised pre-trained models for both audio and text modality with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-24 Krishna D N

Emotion Recognition in Conversation (ERC) is a more challenging task than conventional text emotion recognition. It can be regarded as a personalized and interactive emotion recognition task, which is supposed to consider not only the…

Computation and Language · Computer Science 2021-01-01 Jiangnan Li , Zheng Lin , Peng Fu , Qingyi Si , Weiping Wang

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse activities with…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Zi-Yi Dou , Xitong Yang , Tushar Nagarajan , Huiyu Wang , Jing Huang , Nanyun Peng , Kris Kitani , Fu-Jen Chu

Emotional text-to-speech (TTS) systems sturggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of existing emotion labels. To address this, we propose a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Kun Zhou , You Zhang , Dianwen Ng , Shengkui Zhao , Hao Wang , Bin Ma

In this manuscript, the topic of multi-corpus Speech Emotion Recognition (SER) is approached from a deep transfer learning perspective. A large corpus of emotional speech data, EmoSet, is assembled from a number of existing SER corpora. In…

Sound · Computer Science 2021-03-16 Maurice Gerczuk , Shahin Amiriparian , Sandra Ottl , Björn Schuller

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to-speech (TTS) screen…

Social and Information Networks · Computer Science 2024-10-28 Suparna De , Ionut Bostan , Nishanth Sastry

How much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive language disorders.…

Sound · Computer Science 2025-02-04 Jiaxin Ye , Boyuan Cao , Hongming Shan

Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a novel EVC framework capable of generating high-quality…

Sound · Computer Science 2025-05-21 Yu Pan , Yanni Hu , Yuguang Yang , Jixun Yao , Jianhao Ye , Hongbin Zhou , Lei Ma , Jianjun Zhao

Speech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level labels. However, not all frames in an audio have affective…

Sound · Computer Science 2023-12-29 Qifei Li , Yingming Gao , Cong Wang , Yayue Deng , Jinlong Xue , Yichen Han , Ya Li

Recent text-to-speech models have reached the level of generating natural speech similar to what humans say. But there still have limitations in terms of expressiveness. The existing emotional speech synthesis models have shown…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Yoori Oh , Juheon Lee , Yoseob Han , Kyogu Lee

Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human emotional…

Sound · Computer Science 2024-12-13 Weizhen Bian , Yubo Zhou , Kaitai Zhang , Xiaohan Gu

Generating versatile and appropriate synthetic speech requires control over the output expression separate from the spoken text. Important non-textual speech variation is seldom annotated, in which case output control must be learned in an…

Audio and Speech Processing · Electrical Eng. & Systems 2018-09-11 Gustav Eje Henter , Jaime Lorenzo-Trueba , Xin Wang , Junichi Yamagishi

In this work we introduce NWT, an expressive speech-to-video model. Unlike approaches that use domain-specific intermediate representations such as pose keypoints, NWT learns its own latent representations, with minimal assumptions about…

Sound · Computer Science 2021-06-09 Rayhane Mama , Marc S. Tyndel , Hashiam Kadhim , Cole Clifford , Ragavan Thurairatnam

The increasing use of dialogue agents makes it extremely desirable for them to understand and acknowledge the implied emotions to respond like humans with empathy. Chatbots using traditional techniques analyze emotions based on the context…

Computation and Language · Computer Science 2021-05-27 Akhilesh Ravi , Amit Yadav , Jainish Chauhan , Jatin Dholakia , Naman Jain , Mayank Singh

Current computational-emotion research has focused on applying acoustic properties to analyze how emotions are perceived mathematically or used in natural language processing machine learning models. While recent interest has focused on…

Sound · Computer Science 2021-07-06 Daniel Szelogowski

Automated emotion detection in speech is a challenging task due to the complex interdependence between words and the manner in which they are spoken. It is made more difficult by the available datasets; their small size and incompatible…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-16 Amith Ananthram , Kailash Karthik Saravanakumar , Jessica Huynh , Homayoon Beigi

Speech emotion recognition (SER) is the task of recognising human's emotional states from speech. SER is extremely prevalent in helping dialogue systems to truly understand our emotions and become a trustworthy human conversational partner.…

Sound · Computer Science 2022-10-27 Zhao Ren , Thanh Tam Nguyen , Yi Chang , Björn W. Schuller
‹ Prev 1 8 9 10 Next ›