中文
相关论文

相关论文: Generating Diverse Vocal Bursts with StyleGAN2 and…

200 篇论文

Affective computing faces a major challenge: the lack of high-quality, diverse depth facial datasets for recognizing subtle emotional expressions. We propose a framework for synthetic depth face generation using an optimized GAN with…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Seyed Muhammad Hossein Mousavi , S. Younes Mirinezhad

The ACII Affective Vocal Bursts (A-VB) competition introduces a new topic in affective computing, which is understanding emotional expression using the non-verbal sound of humans. We are familiar with emotion recognition via verbal vocal or…

音频与语音处理 · 电气工程与系统科学 2022-10-04 Dang-Khanh Nguyen , Sudarshan Pant , Ngoc-Huynh Ho , Guee-Sang Lee , Soo-Huyng Kim , Hyung-Jeong Yang

This report provides a brief description of our proposed solution for the Vocal Burst Type classification task of the ACII 2022 Affective Vocal Bursts (A-VB) Competition. We experimented with two approaches as part of our solution for the…

声音 · 计算机科学 2022-10-14 Muhammad Shehram Shah Syed , Zafi Sherhan Syed , Abbas Syed

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing…

计算机视觉与模式识别 · 计算机科学 2023-06-01 Chao Xu , Junwei Zhu , Jiangning Zhang , Yue Han , Wenqing Chu , Ying Tai , Chengjie Wang , Zhifeng Xie , Yong Liu

Affective speech analysis is an ongoing topic of research. A relatively new problem in this field is the analysis of vocal bursts, which are nonverbal vocalisations such as laughs or sighs. Current state-of-the-art approaches to address…

声音 · 计算机科学 2022-09-29 Tobias Hallmen , Silvan Mertes , Dominik Schiller , Elisabeth André

Articulation, emotion, and personality play strong roles in the orofacial movements. To improve the naturalness and expressiveness of virtual agents (VAs), it is important that we carefully model the complex interplay between these factors.…

人机交互 · 计算机科学 2023-05-15 Najmeh Sadoughi , Carlos Busso

Automated emotion detection in speech is a challenging task due to the complex interdependence between words and the manner in which they are spoken. It is made more difficult by the available datasets; their small size and incompatible…

音频与语音处理 · 电气工程与系统科学 2020-11-16 Amith Ananthram , Kailash Karthik Saravanakumar , Jessica Huynh , Homayoon Beigi

Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a novel EVC framework capable of generating high-quality…

声音 · 计算机科学 2025-05-21 Yu Pan , Yanni Hu , Yuguang Yang , Jixun Yao , Jianhao Ye , Hongbin Zhou , Lei Ma , Jianjun Zhao

Emotional voice conversion (EVC) is one way to generate expressive synthetic speech. Previous approaches mainly focused on modeling one-to-one mapping, i.e., conversion from one emotional state to another emotional state, with Mel-cepstral…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Songxiang Liu , Yuewen Cao , Helen Meng

Human conversational styles are measured by the sense of humor, personality, and tone of voice. These characteristics have become essential for conversational intelligent virtual assistants. However, most of the state-of-the-art intelligent…

It remains a significant challenge how to quantitatively control the expressiveness of speech emotion in speech generation. In this work, we present a novel approach for manipulating the rendering of emotions for speech generation. We…

声音 · 计算机科学 2024-10-01 Sho Inoue , Kun Zhou , Shuai Wang , Haizhou Li

Emotion recognition (ER) from speech signals is a robust approach since it cannot be imitated like facial expression or text based sentiment analysis. Valuable information underlying the emotions are significant for human-computer…

声音 · 计算机科学 2023-12-19 David Hason Rudd , Huan Huo , Guandong Xu

Speech emotion recognition systems have high prediction latency because of the high computational requirements for deep learning models and low generalizability mainly because of the poor reliability of emotional measurements across…

声音 · 计算机科学 2023-02-23 Abdul Rehman , Zhen-Tao Liu , Min Wu , Wei-Hua Cao , Cheng-Shan Jiang

This paper introduces WaveGrad, a conditional model for waveform generation which estimates gradients of the data density. The model is built on prior work on score matching and diffusion probabilistic models. It starts from a Gaussian…

音频与语音处理 · 电气工程与系统科学 2020-10-12 Nanxin Chen , Yu Zhang , Heiga Zen , Ron J. Weiss , Mohammad Norouzi , William Chan

Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation. A mel-spectrogram is extracted from the waveform by a simple,…

In this paper, we propose multi-band MelGAN, a much faster waveform generation model targeting to high-quality text-to-speech. Specifically, we improve the original MelGAN by the following aspects. First, we increase the receptive field of…

声音 · 计算机科学 2020-11-18 Geng Yang , Shan Yang , Kai Liu , Peng Fang , Wei Chen , Lei Xie

The ACII Affective Vocal Bursts Workshop & Competition is focused on understanding multiple affective dimensions of vocal bursts: laughs, gasps, cries, screams, and many other non-linguistic vocalizations central to the expression of…

音频与语音处理 · 电气工程与系统科学 2022-10-31 Alice Baird , Panagiotis Tzirakis , Jeffrey A. Brooks , Christopher B. Gregory , Björn Schuller , Anton Batliner , Dacher Keltner , Alan Cowen

Conditional discrete generative models struggle to faithfully compose multiple input conditions. To address this, we derive a theoretically-grounded formulation for composing discrete probabilistic generative processes, with masked…

机器学习 · 计算机科学 2026-04-08 Jamie Stirling , Noura Al-Moubayed , Chris G. Willcocks , Hubert P. H. Shum

Conditional discrete generative models struggle to faithfully compose multiple input conditions. To address this, we derive a theoretically-grounded formulation for composing discrete probabilistic generative processes, with masked…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Jamie Stirling , Noura Al-Moubayed , Chris G. Willcocks , Hubert P. H. Shum

Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech signals, which is then…

声音 · 计算机科学 2020-12-18 Mostafa Sadeghi , Simon Leglaive , Xavier Alameda-PIneda , Laurent Girin , Radu Horaud