English
Related papers

Related papers: Generating Diverse Vocal Bursts with StyleGAN2 and…

200 papers

Most GAN(Generative Adversarial Network)-based approaches towards high-fidelity waveform generation heavily rely on discriminators to improve their performance. However, GAN methods introduce much uncertainty into the generation process and…

Sound · Computer Science 2022-03-22 Shengyuan Xu , Wenxiao Zhao , Jing Guo

This paper proposes a modeling-by-generation (MbG) excitation vocoder for a neural text-to-speech (TTS) system. Recently proposed neural excitation vocoders can realize qualified waveform generation by combining a vocal tract filter with a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Eunwoo Song , Min-Jae Hwang , Ryuichi Yamamoto , Jin-Seob Kim , Ohsung Kwon , Jae-Min Kim

Speech emotion recognition is a challenging task because the emotion expression is complex, multimodal and fine-grained. In this paper, we propose a novel multimodal deep learning approach to perform fine-grained emotion recognition from…

Sound · Computer Science 2021-07-16 Hang Li , Wenbiao Ding , Zhongqin Wu , Zitao Liu

Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of such intensity fluctuations has been largely overlooked by…

Sound · Computer Science 2024-10-01 Jingyi Xu , Hieu Le , Zhixin Shu , Yang Wang , Yi-Hsuan Tsai , Dimitris Samaras

Generative Adversarial Networks (GANs) are powerful models able to synthesize data samples closely resembling the distribution of real data, yet the diversity of those generated samples is limited due to the so-called mode collapse…

Computer Vision and Pattern Recognition · Computer Science 2023-06-26 Jan Dubiński , Kamil Deja , Sandro Wenzel , Przemysław Rokita , Tomasz Trzciński

The advent of Large Models marks a new era in machine learning, significantly outperforming smaller models by leveraging vast datasets to capture and synthesize complex patterns. Despite these advancements, the exploration into scaling,…

Sound · Computer Science 2024-02-05 Shijia Liao , Shiyi Lan , Arun George Zachariah

Classical parametric speech coding techniques provide a compact representation for speech signals. This affords a very low transmission rate but with a reduced perceptual quality of the reconstructed signals. Recently, autoregressive deep…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-02 Ahmed Mustafa , Arijit Biswas , Christian Bergler , Julia Schottenhamml , Andreas Maier

Generating conversational gestures from speech audio is challenging due to the inherent one-to-many mapping between audio and body motions. Conventional CNNs/RNNs assume one-to-one mapping, and thus tend to predict the average of all…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Jing Li , Di Kang , Wenjie Pei , Xuefei Zhe , Ying Zhang , Zhenyu He , Linchao Bao

Decoding speech from brain signals is a challenging research problem. Although existing technologies have made progress in reconstructing the mel spectrograms of auditory stimuli at the word or letter level, there remain core challenges in…

Sound · Computer Science 2025-08-12 Cunhang Fan , Sheng Zhang , Jingjing Zhang , Enrui Liu , Xinhui Li , Gangming Zhao , Zhao Lv

The task of audio-driven portrait animation involves generating a talking head video using an identity image and an audio track of speech. While many existing approaches focus on lip synchronization and video quality, few tackle the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Jian Zhang , Weijian Mai , Zhijun Zhang

Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifically, we design an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-23 Sefik Emre Eskimez , You Zhang , Zhiyao Duan

In this paper, we describe our contribution to Task 2 of the DCASE 2018 Audio Challenge. While it has become ubiquitous to utilize an ensemble of machine learning methods for classification tasks to obtain better predictive performance, the…

Sound · Computer Science 2018-11-28 Marcel Lederle , Benjamin Wilhelm

Deep generative models provide powerful tools for distributions over complicated manifolds, such as those of natural images. But many of these methods, including generative adversarial networks (GANs), can be difficult to train, in part…

Machine Learning · Statistics 2017-11-08 Akash Srivastava , Lazar Valkov , Chris Russell , Michael U. Gutmann , Charles Sutton

Vocal bursts play an important role in communicating affect, making them valuable for improving speech emotion recognition. Here, we present our approach for classifying vocal bursts and predicting their emotional significance in the ACII…

Sound · Computer Science 2022-09-28 Vincent Karas , Andreas Triantafyllopoulos , Meishu Song , Björn W. Schuller

Speech is a rich biometric signal that contains information about the identity, gender and emotional state of the speaker. In this work, we explore its potential to generate face images of a speaker by conditioning a Generative Adversarial…

Recent advances in diffusion models have positioned them as powerful generative frameworks for speech synthesis, demonstrating substantial improvements in audio quality and stability. Nevertheless, their effectiveness in vocoders…

Sound · Computer Science 2025-12-01 Teysir Baoueb , Xiaoyu Bie , Mathieu Fontaine , Gaël Richard

Adversarial waveform generation has been a popular approach as the backend of singing voice conversion (SVC) to generate high-quality singing audio. However, the instability of GAN also leads to other problems, such as pitch jitters and U/V…

Sound · Computer Science 2022-01-26 Haohan Guo , Zhiping Zhou , Fanbo Meng , Kai Liu

Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Guanwen Feng , Haoran Cheng , Yunan Li , Zhiyuan Ma , Chaoneng Li , Zhihao Qian , Qiguang Miao , Chi-Man Pun

Audio diffusion models can synthesize a wide variety of sounds. Existing models often operate on the latent domain with cascaded phase recovery modules to reconstruct waveform. This poses challenges when generating high-fidelity audio. In…

Sound · Computer Science 2023-11-21 Ge Zhu , Yutong Wen , Marc-André Carbonneau , Zhiyao Duan

This paper presents a deep learning-based approach to emotion detection using Conditional Generative Adversarial Networks (cGANs). Unlike traditional unimodal techniques that rely on a single data type, we explore a multimodal framework…

Machine Learning · Computer Science 2025-08-07 Anushka Srivastava
‹ Prev 1 4 5 6 7 8 10 Next ›