English
Related papers

Related papers: Generating Diverse Vocal Bursts with StyleGAN2 and…

200 papers

Diffusion models have recently been shown to be relevant for high-quality speech generation. Most work has been focused on generating spectrograms, and as such, they further require a subsequent model to convert the spectrogram to a…

Sound · Computer Science 2024-03-12 Roi Benita , Michael Elad , Joseph Keshet

Emotional state recognition through speech is being a very interesting research topic nowadays. Using subliminal information of speech, denominated as prosody, it is possible to recognize the emotional state of the person. One of the main…

Computer Vision and Pattern Recognition · Computer Science 2014-03-20 Inma Mohino-Herranz , Roberto Gil-Pita , Sagrario Alonso-Diaz , Manuel Rosa-Zurera

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Lele Chen , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

We explore unsupervised pre-training for speech recognition by learning representations of raw audio. wav2vec is trained on large amounts of unlabeled audio data and the resulting representations are then used to improve acoustic model…

Computation and Language · Computer Science 2019-09-12 Steffen Schneider , Alexei Baevski , Ronan Collobert , Michael Auli

Insufficient recordings and the scarcity of anomalies present significant challenges in developing and validating robust anomaly detection systems for machine sounds. To address these limitations, we propose a novel approach for generating…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-30 Harsh Purohit , Tomoya Nishida , Kota Dohi , Takashi Endo , Yohei Kawaguchi

Emotional language generation is one of the keys to human-like artificial intelligence. Humans use different type of emotions depending on the situation of the conversation. Emotions also play an important role in mediating the engagement…

Computation and Language · Computer Science 2019-11-27 Sashank Santhanam , Samira Shaikh

Evaluating the emotional intelligence (EI) of audio language models (ALMs) is critical. However, existing benchmarks mostly rely on synthesized speech, are limited to single-turn interactions, and depend heavily on open-ended scoring. This…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-27 Shuiyuan Wang , Zhixian Zhao , Hongfei Xue , Chengyou Wang , Shuai Wang , Hui Bu , Xin Xu , Lei Xie

Modern text-to-speech synthesis pipelines typically involve multiple processing stages, each of which is designed or learnt independently from the rest. In this work, we take on the challenging task of learning to synthesise speech from…

Sound · Computer Science 2021-03-18 Jeff Donahue , Sander Dieleman , Mikołaj Bińkowski , Erich Elsen , Karen Simonyan

Automatically assessing emotional valence in human speech has historically been a difficult task for machine learning algorithms. The subtle changes in the voice of the speaker that are indicative of positive or negative emotional states…

Computation and Language · Computer Science 2017-05-09 Jonathan Chang , Stefan Scherer

How much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive language disorders.…

Sound · Computer Science 2025-02-04 Jiaxin Ye , Boyuan Cao , Hongming Shan

Deep generative models have achieved significant progress in speech synthesis to date, while high-fidelity singing voice synthesis is still an open problem for its long continuous pronunciation, rich high-frequency parts, and strong…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-08 Rongjie Huang , Chenye Cui , Feiyang Chen , Yi Ren , Jinglin Liu , Zhou Zhao , Baoxing Huai , Zhefeng Wang

Current state-of-the-art photorealistic generators are computationally expensive, involve unstable training processes, and have real and synthetic distributions that are dissimilar in higher-dimensional spaces. To solve these issues, we…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Badr Belhiti , Justin Milushev , Avinash Gupta , John Breedis , Johnson Dinh , Jesse Pisel , Michael Pyrcz

Existing gesture generation methods primarily focus on upper body gestures based on audio features, neglecting speech content, emotion, and locomotion. These limitations result in stiff, mechanical gestures that fail to convey the true…

Sound · Computer Science 2026-03-10 Yongkang Cheng , Mingjiang Liang , Shaoli Huang , Gaoge Han , Jifeng Ning , Wei Liu

Recent advances in visually-induced audio generation are based on sampling short, low-fidelity, and one-class sounds. Moreover, sampling 1 second of audio from the state-of-the-art model takes minutes on a high-end GPU. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Vladimir Iashin , Esa Rahtu

Recently, mainstream mel-spectrogram-based neural vocoders rely on generative adversarial network (GAN) for high-fidelity speech generation, e.g., HiFi-GAN and BigVGAN. However, the use of GAN restricts training efficiency and model…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Hui-Peng Du , Yang Ai , Rui-Chen Zheng , Ye-Xin Lu , Zhen-Hua Ling

In this paper, we present a Diffusion GAN based approach (Prosodic Diff-TTS) to generate the corresponding high-fidelity speech based on the style description and content text as an input to generate speech samples within only 4 denoising…

Sound · Computer Science 2023-10-30 Neeraj Kumar , Ankur Narang , Brejesh Lall

Automatic speaker verification (ASV) systems are highly vulnerable to presentation attacks, also called spoofing attacks. Replay is among the simplest attacks to mount - yet difficult to detect reliably. The generalization failure of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-24 Bhusan Chettri , Tomi Kinnunen , Emmanouil Benetos

The EMelodyGen system focuses on emotional melody generation in ABC notation controlled by the musical feature template. Owing to the scarcity of well-structured and emotionally labeled sheet music, we designed a template for controlling…

Information Retrieval · Computer Science 2025-05-20 Monan Zhou , Xiaobing Li , Feng Yu , Wei Li

Interest in generative Electrocardiogram-Language Models (ELMs) is growing, as they can produce textual responses conditioned on ECG signals and textual queries. Unlike traditional classifiers that output label probabilities, ELMs are more…

Computation and Language · Computer Science 2025-10-02 Xiaoyu Song , William Han , Tony Chen , Chaojing Duan , Michael A. Rosenberg , Emerson Liu , Ding Zhao

This paper proposes a hierarchical generative model with a multi-grained latent variable to synthesize expressive speech. In recent years, fine-grained latent variables are introduced into the text-to-speech synthesis that enable the fine…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-28 Yukiya Hono , Kazuna Tsuboi , Kei Sawada , Kei Hashimoto , Keiichiro Oura , Yoshihiko Nankaku , Keiichi Tokuda
‹ Prev 1 8 9 10 Next ›