English
Related papers

Related papers: ANGUS: Real-time manipulation of vocal roughness f…

200 papers

Millions of people reach out to digital assistants such as Siri every day, asking for information, making phone calls, seeking assistance, and much more. The expectation is that such assistants should understand the intent of the users…

Computation and Language · Computer Science 2019-07-02 Vikramjit Mitra , Sue Booker , Erik Marchi , David Scott Farrar , Ute Dorothea Peitz , Bridget Cheng , Ermine Teves , Anuj Mehta , Devang Naik

Exploring moral dilemmas allows individuals to navigate moral complexity, where a reversal in decision certainty, shifting toward the opposite of one's initial choice, could reflect open-mindedness and less rigidity. This study probes how…

Human-Computer Interaction · Computer Science 2024-12-30 Chenyi Zhang , Zhenhao Zhang , Wei Zhang , Tian Zeng , Black Sun , Jian Zhao , Pengcheng An

This paper proposes a neural network that performs audio transformations to user-specified sources (e.g., vocals) of a given audio track according to a given description while preserving other sources not mentioned in the description. Audio…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-29 Woosung Choi , Minseok Kim , Marco A. Martínez Ramírez , Jaehwa Chung , Soonyoung Jung

Acoustic environments affect acoustic characteristics of sound to be recognized by physically interacting with sound wave propagation. Thus, training acoustic models for audio and speech tasks requires regularization on various acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-08 Hyeonuk Nam , Seong-Hu Kim , Yong-Hwa Park

The proposed model is only for the audio module. All videos in the OMG Emotion Dataset are converted to WAV files. The proposed model makes use of semi-supervised learning for the emotion recognition. A GAN is trained with unsupervised…

Sound · Computer Science 2018-05-07 Ingryd Pereira , Diego Santos

Conversational agents (CAs) are revolutionizing human-computer interaction by evolving from text-based chatbots to empathetic digital humans (DHs) capable of rich emotional expressions. This paper explores the integration of neural and…

Human-Computer Interaction · Computer Science 2025-08-11 Nastaran Saffaryazdi , Tamil Selvan Gunasekaran , Kate Laveys , Elizabeth Broadbent , Mark Billinghurst

Despite rapid advancement in recent years, current speech enhancement models often produce speech that differs in perceptual quality from real clean speech. We propose a learning objective that formalizes differences in perceptual quality,…

The main motivation for Automatic Speech Recognition (ASR) is efficient interfaces to computers, and for the interfaces to be natural and truly useful, it should provide coverage for a large group of users. The purpose of these tasks is to…

Computation and Language · Computer Science 2013-03-25 Urmila Shrawankar , VM Thakare

Voice user interfaces and digital assistants are rapidly entering our lives and becoming singular touch points spanning our devices. These always-on services capture and transmit our audio data to powerful cloud services for further…

Computation and Language · Computer Science 2022-10-03 Ranya Aloufi , Hamed Haddadi , David Boyle

With the widespread application of automatic speech recognition (ASR) systems, their vulnerability to adversarial attacks has been extensively studied. However, most existing adversarial examples are generated on specific individual models,…

Sound · Computer Science 2025-03-26 Weifei Jin , Junjie Su , Hejia Wang , Yulin Ye , Jie Hao

Realistic talking-head video generation is critical for virtual avatars, film production, and interactive systems. Current methods struggle with nuanced emotional expressions due to the lack of fine-grained emotion control. To address this…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Jiayi Lyu , Leigang Qu , Wenjing Zhang , Hanyu Jiang , Kai Liu , Zhenglin Zhou , Xiaobo Xia , Jian Xue , Tat-Seng Chua

Most automatic emotion recognition systems exploit time-continuous annotations of emotion to provide fine-grained descriptions of spontaneous expressions as observed in real-life interactions. As emotion is rather subjective, its annotation…

Sound · Computer Science 2022-09-22 Sina Alisamir , Fabien Ringeval , Francois Portet

Depression, as a typical mental disorder, has become a prevalent issue significantly impacting public health. However, the prevention and treatment of depression still face multiple challenges, including complex diagnostic procedures,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yu Luo , Nan Huang , Sophie Yu , Hendry Xu , Jerry Wang , Colin Wang , Zhichao Liu , Chen Zeng

Data augmentation has proven to be effective in training neural networks. Recently, a method called RandAug was proposed, randomly selecting data augmentation techniques from a predefined search space. RandAug has demonstrated significant…

A speech emotion recognition algorithm based on multi-feature and Multi-lingual fusion is proposed in order to resolve low recognition accuracy caused by lack of large speech dataset and low robustness of acoustic features in the…

Computation and Language · Computer Science 2020-01-17 Chunyi Wang

Non-Verbal Vocalisations (NVVs) are short `non-word' utterances without proper linguistic (semantic) meaning but conveying connotations -- be this emotions/affects or other paralinguistic information. We start this contribution with a…

Sound · Computer Science 2025-08-05 Anton Batliner , Shahin Amiriparian , Björn W. Schuller

The disparity between the computational demands of deep learning and the capabilities of compute hardware is expanding drastically. Although deep learning achieves remarkable performance in countless tasks, its escalating requirements for…

Machine Learning · Computer Science 2025-09-12 Xiao Wang , Hendrik Borras , Bernhard Klein , Holger Fröning

Animal vocalization denoising is a task similar to human speech enhancement, which is relatively well-studied. In contrast to the latter, it comprises a higher diversity of sound production mechanisms and recording environments, and this…

Hearing aids use dynamic range compression (DRC), a form of automatic gain control, to make quiet sounds louder and loud sounds quieter. Compression can improve listening comfort, but it can also cause distortion in noisy environments. It…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-28 Ryan M. Corey , Andrew C. Singer

This paper is devoted to improve automatic emotion recognition from speech by incorporating rhythm and temporal features. Research on automatic emotion recognition so far has mostly been based on applying features like MFCCs, pitch and…

Computer Vision and Pattern Recognition · Computer Science 2013-03-08 Mayank Bhargava , Tim Polzehl
‹ Prev 1 3 4 5 6 7 10 Next ›