English
Related papers

Related papers: Attention-based Interactive Disentangling Network …

200 papers

Entrainment is a known adaptation mechanism that causes interaction participants to adapt or synchronize their acoustic characteristics. Understanding how interlocutors tend to adapt to each other's speaking style through entrainment…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-15 Md Nasir , Brian Baucom , Shrikanth Narayanan , Panayiotis Georgiou

Network embedding represents nodes in a continuous vector space and preserves structure information from the Network. Existing methods usually adopt a "one-size-fits-all" approach when concerning multi-scale structure information, such as…

Machine Learning · Computer Science 2018-03-28 Lei Sang , Min Xu , Shengsheng Qian , Xindong Wu

Speech emotion recognition (SER) has drawn increasing attention for its applications in human-machine interaction. However, existing SER methods ignore the information gap between the pre-training speech recognition task and the downstream…

Sound · Computer Science 2023-10-03 Dongyuan Li , Yusong Wang , Kotaro Funakoshi , Manabu Okumura

Researchers have developed excellent feed-forward models that learn to map images to desired outputs, such as to the images' latent factors, or to other images, using supervised learning. Learning such mappings from unlabelled data, or…

Computer Vision and Pattern Recognition · Computer Science 2017-11-21 Hsiao-Yu Fish Tung , Adam W. Harley , William Seto , Katerina Fragkiadaki

Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion…

Computation and Language · Computer Science 2018-04-02 Egor Lakomkin , Cornelius Weber , Sven Magg , Stefan Wermter

Emotion recognition is a fundamental component of next-generation human-computer interaction (HCI), enabling machines to perceive, understand, and respond to users' affective states. However, existing systems often rely on single-modality…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Ziwen Zhong , Zhitao Shu , Yue Zhao

This paper proposes a speech emotion recognition method based on speech features and speech transcriptions (text). Speech features such as Spectrogram and Mel-frequency Cepstral Coefficients (MFCC) help retain emotion-related low-level…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-14 Suraj Tripathi , Abhay Kumar , Abhiram Ramesh , Chirag Singh , Promod Yenigalla

Voice conversion for highly expressive speech is challenging. Current approaches struggle with the balancing between speaker similarity, intelligibility and expressiveness. To address this problem, we propose Expressive-VC, a novel…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-10 Ziqian Ning , Qicong Xie , Pengcheng Zhu , Zhichao Wang , Liumeng Xue , Jixun Yao , Lei Xie , Mengxiao Bi

Voice-enabled interactions provide more human-like experiences in many popular IoT systems. Cloud-based speech analysis services extract useful information from voice input using speech recognition techniques. The voice signal is a rich…

Cryptography and Security · Computer Science 2019-08-13 Ranya Aloufi , Hamed Haddadi , David Boyle

Dynamic emotion recognition in the wild remains challenging due to the transient nature of emotional expressions and temporal misalignment of multi-modal cues. Traditional approaches predict valence and arousal and often overlook the…

Machine Learning · Computer Science 2025-05-05 Vrushank Ahire , Kunal Shah , Mudasir Nazir Khan , Nikhil Pakhale , Lownish Rai Sookha , M. A. Ganaie , Abhinav Dhall

In this paper, adaptive mechanisms are applied in deep neural network (DNN) training for x-vector-based text-independent speaker verification. First, adaptive convolutional neural networks (ACNNs) are employed in frame-level embedding…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-18 Bin Gu , Wu Guo , Lirong Dai , Jun Du

Speech emotion recognition is a challenging task because the emotion expression is complex, multimodal and fine-grained. In this paper, we propose a novel multimodal deep learning approach to perform fine-grained emotion recognition from…

Sound · Computer Science 2021-07-16 Hang Li , Wenbiao Ding , Zhongqin Wu , Zitao Liu

Speech-to-text translation (ST), which translates source language speech into target language text, has attracted intensive attention in recent years. Compared to the traditional pipeline system, the end-to-end ST model has potential…

Computation and Language · Computer Science 2019-12-17 Yuchen Liu , Jiajun Zhang , Hao Xiong , Long Zhou , Zhongjun He , Hua Wu , Haifeng Wang , Chengqing Zong

Large language models have revolutionized sign language generation by automatically transforming text into high-quality sign language videos, providing accessible communication for the Deaf community. However, existing LLM-based approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yanchao Zhao , Jihao Zhu , Yu Liu , Weizhuo Chen , Yuling Yang , Kun Peng

Learning effective joint representations has been a central task in multi-modal sentiment analysis. Previous works addressing this task focus on exploring sophisticated fusion techniques to enhance performance. However, the inherent…

Multimedia · Computer Science 2024-08-20 Weichen Dai , Xingyu Li , Zeyu Wang , Pengbo Hu , Ji Qi , Jianlin Peng , Yi Zhou

Recently, voice conversion (VC) has been widely studied. Many VC systems use disentangle-based learning techniques to separate the speaker and the linguistic content information from a speech signal. Subsequently, they convert the voice by…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-03 Yen-Hao Chen , Da-Yi Wu , Tsung-Han Wu , Hung-yi Lee

Expressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disentangle the emotional…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-22 Zongyang Du , Berrak Sisman , Kun Zhou , Haizhou Li

In recent years, deep learning has achieved innovative advancements in various fields, including the analysis of human emotions and behaviors. Initiatives such as the Affective Behavior Analysis in-the-wild (ABAW) competition have been…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Seongjae Min , Junseok Yang , Sangjun Lim , Junyong Lee , Sangwon Lee , Sejoon Lim

Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details of the reference speech. To this end, we propose a novel…

Sound · Computer Science 2025-10-03 Jianing Yang , Sheng Li , Takahiro Shinozaki , Yuki Saito , Hiroshi Saruwatari

The field of Explainable Artificial Intelligence (XAI) aims to build explainable and interpretable machine learning (or deep learning) methods without sacrificing prediction performance. Convolutional Neural Networks (CNNs) have been…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Shaw-Hwa Lo , Yiqiao Yin