English
Related papers

Related papers: Affective social anthropomorphic intelligent syste…

200 papers

This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Puneet Kumar , Sarthak Malik , Balasubramanian Raman

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

Computation and Language · Computer Science 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

Human emotions are essentially molded by lived experiences, from which we construct personalised meaning. The engagement in such meaning-making process has been practiced as an intervention in various psychotherapies to promote wellness.…

Human-Computer Interaction · Computer Science 2024-03-06 Qian Wan , Xin Feng , Yining Bei , Zhiqi Gao , Zhicong Lu

This paper explores predicting suitable prosodic features for fine-grained emotion analysis from the discourse-level text. To obtain fine-grained emotional prosodic features as predictive values for our model, we extract a phoneme-level…

Sound · Computer Science 2023-09-22 Xianhao Wei , Jia Jia , Xiang Li , Zhiyong Wu , Ziyi Wang

To provide consistent emotional interaction with users, dialog systems should be capable to automatically select appropriate emotions for responses like humans. However, most existing works focus on rendering specified emotions in responses…

Computation and Language · Computer Science 2021-07-01 Wen Zhiyuan , Cao Jiannong , Yang Ruosong , Liu Shuaiqi , Shen Jiaxing

Speech emotion recognition (SER) has traditionally relied on categorical or dimensional labels. However, this technique is limited in representing both the diversity and interpretability of emotions. To overcome this limitation, we focus on…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-19 Ryotaro Nagase , Ryoichi Takashima , Yoichi Yamashita

This paper proposes a system capable of recognizing a speaker's utterance-level emotion through multimodal cues in a video. The system seamlessly integrates multiple AI models to first extract and pre-process multimodal information from the…

Human-Computer Interaction · Computer Science 2023-08-29 Sun-Kyung Lee , Jong-Hwan Kim

Attackers may manipulate audio with the intent of presenting falsified reports, changing an opinion of a public figure, and winning influence and power. The prevalence of inauthentic multimedia continues to rise, so it is imperative to…

Sound · Computer Science 2022-05-05 Emily R. Bartusiak , Edward J. Delp

Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Rahil Parikh , Nadee Seneviratne , Ganesh Sivaraman , Shihab Shamma , Carol Espy-Wilson

Emotion prediction is the field of study to understand human emotions. Existing methods focus on modalities like text, audio, facial expressions, etc., which could be private to the user. Emotion can be derived from the subject's…

Signal Processing · Electrical Eng. & Systems 2023-08-25 Dhruv Limbani , Daketi Yatin , Nitish Chaturvedi , Vaishnavi Moorthy , Pushpalatha M , Harichandana BSS , Sumit Kumar

Traditional voice conversion methods rely on parallel recordings of multiple speakers pronouncing the same sentences. For real-world applications however, parallel data is rarely available. We propose MelGAN-VC, a voice conversion method…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-06 Marco Pasini

In human-computer interaction (HCI), Speech Emotion Recognition (SER) is a key technology for understanding human intentions and emotions. Traditional SER methods struggle to effectively capture the long-term temporal correla-tions and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-18 Xincheng Wang , Liejun Wang , Yinfeng Yu , Xinxin Jiao

Accurate emotion recognition is pivotal for nuanced and engaging human-computer interactions, yet remains difficult to achieve, especially in dynamic, conversation-like settings. In this study, we showcase how integrating eye-tracking data,…

Human-Computer Interaction · Computer Science 2025-11-03 Meisam Jamshidi Seikavandi , Jostein Fimland , Maria Barrett , Paolo Burelli

Turn-taking prediction models are essential components in spoken dialogue systems and conversational robots. Recent approaches leverage transformer-based architectures to predict speech activity continuously and in real-time. In this study,…

Computation and Language · Computer Science 2025-07-04 Koji Inoue , Mikey Elmers , Yahui Fu , Zi Haur Pang , Divesh Lala , Keiko Ochi , Tatsuya Kawahara

In human dialogue, nonverbal information such as nodding and facial expressions is as crucial as verbal information, and spoken dialogue systems are also expected to express such nonverbal behaviors. We focus on nodding, which is critical…

Human-Computer Interaction · Computer Science 2025-08-05 Kazushi Kato , Koji Inoue , Divesh Lala , Keiko Ochi , Tatsuya Kawahara

The increasing demand for mental health services has highlighted the need for innovative solutions, particularly in the realm of psychological conversational AI, where the availability of sensitive data is scarce. In this work, we explored…

Human-Computer Interaction · Computer Science 2024-12-31 Alessandro De Grandi , Federico Ravenda , Andrea Raballo , Fabio Crestani

Speech emotion recognition (SER) has been a challenging problem in spoken language processing research, because it is unclear how human emotions are connected to various components of sounds such as pitch, loudness, and energy. This paper…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Tai Vu

The growing prevalence of Large Language Models (LLMs) is reshaping online text-based communication; a transformation that is extensively studied as AI-mediated communication. However, much of the existing research remains bound by…

Human-Computer Interaction · Computer Science 2025-02-26 Shutaro Aoyama , Rintaro Chujo , Ari Hautasaari , Takeshi Naemura

With large language models (LLMs) becoming increasingly prevalent in daily life, so too has the tendency to attribute to them human-like minds and emotions, or anthropomorphize them. Here, we investigate dimensions people use to…

In this paper, we propose a new methodology for emotional speech recognition using visual deep neural network models. We employ the transfer learning capabilities of the pre-trained computer vision deep models to have a mandate for the…

Computer Vision and Pattern Recognition · Computer Science 2022-04-08 Waleed Ragheb , Mehdi Mirzapour , Ali Delfardi , Hélène Jacquenet , Lawrence Carbon
‹ Prev 1 8 9 10 Next ›