中文
相关论文

相关论文: Dyadic Speech-based Affect Recognition using DAMI-…

200 篇论文

Natural conversations between humans often involve a large number of non-verbal nuanced expressions, displayed at key times throughout the conversation. Understanding and being able to model these complex interactions is essential for…

计算机视觉与模式识别 · 计算机科学 2022-07-12 Renke Wang , Ifeoma Nwogu

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Mengyi Shan , Shouchieh Chang , Ziqian Bai , Shichen Liu , Yinda Zhang , Luchuan Song , Rohit Pandey , Sean Fanello , Zeng Huang

Diagnostic procedures for ASD (autism spectrum disorder) involve semi-naturalistic interactions between the child and a clinician. Computational methods to analyze these sessions require an end-to-end speech and language processing pipeline…

音频与语音处理 · 电气工程与系统科学 2019-10-28 Rimita Lahiri , Manoj Kumar , Somer Bishop , Shrikanth Narayanan

Human-human communication is like a delicate dance where listeners and speakers concurrently interact to maintain conversational dynamics. Hence, an effective model for generating listener nonverbal behaviors requires understanding the…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Minh Tran , Di Chang , Maksim Siniukov , Mohammad Soleymani

Different from the emotion recognition in individual utterances, we propose a multimodal learning framework using relation and dependencies among the utterances for conversational emotion analysis. The attention mechanism is applied to the…

计算与语言 · 计算机科学 2019-10-25 Zheng Lian , Jianhua Tao , Bin Liu , Jian Huang

Affect understanding capability is essential for social robots to autonomously interact with a group of users in an intuitive and reciprocal way. However, the challenge of multi-person affect understanding comes from not only the accurate…

计算机视觉与模式识别 · 计算机科学 2023-01-02 Yubin Kim , Huili Chen , Sharifa Alghowinem , Cynthia Breazeal , Hae Won Park

Psychomotor retardation in depression has been associated with speech timing changes from dyadic clinical interviews. In this work, we investigate speech timing features from free-living dyadic interactions. Apart from the possibility of…

声音 · 计算机科学 2022-09-09 Bishal Lamichhane , Nidal Moukaddam , Ankit B. Patel , Ashutosh Sabharwal

When recognizing emotions from speech, we encounter two common problems: how to optimally capture emotion-relevant information from the speech signal and how to best quantify or categorize the noisy subjective emotion labels.…

音频与语音处理 · 电气工程与系统科学 2022-11-04 Sofoklis Kakouros , Themos Stafylakis , Ladislav Mosner , Lukas Burget

The potential of multimodal generative artificial intelligence (mAI) to replicate human grounded language understanding, including the pragmatic, context-rich aspects of communication, remains to be clarified. Humans are known to use…

In speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers.…

声音 · 计算机科学 2024-01-19 Cheng Lu , Yuan Zong , Hailun Lian , Yan Zhao , Björn Schuller , Wenming Zheng

Many applications of speech communication and speaker identification suffer from the problem of co-channel speech. This paper deals with a multi-resolution dyadic wavelet transform method for usable segments of co-channel speech detection…

声音 · 计算机科学 2013-01-03 Wajdi Ghezaiel , Amel Ben Slimane Rahmouni , Ezzedine Ben Braiek

Multimodal dialogue emotion recognition captures emotional cues by fusing text, visual, and audio modalities. However, existing approaches still suffer from notable limitations in modeling emotional dependencies and learning multimodal…

多媒体 · 计算机科学 2026-03-12 Yunsheng Wang , Yuntao Shou , Yilong Tan , Wei Ai , Tao Meng , Keqin Li

Human emotional expression emerges through coordinated vocal, facial, and gestural signals. While speech face alignment is well established, the broader dynamics linking emotionally expressive speech to regional facial and hand motion…

多媒体 · 计算机科学 2025-06-13 Von Ralph Dane Marquez Herbuela , Yukie Nagai

The dyadic reaction generation task involves synthesizing responsive facial reactions that align closely with the behaviors of a conversational partner, enhancing the naturalness and effectiveness of human-like interaction simulations. This…

机器学习 · 计算机科学 2025-05-14 Minh-Duc Nguyen , Hyung-Jeong Yang , Soo-Hyung Kim , Ji-Eun Shin , Seung-Won Kim

This study investigates the efficacy of using multimodal machine learning techniques to detect deception in dyadic interactions, focusing on the integration of data from both the deceiver and the deceived. We compare early and late fusion…

机器学习 · 计算机科学 2025-12-12 Thomas Jack Samuels , Franco Rugolon , Stephan Hau , Lennart Högman

Spoken languages often utilise intonation, rhythm, intensity, and structure, to communicate intention, which can be interpreted differently depending on the rhythm of speech of their utterance. These speech acts provide the foundation of…

Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or employ speech…

计算与语言 · 计算机科学 2020-04-06 Haiyang Xu , Hui Zhang , Kun Han , Yun Wang , Yiping Peng , Xiangang Li

Recent studies have shown that frame-level deep speaker features can be derived from a deep neural network with the training target set to discriminate speakers by a short speech segment. By pooling the frame-level features, utterance-level…

音频与语音处理 · 电气工程与系统科学 2018-11-09 Lantian Li , Zhiyuan Tang , Ying Shi , Dong Wang

Over the past decade, wearable computing devices (``smart glasses'') have undergone remarkable advancements in sensor technology, design, and processing power, ushering in a new era of opportunity for high-density human behavior data.…

Acoustic-prosodic entrainment describes the tendency of humans to align or adapt their speech acoustics to each other in conversation. This alignment of spoken behavior has important implications for conversational success. However,…

音频与语音处理 · 电气工程与系统科学 2018-07-13 Megan M. Willi , Stephanie A. Borrie , Tyson S. Barrett , Ming Tu , Visar Berisha