中文
相关论文

相关论文: EmotionGesture: Audio-Driven Diverse Emotional Co-…

200 篇论文

While the field of co-speech gesture generation has seen significant advances, producing holistic, semantically grounded gestures remains a challenge. Existing approaches rely on external semantic retrieval methods, which limit their…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Lanmiao Liu , Esam Ghaleb , Aslı Özyürek , Zerrin Yumak

Co-speech gestures increase engagement and improve speech understanding. Most data-driven robot systems generate rhythmic beat-like motion, yet few integrate semantic emphasis. To address this, we propose a lightweight transformer that…

Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Xuli Shen , Hua Cai , Dingding Yu , Weilin Shen , Qing Xu , Xiangyang Xue

Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Mengchao Wang , Qiang Wang , Fan Jiang , Yaqi Fan , Yunpeng Zhang , Yonggang Qi , Kun Zhao , Mu Xu

With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face conversation has both spoken and non-verbal aspects (here,…

音频与语音处理 · 电气工程与系统科学 2023-09-15 Shivam Mehta , Siyang Wang , Simon Alexanderson , Jonas Beskow , Éva Székely , Gustav Eje Henter

Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a single audio stream to a single speaker's motion, without…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yichen Peng , Jyun-Ting Song , Siyeol Jung , Ruofan Liu , Haiyang Liu , Xuangeng Chu , Ruicong Liu , Erwin Wu , Hideki Koike , Kris Kitani

State-of-the-art speech synthesis models try to get as close as possible to the human voice. Hence, modelling emotions is an essential part of Text-To-Speech (TTS) research. In our work, we selected FastSpeech2 as the starting point and…

音频与语音处理 · 电气工程与系统科学 2023-07-04 Daria Diatlova , Vitaly Shutov

Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic full-body motion based on text prompts, such as talking while…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Bohong Chen , Yumeng Li , Yao-Xiang Ding , Tianjia Shao , Kun Zhou

Despite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important features on natural…

计算机视觉与模式识别 · 计算机科学 2021-05-21 Xinya Ji , Hang Zhou , Kaisiyuan Wang , Wayne Wu , Chen Change Loy , Xun Cao , Feng Xu

Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches still encounter…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Xingpei Ma , Jiaran Cai , Yuansheng Guan , Shenneng Huang , Qiang Zhang , Shunsi Zhang

We propose a new framework for gesture generation, aiming to allow data-driven approaches to produce more semantically rich gestures. Our approach first predicts whether to gesture, followed by a prediction of the gesture properties. Those…

人机交互 · 计算机科学 2021-09-28 Taras Kucherenko , Rajmund Nagy , Patrik Jonell , Michael Neff , Hedvig Kjellström , Gustav Eje Henter

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Changan Chen , Juze Zhang , Shrinidhi K. Lakshmikanth , Yusu Fang , Ruizhi Shao , Gordon Wetzstein , Li Fei-Fei , Ehsan Adeli

This paper describes a system developed for the GENEA (Generation and Evaluation of Non-verbal Behaviour for Embodied Agents) Challenge 2023. Our solution builds on an existing diffusion-based motion synthesis model. We propose a…

音频与语音处理 · 电气工程与系统科学 2023-09-12 Anna Deichler , Shivam Mehta , Simon Alexanderson , Jonas Beskow

Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced…

计算机视觉与模式识别 · 计算机科学 2024-02-05 Guanwen Feng , Haoran Cheng , Yunan Li , Zhiyuan Ma , Chaoneng Li , Zhihao Qian , Qiguang Miao , Chi-Man Pun

While previous audio-driven talking head generation (THG) methods generate head poses from driving audio, the generated poses or lips cannot match the audio well or are not editable. In this study, we propose \textbf{PoseTalk}, a THG system…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Jun Ling , Yiwen Wang , Han Xue , Rong Xie , Li Song

Several works have developed end-to-end pipelines for generating lip-synced talking faces with various real-world applications, such as teaching and language translation in videos. However, these prior works fail to create realistic-looking…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Sahil Goyal , Shagun Uppal , Sarthak Bhagat , Yi Yu , Yifang Yin , Rajiv Ratn Shah

Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have…

声音 · 计算机科学 2024-10-17 Mithun Manivannan , Vignesh Nethrapalli , Mark Cartwright

This paper presents FaceXHuBERT, a text-less speech-driven 3D facial animation generation method that allows to capture personalized and subtle cues in speech (e.g. identity, emotion and hesitation). It is also very robust to background…

计算机视觉与模式识别 · 计算机科学 2023-03-10 Kazi Injamamul Haque , Zerrin Yumak

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Human speech is often accompanied by body gestures including arm and hand gestures. We present a method that reenacts a high-quality video with gestures matching a target speech audio. The key idea of our method is to split and re-assemble…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Yang Zhou , Jimei Yang , Dingzeyu Li , Jun Saito , Deepali Aneja , Evangelos Kalogerakis