中文
相关论文

相关论文: Beyond Talking -- Generating Holistic 3D Human Dya…

200 篇论文

We introduce a video framework for modeling the association between verbal and non-verbal communication during dyadic conversation. Given the input speech of a speaker, our approach retrieves a video of a listener, who has facial…

计算机视觉与模式识别 · 计算机科学 2023-01-27 Scott Geng , Revant Teotia , Purva Tendulkar , Sachit Menon , Carl Vondrick

Audio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, expressive facial expressions, natural head pose generation, and high video quality. However, no model has yet led…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Xusen Sun , Longhao Zhang , Hao Zhu , Peng Zhang , Bang Zhang , Xinya Ji , Kangneng Zhou , Daiheng Gao , Liefeng Bo , Xun Cao

Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing part-decomposed or…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Xuanmeng Sha , Liyun Zhang , Tomohiro Mashita , Naoya Chiba , Yuki Uranishi

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears,…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

This paper presents a novel framework for speech-driven gesture production, applicable to virtual agents to enhance human-computer interaction. Specifically, we extend recent deep-learning-based, data-driven methods for speech-driven…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Taras Kucherenko , Dai Hasegawa , Naoshi Kaneko , Gustav Eje Henter , Hedvig Kjellström

This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-realistic quality, and…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Xiang Deng , Youxin Pang , Xiaochen Zhao , Chao Xu , Lizhen Wang , Hongjiang Xiao , Shi Yan , Hongwen Zhang , Yebin Liu

The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have explored talking or…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Joon Son Chung , Shinji Watanabe

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an efficient…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Qi Wu , Yubo Zhao , Yifan Wang , Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

This paper reports on the GENEA Challenge 2023, in which participating teams built speech-driven gesture-generation systems using the same speech and motion dataset, followed by a joint evaluation. This year's challenge provided data on…

We present a new research task and a dataset to understand human social interactions via computational methods, to ultimately endow machines with the ability to encode and decode a broad channel of social signals humans use. This research…

计算机视觉与模式识别 · 计算机科学 2019-06-11 Hanbyul Joo , Tomas Simon , Mina Cikara , Yaser Sheikh

This work presents Interactive Conversational 3D Virtual Human (ICo3D), a method for generating an interactive, conversational, and photorealistic 3D human avatar. Based on multi-view captures of a subject, we create an animatable 3D face…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Richard Shaw , Youngkyoon Jang , Athanasios Papaioannou , Arthur Moreau , Helisa Dhamo , Zhensong Zhang , Eduardo Pérez-Pellitero

This paper addresses the problem of generating 3D interactive human motion from text. Given a textual description depicting the actions of different body parts in contact with static objects, we synthesize sequences of 3D body poses that…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Sihan Ma , Qiong Cao , Jing Zhang , Dacheng Tao

Generating human motions from textual descriptions has gained growing research interest due to its wide range of applications. However, only a few works consider human-scene interactions together with text conditions, which is crucial for…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Zhi Cen , Huaijin Pi , Sida Peng , Zehong Shen , Minghui Yang , Shuai Zhu , Hujun Bao , Xiaowei Zhou

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized Variational…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Sohan Anisetty , James Hays

Generating 3D-based body movements from speech shows great potential in extensive downstream applications, while it still suffers challenges in imitating realistic human movements. Predominant research efforts focus on end-to-end generation…

人工智能 · 计算机科学 2025-12-16 Xin Guo , Yifan Zhao , Jia Li

Generating realistic 3D human-human interactions from textual descriptions remains a challenging task. Existing approaches, typically based on diffusion models, often produce results lacking realism and fidelity. In this work, we introduce…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Muhammad Gohar Javed , Chuan Guo , Li Cheng , Xingyu Li

Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high-fidelity physical plausibility. Existing methods remain limited in their ability…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Zerui Chen , Rolandos Alexandros Potamias , Shizhe Chen , Jiankang Deng , Cordelia Schmid , Stefanos Zafeiriou

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

Animating virtual characters with holistic co-speech gestures is a challenging but critical task. Previous systems have primarily focused on the weak correlation between audio and gestures, leading to physically unnatural outcomes that…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yongkang Cheng , Shaoli Huang

Speech-driven 3D motion synthesis seeks to create lifelike animations based on human speech, with potential uses in virtual reality, gaming, and the film production. Existing approaches reply solely on speech audio for motion generation,…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Wenshuo Peng , Kaipeng Zhang , Sai Qian Zhang