中文
相关论文

相关论文: Learning Hierarchical Cross-Modal Association for …

200 篇论文

Human communication combines speech with expressive nonverbal cues such as hand gestures that serve manifold communicative functions. Yet, current generative gesture generation approaches are restricted to simple, repetitive beat gestures…

人机交互 · 计算机科学 2025-10-21 Hendric Voss , Stefan Kopp

Gestures play a key role in human communication. Recent methods for co-speech gesture generation, while managing to generate beat-aligned motions, struggle generating gestures that are semantically aligned with the utterance. Compared to…

计算机视觉与模式识别 · 计算机科学 2024-03-27 Muhammad Hamza Mughal , Rishabh Dabral , Ikhsanul Habibie , Lucia Donatelli , Marc Habermann , Christian Theobalt

Generating co-speech gestures in real time requires both temporal coherence and efficient sampling. We introduce a novel framework for streaming gesture generation that extends Rolling Diffusion models with structured progressive noise…

机器学习 · 计算机科学 2025-11-20 Evgeniia Vu , Andrei Boiarov , Dmitry Vetrov

Generating full-body human gestures based on speech signals remains challenges on quality and speed. Existing approaches model different body regions such as body, legs and hands separately, which fail to capture the spatial interactions…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Pinxin Liu , Luchuan Song , Junhua Huang , Haiyang Liu , Chenliang Xu

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Mengyi Shan , Shouchieh Chang , Ziqian Bai , Shichen Liu , Yinda Zhang , Luchuan Song , Rohit Pandey , Sean Fanello , Zeng Huang

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

In this paper, we propose a novel approach to convert given speech audio to a photo-realistic speaking video of a specific person, where the output video has synchronized, realistic, and expressive rich body dynamics. We achieve this by…

计算机视觉与模式识别 · 计算机科学 2020-10-12 Miao Liao , Sibo Zhang , Peng Wang , Hao Zhu , Xinxin Zuo , Ruigang Yang

The art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the…

人机交互 · 计算机科学 2023-05-09 Sicheng Yang , Zhiyong Wu , Minglei Li , Zhensong Zhang , Lei Hao , Weihong Bao , Ming Cheng , Long Xiao

Body language such as conversational gesture is a powerful way to ease communication. Conversational gestures do not only make a speech more lively but also contain semantic meaning that helps to stress important information in the…

机器人学 · 计算机科学 2022-10-14 Hitoshi Teshima , Naoki Wake , Diego Thomas , Yuta Nakashima , Hiroshi Kawasaki , Katsushi Ikeuchi

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we…

计算机视觉与模式识别 · 计算机科学 2019-05-13 Lele Chen , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

Co-speech gestures are fundamental for communication. The advent of recent deep learning techniques has facilitated the creation of lifelike, synchronous co-speech gestures for Embodied Conversational Agents. "In-the-wild" datasets,…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Téo Guichoux , Laure Soulier , Nicolas Obin , Catherine Pelachaud

We present a multimodal learning-based method to simultaneously synthesize co-speech facial expressions and upper-body gestures for digital characters using RGB video data captured using commodity cameras. Our approach learns from sparse…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Uttaran Bhattacharya , Aniket Bera , Dinesh Manocha

Co-speech gesture generation has significantly advanced human-computer interaction, yet speaker movements remain constrained due to the omission of text-driven non-spontaneous gestures (e.g., bowing while talking). Existing methods face two…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Fengyi Fang , Sicheng Yang , Wenming Yang

This paper presents a novel framework for speech-driven gesture production, applicable to virtual agents to enhance human-computer interaction. Specifically, we extend recent deep-learning-based, data-driven methods for speech-driven…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Taras Kucherenko , Dai Hasegawa , Naoshi Kaneko , Gustav Eje Henter , Hedvig Kjellström

In dyadic speaker-listener interactions, the listener's head reactions along with the speaker's head movements, constitute an important non-verbal semantic expression together. The listener Head generation task aims to synthesize responsive…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Zhigang Chang , Weitai Hu , Qing Yang , Shibao Zheng

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)…

计算机视觉与模式识别 · 计算机科学 2021-07-21 Suzhen Wang , Lincheng Li , Yu Ding , Changjie Fan , Xin Yu

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

Given a piece of text, a video clip and a reference audio, the movie dubbing (also known as visual voice clone V2C) task aims to generate speeches that match the speaker's emotion presented in the video using the desired speaker voice as…

计算与语言 · 计算机科学 2023-04-05 Gaoxiang Cong , Liang Li , Yuankai Qi , Zhengjun Zha , Qi Wu , Wenyu Wang , Bin Jiang , Ming-Hsuan Yang , Qingming Huang

Human communication seamlessly integrates speech and bodily motion, where hand gestures naturally complement vocal prosody to express intent, emotion, and emphasis. While recent text-to-speech (TTS) systems have begun incorporating…

音频与语音处理 · 电气工程与系统科学 2026-03-23 Lokesh Kumar , Nirmesh Shah , Ashishkumar P. Gudmalwar , Pankaj Wasnik

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Zhen Ye , Xu Tan , Aoxiong Yin , Hongzhan Lin , Guangyan Zhang , Peiwen Sun , Yiming Li , Chi-Min Chan , Wei Ye , Shikun Zhang , Wei Xue