中文
相关论文

相关论文: Learning Hierarchical Cross-Modal Association for …

200 篇论文

Co-Speech Gesture Video Generation aims to generate vivid speech videos from audio-driven still images, which is challenging due to the diversity of body parts in terms of motion amplitude, audio relevance, and detailed features. Relying…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Siyuan Wang , Jiawei Liu , Wei Wang , Yeying Jin , Jinsong Du , Zhi Han

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap…

音频与语音处理 · 电气工程与系统科学 2025-03-24 Ji-Hoon Kim , Jeongsoo Choi , Jaehun Kim , Chaeyoung Jung , Joon Son Chung

Embodied agents, in the form of virtual agents or social robots, are rapidly becoming more widespread. In human-human interactions, humans use nonverbal behaviours to convey their attitudes, feelings, and intentions. Therefore, this…

人工智能 · 计算机科学 2026-04-30 Carson Yu Liu , Gelareh Mohammadi , Yang Song , Wafa Johal

Generating vivid and emotional 3D co-speech gestures is crucial for virtual avatar animation in human-machine interaction applications. While the existing methods enable generating the gestures to follow a single emotion label, they…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Xingqun Qi , Jiahao Pan , Peng Li , Ruibin Yuan , Xiaowei Chi , Mengfei Li , Wenhan Luo , Wei Xue , Shanghang Zhang , Qifeng Liu , Yike Guo

When humans speak, gestures help convey communicative intentions, such as adding emphasis or describing concepts. However, current co-speech gesture generation methods rely solely on superficial linguistic cues (e.g. speech audio or text…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Pinxin Liu , Haiyang Liu , Luchuan Song , Jason J. Corso , Chenliang Xu

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio, we output multiple possibilities of gestural motion for an…

计算机视觉与模式识别 · 计算机科学 2024-01-04 Evonne Ng , Javier Romero , Timur Bagautdinov , Shaojie Bai , Trevor Darrell , Angjoo Kanazawa , Alexander Richard

Synthesizing realistic co-speech gestures is an important and yet unsolved problem for creating believable motions that can drive a humanoid robot to interact and communicate with human users. Such capability will improve the impressions of…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Shuhong Lu , Youngwoo Yoon , Andrew Feng

Due to their significance in human communication, the automatic generation of co-speech gestures in artificial embodied agents has received a lot of attention. Although modern deep learning approaches can generate realistic-looking…

人机交互 · 计算机科学 2023-07-20 Hendric Voß , Stefan Kopp

Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing part-decomposed or…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Xuanmeng Sha , Liyun Zhang , Tomohiro Mashita , Naoya Chiba , Yuki Uranishi

The automatic co-speech gesture generation draws much attention in computer animation. Previous works designed network structures on individual datasets, which resulted in a lack of data volume and generalizability across different motion…

We propose a real-time system for synthesizing gestures directly from speech. Our data-driven approach is based on Generative Adversarial Neural Networks to model the speech-gesture relationship. We utilize the large amount of speaker video…

计算机视觉与模式识别 · 计算机科学 2022-08-08 Manuel Rebol , Christian Gütl , Krzysztof Pietroszek

A key component of dyadic spoken interactions is the contextually relevant non-verbal gestures, such as head movements that reflect a listener's response to the interlocutor's speech. Although significant progress has been made in the…

机器人学 · 计算机科学 2024-10-01 Bishal Ghosh , Emma Li , Tanaya Guha

While previous audio-driven talking head generation (THG) methods generate head poses from driving audio, the generated poses or lips cannot match the audio well or are not editable. In this study, we propose \textbf{PoseTalk}, a THG system…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Jun Ling , Yiwen Wang , Han Xue , Rong Xie , Li Song

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader…

声音 · 计算机科学 2025-03-25 Yong Ren , Chenxing Li , Manjie Xu , Wei Liang , Yu Gu , Rilin Chen , Dong Yu

Speech-driven gesture generation aims at synthesizing a gesture sequence synchronized with the input speech signal. Previous methods leverage neural networks to directly map a compact audio representation to the gesture sequence, ignoring…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Fengqi Liu , Hexiang Wang , Jingyu Gong , Ran Yi , Qianyu Zhou , Xuequan Lu , Jiangbo Lu , Lizhuang Ma

Generating full-body human gestures encompassing face, body, hands, and global movements from audio is a valuable yet challenging task in virtual avatar creation. Previous systems focused on tokenizing the human gestures framewisely and…

图形学 · 计算机科学 2025-05-20 Zhizhuo Yin , Yuk Hang Tsui , Pan Hui

Recent increase of remote-work, online meeting and tele-operation task makes people find that gesture for avatars and communication robots is more important than we have thought. It is one of the key factors to achieve smooth and natural…

人机交互 · 计算机科学 2023-09-29 Hitoshi Teshima , Naoki Wake , Diego Thomas , Yuta Nakashima , Hiroshi Kawasaki , Katsushi Ikeuchi

How can we teach robots or virtual assistants to gesture naturally? Can we go further and adapt the gesturing style to follow a specific speaker? Gestures that are naturally timed with corresponding speech during human communication are…

计算机视觉与模式识别 · 计算机科学 2020-07-27 Chaitanya Ahuja , Dong Won Lee , Yukiko I. Nakano , Louis-Philippe Morency

The generation of co-speech gestures for digital humans is an emerging area in the field of virtual human creation. Prior research has made progress by using acoustic and semantic information as input and adopting classify method to…

声音 · 计算机科学 2024-04-16 Fan Zhang , Naye Ji , Fuxing Gao , Siyuan Zhao , Zhaohan Wang , Shunman Li

Audio-driven co-speech human gesture generation has made remarkable advancements recently. However, most previous works only focus on single person audio-driven gesture generation. We aim at solving the problem of conversational co-speech…

人机交互 · 计算机科学 2024-01-12 Haiwei Xue , Sicheng Yang , Zhensong Zhang , Zhiyong Wu , Minglei Li , Zonghong Dai , Helen Meng