中文
相关论文

相关论文: LivelySpeaker: Towards Semantic-Aware Co-Speech Ge…

200 篇论文

Human speech is often accompanied by hand and arm gestures. Given audio speech input, we generate plausible gestures to go along with the sound. Specifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a…

计算机视觉与模式识别 · 计算机科学 2019-06-11 Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , Jitendra Malik

Co-speech gesture generation enhances human-computer interaction realism through speech-synchronized gesture synthesis. However, generating semantically meaningful gestures remains a challenging problem. We propose SARGes, a novel framework…

计算与语言 · 计算机科学 2025-03-27 Nan Gao , Yihua Bao , Dongdong Weng , Jiayi Zhao , Jia Li , Yan Zhou , Pengfei Wan , Di Zhang

Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjusting the style while…

音频与语音处理 · 电气工程与系统科学 2025-05-01 Hanzhao Li , Yuke Li , Xinsheng Wang , Jingbin Hu , Qicong Xie , Shan Yang , Lei Xie

Gesture behavior is a natural part of human conversation. Much work has focused on removing the need for tedious hand-animation to create embodied conversational agents by designing speech-driven gesture generators. However, these…

人机交互 · 计算机科学 2020-10-05 Ylva Ferstl , Michael Neff , Rachel McDonnell

Co-speech gesture generation is a critical area of research aimed at synthesizing speech-synchronized human-like gestures. Existing methods often suffer from issues such as rhythmic inconsistency, motion jitter, foot sliding and limited…

声音 · 计算机科学 2026-01-09 Yujiao Jiang , Qingmin Liao , Zongqing Lu

This study explores two frameworks for co-speech gesture generation, AQ-GT and its semantically-augmented variant AQ-GT-a, to evaluate their ability to convey meaning through gestures and how humans perceive the resulting movements. Using…

人机交互 · 计算机科学 2025-10-21 Hendric Voss , Lisa Michelle Bohnenkamp , Stefan Kopp

In recent years because of the advances in computer vision research, free hand gestures have been explored as means of human-computer interaction (HCI). Together with improved speech processing technology it is an important step toward…

计算机视觉与模式识别 · 计算机科学 2007-05-23 S. Kettebekov , R. Sharma

Automatic gesture generation from speech generally relies on implicit modelling of the nondeterministic speech-gesture relationship and can result in averaged motion lacking defined form. Here, we propose a database-driven approach of…

人机交互 · 计算机科学 2021-03-05 Ylva Ferstl , Michael Neff , Rachel McDonnell

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Jiaxin Ye , Hongming Shan

Speech-driven gesture synthesis is a field of growing interest in virtual human creation. However, a critical challenge is the inherent intricate one-to-many mapping between speech and gestures. Previous studies have explored and achieved…

图形学 · 计算机科学 2023-02-03 Fan Zhang , Naye Ji , Fuxing Gao , Yongping Li

We present Social Agent, a novel framework for synthesizing realistic and contextually appropriate co-speech nonverbal behaviors in dyadic conversations. In this framework, we develop an agentic system driven by a Large Language Model (LLM)…

图形学 · 计算机科学 2025-10-07 Zeyi Zhang , Yanju Zhou , Heyuan Yao , Tenglong Ao , Xiaohang Zhan , Libin Liu

This paper describes a system developed for the GENEA (Generation and Evaluation of Non-verbal Behaviour for Embodied Agents) Challenge 2023. Our solution builds on an existing diffusion-based motion synthesis model. We propose a…

音频与语音处理 · 电气工程与系统科学 2023-09-12 Anna Deichler , Shivam Mehta , Simon Alexanderson , Jonas Beskow

We introduce a real-time, human-in-the-loop gesture control framework that can dynamically adapt audio and music based on human movement by analyzing live video input. By creating a responsive connection between visual and auditory stimuli,…

人机交互 · 计算机科学 2025-04-29 Mahya Khazaei , Ali Bahrani , George Tzanetakis

Recent advances in interactive technologies have highlighted the prominence of audio signals for semantic encoding. This paper explores a new task, where audio signals are used as conditioning inputs to generate motions that align with the…

声音 · 计算机科学 2025-05-30 Zi-An Wang , Shihao Zou , Shiyao Yu , Mingyuan Zhang , Chao Dong

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to…

音频与语音处理 · 电气工程与系统科学 2024-12-17 Leying Zhang , Yao Qian , Long Zhou , Shujie Liu , Dongmei Wang , Xiaofei Wang , Midia Yousefi , Yanmin Qian , Jinyu Li , Lei He , Sheng Zhao , Michael Zeng

Co-speech gesture video generation aims to synthesize realistic, audio-aligned videos of speakers, complete with synchronized facial expressions and body gestures. This task presents challenges due to the significant one-to-many mapping…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Xu Yang , Shaoli Huang , Shenbo Xie , Xuelin Chen , Yifei Liu , Changxing Ding

Lip-to-speech involves generating a natural-sounding speech synchronized with a soundless video of a person talking. Despite recent advances, current methods still cannot produce high-quality speech with high levels of intelligibility for…

音频与语音处理 · 电气工程与系统科学 2024-03-29 Yochai Yemini , Aviv Shamsian , Lior Bracha , Sharon Gannot , Ethan Fetaya

Textless spoken language models (SLMs) are generative models of speech that do not rely on text supervision. Most textless SLMs learn to predict the next semantic token, a discrete representation of linguistic content, and rely on a…

计算与语言 · 计算机科学 2025-10-23 Ju-Chieh Chou , Jiawei Zhou , Karen Livescu

In this paper we introduce a new synchronisation task, Gesture-Sync: determining if a person's gestures are correlated with their speech or not. In comparison to Lip-Sync, Gesture-Sync is far more challenging as there is a far looser…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Sindhu B Hegde , Andrew Zisserman

Embodied conversational agents benefit from being able to accompany their speech with gestures. Although many data-driven approaches to gesture generation have been proposed in recent years, it is still unclear whether such systems can…

人机交互 · 计算机科学 2022-01-17 Taras Kucherenko , Rajmund Nagy , Michael Neff , Hedvig Kjellström , Gustav Eje Henter