中文
相关论文

相关论文: LiveGesture Streamable Co-Speech Gesture Generatio…

200 篇论文

Audio-driven talking head generation aims to create vivid and realistic videos from a static portrait and speech. Existing AR-based methods rely on intermediate facial representations, which limit their expressiveness and realism.…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Yuzhe Weng , Haotian Wang , Yuanhong Yu , Jun Du , Shan He , Xiaoyan Wu , Haoran Xu

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Jiaxin Ye , Hongming Shan

Masked-based autoregressive models have demonstrated promising image generation capability in continuous space. However, their potential for video generation remains under-explored. In this paper, we propose \textbf{VideoMAR}, a concise and…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Hu Yu , Biao Gong , Hangjie Yuan , DanDan Zheng , Weilong Chai , Jingdong Chen , Kecheng Zheng , Feng Zhao

Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking the vital role of…

图形学 · 计算机科学 2025-05-09 Jinhe Huang , Yongkang Cheng , Yuming Hang , Gaoge Han , Jinewei Li , Jing Zhang , Xingjian Gu

People naturally conduct spontaneous body motions to enhance their speeches while giving talks. Body motion generation from speech is inherently difficult due to the non-deterministic mapping from speech to body motions. Most existing works…

计算机视觉与模式识别 · 计算机科学 2022-03-07 Jing Xu , Wei Zhang , Yalong Bai , Qibin Sun , Tao Mei

Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing part-decomposed or…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Xuanmeng Sha , Liyun Zhang , Tomohiro Mashita , Naoya Chiba , Yuki Uranishi

Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic full-body motion based on text prompts, such as talking while…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Bohong Chen , Yumeng Li , Yao-Xiang Ding , Tianjia Shao , Kun Zhou

Text-to-motion generation has advanced rapidly, yet two challenges persist. First, existing motion autoencoders compress each frame into a single monolithic latent vector, entangling trajectory and per-joint rotations in an unstructured…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Zeyu Ling , Qing Shuai , Teng Zhang , Shiyang Li , Bo Han , Changqing Zou

We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Sicheng Xu , Guojun Chen , Yu-Xiao Guo , Jiaolong Yang , Chong Li , Zhenyu Zang , Yizhong Zhang , Xin Tong , Baining Guo

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in virtual interaction…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Xi Liu , Ying Guo , Cheng Zhen , Tong Li , Yingying Ao , Pengfei Yan

A good co-speech motion generation cannot be achieved without a careful integration of common rhythmic motion and rare yet essential semantic motion. In this work, we propose SemTalk for holistic co-speech motion generation with frame-level…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Xiangyue Zhang , Jianfang Li , Jiaxu Zhang , Ziqiang Dang , Jianqiang Ren , Liefeng Bo , Zhigang Tu

Gaze and head movements play a central role in expressive 3D media, human-agent interaction, and immersive communication. Existing works often model facial components in isolation and lack mechanisms for generating personalized, style-aware…

图形学 · 计算机科学 2026-01-05 Chengwei Shi , Chong Cao

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to…

音频与语音处理 · 电气工程与系统科学 2024-12-17 Leying Zhang , Yao Qian , Long Zhou , Shujie Liu , Dongmei Wang , Xiaofei Wang , Midia Yousefi , Yanmin Qian , Jinyu Li , Lei He , Sheng Zhao , Michael Zeng

We present TANGO, a framework for generating co-speech body-gesture videos. Given a few-minute, single-speaker reference video and target speech audio, TANGO produces high-fidelity videos with synchronized body gestures. TANGO builds on…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Haiyang Liu , Xingchao Yang , Tomoya Akiyama , Yuantian Huang , Qiaoge Li , Shigeru Kuriyama , Takafumi Taketomi

Text-to-motion generation is an emerging and challenging problem, which aims to synthesize motion with the same semantics as the input text. However, due to the lack of diverse labeled training data, most approaches either limit to specific…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Junfan Lin , Jianlong Chang , Lingbo Liu , Guanbin Li , Liang Lin , Qi Tian , Chang Wen Chen

Generative models are reshaping the live-streaming industry by redefining how content is created, styled, and delivered. Previous image-based streaming diffusion models have powered efficient and creative live streaming products but have…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Tianrui Feng , Zhi Li , Shuo Yang , Haocheng Xi , Muyang Li , Xiuyu Li , Lvmin Zhang , Keting Yang , Kelly Peng , Song Han , Maneesh Agrawala , Kurt Keutzer , Akio Kodaira , Chenfeng Xu

For realistic talking head generation, creating natural head motion while maintaining accurate lip synchronization is essential. To fulfill this challenging task, we propose DisCoHead, a novel method to disentangle and control head pose and…

计算机视觉与模式识别 · 计算机科学 2023-03-15 Geumbyeol Hwang , Sunwon Hong , Seunghyun Lee , Sungwoo Park , Gyeongsu Chae

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic videos are disregarded.…

计算机视觉与模式识别 · 计算机科学 2023-06-16 Jiawei Liu , Weining Wang , Sihan Chen , Xinxin Zhu , Jing Liu

Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite their photorealistic rendering capability, still frequently…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Xiangyang Luo , Xiaozhe Xin , Tao Feng , Xu Guo , Meiguang Jin , Junfeng Ma

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Yue Gao , Yuan Zhou , Jinglu Wang , Xiao Li , Xiang Ming , Yan Lu