English
Related papers

Related papers: GestureLSM: Latent Shortcut based Co-Speech Gestur…

200 papers

Deriving co-speech 3D gestures has seen tremendous progress in virtual avatar animation. Yet, the existing methods often produce stiff and unreasonable gestures with unseen human speech inputs due to the limited 3D speech-gesture data. In…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Xingqun Qi , Hengyuan Zhang , Yatian Wang , Jiahao Pan , Chen Liu , Peng Li , Xiaowei Chi , Mengfei Li , Wei Xue , Shanghang Zhang , Wenhan Luo , Qifeng Liu , Yike Guo

Unconstrained lip-to-speech synthesis aims to generate corresponding speeches from silent videos of talking faces with no restriction on head poses or vocabulary. Current works mainly use sequence-to-sequence models to solve this problem,…

Sound · Computer Science 2022-07-14 Yongqi Wang , Zhou Zhao

Spatio-temporal coherency is a major challenge in synthesizing high quality videos, particularly in synthesizing human videos that contain rich global and local deformations. To resolve this challenge, previous approaches have resorted to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Yaohui Wang , Xin Ma , Xinyuan Chen , Cunjian Chen , Antitza Dantcheva , Bo Dai , Yu Qiao

People naturally conduct spontaneous body motions to enhance their speeches while giving talks. Body motion generation from speech is inherently difficult due to the non-deterministic mapping from speech to body motions. Most existing works…

Computer Vision and Pattern Recognition · Computer Science 2022-03-07 Jing Xu , Wei Zhang , Yalong Bai , Qibin Sun , Tao Mei

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling due to its iterative…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Taekyung Ki , Dongchan Min , Gyeongsu Chae

Co-speech gesture video generation aims to synthesize realistic, audio-aligned videos of speakers, complete with synchronized facial expressions and body gestures. This task presents challenges due to the significant one-to-many mapping…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Xu Yang , Shaoli Huang , Shenbo Xie , Xuelin Chen , Yifei Liu , Changxing Ding

The emergence of large language models (LLMs) has sparked significant interest in extending their remarkable language capabilities to speech. However, modality alignment between speech and text still remains an open problem. Current…

Computation and Language · Computer Science 2024-05-29 Chen Wang , Minpeng Liao , Zhongqiang Huang , Jinliang Lu , Junhong Wu , Yuchen Liu , Chengqing Zong , Jiajun Zhang

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Zexu Pan , Xinyuan Qian , Shengkui Zhao , Kun Zhou , Bin Ma

Recently, the text-to-3D task has developed rapidly due to the appearance of the SDS method. However, the SDS method always generates 3D objects with poor quality due to the over-smooth issue. This issue is attributed to two factors: 1) the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Yiming Zhong , Xiaolin Zhang , Yao Zhao , Yunchao Wei

In modern society, people should not be identified based on their disability, rather, it is environments that can disable people with impairments. Improvements to automatic Sign Language Recognition (SLR) will lead to more enabling…

Computer Vision and Pattern Recognition · Computer Science 2022-02-23 Jordan J. Bird

Human motion stylization aims to revise the style of an input motion while keeping its content unaltered. Unlike existing works that operate directly in pose space, we leverage the latent space of pretrained autoencoders as a more…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Chuan Guo , Yuxuan Mu , Xinxin Zuo , Peng Dai , Youliang Yan , Juwei Lu , Li Cheng

Continuous sign language recognition (CSLR) requires precise spatio-temporal modeling to accurately recognize sequences of gestures in videos. Existing frameworks often rely on CNN-based spatial backbones combined with temporal convolution…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Ahmed Abul Hasanaath , Hamzah Luqman

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

Hand gesture recognition using multichannel surface electromyography (sEMG) is challenging due to unstable predictions and inefficient time-varying feature enhancement. To overcome the lack of signal based time-varying feature problems, we…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Jungpil Shin , Abu Saleh Musa Miah , Sota Konnai , Shu Hoshitaka , Pankoo Kim

Recent advancements in large language models (LLMs) have led to significant progress in text-based dialogue systems. These systems can now generate high-quality responses that are accurate and coherent across a wide range of topics and…

Computation and Language · Computer Science 2025-01-10 Long Mai , Julie Carson-Berndsen

Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Peng Jin , Hao Li , Zesen Cheng , Kehan Li , Runyi Yu , Chang Liu , Xiangyang Ji , Li Yuan , Jie Chen

Cued Speech (CS) is an advanced visual phonetic encoding system that integrates lip reading with hand codings, enabling people with hearing impairments to communicate efficiently. CS video generation aims to produce specific lip and gesture…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Wentao Lei , Li Liu , Jun Wang

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, existing VLA systems primarily rely on textual instructions and struggle to resolve spatial…

Robotics · Computer Science 2026-05-22 Wenxuan Guo , Ziyuan Li , Meng Zhang , Yichen Liu , Yimeng Dong , Chuxi Xu , Yunfei Wei , Ze Chen , Erjin Zhou , Jianjiang Feng

While the community keeps promoting end-to-end models over conventional hybrid models, which usually are long short-term memory (LSTM) models trained with a cross entropy criterion followed by a sequence discriminative training criterion,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-18 Jinyu Li , Rui Zhao , Eric Sun , Jeremy H. M. Wong , Amit Das , Zhong Meng , Yifan Gong

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Qi Wu , Yubo Zhao , Yifan Wang , Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang