English
Related papers

Related papers: SIGGesture: Generalized Co-Speech Gesture Synthesi…

200 papers

Scene text recognition (STR) suffers from challenges of either less realistic synthetic training data or the difficulty of collecting sufficient high-quality real-world data, limiting the effectiveness of trained models. Meanwhile, despite…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Xingsong Ye , Yongkun Du , Yunbo Tao , Zhineng Chen

Gestures are pivotal in enhancing co-speech communication. While recent works have mostly focused on point-level motion transformation or fully supervised motion representations through data-driven approaches, we explore the representation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Huan Yang , Jiahui Chen , Chaofan Ding , Runhua Shi , Siyu Xiong , Qingqi Hong , Xiaoqi Mo , Xinhan Di

Text-to-motion generation has gained increasing attention, but most existing methods are limited to generating short-term motions that correspond to a single sentence describing a single action. However, when a text stream describes a…

Computer Vision and Pattern Recognition · Computer Science 2023-08-04 Zhao Yang , Bing Su , Ji-Rong Wen

Diffusion models have become central to various image editing tasks, yet they often fail to fully adhere to physical laws, particularly with effects like shadows, reflections, and occlusions. In this work, we address the challenge of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-23 Ankit Dhiman , Manan Shah , R Venkatesh Babu

Recent successes in image synthesis are powered by large-scale diffusion models. However, most methods are currently limited to either text- or image-conditioned generation for synthesizing an entire image, texture transfer or inserting…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Yufei Ye , Xueting Li , Abhinav Gupta , Shalini De Mello , Stan Birchfield , Jiaming Song , Shubham Tulsiani , Sifei Liu

We present a generative adversarial network to synthesize 3D pose sequences of co-speech upper-body gestures with appropriate affective expressions. Our network consists of two components: a generator to synthesize gestures from a joint…

Multimedia · Computer Science 2024-11-26 Uttaran Bhattacharya , Elizabeth Childs , Nicholas Rewkowski , Dinesh Manocha

Text-guided image generation has advanced rapidly with large-scale diffusion models, yet achieving precise stylization with visual exemplars remains difficult. Existing approaches often depend on task-specific retraining or expensive…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Yingying Deng , Xiangyu He , Fan Tang , Weiming Dong , Xucheng Yin

Speech-driven 3D facial animation synthesis has been a challenging task both in industry and research. Recent methods mostly focus on deterministic deep learning methods meaning that given a speech input, the output is always the same.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Stefan Stan , Kazi Injamamul Haque , Zerrin Yumak

Generating full-body human gestures based on speech signals remains challenges on quality and speed. Existing approaches model different body regions such as body, legs and hands separately, which fail to capture the spatial interactions…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Pinxin Liu , Luchuan Song , Junhua Huang , Haiyang Liu , Chenliang Xu

We present TexFusion (Texture Diffusion), a new method to synthesize textures for given 3D geometries, using large-scale text-guided image diffusion models. In contrast to recent works that leverage 2D text-to-image diffusion models to…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Tianshi Cao , Karsten Kreis , Sanja Fidler , Nicholas Sharp , Kangxue Yin

We introduce DiffuseST, a low-latency, direct speech-to-speech translation system capable of preserving the input speaker's voice zero-shot while translating from multiple source languages into English. We experiment with the synthesizer…

Generative AI has demonstrated impressive performance in various fields, among which speech synthesis is an interesting direction. With the diffusion model as the most popular generative model, numerous works have attempted two active…

We propose a real-time system for synthesizing gestures directly from speech. Our data-driven approach is based on Generative Adversarial Neural Networks to model the speech-gesture relationship. We utilize the large amount of speaker video…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Manuel Rebol , Christian Gütl , Krzysztof Pietroszek

Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing models typically specialize either in descriptive motion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Anna Deichler , Jim O'Regan , Teo Guichoux , David Johansson , Jonas Beskow

Co-speech gestures convey a wide variety of meanings and play an important role in face-to-face human interactions. These gestures significantly influence the addressee's engagement, recall, comprehension, and attitudes toward the speaker.…

Human-Computer Interaction · Computer Science 2025-03-19 Parisa Ghanad Torshizi , Laura B. Hensel , Ari Shapiro , Stacy C. Marsella

Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic full-body motion based on text prompts, such as talking while…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Bohong Chen , Yumeng Li , Yao-Xiang Ding , Tianjia Shao , Kun Zhou

Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Peng Jin , Hao Li , Zesen Cheng , Kehan Li , Runyi Yu , Chang Liu , Xiangyang Ji , Li Yuan , Jie Chen

For human-like agents, including virtual avatars and social robots, making proper gestures while speaking is crucial in human--agent interaction. Co-speech gestures enhance interaction experiences and make the agents look alive. However, it…

Graphics · Computer Science 2020-09-07 Youngwoo Yoon , Bok Cha , Joo-Haeng Lee , Minsu Jang , Jaeyeon Lee , Jaehong Kim , Geehyuk Lee

Building generic robotic manipulation systems often requires large amounts of real-world data, which can be dificult to collect. Synthetic data generation offers a promising alternative, but limiting the sim-to-real gap requires significant…

Robotics · Computer Science 2024-11-18 Thomas Lips , Francis wyffels

Semantic communication is expected to be one of the cores of next-generation AI-based communications. One of the possibilities offered by semantic communication is the capability to regenerate, at the destination side, images or videos…

Artificial Intelligence · Computer Science 2026-05-18 Eleonora Grassucci , Sergio Barbarossa , Danilo Comminiello
‹ Prev 1 3 4 5 6 7 10 Next ›