English
Related papers

Related papers: Towards Streaming Speech-to-Avatar Synthesis

200 papers

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Mengyi Shan , Shouchieh Chang , Ziqian Bai , Shichen Liu , Yinda Zhang , Luchuan Song , Rohit Pandey , Sean Fanello , Zeng Huang

Most organisms including humans function by coordinating and integrating sensory signals with motor actions to survive and accomplish desired tasks. Learning these complex sensorimotor mappings proceeds simultaneously and often in an…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-26 Yashish M. Siriwardena , Carol Espy-Wilson , Shihab Shamma

Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven…

Image and Video Processing · Electrical Eng. & Systems 2024-09-25 Hong Nguyen , Sean Foley , Kevin Huang , Xuan Shi , Tiantian Feng , Shrikanth Narayanan

Text-to-speech synthesis (TTS) has witnessed rapid progress in recent years, where neural methods became capable of producing audios with high naturalness. However, these efforts still suffer from two types of latencies: (a) the {\em…

Computation and Language · Computer Science 2020-10-08 Mingbo Ma , Baigong Zheng , Kaibo Liu , Renjie Zheng , Hairong Liu , Kainan Peng , Kenneth Church , Liang Huang

We propose a method for synthesizing edited photo-realistic digital avatars with text instructions. Given a short monocular RGB video and text instructions, our method uses an image-conditioned diffusion model to edit one head image and…

Computer Vision and Pattern Recognition · Computer Science 2023-06-06 Shaoxu Li

Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital…

Sound · Computer Science 2025-10-15 Tianbao Zhang , Jian Zhao , Yuer Li , Zheng Zhu , Ping Hu , Zhaoxin Fan , Wenjun Wu , Xuelong Li

Video-to-video synthesis is a challenging problem aiming at learning a translation function between a sequence of semantic maps and a photo-realistic video depicting the characteristics of a driving video. We propose a head-to-head system…

Computer Vision and Pattern Recognition · Computer Science 2020-06-19 Mohammad Rami Koujan , Michail Christos Doukas , Anastasios Roussos , Stefanos Zafeiriou

The goal of this work is zero-shot text-to-speech synthesis, with speaking styles and voices learnt from facial characteristics. Inspired by the natural fact that people can imagine the voice of someone when they look at his or her face, we…

Machine Learning · Computer Science 2023-02-28 Jiyoung Lee , Joon Son Chung , Soo-Whan Chung

Existing end-to-end sign-language animation systems suffer from low naturalness, limited facial/body expressivity, and no user control. We propose a human-centered, real-time speech-to-sign animation framework that integrates (1) a…

Human-Computer Interaction · Computer Science 2025-06-25 Yingchao Li

Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current methods still face challenges in maintaining lip-sync…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Bin Liu , Zhixiang Xiong , Zhifen He , Bo Li

Acoustic-to-articulatory inversion (AAI) is to convert audio into articulator movements, such as ultrasound tongue imaging (UTI) data. An issue of existing AAI methods is only using the personalized acoustic information to derive the…

Sound · Computer Science 2024-03-13 Yudong Yang , Rongfeng Su , Xiaokang Liu , Nan Yan , Lan Wang

Controllability, generalizability and efficiency are the major objectives of constructing face avatars represented by neural implicit field. However, existing methods have not managed to accommodate the three requirements simultaneously.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Zhiyuan Ma , Xiangyu Zhu , Guojun Qi , Zhen Lei , Lei Zhang

The capacity to create realistic virtual humans has progressed significantly, and such characters can be found in many applications across entertainment, education and health. As an essential element of interactive virtual humans,…

Graphics · Computer Science 2026-05-12 Haoyang Du , Yinghan Xu , John Dingliana , Brian Keegan , Rachel McDonnell , Cathy Ennis

Voice impersonation is not the same as voice transformation, although the latter is an essential element of it. In voice impersonation, the resultant voice must convincingly convey the impression of having been naturally produced by the…

Sound · Computer Science 2018-02-21 Yang Gao , Rita Singh , Bhiksha Raj

Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio-visual diffusion models remain too slow for interactive use…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Chunyu Li , Jiaye Li , Ruiqiao Mei , Haoyuan Xia , Hao Zhu , Jingdong Wang , Siyu Zhu

Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. This approach circumvents delays and cascading errors associated with model cascading.…

Acoustic-to-articulatory inversion (AAI) is to obtain the movement of articulators from speech signals. Until now, achieving a speaker-independent AAI remains a challenge given the limited data. Besides, most current works only use audio…

Sound · Computer Science 2022-04-05 Jianrong Wang , Jinyu Liu , Longxuan Zhao , Shanyu Wang , Ruiguo Yu , Li Liu

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and fluidity. They also…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Qijun Gan , Ruizi Yang , Jianke Zhu , Shaofei Xue , Steven Hoi

We propose a method for synthesizing photo-realistic digital avatars from only one portrait as the reference. Given a portrait, our method synthesizes a coarse talking head video using driving keypoints features. And with the coarse video,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Shaoxu Li

Talking face generation aims to synthesize realistic speaking portraits from a single image, yet existing methods often rely on explicit optical flow and local warping, which fail to model complex global motions and cause identity drift. We…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Bo Chen , Tao Liu , Qi Chen , Xie Chen , Zilong Zheng
‹ Prev 1 3 4 5 6 7 10 Next ›