English
Related papers

Related papers: CustomListener: Text-guided Responsive Interaction…

200 papers

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Balamurugan Thambiraja , Ikhsanul Habibie , Sadegh Aliakbarian , Darren Cosker , Christian Theobalt , Justus Thies

Generating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for training or fine-tuning. Existing person-generic methods have…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Weizhi Zhong , Chaowei Fang , Yinqi Cai , Pengxu Wei , Gangming Zhao , Liang Lin , Guanbin Li

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Ziqi Zhou , Weize Quan , Hailin Shi , Wei Li , Lili Wang , Dong-Ming Yan

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

Talking face generation aims to synthesize realistic speaking portraits from a single image, yet existing methods often rely on explicit optical flow and local warping, which fail to model complex global motions and cause identity drift. We…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Bo Chen , Tao Liu , Qi Chen , Xie Chen , Zilong Zheng

Cross-modality generation is an emerging topic that aims to synthesize data in one modality based on information in a different modality. In this paper, we consider a task of such: given an arbitrary audio speech and one lip image of…

Computer Vision and Pattern Recognition · Computer Science 2018-05-23 Lele Chen , Zhiheng Li , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

Generating vivid and diverse 3D co-speech gestures is crucial for various applications in animating virtual avatars. While most existing methods can generate gestures from audio directly, they usually overlook that emotion is one of the key…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Xingqun Qi , Chen Liu , Lincheng Li , Jie Hou , Haoran Xin , Xin Yu

The creation of listener facial responses aims to simulate interactive communication feedback from a listener during a face-to-face conversation. Our goal is to generate believable videos of listeners' heads that respond authentically to a…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Minh-Duc Nguyen , Hyung-Jeong Yang , Seung-Won Kim , Ji-Eun Shin , Soo-Hyung Kim

Audio-driven talking head generation faces a fundamental trade-off between personalization and generalization, limiting its practical application. Implicit models often achieve generalization at the cost of structural incoherence, resulting…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Shiyu Liu , Kui Jiang , Junjun Jiang , Xianming Liu , Xiaocheng Feng , Hongxun Yao , Qi Tian

In recent years, Text-to-Audio Generation has achieved remarkable progress, offering sound creators powerful tools to transform textual inspirations into vivid audio. However, existing models predominantly operate directly in the acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Zheqi Dai , Guangyan Zhang , Haolin He , Xiquan Li , Jingyu Li , Chunyat Wu , Yiwen Guo , Qiuqiang Kong

The generation of personalized dialogue is vital to natural and human-like conversation. Typically, personalized dialogue generation models involve conditioning the generated response on the dialogue history and a representation of the…

Computation and Language · Computer Science 2021-11-23 Jing Yang Lee , Kong Aik Lee , Woon Seng Gan

Recent advances have demonstrated compelling capabilities in synthesizing real individuals into generated videos, reflecting the growing demand for identity-aware content creation. Nevertheless, an openly accessible framework enabling…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Yingjie Chen , Shilun Lin , Cai Xing , Binxin Yang , Long Zhou , Qixin Yan , Wenjing Wang , Dingming Liu , Hao Liu , Chen Li , Jing Lyu

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Chao Xu , Junwei Zhu , Jiangning Zhang , Yue Han , Wenqing Chu , Ying Tai , Chengjie Wang , Zhifeng Xie , Yong Liu

Controllable generation using StyleGANs is usually achieved by training the model using labeled data. For audio textures, however, there is currently a lack of large semantically labeled datasets. Therefore, to control generation, we…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Purnima Kamath , Chitralekha Gupta , Lonce Wyse , Suranga Nanayakkara

Text-to-motion (T2M) generation is becoming a practical tool for animation and interactive avatars. However, modifying specific body parts while maintaining overall motion coherence remains challenging. Existing methods typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Minyue Dai , Ke Fan , Anyi Rao , Jingbo Wang , Bo Dai

Co-speech gesture generation is to synthesize a gesture sequence that not only looks real but also matches with the input speech audio. Our method generates the movements of a complete upper body, including arms, hands, and the head.…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Shenhan Qian , Zhi Tu , Yihao Zhi , Wen Liu , Shenghua Gao

All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise to the utterance or…

Computer Vision and Pattern Recognition · Computer Science 2019-10-03 Gaurav Mittal , Baoyuan Wang

This paper presents an innovative approach to enhance control over audio generation by emphasizing the alignment between audio and text representations during model training. In the context of language model-based audio generation, the…

In this paper, we abstract the process of people hearing speech, extracting meaningful cues, and creating various dynamically audio-consistent talking faces, termed Listening and Imagining, into the task of high-fidelity diverse talking…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Chao Xu , Yang Liu , Jiazheng Xing , Weida Wang , Mingze Sun , Jun Dan , Tianxin Huang , Siyuan Li , Zhi-Qi Cheng , Ying Tai , Baigui Sun

Customized video generation aims to generate high-quality videos guided by text prompts and subject's reference images. However, since it is only trained on static images, the fine-tuning process of subject learning disrupts abilities of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Tao Wu , Yong Zhang , Xintao Wang , Xianpan Zhou , Guangcong Zheng , Zhongang Qi , Ying Shan , Xi Li