English
Related papers

Related papers: DiffusionTalker: Efficient and Compact Speech-Driv…

200 papers

By leveraging the text-to-image diffusion priors, score distillation can synthesize 3D contents without paired text-3D training data. Instead of spending hours of online optimization per text prompt, recent studies have been focused on…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Zhiyuan Ma , Yuxiang Wei , Yabin Zhang , Xiangyu Zhu , Zhen Lei , Lei Zhang

Diffusion distillation models effectively accelerate reverse sampling by compressing the process into fewer steps. However, these models still exhibit a performance gap compared to their pre-trained diffusion model counterparts, exacerbated…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Geon Yeong Park , Sang Wan Lee , Jong Chul Ye

This paper presents FaceXHuBERT, a text-less speech-driven 3D facial animation generation method that allows to capture personalized and subtle cues in speech (e.g. identity, emotion and hesitation). It is also very robust to background…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Kazi Injamamul Haque , Zerrin Yumak

The art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the…

Human-Computer Interaction · Computer Science 2023-05-09 Sicheng Yang , Zhiyong Wu , Minglei Li , Zhensong Zhang , Lei Hao , Weihong Bao , Ming Cheng , Long Xiao

3D head stylization transforms realistic facial features into artistic representations, enhancing user engagement across gaming and virtual reality applications. While 3D-aware generators have made significant advancements, many 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Bahri Batuhan Bilecen , Ahmet Berke Gokmen , Furkan Guzelant , Aysegul Dundar

Representation knowledge distillation aims at transferring rich information from one model to another. Common approaches for representation distillation mainly focus on the direct minimization of distance metrics between the models'…

Computer Vision and Pattern Recognition · Computer Science 2022-04-06 Emanuel Ben-Baruch , Matan Karklinsky , Yossi Biton , Avi Ben-Cohen , Hussam Lawen , Nadav Zamir

Recent deep learning models demand larger datasets, driving the need for dataset distillation to create compact, cost-efficient datasets while maintaining performance. Due to the powerful image generation capability of diffusion, it has…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Lin Zhao , Yushu Wu , Xinru Jiang , Jianyang Gu , Yanzhi Wang , Xiaolin Xu , Pu Zhao , Xue Lin

Leveraging Stable Diffusion for the generation of personalized portraits has emerged as a powerful and noteworthy tool, enabling users to create high-fidelity, custom character avatars based on their specific prompts. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Siying Cui , Jia Guo , Xiang An , Jiankang Deng , Yongle Zhao , Xinyu Wei , Ziyong Feng

Audio-driven talking head generation is crucial for applications in virtual reality, digital avatars, and film production. While NeRF-based methods enable high-fidelity reconstruction, they suffer from low rendering efficiency and…

Sound · Computer Science 2025-09-23 Tianheng Zhu , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng

We present a novel approach for synthesizing 3D talking heads with controllable emotion, featuring enhanced lip synchronization and rendering quality. Despite significant progress in the field, prior methods still suffer from multi-view…

Computer Vision and Pattern Recognition · Computer Science 2024-08-02 Qianyun He , Xinya Ji , Yicheng Gong , Yuanxun Lu , Zhengyu Diao , Linjia Huang , Yao Yao , Siyu Zhu , Zhan Ma , Songcen Xu , Xiaofei Wu , Zixiao Zhang , Xun Cao , Hao Zhu

Recently, multi-person video generation has started to gain prominence. While a few preliminary works have explored audio-driven multi-person talking video generation, they often face challenges due to the high costs of diverse multi-person…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zhizhou Zhong , Yicheng Ji , Zhe Kong , Yiying Liu , Jiarui Wang , Jiasun Feng , Lupeng Liu , Xiangyi Wang , Yanjia Li , Yuqing She , Ying Qin , Huan Li , Shuiyang Mao , Wei Liu , Wenhan Luo

We introduce FactorPortrait, a video diffusion method for controllable portrait animation that enables lifelike synthesis from disentangled control signals of facial expressions, head movement, and camera viewpoints. Given a single portrait…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Jiapeng Tang , Kai Li , Chengxiang Yin , Liuhao Ge , Fei Jiang , Jiu Xu , Matthias Nießner , Christian Häne , Timur Bagautdinov , Egor Zakharov , Peihong Guo

Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Yuchi Wang , Junliang Guo , Jianhong Bai , Runyi Yu , Tianyu He , Xu Tan , Xu Sun , Jiang Bian

This paper presents a simple yet effective framework MaskCLIP, which incorporates a newly proposed masked self-distillation into contrastive language-image pretraining. The core idea of masked self-distillation is to distill representation…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Xiaoyi Dong , Jianmin Bao , Yinglin Zheng , Ting Zhang , Dongdong Chen , Hao Yang , Ming Zeng , Weiming Zhang , Lu Yuan , Dong Chen , Fang Wen , Nenghai Yu

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

Lipreading has witnessed a lot of progress due to the resurgence of neural networks. Recent works have placed emphasis on aspects such as improving performance by finding the optimal architecture or improving generalization. However, there…

Computer Vision and Pattern Recognition · Computer Science 2021-06-03 Pingchuan Ma , Brais Martinez , Stavros Petridis , Maja Pantic

Recent advancements in human image animation have been propelled by video diffusion models, yet their reliance on numerous iterative denoising steps results in high inference costs and slow speeds. An intuitive solution involves adopting…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Xiang Wang , Shiwei Zhang , Hangjie Yuan , Yujie Wei , Yingya Zhang , Changxin Gao , Yuehuan Wang , Nong Sang

Voice interfaces integral to the human-computer interaction systems can benefit from speech emotion recognition (SER) to customize responses based on user emotions. Since humans convey emotions through multi-modal audio-visual cues,…

Machine Learning · Computer Science 2025-07-02 Varsha Pendyala , Pedro Morgado , William Sethares

Knowledge distillation in neural networks refers to compressing a large model or dataset into a smaller version of itself. We introduce Privacy Distillation, a framework that allows a text-to-image generative model to teach another model…

Animating virtual avatars to make co-speech gestures facilitates various applications in human-machine interaction. The existing methods mainly rely on generative adversarial networks (GANs), which typically suffer from notorious mode…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Lingting Zhu , Xian Liu , Xuanyu Liu , Rui Qian , Ziwei Liu , Lequan Yu