English
Related papers

Related papers: DiffusionTalker: Efficient and Compact Speech-Driv…

200 papers

Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model based on a conditional…

Sound · Computer Science 2024-12-05 Jiaxuan Liu , Zhaoci Liu , Yajun Hu , Yingying Gao , Shilei Zhang , Zhenhua Ling

Diffusion Language Models (DLMs) have recently achieved strong results in text generation. However, their multi-step sampling leads to slow inference, limiting practical use. To address this, we extend Inverse Distillation, a technique…

Machine Learning · Computer Science 2026-02-24 David Li , Nikita Gushchin , Dmitry Abulkhanov , Eric Moulines , Ivan Oseledets , Maxim Panov , Alexander Korotin

Portrait animation has witnessed tremendous quality improvements thanks to recent advances in video diffusion models. However, these 2D methods often compromise 3D consistency and speed, limiting their applicability in real-world scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Kaiwen Jiang , Xueting Li , Seonwook Park , Ravi Ramamoorthi , Shalini De Mello , Koki Nagano

Training machine learning models on massive datasets is expensive and time-consuming. Dataset distillation addresses this by creating a small synthetic dataset that achieves the same performance as the full dataset. Recent methods use…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Jeffrey A. Chan-Santiago , Mubarak Shah

Knowledge distillation methods have recently shown to be a promising direction to speedup the synthesis of large-scale diffusion models by requiring only a few inference steps. While several powerful distillation methods were recently…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Nikita Starodubcev , Artem Fedorov , Artem Babenko , Dmitry Baranchuk

Semantic correspondence, the task of determining relationships between different parts of images, underpins various applications including 3D reconstruction, image-to-image translation, object tracking, and visual place recognition. Recent…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Frank Fundel , Johannes Schusterbauer , Vincent Tao Hu , Björn Ommer

Talking face generation aims to synthesize realistic speaking portraits from a single image, yet existing methods often rely on explicit optical flow and local warping, which fail to model complex global motions and cause identity drift. We…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Bo Chen , Tao Liu , Qi Chen , Xie Chen , Zilong Zheng

Existing Knowledge Distillation (KD) methods often align feature information between teacher and student by exploring meaningful feature processing and loss functions. However, due to the difference in feature distributions between the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Yu Wang , Chuanguang Yang , Zhulin An , Weilun Feng , Jiarui Zhao , Chengqing Yu , Libo Huang , Boyu Diao , Yongjun Xu

The representation gap between teacher and student is an emerging topic in knowledge distillation (KD). To reduce the gap and improve the performance, current methods often resort to complicated training schemes, loss functions, and feature…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Tao Huang , Yuan Zhang , Mingkai Zheng , Shan You , Fei Wang , Chen Qian , Chang Xu

3D Gaussian splatting-based talking head synthesis has recently gained attention for its ability to render high-fidelity images with real-time inference speed. However, since it is typically trained on only a short video that lacks the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Junuk Cha , Seongro Yoon , Valeriya Strizhkova , Francois Bremond , Seungryul Baek

In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Xingyu Miao , Haoran Duan , Yang Bai , Tejal Shah , Jun Song , Yang Long , Rajiv Ranjan , Ling Shao

Audio-driven 3D facial animation aims to map input audio to realistic facial motion. Despite significant progress, limitations arise from inconsistent 3D annotations, restricting previous models to training on specific annotations and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-02 Xiangyu Fan , Jiaqi Li , Zhiqian Lin , Weiye Xiao , Lei Yang

Current talking avatars mostly generate co-speech gestures based on audio and text of the utterance, without considering the non-speaking motion of the speaker. Furthermore, previous works on co-speech gesture generation have designed…

Multimedia · Computer Science 2024-01-09 Sicheng Yang , Zunnan Xu , Haiwei Xue , Yongkang Cheng , Shaoli Huang , Mingming Gong , Zhiyong Wu

Generating talking head videos through a face image and a piece of speech audio still contains many challenges. ie, unnatural head movement, distorted expression, and identity modification. We argue that these issues are mainly because of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Wenxuan Zhang , Xiaodong Cun , Xuan Wang , Yong Zhang , Xi Shen , Yu Guo , Ying Shan , Fei Wang

The paper introduces AniTalker, an innovative framework designed to generate lifelike talking faces from a single portrait. Unlike existing models that primarily focus on verbal cues such as lip synchronization and fail to capture the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Tao Liu , Feilong Chen , Shuai Fan , Chenpeng Du , Qi Chen , Xie Chen , Kai Yu

Standard Latent Diffusion Models rely on a complex, three-part architecture consisting of a separate encoder, decoder, and diffusion network, which are trained in multiple stages. This modular design is computationally inefficient, leads to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Xiyuan Wang , Muhan Zhang

Knowledge Distillation (KD) refers to transferring knowledge from a large model to a smaller one, which is widely used to enhance model performance in machine learning. It tries to align embedding spaces generated from the teacher and the…

Computer Vision and Pattern Recognition · Computer Science 2020-11-03 Weidong Shi , Guanghui Ren , Yunpeng Chen , Shuicheng Yan

Diffusion models have shown impressive potential on talking head generation. While plausible appearance and talking effect are achieved, these methods still suffer from temporal, 3D or expression inconsistency due to the error accumulation…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Haijie Yang , Zhenyu Zhang , Hao Tang , Jianjun Qian , Jian Yang

Speech-driven 3D facial animation has garnered lots of attention thanks to its broad range of applications. Despite recent advancements in achieving realistic lip motion, current methods fail to capture the nuanced emotional undertones…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Jisoo Kim , Jungbin Cho , Joonho Park , Soonmin Hwang , Da Eun Kim , Geon Kim , Youngjae Yu

Current diffusion-based portrait animation models predominantly focus on enhancing visual quality and expression realism, while overlooking generation latency and real-time performance, which restricts their application range in the live…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Zhiyuan Li , Chi-Man Pun , Chen Fang , Jue Wang , Xiaodong Cun