English
Related papers

Related papers: KSDiff: Keyframe-Augmented Speech-Aware Dual-Path …

200 papers

Diffusion-based audio-driven talking-head generation enables realistic portrait animation, but also introduces risks of misuse, such as fraud and misinformation. Existing protection methods are largely limited to a single modality, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Wenli Zhang , Xianglong Shi , Sirui Zhao , Xinqi Chen , Guo Cheng , Yifan Xu , Tong Xu , Yong Liao

Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hindered their applications to speech synthesis. This paper…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-22 Rongjie Huang , Max W. Y. Lam , Jun Wang , Dan Su , Dong Yu , Yi Ren , Zhou Zhao

Recent advancements in diffusion models have greatly improved the quality and diversity of synthesized content. To harness the expressive power of diffusion models, researchers have explored various controllable mechanisms that allow users…

Computer Vision and Pattern Recognition · Computer Science 2023-04-28 Tsai-Shien Chen , Chieh Hubert Lin , Hung-Yu Tseng , Tsung-Yi Lin , Ming-Hsuan Yang

The human brain has the capability to associate the unknown person's voice and face by leveraging their general relationship, referred to as ``cross-modal speaker verification''. This task poses significant challenges due to the complex…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-26 Ruijie Tao , Zhan Shi , Yidi Jiang , Duc-Tuan Truong , Eng-Siong Chng , Massimo Alioto , Haizhou Li

Visuomotor imitation learning policies enable robots to efficiently acquire manipulation skills from visual demonstrations. However, as scene complexity and visual distractions increase, policies that perform well in simple settings often…

Artificial Intelligence · Computer Science 2025-11-11 Yuhang Dong , Haizhou Ge , Yupei Zeng , Jiangning Zhang , Beiwen Tian , Hongrui Zhu , Yufei Jia , Ruixiang Wang , Zhucun Xue , Guyue Zhou , Longhua Ma , Guanzhong Tian

With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which tends to introduce…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-31 Nan Xu , Zhaolong Huang , Xiaonan Zhi

Collaborative 3D object detection holds significant importance in the field of autonomous driving, as it greatly enhances the perception capabilities of each individual agent by facilitating information exchange among multiple agents.…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Zhe Huang , Shuo Wang , Yongcai Wang , Lei Wang

The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer from limited realism…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Qingcheng Zhao , Pengyu Long , Qixuan Zhang , Dafei Qin , Han Liang , Longwen Zhang , Yingliang Zhang , Jingyi Yu , Lan Xu

Discrete diffusion models generate sequences by iteratively denoising samples corrupted by categorical noise, offering an appealing alternative to autoregressive decoding for structured and symbolic generation. However, standard training…

Machine Learning · Computer Science 2026-02-04 Huu Binh Ta , Michael Cardei , Alvaro Velasquez , Ferdinando Fioretto

Audio-driven talking face generation has garnered significant interest within the domain of digital human research. Existing methods are encumbered by intricate model architectures that are intricately dependent on each other, complicating…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Dong Zhao , Jiaying Shi , Wenjun Li , Shudong Wang , Shenghui Xu , Zhaoming Pan

Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over facial animation such as speaking style and emotional expression,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Baiqin Wang , Xiangyu Zhu , Fan Shen , Hao Xu , Zhen Lei

Event-based cameras are bio-inspired sensors that asynchronously capture pixel intensity changes with microsecond latency, high temporal resolution, and high dynamic range, providing information on the spatiotemporal dynamics of a scene. We…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Rodrigo Verschae , Ignacio Bugueno-Cordova

Self-supervised representation learning has gained increasing attention for strong generalization ability without relying on paired datasets. However, it has not been explored sufficiently for facial representation. Self-supervised facial…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Ruian He , Zhen Xing , Weimin Tan , Bo Yan

Online Speech Enhancement was mainly reserved for predictive models. A key advantage of these models is that for an incoming signal frame from a stream of data, the model is called only once for enhancement. In contrast, generative Speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-22 Bunlong Lay , Rostislav Makarov , Simon Welker , Maris Hillemann , Timo Gerkmann

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Suzhen Wang , Lincheng Li , Yu Ding , Changjie Fan , Xin Yu

Multimodal neuroimaging provides complementary insights for Alzheimer's disease diagnosis, yet clinical datasets frequently suffer from missing modalities. We propose ACADiff, a framework that synthesizes missing brain imaging modalities…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Rong Zhou , Houliang Zhou , Yao Su , Brian Y. Chen , Yu Zhang , Lifang He , Alzheimer's Disease Neuroimaging Initiative

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between current results and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Xiaozhong Ji , Chuming Lin , Zhonggan Ding , Ying Tai , Junwei Zhu , Xiaobin Hu , Donghao Luo , Yanhao Ge , Chengjie Wang

We present AnaMoDiff, a novel diffusion-based method for 2D motion analogies that is applied to raw, unannotated videos of articulated characters. Our goal is to accurately transfer motions from a 2D driving video onto a source character,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Maham Tanveer , Yizhi Wang , Ruiqi Wang , Nanxuan Zhao , Ali Mahdavi-Amiri , Hao Zhang

Speech-driven 3D facial animation has been an attractive task in both academia and industry. Traditional methods mostly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consider the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Peng Chen , Xiaobao Wei , Ming Lu , Yitong Zhu , Naiming Yao , Xingyu Xiao , Hui Chen

Portrait animation aims to synthesize talking videos from a static reference face, conditioned on audio and style frame cues (e.g., emotion and head poses), while ensuring precise lip synchronization and faithful reproduction of speaking…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 He Feng , Yongjia Ma , Donglin Di , Lei Fan , Tonghua Su , Xiangqian Wu