中文
相关论文

相关论文: Dual Audio-Centric Modality Coupling for Talking H…

200 篇论文

We introduce Diffusion-based Audio Captioning (DAC), a non-autoregressive diffusion model tailored for diverse and efficient audio captioning. Although existing captioning models relying on language backbones have achieved remarkable…

计算与语言 · 计算机科学 2025-06-03 Manjie Xu , Chenxing Li , Xinyi Tu , Yong Ren , Ruibo Fu , Wei Liang , Dong Yu

Large Language Models (LLMs) have demonstrated remarkable success across diverse fields, establishing a powerful paradigm for complex information processing. This has inspired the integration of speech into LLM frameworks, often by…

音频与语音处理 · 电气工程与系统科学 2025-12-30 Xiangyu Zhang , Fuming Fang , Peng Gao , Bin Qin , Beena Ahmed , Julien Epps

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

Imagine having a conversation with a socially intelligent agent. It can attentively listen to your words and offer visual and linguistic feedback promptly. This seamless interaction allows for multiple rounds of conversation to flow…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Yongming Zhu , Longhao Zhang , Zhengkun Rong , Tianshu Hu , Shuang Liang , Zhipeng Ge

Deepfakes are AI-generated media in which the original content is digitally altered to create convincing but manipulated images, videos, or audio. Among the various types of deepfakes, lip-syncing deepfakes are one of the most challenging…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Soumyya Kanti Datta , Shan Jia , Siwei Lyu

The talking head generation recently attracted considerable attention due to its widespread application prospects, especially for digital avatars and 3D animation design. Inspired by this practical demand, several works explored Neural…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Tianyong Wang , Xiangyu Liang , Wangguandong Zheng , Dan Niu , Haifeng Xia , Siyu Xia

We present a novel approach for generating realistic speaking and talking faces by synthesizing a person's voice and facial movements from a static image, a voice profile, and a target text. The model encodes the prompt/driving text, the…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Aashish Chandra , Aashutosh A , Abhijit Das

Generating realistic, dyadic talking head video requires ultra-low latency. Existing chunk-based methods require full non-causal context windows, introducing significant delays. This high latency critically prevents the immediate,…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Bohong Chen , Haiyang Liu

Audio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to previous studies.…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Dongze Li , Kang Zhao , Wei Wang , Bo Peng , Yingya Zhang , Jing Dong , Tieniu Tan

Body-conduction microphone signals (BMS) bypass airborne sound, providing strong noise resistance. However, a complementary modality is required to compensate for the inherent loss of high-frequency information. In this study, we propose a…

声音 · 计算机科学 2025-08-29 Yunsik Kim , Yoonyoung Chung

Engagement estimation plays a crucial role in understanding human social behaviors, attracting increasing research interests in fields such as affective computing and human-computer interaction. In this paper, we propose a Dialogue-Aware…

人机交互 · 计算机科学 2024-10-14 Jia Li , Yangchen Yu , Yin Chen , Yu Zhang , Peng Jia , Yunbo Xu , Ziqiang Li , Meng Wang , Richang Hong

Recent advances in deep learning for sequential data have given rise to fast and powerful models that produce realistic videos of talking humans. The state of the art in talking face generation focuses mainly on lip-syncing, being…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Georgios Milis , Panagiotis P. Filntisis , Anastasios Roussos , Petros Maragos

Talking head video compression has advanced with neural rendering and keypoint-based methods, but challenges remain, especially at low bit rates, including handling large head movements, suboptimal lip synchronization, and distorted facial…

图像与视频处理 · 电气工程与系统科学 2025-06-17 Riku Takahashi , Ryugo Morita , Jinjia Zhou

In dyadic speaker-listener interactions, the listener's head reactions along with the speaker's head movements, constitute an important non-verbal semantic expression together. The listener Head generation task aims to synthesize responsive…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Zhigang Chang , Weitai Hu , Qing Yang , Shibao Zheng

Audio-driven talking head generation is a significant and challenging task applicable to various fields such as virtual avatars, film production, and online conferences. However, the existing GAN-based models emphasize generating…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Jintao Tan , Xize Cheng , Lingyu Xiong , Lei Zhu , Xiandong Li , Xianjia Wu , Kai Gong , Minglei Li , Yi Cai

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Hui Fu , Zeqing Wang , Ke Gong , Keze Wang , Tianshui Chen , Haojie Li , Haifeng Zeng , Wenxiong Kang

For realistic talking head generation, creating natural head motion while maintaining accurate lip synchronization is essential. To fulfill this challenging task, we propose DisCoHead, a novel method to disentangle and control head pose and…

计算机视觉与模式识别 · 计算机科学 2023-03-15 Geumbyeol Hwang , Sunwon Hong , Seunghyun Lee , Sungwoo Park , Gyeongsu Chae

Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, producing both modalities simultaneously still typically relies on…

Audio-Driven Talking Face Generation aims at generating realistic videos of talking faces, focusing on accurate audio-lip synchronization without deteriorating any identity-related visual details. Recent state-of-the-art methods are based…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazım Kemal Ekenel , Alexander Waibel

Although the semantic communication with joint semantic-channel coding design has shown promising performance in transmitting data of different modalities over physical layer channels, the synchronization and packet-level forward error…

图像与视频处理 · 电气工程与系统科学 2024-08-13 Yun Tian , Jingkai Ying , Zhijin Qin , Ye Jin , Xiaoming Tao