English
Related papers

Related papers: OmniTalker: One-shot Real-time Text-Driven Talking…

200 papers

Speech-driven facial animation is the process which uses speech signals to automatically synthesize a talking character. The majority of work in this domain creates a mapping from audio features to visual features. This often requires…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-20 Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

Speech-driven 3D facial animation with accurate lip synchronization has been widely studied. However, synthesizing realistic motions for the entire face during speech has rarely been explored. In this work, we present a joint audio-text…

Computer Vision and Pattern Recognition · Computer Science 2021-12-08 Yingruo Fan , Zhaojiang Lin , Jun Saito , Wenping Wang , Taku Komura

Recent works on audio-driven talking head synthesis using Neural Radiance Fields (NeRF) have achieved impressive results. However, due to inadequate pose and expression control caused by NeRF implicit representation, these methods still…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Hongyun Yu , Zhan Qu , Qihang Yu , Jianchuan Chen , Zhonghua Jiang , Zhiwen Chen , Shengyu Zhang , Jimin Xu , Fei Wu , Chengfei Lv , Gang Yu

Generative modeling has recently achieved remarkable success across text, image, and audio domains, demonstrating powerful capabilities for unified representation learning. However, audio generation models still face challenges in terms of…

Sound · Computer Science 2025-10-31 Chengwei Liu , Haoyin Yan , Shaofei Xue , Xiaotao Liang , Yinghao Liu , Zheng Xue , Gang Song , Boyang Zhou

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Jiaben Chen , Zixin Wang , Ailing Zeng , Yang Fu , Xueyang Yu , Siyuan Cen , Julian Tanke , Yihang Chen , Koichi Saito , Yuki Mitsufuji , Chuang Gan

Audio-Driven Talking Face Generation aims at generating realistic videos of talking faces, focusing on accurate audio-lip synchronization without deteriorating any identity-related visual details. Recent state-of-the-art methods are based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazım Kemal Ekenel , Alexander Waibel

Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing audio-video generation pipelines typically decompose the…

This paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture. It adopts a unified transformer-based visual encoder for both image and video inputs, and thus can…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Junke Wang , Dongdong Chen , Zuxuan Wu , Chong Luo , Luowei Zhou , Yucheng Zhao , Yujia Xie , Ce Liu , Yu-Gang Jiang , Lu Yuan

We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation methods shows that crucial information is the movement of the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-23 Ander Arriandiaga , Giovanni Morrone , Luca Pasa , Leonardo Badino , Chiara Bartolozzi

Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelity lip-sync with different techniques, their task-orientated…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Jiazhi Guan , Zhiliang Xu , Hang Zhou , Kaisiyuan Wang , Shengyi He , Zhanwang Zhang , Borong Liang , Haocheng Feng , Errui Ding , Jingtuo Liu , Jingdong Wang , Youjian Zhao , Ziwei Liu

Audio-driven portrait animation aims to synthesize portrait videos that are conditioned by given audio. Animating high-fidelity and multimodal video portraits has a variety of applications. Previous methods have attempted to capture…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Yunfei Liu , Lijian Lin , Fei Yu , Changyin Zhou , Yu Li

Talking face generation technology creates talking videos from arbitrary appearance and motion signal, with the "arbitrary" offering ease of use but also introducing challenges in practical applications. Existing methods work well with…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Chao Liang , Jianwen Jiang , Tianyun Zhong , Gaojie Lin , Zhengkun Rong , Jiaqi Yang , Yongming Zhu

Speech-driven facial animation is the process that automatically synthesizes talking characters based on speech signals. The majority of work in this domain creates a mapping from audio features to visual features. This approach often…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architecture as their backbone, leaving significant room to explore effective and efficient…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Lijiang Li , Zuwei Long , Yunhang Shen , Heting Gao , Haoyu Cao , Xing Sun , Caifeng Shan , Ran He , Chaoyou Fu

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an avatar chatbot system…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Se Jin Park , Chae Won Kim , Hyeongseop Rha , Minsu Kim , Joanna Hong , Jeong Hun Yeo , Yong Man Ro

Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Zhe Li , Weihao Yuan , Weichao Shen , Siyu Zhu , Zilong Dong , Chang Xu

In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ziqiao Peng , Yanbo Fan , Haoyu Wu , Xuan Wang , Hongyan Liu , Jun He , Zhaoxin Fan

Real-time speech-driven 3D facial animation has been attractive in academia and industry. Traditional methods mainly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consider the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Peng Chen , Xiaobao Wei , Ming Lu , Hui Chen , Feng Tian

Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed. However, existing generation models often entangle these factors, limiting independent control. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Yukun Wang , Ruihuang Li , Jiale Tao , Shiyuan Yang , Liyi Chen , Zhantao Yang , Handz , Yulan Guo , Shuai Shao , Qinglin Lu

We present MGM-Omni, a unified Omni LLM for omni-modal understanding and expressive, long-horizon speech generation. Unlike cascaded pipelines that isolate speech synthesis, MGM-Omni adopts a "brain-mouth" design with a dual-track,…

Sound · Computer Science 2025-09-30 Chengyao Wang , Zhisheng Zhong , Bohao Peng , Senqiao Yang , Yuqi Liu , Haokun Gui , Bin Xia , Jingyao Li , Bei Yu , Jiaya Jia