中文
相关论文

相关论文: GenSync: A Generalized Talking Head Framework for …

200 篇论文

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Hebeizi Li , Zihao Liang , Benyuan Sun , Zihao Yin , Xiao Sha , Chenliang Wang , Yi Yang

The goal of this paper is to synthesise talking faces with controllable facial motions. To achieve this goal, we propose two key ideas. The first is to establish a canonical space where every face has the same motion patterns but different…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Youngjoon Jang , Kyeongha Rho , Jong-Bin Woo , Hyeongkeun Lee , Jihwan Park , Youshin Lim , Byeong-Yeol Kim , Joon Son Chung

Achieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. Traditional Generative Adversarial Networks (GAN) struggle to maintain consistent facial identity, while…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Ziqiao Peng , Wentao Hu , Yue Shi , Xiangyu Zhu , Xiaomei Zhang , Hao Zhao , Jun He , Hongyan Liu , Zhaoxin Fan

In this paper, we introduce ConversaSynth, a framework designed to generate synthetic conversation audio using large language models (LLMs) with multiple persona settings. The framework first creates diverse and coherent text-based…

声音 · 计算机科学 2025-07-08 Kaung Myat Kyaw , Jonathan Hoyin Chan

We introduce MIGS (Multi-Identity Gaussian Splatting), a novel method that learns a single neural representation for multiple identities, using only monocular videos. Recent 3D Gaussian Splatting (3DGS) approaches for human avatars require…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Aggelina Chatziagapi , Grigorios G. Chrysos , Dimitris Samaras

Novel-view synthesis aims to generate novel views of a scene from multiple input images or videos, and recent advancements like 3D Gaussian splatting (3DGS) have achieved notable success in producing photorealistic renderings with efficient…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Xi Liu , Chaoyi Zhou , Siyu Huang

3D Gaussian Splatting (3DGS) has demonstrated impressive novel view synthesis performance. While conventional methods require per-scene optimization, more recently several feed-forward methods have been proposed to generate pixel-aligned…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Shengjun Zhang , Xin Fei , Fangfu Liu , Haixu Song , Yueqi Duan

Animating still face images with deep generative models using a speech input signal is an active research topic and has seen important recent progress.However, much of the effort has been put into lip syncing and rendering quality while the…

图形学 · 计算机科学 2024-12-13 Louis Airale , Dominique Vaufreydaz , Xavier Alameda-Pineda

3D Gaussian Splatting (3DGS) has become a state-of-the-art framework for real-time, high-fidelity novel view synthesis. However, its substantial storage requirements and inherently unstructured representation pose challenges for deployment…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Yuqin Lu , Yang Zhou , Yihua Dai , Guiqing Li , Shengfeng He

Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual datasets, such as GRID,…

声音 · 计算机科学 2023-03-02 Zhe Niu , Brian Mak

While 3D Gaussian Splatting enables high-quality real-time rendering, existing Gaussian-based frameworks for 3D semantic segmentation still face significant challenges in boundary recognition accuracy. To address this, we propose a novel…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Zehao Li , Wenwei Han , Yujun Cai , Hao Jiang , Baolong Bi , Shuqin Gao , Honglong Zhao , Zhaoqi Wang

Audio-driven lip sync has recently drawn significant attention due to its widespread application in the multimedia domain. Individuals exhibit distinct lip shapes when speaking the same utterance, attributed to the unique speaking styles of…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Weizhi Zhong , Jichang Li , Yinqi Cai , Ming Li , Feng Gao , Liang Lin , Guanbin Li

We present FlashLips, a two-stage, mask-free lip-sync system that decouples lips control from rendering and achieves real-time performance, with our U-Net variant running at over 100 FPS on a single GPU, while matching the visual quality of…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Andreas Zinonos , Michał Stypułkowski , Antoni Bigata , Stavros Petridis , Maja Pantic , Nikita Drobyshev

This paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face animation, fail to produce accurate and plausible…

计算机视觉与模式识别 · 计算机科学 2022-05-23 Alexander Richard , Michael Zollhoefer , Yandong Wen , Fernando de la Torre , Yaser Sheikh

We introduce HyperGS, a novel framework for Hyperspectral Novel View Synthesis (HNVS), based on a new latent 3D Gaussian Splatting (3DGS) technique. Our approach enables simultaneous spatial and spectral renderings by encoding material…

计算机视觉与模式识别 · 计算机科学 2024-12-18 Christopher Thirgood , Oscar Mendez , Erin Chao Ling , Jon Storey , Simon Hadfield

The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Chenxu Zhang , Chao Wang , Jianfeng Zhang , Hongyi Xu , Guoxian Song , You Xie , Linjie Luo , Yapeng Tian , Xiaohu Guo , Jiashi Feng

To enable AI agents to interact seamlessly with both humans and 3D environments, they must not only perceive the 3D world accurately but also align human language with 3D spatial representations. While prior work has made significant…

人工智能 · 计算机科学 2025-09-26 Saimouli Katragadda , Cho-Ying Wu , Yuliang Guo , Xinyu Huang , Guoquan Huang , Liu Ren

We tackle the challenge of generating dynamic 4D scenes from monocular, multi-object videos with heavy occlusions, and introduce GenMOJO, a novel approach that integrates rendering-based deformable 3D Gaussian optimization with generative…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Wen-Hsuan Chu , Lei Ke , Jianmeng Liu , Mingxiao Huo , Pavel Tokmakov , Katerina Fragkiadaki

Novel view synthesis has seen significant advancements with 3D Gaussian Splatting (3DGS), enabling real-time photorealistic rendering. However, the inherent fuzziness of Gaussian Splatting presents challenges for 3D scene understanding,…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Abdalla Arafa , Didier Stricker

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input…

计算机视觉与模式识别 · 计算机科学 2023-01-03 Balamurugan Thambiraja , Ikhsanul Habibie , Sadegh Aliakbarian , Darren Cosker , Christian Theobalt , Justus Thies