中文
相关论文

相关论文: VisualSpeaker: Visually-Guided 3D Avatar Lip Synth…

200 篇论文

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

计算机视觉与模式识别 · 计算机科学 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

Speech-driven 3D facial animation has advanced rapidly, yet most approaches remain tied to registered template meshes, preventing effective deployment on raw 3D scans with arbitrary topology. At the same time, modeling controllable…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Federico Nocentini , Thomas Besnier , Claudio Ferrari , Stefano Berretti , Mohamed Daoudi

Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper, we start by…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazim Kemal Ekenel , Alexander Waibel

Visual speech recognition is a technique to identify spoken content in silent speech videos, which has raised significant attention in recent years. Advancements in data-driven deep learning methods have significantly improved both the…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Lei Yang , Junshan Jin , Mingyuan Zhang , Yi He , Bofan Chen , Shilin Wang

With the rapid advancement of 3D representation techniques and generative models, substantial progress has been made in reconstructing full-body 3D avatars from a single image. However, this task remains fundamentally ill-posedness due to…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Gaofeng Liu , Hengsen Li , Ruoyu Gao , Xuetong Li , Zhiyuan Ma , Tao Fang

We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Sicheng Xu , Guojun Chen , Yu-Xiao Guo , Jiaolong Yang , Chong Li , Zhenyu Zang , Yizhong Zhang , Xin Tong , Baining Guo

Lipreading is understanding speech from observed lip movements. An observed series of lip motions is an ordered sequence of visual lip gestures. These gestures are commonly known, but as yet are not formally defined, as `visemes'. In this…

图像与视频处理 · 电气工程与系统科学 2019-09-17 Helen Bear , Richard Harvey

The ability to accurately capture and express emotions is a critical aspect of creating believable characters in video games and other forms of entertainment. Traditionally, this animation has been achieved with artistic effort or…

图形学 · 计算机科学 2023-07-19 Jack Saunders , Steven Caulkin , Vinay Namboodiri

We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We address these limitations with a flow matching based framework.…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Haiyang Liu , Xiaolin Hong , Xuancheng Yang , Yudi Ruan , Xiang Lian , Michael Lingelbach , Hongwei Yi , Wei Li

Creating realistic 3D facial animation is crucial for various applications in the movie production and gaming industry, especially with the burgeoning demand in the metaverse. However, prevalent methods such as blendshape-based approaches…

计算机视觉与模式识别 · 计算机科学 2023-08-14 Haoyu Wang , Haozhe Wu , Junliang Xing , Jia Jia

When we speak, the prosody and content of the speech can be inferred from the movement of our lips. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate speech given only the lip movements of a speaker…

计算机视觉与模式识别 · 计算机科学 2022-06-29 Christen Millerdurai , Lotfy Abdel Khaliq , Timon Ulrich

Recent advances in diffusion-based lip-syncing generative models have demonstrated their ability to produce highly synchronized talking face videos for visual dubbing. Although these models excel at lip synchronization, they often struggle…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Yanyu Zhu , Lichen Bai , Jintao Xu , Hai-tao Zheng

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body or half-body poses,…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Linrui Tian , Siqi Hu , Qi Wang , Bang Zhang , Liefeng Bo

Digital humans and, especially, 3D facial avatars have raised a lot of attention in the past years, as they are the backbone of several applications like immersive telepresence in AR or VR. Despite the progress, facial avatars reconstructed…

计算机视觉与模式识别 · 计算机科学 2023-11-27 Berna Kabadayi , Wojciech Zielonka , Bharat Lal Bhatnagar , Gerard Pons-Moll , Justus Thies

The paper introduces AniTalker, an innovative framework designed to generate lifelike talking faces from a single portrait. Unlike existing models that primarily focus on verbal cues such as lip synchronization and fail to capture the…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Tao Liu , Feilong Chen , Shuai Fan , Chenpeng Du , Qi Chen , Xie Chen , Kai Yu

Generating talking avatar driven by audio remains a significant challenge. Existing methods typically require high computational costs and often lack sufficient facial detail and realism, making them unsuitable for applications that demand…

计算机视觉与模式识别 · 计算机科学 2025-01-27 Yujian Liu , Shidang Xu , Jing Guo , Dingbin Wang , Zairan Wang , Xianfeng Tan , Xiaoli Liu

Vision-guided speech generation aims to produce authentic speech from facial appearance or lip motions without relying on auditory signals, offering significant potential for applications such as dubbing in filmmaking and assisting…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Jiaxin Ye , Hongming Shan

Creating digital avatars from textual prompts has long been a desirable yet challenging task. Despite the promising results achieved with 2D diffusion priors, current methods struggle to create high-quality and consistent animated avatars…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Zhenglin Zhou , Fan Ma , Hehe Fan , Zongxin Yang , Yi Yang

We introduce AvatarBooth, a novel method for generating high-quality 3D avatars using text prompts or specific images. Unlike previous approaches that can only synthesize avatars based on simple text descriptions, our method enables the…

计算机视觉与模式识别 · 计算机科学 2023-06-19 Yifei Zeng , Yuanxun Lu , Xinya Ji , Yao Yao , Hao Zhu , Xun Cao

Speech-driven facial animation is the process which uses speech signals to automatically synthesize a talking character. The majority of work in this domain creates a mapping from audio features to visual features. This often requires…

音频与语音处理 · 电气工程与系统科学 2018-07-20 Konstantinos Vougioukas , Stavros Petridis , Maja Pantic
‹ 上一页 1 8 9 10 下一页 ›