English
Related papers

Related papers: FaceXHuBERT: Text-less Speech-driven E(X)pressive …

200 papers

We present SentiAvatar, a framework for building expressive interactive 3D digital humans, and use it to create SuSu, a virtual character that speaks, gestures, and emotes in real time. Achieving such a system remains challenging, as it…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Chuhao Jin , Rui Zhang , Qingzhe Gao , Haoyu Shi , Dayu Wu , Yichen Jiang , Yihan Wu , Ruihua Song

Existing methodologies for animating portrait images face significant challenges, particularly in handling non-frontal perspectives, rendering dynamic objects around the portrait, and generating immersive, realistic backgrounds. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Jiahao Cui , Hui Li , Yun Zhan , Hanlin Shang , Kaihui Cheng , Yuqi Ma , Shan Mu , Hang Zhou , Jingdong Wang , Siyu Zhu

For realistic talking head generation, creating natural head motion while maintaining accurate lip synchronization is essential. To fulfill this challenging task, we propose DisCoHead, a novel method to disentangle and control head pose and…

Computer Vision and Pattern Recognition · Computer Science 2023-03-15 Geumbyeol Hwang , Sunwon Hong , Seunghyun Lee , Sungwoo Park , Gyeongsu Chae

3D facial animation is often produced by manipulating facial deformation models (or rigs), that are traditionally parameterized by expression controls. A key component that is usually overlooked is expression 'style', as in, how a…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Lingchen Yang , Gaspard Zoss , Prashanth Chandran , Paulo Gotardo , Markus Gross , Barbara Solenthaler , Eftychios Sifakis , Derek Bradley

Audio-driven talking face generation has garnered significant interest within the domain of digital human research. Existing methods are encumbered by intricate model architectures that are intricately dependent on each other, complicating…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Dong Zhao , Jiaying Shi , Wenjun Li , Shudong Wang , Shenghui Xu , Zhaoming Pan

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio, we output multiple possibilities of gestural motion for an…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Evonne Ng , Javier Romero , Timur Bagautdinov , Shaojie Bai , Trevor Darrell , Angjoo Kanazawa , Alexander Richard

While 3D facial animation has made impressive progress, challenges still exist in realizing fine-grained stylized 3D facial expression manipulation due to the lack of appropriate datasets. In this paper, we introduce the AUBlendSet, a 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Hao Li , Ju Dai , Feng Zhou , Kaida Ning , Lei Li , Junjun Pan

The creation of increasingly vivid 3D talking face has become a hot topic in recent years. Currently, most speech-driven works focus on lip synchronisation but neglect to effectively capture the correlations between emotions and facial…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Yihong Lin , Liang Peng , Zhaoxin Fan , Xianjia Wu , Jianqiao Hu , Xiandong Li , Wenxiong Kang , Songju Lei

Recent years have witnessed significant advancements in self-supervised learning (SSL) methods for speech-processing tasks. Various speech-based SSL models have been developed and present promising performance on a range of downstream tasks…

Computation and Language · Computer Science 2023-10-02 Guanrou Yang , Ziyang Ma , Zhisheng Zheng , Yakun Song , Zhikang Niu , Xie Chen

The generation of stylistic 3D facial animations driven by speech presents a significant challenge as it requires learning a many-to-many mapping between speech, style, and the corresponding natural facial motion. However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Zhiyao Sun , Tian Lv , Sheng Ye , Matthieu Lin , Jenny Sheng , Yu-Hui Wen , Minjing Yu , Yong-Jin Liu

In the task of talking face generation, the objective is to generate a face video with lips synchronized to the corresponding audio while preserving visual details and identity information. Current methods face the challenge of learning…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Seymanur Aktı , Hazım Kemal Ekenel , Alexander Waibel

Although automatically animating audio-driven talking heads has recently received growing interest, previous efforts have mainly concentrated on achieving lip synchronization with the audio, neglecting two crucial elements for generating…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Shuai Tan , Bin Ji , Ye Pan

This study explores a streamlined facial data collection method for conversational contexts, addressing the limitations of existing approaches that often require extensive datasets and prioritize technical metrics over user perception and…

Human-Computer Interaction · Computer Science 2026-02-03 Seoyoung Kang , Seokhwan Yang , Hail Song , Boram Yoon , Jinwook Kim , Kangsoo Kim , Woontack Woo

Generating emotional talking faces is a practical yet challenging endeavor. To create a lifelike avatar, we draw upon two critical insights from a human perspective: 1) The connection between audio and the non-deterministic facial dynamics,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Shuai Tan , Bin Ji , Ye Pan

This paper presents a novel approach for generating 3D talking heads from raw audio inputs. Our method grounds on the idea that speech related movements can be comprehensively and efficiently described by the motion of a few control points…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Federico Nocentini , Claudio Ferrari , Stefano Berretti

The recent state of the art on monocular 3D face reconstruction from image data has made some impressive advancements, thanks to the advent of Deep Learning. However, it has mostly focused on input coming from a single RGB image,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Panagiotis P. Filntisis , George Retsinas , Foivos Paraperas-Papantoniou , Athanasios Katsamanis , Anastasios Roussos , Petros Maragos

Generating speech-driven 3D talking heads presents numerous challenges; among those is dealing with varying mesh topologies where no point-wise correspondence exists across the meshes the model can animate. While previous literature works…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Federico Nocentini , Thomas Besnier , Claudio Ferrari , Sylvain Arguillere , Mohamed Daoudi , Stefano Berretti

We present FlexAvatar, a flexible large reconstruction model for high-fidelity 3D head avatars with detailed dynamic deformation from single or sparse images, without requiring camera poses or expression labels. It leverages a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Cheng Peng , Zhuo Su , Liao Wang , Chen Guo , Zhaohu Li , Chengjiang Long , Zheng Lv , Jingxiang Sun , Chenyangguang Zhang , Yebin Liu

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper introduces JoyGen, a…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Qili Wang , Dajiang Wu , Zihang Xu , Junshi Huang , Jun Lv

Existing self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However, many current models are trained on a clean corpus from a…

Sound · Computer Science 2023-03-01 Dianwen Ng , Ruixi Zhang , Jia Qi Yip , Zhao Yang , Jinjie Ni , Chong Zhang , Yukun Ma , Chongjia Ni , Eng Siong Chng , Bin Ma
‹ Prev 1 8 9 10 Next ›