English
Related papers

Related papers: Personalized Speech-driven Expressive 3D Facial An…

200 papers

This paper presents a novel framework for automatic speech-driven gesture generation, applicable to human-agent interaction including both virtual agents and robots. Specifically, we extend recent deep-learning-based, data-driven methods…

Human-Computer Interaction · Computer Science 2019-06-12 Taras Kucherenko , Dai Hasegawa , Gustav Eje Henter , Naoshi Kaneko , Hedvig Kjellström

Speech synthesis has significantly advanced from statistical methods to deep neural network architectures, leading to various text-to-speech (TTS) models that closely mimic human speech patterns. However, capturing nuances such as emotion…

Sound · Computer Science 2025-01-14 Shaozuo Zhang , Ambuj Mehrish , Yingting Li , Soujanya Poria

Humans often speak in a continuous manner which leads to coherent and consistent prosody properties across neighboring utterances. However, most state-of-the-art speech synthesis systems only consider the information within each sentence…

Sound · Computer Science 2023-05-19 Ya-Jie Zhang , Wei Song , Yanghao Yue , Zhengchen Zhang , Youzheng Wu , Xiaodong He

This study delves into the intricacies of synchronizing facial dynamics with multilingual audio inputs, focusing on the creation of visually compelling, time-synchronized animations through diffusion-based techniques. Diverging from…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Rui Zhang , Yixiao Fang , Zhengnan Lu , Pei Cheng , Zebiao Huang , Bin Fu

Expressive speech synthesis aims to generate speech that captures a wide range of para-linguistic features, including emotion and articulation, though current research primarily emphasizes emotional aspects over the nuanced articulatory…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-24 Zehua Kcriss Li , Meiying Melissa Chen , Yi Zhong , Pinxin Liu , Zhiyao Duan

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Linrui Tian , Qi Wang , Bang Zhang , Liefeng Bo

Modeling virtual agents with behavior style is one factor for personalizing human agent interaction. We propose an efficient yet effective machine learning approach to synthesize gestures driven by prosodic features and text in the style of…

Sound · Computer Science 2022-08-04 Mireille Fares , Michele Grimaldi , Catherine Pelachaud , Nicolas Obin

This work proposes to explore a new area of dynamic speech emotion recognition. Unlike traditional methods, we assume that each audio track is associated with a sequence of emotions active at different moments in time. The study…

Sound · Computer Science 2025-08-22 Ilya Fedorov , Dmitry Korobchenko

In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Chenxu Zhang , Yifan Zhao , Yifei Huang , Ming Zeng , Saifeng Ni , Madhukar Budagavi , Xiaohu Guo

Critical obstacles in training classifiers to detect facial actions are the limited sizes of annotated video databases and the relatively low frequencies of occurrence of many actions. To address these problems, we propose an approach that…

Computer Vision and Pattern Recognition · Computer Science 2020-10-22 Koichiro Niinuma , Itir Onal Ertugrul , Jeffrey F Cohn , László A Jeni

Facial animation is a core component for creating digital characters in Computer Graphics (CG) industry. A typical production workflow relies on sparse, semantically meaningful keyframes to precisely control facial expressions. Enabling…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Jingchao Wu , Zejian Kang , Haibo Liu , Yuanchen Fei , Xiangru Huang

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Ziqi Zhang , Cheng Deng

In this work, we address the problem of 4D facial expressions generation. This is usually addressed by animating a neutral 3D face to reach an expression peak, and then get back to the neutral state. In the real world though, people show…

Computer Vision and Pattern Recognition · Computer Science 2023-05-19 Naima Otberdout , Claudio Ferrari , Mohamed Daoudi , Stefano Berretti , Alberto Del Bimbo

Audio-driven talking head generation is advancing from 2D to 3D content. Notably, Neural Radiance Field (NeRF) is in the spotlight as a means to synthesize high-quality 3D talking head outputs. Unfortunately, this NeRF-based approach…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Gihoon Kim , Kwanggyoon Seo , Sihun Cha , Junyong Noh

3D Gaussian splatting-based talking head synthesis has recently gained attention for its ability to render high-fidelity images with real-time inference speed. However, since it is typically trained on only a short video that lacks the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Junuk Cha , Seongro Yoon , Valeriya Strizhkova , Francois Bremond , Seungryul Baek

The diversity of facial shapes and motions among persons is one of the greatest challenges for automatic analysis of facial expressions. In this paper, we propose a feature describing expression intensity over time, while being invariant to…

Computer Vision and Pattern Recognition · Computer Science 2018-05-04 Maren Awiszus , Stella Graßhof , Felix Kuhnke , Jörn Ostermann

When we speak, the prosody and content of the speech can be inferred from the movement of our lips. In this work, we explore the task of lip to speech synthesis, i.e., learning to generate speech given only the lip movements of a speaker…

Computer Vision and Pattern Recognition · Computer Science 2022-06-29 Christen Millerdurai , Lotfy Abdel Khaliq , Timon Ulrich

Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchronization, often overlooking aspects such as visual quality,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Fatemeh Nazarieh , Zhenhua Feng , Diptesh Kanojia , Muhammad Awais , Josef Kittler

We present techniques for improving performance driven facial animation, emotion recognition, and facial key-point or landmark prediction using learned identity invariant representations. Established approaches to these problems can work…

Computer Vision and Pattern Recognition · Computer Science 2016-05-24 David Rim , Sina Honari , Md Kamrul Hasan , Chris Pal

Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to…

Sound · Computer Science 2025-05-20 Yifan Hu , Rui Liu , Yi Ren , Xiang Yin , Haizhou Li