English
Related papers

Related papers: Learning to Listen: Modeling Non-Deterministic Dya…

200 papers

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

Conversation is ubiquitous in social life, but the empirical study of this interactive process has been thwarted by tools that are insufficiently modular and unadaptive to researcher needs. To relieve many constraints in conversation…

Human-Computer Interaction · Computer Science 2026-03-24 David M. Markowitz

The domain of 3D talking head generation has witnessed significant progress in recent years. A notable challenge in this field consists in blending speech-related motions with expression dynamics, which is primarily caused by the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Federico Nocentini , Claudio Ferrari , Stefano Berretti

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been…

Sound · Computer Science 2022-07-08 Junwen Xiong , Yu Zhou , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or employ speech…

Computation and Language · Computer Science 2020-04-06 Haiyang Xu , Hui Zhang , Kun Han , Yun Wang , Yiping Peng , Xiangang Li

Speech-driven 3D facial animation has been widely studied, yet there is still a gap to achieving realism and vividness due to the highly ill-posed nature and scarcity of audio-visual data. Existing works typically formulate the cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Jinbo Xing , Menghan Xia , Yuechen Zhang , Xiaodong Cun , Jue Wang , Tien-Tsin Wong

Most of the prevalent approaches in speech prosody modeling rely on learning global style representations in a continuous latent space which encode and transfer the attributes of reference speech. However, recent work on neural codecs which…

Neural latent variable models enable the discovery of interesting structure in speech audio data. This paper presents a comparison of two different approaches which are broadly based on predicting future time-steps or auto-encoding the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-28 Henry Zhou , Alexei Baevski , Michael Auli

Emotion is a critical component of artificial social intelligence. However, while current methods excel in lip synchronization and image quality, they often fail to generate accurate and controllable emotional expressions while preserving…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Wenqing Wang , Yun Fu

In this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextual sentiments as well as speech rhythm and pauses. To be…

Computer Vision and Pattern Recognition · Computer Science 2021-05-10 Lincheng Li , Suzhen Wang , Zhimeng Zhang , Yu Ding , Yixing Zheng , Xin Yu , Changjie Fan

It is common in everyday spoken communication that we look at the turning head of a talker to listen to his/her voice. Humans see the talker to listen better, so do machines. However, previous studies on audio-visual speaker extraction have…

Sound · Computer Science 2023-09-14 Qinghua Liu , Meng Ge , Zhizheng Wu , Haizhou Li

Personality computing has become an emerging topic in computer vision, due to the wide range of applications it can be used for. However, most works on the topic have focused on analyzing the individual, even when applied to interaction…

Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Munender Varshney , Ravindra Yadav , Vinay P. Namboodiri , Rajesh M Hegde

Audio-driven talking head animation is a challenging research topic with many real-world applications. Recent works have focused on creating photo-realistic 2D animation, while learning different talking or singing styles remains an open…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Trong-Thang Pham , Nhat Le , Tuong Do , Hung Nguyen , Erman Tjiputra , Quang D. Tran , Anh Nguyen

In this paper, we introduce a new task, Reactive Listener Motion Generation from Speaker Utterance, which aims to generate naturalistic listener body motions that appropriately respond to a speaker's utterance. However, modeling such…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Cheng Luo , Bizhu Wu , Bing Li , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen , Bernard Ghanem

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Wenqing Wang , Yun Fu

Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a textdependent…

Computer Vision and Pattern Recognition · Computer Science 2020-10-14 Jianrong Wang , Tong Wu , Shanyu Wang , Mei Yu , Qiang Fang , Ju Zhang , Li Liu

In order to be widely applicable, speech-driven 3D head avatars must articulate their lips in accordance with speech, while also conveying the appropriate emotions with dynamically changing facial expressions. The key problem is that…

Graphics · Computer Science 2026-01-28 Radek Daněček , Carolin Schmitt , Senya Polikovsky , Michael J. Black

Speech-driven 3D facial animation has improved a lot recently while most related works only utilize acoustic modality and neglect the influence of visual and textual cues, leading to unsatisfactory results in terms of precision and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Tianshun Han , Shengnan Gui , Yiqing Huang , Baihui Li , Lijian Liu , Benjia Zhou , Ning Jiang , Quan Lu , Ruicong Zhi , Yanyan Liang , Du Zhang , Jun Wan

Speech-driven facial animation aims to synthesize lip-synchronized 3D talking faces following the given speech signal. Prior methods to this task mostly focus on pursuing realism with deterministic systems, yet characterizing the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Chunzhi Gu , Shigeru Kuriyama , Katsuya Hotta