English
Related papers

Related papers: Voice2Mesh: Cross-Modal 3D Face Model Generation f…

200 papers

In this work, we address the lack of 3D understanding of generative neural networks by introducing a persistent 3D feature embedding for view synthesis. To this end, we propose DeepVoxels, a learned representation that encodes the…

Computer Vision and Pattern Recognition · Computer Science 2019-04-12 Vincent Sitzmann , Justus Thies , Felix Heide , Matthias Nießner , Gordon Wetzstein , Michael Zollhöfer

Existing generative models for 3D shapes are typically trained on a large 3D dataset, often of a specific object category. In this paper, we investigate the deep generative model that learns from only a single reference 3D shape.…

Graphics · Computer Science 2022-12-19 Rundi Wu , Changxi Zheng

Recent advances in deep learning have significantly increased the performance of face recognition systems. The performance and reliability of these models depend heavily on the amount and quality of the training data. However, the…

Computer Vision and Pattern Recognition · Computer Science 2018-02-19 Adam Kortylewski , Andreas Schneider , Thomas Gerig , Bernhard Egger , Andreas Morel-Forster , Thomas Vetter

Generating speech from a face image is crucial for developing virtual humans capable of interacting using their unique voices, without relying on pre-recorded human speech. In this paper, we propose Face-StyleSpeech, a zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Minki Kang , Wooseok Han , Eunho Yang

Despite recent advances in appearance-based gaze estimation techniques, the need for training data that covers the target head pose and gaze distribution remains a crucial challenge for practical deployment. This work examines a novel…

Computer Vision and Pattern Recognition · Computer Science 2022-04-28 Jiawei Qin , Takuru Shimoyama , Yusuke Sugano

Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a textdependent…

Computer Vision and Pattern Recognition · Computer Science 2020-10-14 Jianrong Wang , Tong Wu , Shanyu Wang , Mei Yu , Qiang Fang , Ju Zhang , Li Liu

In this paper, we present an improved model for voicing silent speech, where audio is synthesized from facial electromyography (EMG) signals. To give our model greater flexibility to learn its own input features, we directly use EMG signals…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-22 David Gaddy , Dan Klein

Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Jinting Wang , Jun Wang , Hei Victor Cheng , Li Liu

We propose a novel learned deep prior of body motion for 3D hand shape synthesis and estimation in the domain of conversational gestures. Our model builds upon the insight that body motion and hand gestures are strongly correlated in…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Evonne Ng , Shiry Ginosar , Trevor Darrell , Hanbyul Joo

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

We introduce a seemingly impossible task: given only an audio clip of someone speaking, decide which of two face images is the speaker. In this paper we study this, and a number of related cross-modal tasks, aimed at answering the question:…

Computer Vision and Pattern Recognition · Computer Science 2018-04-04 Arsha Nagrani , Samuel Albanie , Andrew Zisserman

Recently, a lot of attention has been focused on the incorporation of 3D data into face analysis and its applications. Despite providing a more accurate representation of the face, 3D facial images are more complex to acquire than 2D…

Computer Vision and Pattern Recognition · Computer Science 2021-03-01 Araceli Morales , Gemma Piella , Federico M. Sukno

Existing methods for 3D face reconstruction from a few casually captured images employ deep learning based models along with a 3D Morphable Model(3DMM) as face geometry prior. Structure From Motion(SFM), followed by Multi-View Stereo (MVS),…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Raja Kumar , Jiahao Luo , Alex Pang , James Davis

Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content, accents, and prosody. In this paper, we propose RESOUND, a…

Sound · Computer Science 2025-05-29 Long-Khanh Pham , Thanh V. T. Tran , Minh-Tan Pham , Van Nguyen

Generating photo-realistic video portrait with arbitrary speech audio is a crucial problem in film-making and virtual reality. Recently, several works explore the usage of neural radiance field in this task to improve 3D realness and image…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Zhenhui Ye , Ziyue Jiang , Yi Ren , Jinglin Liu , JinZheng He , Zhou Zhao

Automatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is difficult due to the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Kunkun Pang , Dafei Qin , Yingruo Fan , Julian Habekost , Takaaki Shiratori , Junichi Yamagishi , Taku Komura

The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal…

Sound · Computer Science 2020-11-05 Soo-Whan Chung , Hong Goo Kang , Joon Son Chung

We introduce FaceGPT, a self-supervised learning framework for Large Vision-Language Models (VLMs) to reason about 3D human faces from images and text. Typical 3D face reconstruction methods are specialized algorithms that lack semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Haoran Wang , Mohit Mendiratta , Christian Theobalt , Adam Kortylewski

This paper delves into the emerging field of face-based voice conversion, leveraging the unique relationship between an individual's facial features and their vocal characteristics. We present a novel face-based voice conversion framework…

Sound · Computer Science 2024-08-20 Jaejun Lee , Yoori Oh , Injune Hwang , Kyogu Lee

Reconstructing 3D face from a single unconstrained image remains a challenging problem due to diverse conditions in unconstrained environments. Recently, learning-based methods have achieved notable results by effectively capturing complex…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Danling Cao
‹ Prev 1 3 4 5 6 7 10 Next ›