中文
相关论文

相关论文: Rethinking Voice-Face Correlation: A Geometry View

200 篇论文

We introduce GaussianSpeech, a novel approach that synthesizes high-fidelity animation sequences of photo-realistic, personalized 3D human head avatars from spoken audio. To capture the expressive, detailed nature of human heads, including…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Shivangi Aneja , Artem Sevastopolsky , Tobias Kirschstein , Justus Thies , Angela Dai , Matthias Nießner

While deep learning technologies are now capable of generating realistic images confusing humans, the research efforts are turning to the synthesis of images for more concrete and application-specific purposes. Facial image generation based…

计算机视觉与模式识别 · 计算机科学 2020-06-11 Yeqi Bai , Tao Ma , Lipo Wang , Zhenjie Zhang

Recovering the spatial layout of the cameras and the geometry of the scene from extreme-view images is a longstanding challenge in computer vision. Prevailing 3D reconstruction algorithms often adopt the image matching paradigm and presume…

计算机视觉与模式识别 · 计算机科学 2022-06-17 Wei-Chiu Ma , Anqi Joyce Yang , Shenlong Wang , Raquel Urtasun , Antonio Torralba

We introduce FaceGPT, a self-supervised learning framework for Large Vision-Language Models (VLMs) to reason about 3D human faces from images and text. Typical 3D face reconstruction methods are specialized algorithms that lack semantic…

计算机视觉与模式识别 · 计算机科学 2024-06-12 Haoran Wang , Mohit Mendiratta , Christian Theobalt , Adam Kortylewski

Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a textdependent…

计算机视觉与模式识别 · 计算机科学 2020-10-14 Jianrong Wang , Tong Wu , Shanyu Wang , Mei Yu , Qiang Fang , Ju Zhang , Li Liu

The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model was not trained. Our approach consists of separate uni-modal…

声音 · 计算机科学 2025-12-05 Christopher Simic , Korbinian Riedhammer , Tobias Bocklet

This paper presents a novel task, zero-shot voice conversion based on face images (zero-shot FaceVC), which aims at converting the voice characteristics of an utterance from any source speaker to a newly coming target speaker, solely…

声音 · 计算机科学 2023-09-19 Zheng-Yan Sheng , Yang Ai , Yan-Nian Chen , Zhen-Hua Ling

We present a novel approach for synthesizing 3D facial motions from audio sequences using key motion embeddings. Despite recent advancements in data-driven techniques, accurately mapping between audio signals and 3D facial meshes remains…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Zhihao Xu , Shengjie Gong , Jiapeng Tang , Lingyu Liang , Yining Huang , Haojie Li , Shuangping Huang

We present to recover the complete 3D facial geometry from a single depth view by proposing an Attention Guided Generative Adversarial Networks (AGGAN). In contrast to existing work which normally requires two or more depth views to recover…

计算机视觉与模式识别 · 计算机科学 2020-09-03 Xiaoxu Cai , Hui Yu , Jianwen Lou , Xuguang Zhang , Gongfa Li , Junyu Dong

Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthesis models currently rely on annotated audio data, but it is…

音频与语音处理 · 电气工程与系统科学 2023-05-31 Rongjie Huang , Chunlei Zhang , Yongqi Wang , Dongchao Yang , Luping Liu , Zhenhui Ye , Ziyue Jiang , Chao Weng , Zhou Zhao , Dong Yu

Face reenactment is challenging due to the need to establish dense correspondence between various face representations for motion transfer. Recent studies have utilized Neural Radiance Field (NeRF) as fundamental representation, which…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Songlin Yang , Wei Wang , Yushi Lan , Xiangyu Fan , Bo Peng , Lei Yang , Jing Dong

We present 3DiFACE, a novel method for personalized speech-driven 3D facial animation and editing. While existing methods deterministically predict facial animations from speech, they overlook the inherent one-to-many relationship between…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Balamurugan Thambiraja , Sadegh Aliakbarian , Darren Cosker , Justus Thies

Recently, 3D face reconstruction and face alignment tasks are gradually combined into one task: 3D dense face alignment. Its goal is to reconstruct the 3D geometric structure of face with pose information. In this paper, we propose a graph…

计算机视觉与模式识别 · 计算机科学 2019-04-12 Huawei Wei , Shuang Liang , Yichen Wei

In this work, we address the challenge of generalizable audio deepfake detection (ADD) across diverse speech synthesis paradigms-including conventional text-to-speech (TTS) systems and modern diffusion or flow-matching (FM) based…

音频与语音处理 · 电气工程与系统科学 2025-11-17 Farhan Sheth , Girish , Mohd Mujtaba Akhtar , Muskaan Singh

We present an unsupervised approach for learning to estimate three dimensional (3D) facial structure from a single image while also predicting 3D viewpoint transformations that match a desired pose and facial geometry. We achieve this by…

计算机视觉与模式识别 · 计算机科学 2018-12-27 Joel Ruben Antony Moniz , Christopher Beckham , Simon Rajotte , Sina Honari , Christopher Pal

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Jiadong Liang , Feng Lu

Reconstructing 3D humans from images captured at multiple perspectives typically requires pre-calibration, like using checkerboards or MVS algorithms, which limits scalability and applicability in diverse real-world scenarios. In this work,…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Xiaozhen Qiao , Wenjia Wang , Zhiyuan Zhao , Jiacheng Sun , Ping Luo , Hongyuan Zhang , Xuelong Li

Recently, a lot of attention has been focused on the incorporation of 3D data into face analysis and its applications. Despite providing a more accurate representation of the face, 3D facial images are more complex to acquire than 2D…

计算机视觉与模式识别 · 计算机科学 2021-03-01 Araceli Morales , Gemma Piella , Federico M. Sukno

Face frontalization consists of synthesizing a frontally-viewed face from an arbitrarily-viewed one. The main contribution of this paper is a robust face alignment method that enables pixel-to-pixel warping. The method simultaneously…

计算机视觉与模式识别 · 计算机科学 2021-03-11 Zhiqi Kang , Mostafa Sadeghi , Radu Horaud

This paper introduces a novel pipeline to reconstruct the geometry of interacting multi-person in clothing on a globally coherent scene space from a single image. The main challenge arises from the occlusion: a part of a human body is not…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Junuk Cha , Hansol Lee , Jaewon Kim , Nhat Nguyen Bao Truong , Jae Shin Yoon , Seungryul Baek