中文
相关论文

相关论文: Controlled AutoEncoders to Generate Faces from Voi…

200 篇论文

Despite rapid progress in text-to-speech (TTS), open-source systems still lack truly instruction-following, fine-grained control over core speech attributes (e.g., pitch, speaking rate, age, emotion, and style). We present VoiceSculptor, an…

Audio-driven facial reenactment is a crucial technique that has a range of applications in film-making, virtual avatars and video conferences. Existing works either employ explicit intermediate face representations (e.g., 2D facial…

计算机视觉与模式识别 · 计算机科学 2023-06-14 Ricong Huang , Peiwen Lai , Yipeng Qin , Guanbin Li

Previous works on voice-face matching and voice-guided face synthesis demonstrate strong correlations between voice and face, but mainly rely on coarse semantic cues such as gender, age, and emotion. In this paper, we aim to investigate the…

计算机视觉与模式识别 · 计算机科学 2023-07-27 Xiang Li , Yandong Wen , Muqiao Yang , Jinglu Wang , Rita Singh , Bhiksha Raj

Generating face image with specific gaze information has attracted considerable attention. Existing approaches typically input gaze values directly for face generation, which is unnatural and requires annotated gaze datasets for training,…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Hengfei Wang , Zhongqun Zhang , Yihua Cheng , Hyung Jin Chang

How much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive language disorders.…

声音 · 计算机科学 2025-02-04 Jiaxin Ye , Boyuan Cao , Hongming Shan

Human social behaviors are inherently multimodal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Hugo Bohy , Minh Tran , Kevin El Haddad , Thierry Dutoit , Mohammad Soleymani

Animating human face images aims to synthesize a desired source identity in a natural-looking way mimicking a driving video's facial movements. In this context, Generative Adversarial Networks have demonstrated remarkable potential in…

计算机视觉与模式识别 · 计算机科学 2024-08-26 Alireza Javanmardi , Alain Pagani , Didier Stricker

The great advancements of generative adversarial networks and face recognition models in computer vision have made it possible to swap identities on images from single sources. Although a lot of studies seems to have proposed almost…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Kaede Shiohara , Xingchao Yang , Takafumi Taketomi

In facial image generation, current text-to-image models often suffer from facial attribute leakage and insufficient physical consistency when responding to local semantic instructions. In this study, we propose Face-MakeUpV2, a facial…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Dawei Dai , Yinxiu Zhou , Chenghang Li , Guolai Jiang , Chengfang Zhang

Our ability to sample realistic natural images, particularly faces, has advanced by leaps and bounds in recent years, yet our ability to exert fine-tuned control over the generative process has lagged behind. If this new technology is to…

计算机视觉与模式识别 · 计算机科学 2020-10-20 Marek Kowalski , Stephan J. Garbin , Virginia Estellers , Tadas Baltrušaitis , Matthew Johnson , Jamie Shotton

3D-controllable portrait synthesis has significantly advanced, thanks to breakthroughs in generative adversarial networks (GANs). However, it is still challenging to manipulate existing face images with precise 3D control. While…

计算机视觉与模式识别 · 计算机科学 2022-08-25 Yuchen Liu , Zhixin Shu , Yijun Li , Zhe Lin , Richard Zhang , S. Y. Kung

The task of face attribute manipulation has found increasing applications, but still remains challenging with the requirement of editing the attributes of a face image while preserving its unique details. In this paper, we choose to combine…

计算机视觉与模式识别 · 计算机科学 2021-05-07 Ruoqi Sun , Chen Huang , Jianping Shi , Lizhuang Ma

Image-to-image translation and voice conversion enable the generation of a new facial image and voice while maintaining some of the semantics such as a pose in an image and linguistic content in audio, respectively. They can aid in the…

计算机视觉与模式识别 · 计算机科学 2023-03-02 Naoya Takahashi , Mayank K. Singh , Yuki Mitsufuji

The primary objective of this work is to present an alternative approach aimed at reducing the dependency on labeled data. Our proposed method involves utilizing autoencoder pre-training within a face image recognition task with two step…

计算机视觉与模式识别 · 计算机科学 2024-02-09 Enoch Solomon , Abraham Woubie , Eyael Solomon Emiru

There is a growing demand for the accessible creation of high-quality 3D avatars that are animatable and customizable. Although 3D morphable models provide intuitive control for editing and animation, and robustness for single-view face…

计算机视觉与模式识别 · 计算机科学 2023-05-05 Connor Z. Lin , Koki Nagano , Jan Kautz , Eric R. Chan , Umar Iqbal , Leonidas Guibas , Gordon Wetzstein , Sameh Khamis

Neural networks have recently become good at engaging in dialog. However, current approaches are based solely on verbal text, lacking the richness of a real face-to-face conversation. We propose a neural conversation model that aims to read…

计算机视觉与模式识别 · 计算机科学 2018-12-05 Hang Chu , Daiqing Li , Sanja Fidler

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies…

多媒体 · 计算机科学 2023-05-25 Zheng-Yan Sheng , Yang Ai , Zhen-Hua Ling

Speech-to-face generation is an intriguing area of research that focuses on generating realistic facial images based on a speaker's audio speech. However, state-of-the-art methods employing GAN-based architectures lack stability and cannot…

计算机视觉与模式识别 · 计算机科学 2023-10-06 Jinting Wang , Li Liu , Jun Wang , Hei Victor Cheng

Learned 3D representations of human faces are useful for computer vision problems such as 3D face tracking and reconstruction from images, as well as graphics applications such as character generation and animation. Traditional models learn…

计算机视觉与模式识别 · 计算机科学 2018-08-02 Anurag Ranjan , Timo Bolkart , Soubhik Sanyal , Michael J. Black

In this paper, we present a method for reprogramming pre-trained audio-driven talking face synthesis models to operate in a text-driven manner. Consequently, we can easily generate face videos that articulate the provided textual sentences,…

图形学 · 计算机科学 2024-01-19 Jeongsoo Choi , Minsu Kim , Se Jin Park , Yong Man Ro