中文
相关论文

相关论文: A Neural Virtual Anchor Synthesizer based on Seq2S…

200 篇论文

Current one-stage methods for visual grounding encode the language query as one holistic sentence embedding before fusion with visual feature. Such a formulation does not treat each word of a query sentence on par when modeling language to…

计算机视觉与模式识别 · 计算机科学 2021-08-03 Heng Zhao , Joey Tianyi Zhou , Yew-Soon Ong

Current text-to-avatar methods often rely on implicit representations (e.g., NeRF, SDF, and DMTet), leading to 3D content that artists cannot easily edit and animate in graphics software. This paper introduces a novel framework for…

图形学 · 计算机科学 2025-05-01 Duotun Wang , Hengyu Meng , Zeyu Cai , Zhijing Shao , Qianxi Liu , Lin Wang , Mingming Fan , Xiaohang Zhan , Zeyu Wang

Conversation is an essential component of virtual avatar activities in the metaverse. With the development of natural language processing, textual and vocal conversation generation has achieved a significant breakthrough. However,…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Yichao Yan , Zanwei Zhou , Zi Wang , Jingnan Gao , Xiaokang Yang

Current expressive avatar systems rely heavily on visual cues, failing when faces are occluded or when emotions remain internal. We present Mind-to-Face, the first framework that decodes non-invasive electroencephalogram (EEG) signals…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Haolin Xiong , Tianwen Fu , Pratusha Bhuvana Prasad , Yunxuan Cai , Haiwei Chen , Wenbin Teng , Hanyuan Xiao , Yajie Zhao

Recent advances in neural network -based text-to-speech have reached human level naturalness in synthetic speech. The present sequence-to-sequence models can directly map text to mel-spectrogram acoustic features, which are convenient for…

音频与语音处理 · 电气工程与系统科学 2019-06-27 Lauri Juvela , Bajibabu Bollepalli , Junichi Yamagishi , Paavo Alku

Due to the emergence of Generative Adversarial Networks, video synthesis has witnessed exceptional breakthroughs. However, existing methods lack a proper representation to explicitly control the dynamics in videos. Human pose, on the other…

计算机视觉与模式识别 · 计算机科学 2018-07-31 Ceyuan Yang , Zhe Wang , Xinge Zhu , Chen Huang , Jianping Shi , Dahua Lin

We present a novel learning-based framework for face reenactment. The proposed method, known as ReenactGAN, is capable of transferring facial movements and expressions from monocular video input of an arbitrary person to a target person.…

计算机视觉与模式识别 · 计算机科学 2018-07-31 Wayne Wu , Yunxuan Zhang , Cheng Li , Chen Qian , Chen Change Loy

Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Although prior methods show promising quality, their reliance on the mesh domain limits their ability…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Alexandre Symeonidis-Herzig , Özge Mercanoğlu Sincan , Richard Bowden

We present a novel approach to generating photo-realistic images of a face with accurate lip sync, given an audio input. By using a recurrent neural network, we achieved mouth landmarks based on audio features. We exploited the power of…

计算机视觉与模式识别 · 计算机科学 2018-03-21 Seyed Ali Jalalifar , Hosein Hasani , Hamid Aghajan

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic method to generate…

计算机视觉与模式识别 · 计算机科学 2021-08-29 Xinsheng Wang , Qicong Xie , Jihua Zhu , Lei Xie , Scharenborg

Deepfakes are AI-generated media in which the original content is digitally altered to create convincing but manipulated images, videos, or audio. Among the various types of deepfakes, lip-syncing deepfakes are one of the most challenging…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Soumyya Kanti Datta , Shan Jia , Siwei Lyu

Face sketch synthesis and reputation have wide range of packages in law enforcement. Despite the amazing progresses had been made in faces cartoon and reputation, maximum current researches regard them as separate responsibilities. On this…

计算机视觉与模式识别 · 计算机科学 2023-05-03 Chandana Navuluri , Sandhya Jukanti , Raghupathi Reddy Allapuram

In facial image generation, current text-to-image models often suffer from facial attribute leakage and insufficient physical consistency when responding to local semantic instructions. In this study, we propose Face-MakeUpV2, a facial…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Dawei Dai , Yinxiu Zhou , Chenghang Li , Guolai Jiang , Chengfang Zhang

Nowadays, it is possible to scan faces and automatically register them with high quality. However, the resulting face meshes often need further processing: we need to stabilize them to remove unwanted head movement. Stabilization is…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Jan Bednarik , Erroll Wood , Vasileios Choutas , Timo Bolkart , Daoye Wang , Chenglei Wu , Thabo Beeler

Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Shuling Zhao , Fa-Ting Hong , Xiaoshui Huang , Dan Xu

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Ashkan Taghipour , Morteza Ghahremani , Zinuo Li , Hamid Laga , Farid Boussaid , Mohammed Bennamoun

We present Neural Head Avatars, a novel neural representation that explicitly models the surface geometry and appearance of an animatable human avatar that can be used for teleconferencing in AR/VR or other applications in the movie or…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Philip-William Grassal , Malte Prinzler , Titus Leistner , Carsten Rother , Matthias Nießner , Justus Thies

While head-mounted displays (HMDs) for Virtual Reality (VR) have become widely available in the consumer market, they pose a considerable obstacle for a realistic face-to-face conversation in VR since HMDs hide a significant portion of the…

计算机视觉与模式识别 · 计算机科学 2024-09-20 Philipp Ladwig , Rene Ebertowski , Alexander Pech , Ralf Dörner , Christian Geiger

This paper presents a generative adversarial learning-based human upper body video synthesis approach to generate an upper body video of target person that is consistent with the body motion, face expression, and pose of the person in…

计算机视觉与模式识别 · 计算机科学 2019-09-13 Zhaoxiang Liu , Huan Hu , Zipeng Wang , Kai Wang , Jinqiang Bai , Shiguo Lian

Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of…

计算机视觉与模式识别 · 计算机科学 2022-06-16 Valentin Gabeur , Paul Hongsuck Seo , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid