English
Related papers

Related papers: Cross-Modal Perceptionist: Can Face Geometry be Gl…

200 papers

Self-supervised learning is showing great promise for monocular depth estimation, using geometry as the only source of supervision. Depth networks are indeed capable of learning representations that relate visual appearance to 3D properties…

Computer Vision and Pattern Recognition · Computer Science 2020-02-28 Vitor Guizilini , Rui Hou , Jie Li , Rares Ambrus , Adrien Gaidon

We present a method to learn the 3D surface of objects directly from a collection of images. Previous work achieved this capability by exploiting additional manual annotation, such as object pose, 3D surface templates, temporal continuity…

Computer Vision and Pattern Recognition · Computer Science 2018-11-28 Attila Szabó , Paolo Favaro

Perceptual geometry refers to the interdisciplinary research whose objectives focuses on study of geometry from the perspective of visual perception, and in turn, applies such geometric findings to the ecological study of vision. Perceptual…

Neurons and Cognition · Quantitative Biology 2011-11-29 Arash Sangari , Hasti Mirkia , Amir H. Assadi

Recent learning-based approaches, in which models are trained by single-view images have shown promising results for monocular 3D face reconstruction, but they suffer from the ill-posed face pose and depth ambiguity issue. In contrast to…

Computer Vision and Pattern Recognition · Computer Science 2020-07-27 Jiaxiang Shang , Tianwei Shen , Shiwei Li , Lei Zhou , Mingmin Zhen , Tian Fang , Long Quan

What is the interplay between semantic representations learned by language models (LM) from surface form alone to those learned from more grounded evidence? We study this question for a scenario where part of the input comes from a…

Computation and Language · Computer Science 2026-04-23 Tianyang Xu , Marcelo Sandoval-Castaneda , Karen Livescu , Greg Shakhnarovich , Kanishka Misra

While progress in 2D generative models of human appearance has been rapid, many applications require 3D avatars that can be animated and rendered. Unfortunately, most existing methods for learning generative models of 3D humans with diverse…

Computer Vision and Pattern Recognition · Computer Science 2023-05-04 Zijian Dong , Xu Chen , Jinlong Yang , Michael J. Black , Otmar Hilliges , Andreas Geiger

We present an algorithm that learns a coarse 3D representation of objects from unposed multi-view 2D mask supervision, then uses it to generate detailed mask and image texture. In contrast to existing voxel-based methods for unposed object…

Computer Vision and Pattern Recognition · Computer Science 2021-06-25 Youssef A. Mejjati , Isa Milefchik , Aaron Gokaslan , Oliver Wang , Kwang In Kim , James Tompkin

We present a minimalistic but effective neural network that computes dense facial correspondences in highly unconstrained RGB images. Our network learns a per-pixel flow and a matchability mask between 2D input photographs of a person and…

Computer Vision and Pattern Recognition · Computer Science 2017-09-05 Ronald Yu , Shunsuke Saito , Haoxiang Li , Duygu Ceylan , Hao Li

When interacting in a three dimensional world, humans must estimate 3D structure from visual inputs projected down to two dimensional retinal images. It has been shown that humans use the persistence of object shape over motion-induced…

Neurons and Cognition · Quantitative Biology 2023-04-03 Marissa Connor , Bruno Olshausen , Christopher Rozell

Previous animatable 3D-aware GANs for human generation have primarily focused on either the human head or full body. However, head-only videos are relatively uncommon in real life, and full body generation typically does not deal with…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Yue Wu , Sicheng Xu , Jianfeng Xiang , Fangyun Wei , Qifeng Chen , Jiaolong Yang , Xin Tong

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work that studies the interaction between audio and visual…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-19 Abhinav Shukla , Stavros Petridis , Maja Pantic

Neural front-ends represent a promising approach to feature extraction for automatic speech recognition (ASR) systems as they enable to learn specifically tailored features for different tasks. Yet, many of the existing techniques remain…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-15 Peter Vieting , Benedikt Hilmes , Ralf Schlüter , Hermann Ney

Recent years have seen the development of mature solutions for reconstructing deformable surfaces from a single image, provided that they are relatively well-textured. By contrast, recovering the 3D shape of texture-less surfaces remains an…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Jan Bednařík , Pascal Fua , Mathieu Salzmann

Some self-supervised cross-modal learning approaches have recently demonstrated the potential of image signals for enhancing point cloud representation. However, it remains a question on how to directly model cross-modal local and global…

Computer Vision and Pattern Recognition · Computer Science 2022-11-24 Honggu Zhou , Xiaogang Peng , Jiawei Mao , Zizhao Wu , Ming Zeng

Medical vision-and-language pre-training provides a feasible solution to extract effective vision-and-language representations from medical images and texts. However, few studies have been dedicated to this field to facilitate medical…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Zhihong Chen , Yuhao Du , Jinpeng Hu , Yang Liu , Guanbin Li , Xiang Wan , Tsung-Hui Chang

In this paper, we propose a multi-speaker face-to-speech waveform generation model that also works for unseen speaker conditions. Using a generative adversarial network (GAN) with linguistic and speaker characteristic features as auxiliary…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Se-Yun Um , Jihyun Kim , Jihyun Lee , Hong-Goo Kang

Creating realistic 3D head assets for virtual characters that match a precise artistic vision remains labor-intensive. We present a novel framework that streamlines this process by providing artists with intuitive control over generated 3D…

In this paper, we propose a framework for disentangling the appearance and geometry representations in the face recognition task. To provide supervision for this aim, we generate geometrically identical faces by incorporating spatial…

Computer Vision and Pattern Recognition · Computer Science 2020-01-15 Ali Dabouei , Fariborz Taherkhani , Sobhan Soleymani , Jeremy Dawson , Nasser M. Nasrabadi

The appearances of children are inherited from their parents, which makes it feasible to predict them. Predicting realistic children's faces may help settle many social problems, such as age-invariant face recognition, kinship verification,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-22 Yuzhi Zhao , Lai-Man Po , Xuehui Wang , Qiong Yan , Wei Shen , Yujia Zhang , Wei Liu , Chun-Kit Wong , Chiu-Sing Pang , Weifeng Ou , Wing-Yin Yu , Buhua Liu

Accurately reconstructing a 3D scene including explicit geometry information is both attractive and challenging. Geometry reconstruction can benefit from incorporating differentiable appearance models, such as Neural Radiance Fields and 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Ancheng Lin , Yusheng Xiang , Paul Kennedy , Jun Li