English
Related papers

Related papers: ShapeClipper: Scalable 3D Shape Learning from Sing…

200 papers

Multimodal models are becoming increasingly effective, in part due to unified components, such as the Transformer architecture. However, multimodal models still often consist of many task- and modality-specific pieces and training…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Michael Tschannen , Basil Mustafa , Neil Houlsby

3D Morphable Model (3DMM) fitting has widely benefited face analysis due to its strong 3D priori. However, previous reconstructed 3D faces suffer from degraded visual verisimilitude due to the loss of fine-grained geometry, which is…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Xiangyu Zhu , Chang Yu , Di Huang , Zhen Lei , Hao Wang , Stan Z. Li

We propose a novel framework to reconstruct super-resolution human shape from a single low-resolution input image. The approach overcomes limitations of existing approaches that reconstruct 3D human shape from a single image, which require…

Computer Vision and Pattern Recognition · Computer Science 2022-08-24 Marco Pesavento , Marco Volino , Adrian Hilton

We study how to synthesize novel views of human body from a single image. Though recent deep learning based methods work well for rigid objects, they often fail on objects with large articulation, like human bodies. The core step of…

Computer Vision and Pattern Recognition · Computer Science 2018-04-13 Hao Zhu , Hao Su , Peng Wang , Xun Cao , Ruigang Yang

3D shape captioning is a challenging application in 3D shape understanding. Captions from recent multi-view based methods reveal that they cannot capture part-level characteristics of 3D shapes. This leads to a lack of detailed part-level…

Computer Vision and Pattern Recognition · Computer Science 2019-08-02 Zhizhong Han , Chao Chen , Yu-Shen Liu , Matthias Zwicker

Semantic aware reconstruction is more advantageous than geometric-only reconstruction for future robotic and AR/VR applications because it represents not only where things are, but also what things are. Object-centric mapping is a task to…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Kejie Li , Hamid Rezatofighi , Ian Reid

In this paper, we present a new text-guided 3D shape generation approach DreamStone that uses images as a stepping stone to bridge the gap between text and shape modalities for generating 3D shapes without requiring paired text and 3D data.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Zhengzhe Liu , Peng Dai , Ruihui Li , Xiaojuan Qi , Chi-Wing Fu

Understanding 3D scenes from a single image is fundamental to a wide variety of tasks, such as for robotics, motion planning, or augmented reality. Existing works in 3D perception from a single RGB image tend to focus on geometric…

Computer Vision and Pattern Recognition · Computer Science 2022-05-17 Manuel Dahnert , Ji Hou , Matthias Nießner , Angela Dai

Despite the recent success of image-text contrastive models like CLIP and SigLIP, these models often struggle with vision-centric tasks that demand high-fidelity image understanding, such as counting, depth estimation, and fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Zineng Tang , Long Lian , Seun Eisape , XuDong Wang , Roei Herzig , Adam Yala , Alane Suhr , Trevor Darrell , David M. Chan

Spatial understanding remains a key challenge in vision-language models. Yet it is still unclear whether such understanding is truly acquired, and if so, through what mechanisms. We present a controllable 1D image-text testbed to probe how…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Takaki Yamamoto , Chihiro Noguchi , Toshihiro Tanizawa

Existing single view, 3D face reconstruction methods can produce beautifully detailed 3D results, but typically only for near frontal, unobstructed viewpoints. We describe a system designed to provide detailed 3D reconstructions of faces…

Computer Vision and Pattern Recognition · Computer Science 2018-04-02 Anh Tuan Tran , Tal Hassner , Iacopo Masi , Eran Paz , Yuval Nirkin , Gerard Medioni

Most 3D face reconstruction methods rely on 3D morphable models, which disentangle the space of facial deformations into identity geometry, expressions and skin reflectance. These models are typically learned from a limited number of 3D…

Computer Vision and Pattern Recognition · Computer Science 2020-10-06 Mallikarjun B R , Ayush Tewari , Hans-Peter Seidel , Mohamed Elgharib , Christian Theobalt

Dynamic scene reconstruction from casual videos has seen recent remarkable progress. Numerous approaches have attempted to overcome the ill-posedness of the task by distilling priors from 2D foundational models and by imposing hand-crafted…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Narek Tumanyan , Samuel Rota Bulò , Denis Rozumny , Lorenzo Porzi , Adam Harley , Tali Dekel , Peter Kontschieder , Jonathon Luiten

Large-scale pre-trained models have demonstrated impressive performance in vision and language tasks within open-world scenarios. Due to the lack of comparable pre-trained models for 3D shapes, recent methods utilize language-image…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Dan Song , Xinwei Fu , Ning Liu , Weizhi Nie , Wenhui Li , Lanjun Wang , You Yang , Anan Liu

Accurately predicting the 3D shape of any arbitrary object in any pose from a single image is a key goal of computer vision research. This is challenging as it requires a model to learn a representation that can infer both the visible and…

Computer Vision and Pattern Recognition · Computer Science 2021-09-03 Anh Thai , Stefan Stojanov , Vijay Upadhya , James M. Rehg

We present Distill CLIP (DCLIP), a fine-tuned variant of the CLIP model that enhances multimodal image-text retrieval while preserving the original model's strong zero-shot classification capabilities. CLIP models are typically constrained…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Daniel Csizmadia , Andrei Codreanu , Victor Sim , Vighnesh Prabhu , Michael Lu , Kevin Zhu , Sean O'Brien , Vasu Sharma

Given a single image of a general object such as a chair, could we also restore its articulated 3D shape similar to human modeling, so as to animate its plausible articulations and diverse motions? This is an interesting new question that…

Computer Vision and Pattern Recognition · Computer Science 2022-07-07 Ji Yang , Xinxin Zuo , Sen Wang , Zhenbo Yu , Xingyu Li , Bingbing Ni , Minglun Gong , Li Cheng

Monocular 3D reconstruction of articulated object categories is challenging due to the lack of training data and the inherent ill-posedness of the problem. In this work we use video self-supervision, forcing the consistency of consecutive…

Computer Vision and Pattern Recognition · Computer Science 2021-04-28 Filippos Kokkinos , Iasonas Kokkinos

Estimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture, lighting, etc. We propose ARTIC3D, a self-supervised framework to reconstruct…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Chun-Han Yao , Amit Raj , Wei-Chih Hung , Yuanzhen Li , Michael Rubinstein , Ming-Hsuan Yang , Varun Jampani

Zero-shot 3D Anomaly Detection is an emerging task that aims to detect anomalies in a target dataset without any target training data, which is particularly important in scenarios constrained by sample scarcity and data privacy concerns.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Zehao Deng , An Liu , Yan Wang
‹ Prev 1 8 9 10 Next ›