English
Related papers

Related papers: From Part to Whole: 3D Generative World Model with…

200 papers

3D generation has witnessed significant advancements, yet efficiently producing high-quality 3D assets from a single image remains challenging. In this paper, we present a triplane autoencoder, which encodes 3D models into a compact…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Bowen Zhang , Tianyu Yang , Yu Li , Lei Zhang , Xi Zhao

Inspired by recent findings that generative diffusion models learn semantically meaningful representations, we use them to discover the intrinsic hierarchical structure in biomedical 3D images using unsupervised segmentation. We show that…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Nurislam Tursynbek , Marc Niethammer

A world model is an internal model that simulates how the world evolves. Given past observations and actions, it predicts the future physical state of both the embodied agent and its environment. Accurate world models are essential for…

Machine Learning · Computer Science 2026-04-22 Zaishuo Xia , Yukuan Lu , Xinyi Li , Yifan Xu , Yubei Chen

Recent advances in 2D image generation have achieved remarkable quality,largely driven by the capacity of diffusion models and the availability of large-scale datasets. However, direct 3D generation is still constrained by the scarcity and…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Xuyi Meng , Chen Wang , Jiahui Lei , Kostas Daniilidis , Jiatao Gu , Lingjie Liu

We propose a new representation for encoding 3D shapes as neural fields. The representation is designed to be compatible with the transformer architecture and to benefit both shape reconstruction and shape generation. Existing works on…

Computer Vision and Pattern Recognition · Computer Science 2022-10-19 Biao Zhang , Matthias Nießner , Peter Wonka

This research aims to study a self-supervised 3D clothing reconstruction method, which recovers the geometry shape and texture of human clothing from a single image. Compared with existing methods, we observe that three primary challenges…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Zhedong Zheng , Jiayin Zhu , Wei Ji , Yi Yang , Tat-Seng Chua

We present a learning framework for recovering the 3D shape, camera, and texture of an object from a single image. The shape is represented as a deformable 3D mesh model of an object category where a shape is parameterized by a learned mean…

Computer Vision and Pattern Recognition · Computer Science 2018-08-01 Angjoo Kanazawa , Shubham Tulsiani , Alexei A. Efros , Jitendra Malik

We study the problem of single-image 3D object reconstruction. Recent works have diverged into two directions: regression-based modeling and generative modeling. Regression methods efficiently infer visible surfaces, but struggle with…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Zixuan Huang , Mark Boss , Aaryaman Vasishta , James M. Rehg , Varun Jampani

Correctly capturing the symmetry transformations of data can lead to efficient models with strong generalization capabilities, though methods incorporating symmetries often require prior knowledge. While recent advancements have been made…

Sequential assembly with geometric primitives has drawn attention in robotics and 3D vision since it yields a practical blueprint to construct a target shape. However, due to its combinatorial property, a greedy method falls short of…

Computer Vision and Pattern Recognition · Computer Science 2020-11-26 Jungtaek Kim , Hyunsoo Chung , Jinhwi Lee , Minsu Cho , Jaesik Park

We present a novel framework for training 3D image-conditioned diffusion models using only 2D supervision. Recovering 3D structure from 2D images is inherently ill-posed due to the ambiguity of possible reconstructions, making generative…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Chensheng Peng , Ido Sobol , Masayoshi Tomizuka , Kurt Keutzer , Chenfeng Xu , Or Litany

World models represent a paradigm shift in generative AI, pursuing predictive understanding and controllable simulation of environments in a structured and generalizable way. We present World Machine, a generative world-modeling…

Detecting and localizing glass in 3D environments poses significant challenges for visual perception systems, as the optical properties of glass often hinder conventional sensors from accurately distinguishing glass surfaces. The lack of…

Robotics · Computer Science 2025-09-09 Kai Zhang , Guoyang Zhao , Jianxing Shi , Bonan Liu , Weiqing Qi , Jun Ma

Autonomous assembly of objects is an essential task in robotics and 3D computer vision. It has been studied extensively in robotics as a problem of motion planning, actuator control and obstacle avoidance. However, the task of developing a…

Computer Vision and Pattern Recognition · Computer Science 2021-08-02 Abhinav Narayan Harish , Rajendra Nagar , Shanmuganathan Raman

Accurately predicting 3D occupancy grids from visual inputs is critical for autonomous driving, but current discriminative methods struggle with noisy data, incomplete observations, and the complex structures inherent in 3D scenes. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Yunshen Wang , Yicheng Liu , Tianyuan Yuan , Yingshi Liang , Xiuyu Yang , Honggang Zhang , Hang Zhao

A crucial ability of human intelligence is to build up models of individual 3D objects from partial scene observations. Recent works achieve object-centric generation but without the ability to infer the representation, or achieve 3D scene…

Machine Learning · Computer Science 2021-07-05 Chang Chen , Fei Deng , Sungjin Ahn

Articulated objects are central to interactive 3D applications, including embodied AI, robotics, and VR/AR, where functional part decomposition and kinematic motion are essential. Yet producing high-fidelity articulated assets remains…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Qingming Liu , Xinyue Yao , Shuyuan Zhang , Yueci Deng , Guiliang Liu , Zhen Liu , Kui Jia

Zero-shot novel view synthesis (NVS) from a single image is an essential problem in 3D object understanding. While recent approaches that leverage pre-trained generative models can synthesize high-quality novel views from in-the-wild…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Jianglong Ye , Peng Wang , Kejie Li , Yichun Shi , Heng Wang

In this work, we present a novel framework built to simplify 3D asset generation for amateur users. To enable interactive generation, our method supports a variety of input modalities that can be easily provided by a human, including…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Yen-Chi Cheng , Hsin-Ying Lee , Sergey Tulyakov , Alexander Schwing , Liangyan Gui

Given a single image of a target object, image-to-3D generation aims to reconstruct its texture and geometric shape. Recent methods often utilize intermediate media, such as multi-view images or videos, to bridge the gap between input image…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Jiacheng Wang , Zhedong Zheng , Wei Xu , Ping Liu