English
Related papers

Related papers: Structured 3D Latents for Scalable and Versatile 3…

200 papers

We propose a framework to learn a structured latent space to represent 4D human body motion, where each latent vector encodes a full motion of the whole 3D human shape. On one hand several data-driven skeletal animation models exist…

Computer Vision and Pattern Recognition · Computer Science 2022-09-02 Mathieu Marsot , Stefanie Wuhrer , Jean-Sebastien Franco , Stephane Durocher

We have introduced SegSplat, a novel framework designed to bridge the gap between rapid, feed-forward 3D reconstruction and rich, open-vocabulary semantic understanding. By constructing a compact semantic memory bank from multi-view 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Peter Siegel , Federico Tombari , Marc Pollefeys , Daniel Barath

Creating interactive digital environments for gaming, robotics, and simulation relies on articulated 3D objects whose functionality emerges from their part geometry and kinematic structure. However, existing approaches remain fundamentally…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Penghao Wang , Siyuan Xie , Hongyu Yan , Xianghui Yang , Jingwei Huang , Chunchao Guo , Jiayuan Gu

Despite the remarkable developments achieved by recent 3D generation works, scaling these methods to geographic extents, such as modeling thousands of square kilometers of Earth's surface, remains an open challenge. We address this through…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Shang Liu , Chenjie Cao , Chaohui Yu , Wen Qian , Jing Wang , Fan Wang

Understanding and generating 3D objects as compositions of meaningful parts is fundamental to human perception and reasoning. However, most text-to-3D methods overlook the semantic and functional structure of parts. While recent part-aware…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Tianjiao Yu , Xinzhuo Li , Muntasir Wahed , Jerry Xiong , Yifan Shen , Ying Shen , Ismini Lourentzou

Rigged 3D assets are fundamental to 3D deformation and animation. However, existing 3D generation methods face challenges in generating animatable geometry, while rigging techniques lack fine-grained structural control over skeleton…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Ruisi Zhao , Haoren Zheng , Zongxin Yang , Hehe Fan , Yi Yang

Structured outputs are essential for large language models (LLMs) in critical applications like agents and information extraction. Despite their capabilities, LLMs often generate outputs that deviate from predefined schemas, significantly…

Computation and Language · Computer Science 2025-05-08 Darren Yow-Bang Wang , Zhengyuan Shen , Soumya Smruti Mishra , Zhichao Xu , Yifei Teng , Haibo Ding

We introduce Elastic Looped Transformers (ELT), a highly parameter-efficient class of visual generative models based on a recurrent transformer architecture. While conventional generative models rely on deep stacks of unique transformer…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Sahil Goyal , Swayam Agrawal , Gautham Govind Anil , Prateek Jain , Sujoy Paul , Aditya Kusupati

With the advancement of computer vision, the recently emerged 3D Gaussian Splatting (3DGS) has increasingly become a popular scene reconstruction algorithm due to its outstanding performance. Distributed 3DGS can efficiently utilize edge…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Haosong Peng , Tianyu Qi , Yufeng Zhan , Hao Li , Yalun Dai , Yuanqing Xia

Recently, the multi-modal fusion of RGB, depth, and semantics has shown great potential in dense Simultaneous Localization and Mapping (SLAM). However, a prerequisite for generating consistent semantic maps is the availability of dense,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Linfei Li , Lin Zhang , Zhong Wang , Ying Shen

Conventional geometry-based SLAM systems lack dense 3D reconstruction capabilities since their data association usually relies on feature correspondences. Additionally, learning-based SLAM systems often fall short in terms of real-time…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Zhongche Qu , Zhi Zhang , Cong Liu , Jianhua Yin

This paper presents a novel method for building scalable 3D generative models utilizing pre-trained video diffusion models. The primary obstacle in developing foundation 3D generative models is the limited availability of 3D data. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Junlin Han , Filippos Kokkinos , Philip Torr

We present a novel framework for dynamic 3D scene reconstruction that integrates three key components: an explicit tri-plane deformation field, a view-conditioned canonical radiance field with spherical harmonics (SH) attention, and a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Asrar Alruwayqi

We propose L3DG, the first approach for generative 3D modeling of 3D Gaussians through a latent 3D Gaussian diffusion formulation. This enables effective generative 3D modeling, scaling to generation of entire room-scale scenes which can be…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Barbara Roessle , Norman Müller , Lorenzo Porzi , Samuel Rota Bulò , Peter Kontschieder , Angela Dai , Matthias Nießner

We present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Zibo Zhao , Wen Liu , Xin Chen , Xianfang Zeng , Rui Wang , Pei Cheng , Bin Fu , Tao Chen , Gang Yu , Shenghua Gao

Traditional Simultaneous Localization and Mapping (SLAM) systems often face limitations including coarse rendering quality, insufficient recovery of scene details, and poor robustness in dynamic environments. 3D Gaussian Splatting (3DGS),…

Robotics · Computer Science 2026-02-05 Li Wang , Ruixuan Gong , Yumo Han , Lei Yang , Lu Yang , Ying Li , Bin Xu , Huaping Liu , Rong Fu

We propose the Variational Shape Learner (VSL), a generative model that learns the underlying structure of voxelized 3D shapes in an unsupervised fashion. Through the use of skip-connections, our model can successfully learn and infer a…

Computer Vision and Pattern Recognition · Computer Science 2018-08-07 Shikun Liu , C. Lee Giles , Alexander G. Ororbia

We propose a novel Auto-Regressive (AR) image generation approach that models images as hierarchical compositions of interpretable visual layers. While AR models have achieved transformative success in language modeling, replicating this…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Siddharth Roheda , Rohit Chowdhury , Aniruddha Bala , Rohan Jaiswal

A major obstacle to establishing reliable structure-property (SP) linkages in materials engineering is the scarcity of diverse 3D microstructure datasets. Limited dataset availability and insufficient control over the analysis and design…

Materials Science · Physics 2025-11-14 Kang-Hyun Lee , Faez Ahmed

3D generative models have been recently successful in generating realistic 3D objects in the form of point clouds. However, most models do not offer controllability to manipulate the shape semantics of component object parts without…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Amaya Dharmasiri , Dinithi Dissanayake , Mohamed Afham , Isuru Dissanayake , Ranga Rodrigo , Kanchana Thilakarathna