English
Related papers

Related papers: AssetFormer: Modular 3D Assets Generation with Aut…

200 papers

Driven by powerful image diffusion models, recent research has achieved the automatic creation of 3D objects from textual or visual guidance. By performing score distillation sampling (SDS) iteratively across different views, these methods…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Zeyu Li , Ruitong Gan , Chuanchen Luo , Yuxi Wang , Jiaheng Liu , Ziwei Zhu Man Zhang , Qing Li , Xucheng Yin , Zhaoxiang Zhang , Junran Peng

Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before real-world deployment. Neural scene reconstruction converts driving logs into interactive…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Tianshi Cao , Jiawei Ren , Yuxuan Zhang , Jaewoo Seo , Jiahui Huang , Shikhar Solanki , Haotian Zhang , Mingfei Guo , Haithem Turki , Muxingzi Li , Yue Zhu , Sipeng Zhang , Zan Gojcic , Sanja Fidler , Kangxue Yin

Structured layouts are preferable in many 2D visual contents (\eg, GUIs, webpages) since the structural information allows convenient layout editing. Computational frameworks can help create structured layouts but require heavy labor input.…

Graphics · Computer Science 2025-10-31 Xin Hu , Pengfei Xu , Jin Zhou , Hongbo Fu , Hui Huang

The creation of manufacturable and editable 3D shapes through Computer-Aided Design (CAD) remains a highly manual and time-consuming task, hampered by the complex topology of boundary representations of 3D solids and unintuitive design…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Md Ferdous Alam , Faez Ahmed

Autoregressive models have driven remarkable progress in language modeling. Their foundational reliance on discrete tokens, unidirectional context, and single-pass decoding, while central to their success, also inspires the exploration of a…

Articulated objects are fundamental for robotics, simulation of physics, and interactive virtual environments. However, reconstructing them from visual input remains challenging, as it requires jointly inferring both part geometry and…

Robotics · Computer Science 2026-03-17 Zhuangzhe Wu , Yue Xin , Chengkai Hou , Minghao Chen , Yaoxu Lyu , Jieyu Zhang , Shanghang Zhang

Autoregressive video models offer distinct advantages over bidirectional diffusion models in creating interactive video content and supporting streaming applications with arbitrary duration. In this work, we present Next-Frame Diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Xinle Cheng , Tianyu He , Jiayi Xu , Junliang Guo , Di He , Jiang Bian

Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D compositional objects from single images--particularly under occlusions--remains challenging. Existing methods often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Hui Shan , Keyang Luo , Ming Li , Sizhe Zheng , Yanwei Fu , Zhen Chen , Xiangru Huang

Digital modeling and reconstruction of human faces serve various applications. However, its availability is often hindered by the requirements of data capturing devices, manual labor, and suitable actors. This situation restricts the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Yunxuan Cai , Sitao Xiang , Zongjian Li , Haiwei Chen , Yajie Zhao

Recent advancements in text-to-image (T2I) generation have spurred the development of text-to-3D asset (T23DA) generation, leveraging pretrained 2D text-to-image diffusion models for text-to-3D asset synthesis. Despite the growing…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Kang Fu , Huiyu Duan , Zicheng Zhang , Xiaohong Liu , Xiongkuo Min , Jia Wang , Guangtao Zhai

We introduce Infinigen-Articulated, a toolkit for generating realistic, procedurally generated articulated assets for robotics simulation. We include procedural generators for 18 common articulated object categories along with high-level…

With the capacity of modeling long-range dependencies in sequential data, transformers have shown remarkable performances in a variety of generative tasks such as image, audio, and text generation. Yet, taming them in generating less…

Computer Vision and Pattern Recognition · Computer Science 2022-04-06 An-Chieh Cheng , Xueting Li , Sifei Liu , Min Sun , Ming-Hsuan Yang

We introduce Meta 3D Gen (3DGen), a new state-of-the-art, fast pipeline for text-to-3D asset generation. 3DGen offers 3D asset creation with high prompt fidelity and high-quality 3D shapes and textures in under a minute. It supports…

Recent text-to-scene generation approaches largely reduced the manual efforts required to create 3D scenes. However, their focus is either to generate a scene layout or to generate objects, and few generate both. The generated scene layout…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Zhenggang Tang , Yuehao Wang , Yuchen Fan , Jun-Kun Chen , Yu-Ying Yeh , Kihyuk Sohn , Zhangyang Wang , Qixing Huang , Alexander Schwing , Rakesh Ranjan , Dilin Wang , Zhicheng Yan

Understanding dynamics from visual observations is a challenging problem that requires disentangling individual objects from the scene and learning their interactions. While recent object-centric models can successfully decompose a scene…

Computer Vision and Pattern Recognition · Computer Science 2023-01-24 Ziyi Wu , Nikita Dvornik , Klaus Greff , Thomas Kipf , Animesh Garg

3D occupancy, an advanced perception technology for driving scenarios, represents the entire scene without distinguishing between foreground and background by quantifying the physical space into a grid map. The widely adopted…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Jinke Li , Xiao He , Chonghua Zhou , Xiaoqiang Cheng , Yang Wen , Dan Zhang

Recent advances in subject-driven image generation using diffusion models have attracted considerable attention for their remarkable capabilities in producing high-quality images. Nevertheless, the potential of Visual Autoregressive (VAR)…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Xin Jiang , Jingwen Chen , Yehao Li , Yingwei Pan , Kezhou Chen , Zechao Li , Ting Yao , Tao Mei

Recently, 3D Gaussian splatting (3D-GS) has achieved great success in reconstructing and rendering real-world scenes. To transfer the high rendering quality to generation tasks, a series of research works attempt to generate 3D-Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Taoran Yi , Jiemin Fang , Zanwei Zhou , Junjie Wang , Guanjun Wu , Lingxi Xie , Xiaopeng Zhang , Wenyu Liu , Xinggang Wang , Qi Tian

Generative 3D models are deployed in gaming, robotics, and immersive creation, making source attribution critical: given a 3D asset, can we identify whether and which generative model created it? This problem faces two core challenges:…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Sihan Ma , Siyuan Liang , Dacheng Tao

3D world models (i.e., learning-based 3D dynamics models) offer a promising approach to generalizable robotic manipulation by capturing the underlying physics of environment evolution conditioned on robot actions. However, existing 3D world…

Robotics · Computer Science 2025-08-27 Suning Huang , Qianzhong Chen , Xiaohan Zhang , Jiankai Sun , Mac Schwager