English
Related papers

Related papers: AssetFormer: Modular 3D Assets Generation with Aut…

200 papers

Speech-driven 3D facial animation is challenging due to the complex geometry of human faces and the limited availability of 3D audio-visual data. Prior works typically focus on learning phoneme-level features of short audio windows with…

Computer Vision and Pattern Recognition · Computer Science 2022-03-18 Yingruo Fan , Zhaojiang Lin , Jun Saito , Wenping Wang , Taku Komura

This paper addresses the challenge of text-conditioned streaming motion generation, which requires us to predict the next-step human pose based on variable-length historical motions and incoming texts. Existing methods struggle to achieve…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Lixing Xiao , Shunlin Lu , Huaijin Pi , Ke Fan , Liang Pan , Yueer Zhou , Ziyong Feng , Xiaowei Zhou , Sida Peng , Jingbo Wang

Dynamic scene reconstruction is a long-term challenge in 3D vision. Recent methods extend 3D Gaussian Splatting to dynamic scenes via additional deformation fields and apply explicit constraints like motion flow to guide the deformation.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 DaDong Jiang , Zhihui Ke , Xiaobo Zhou , Zhi Hou , Xianghui Yang , Wenbo Hu , Tie Qiu , Chunchao Guo

This paper focuses on the task of speech-driven 3D facial animation, which aims to generate realistic and synchronized facial motions driven by speech inputs. Recent methods have employed audio-conditioned diffusion models for 3D facial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yifan Yang , Zhi Cen , Sida Peng , Xiangwei Chen , Yifu Deng , Xinyu Zhu , Fan Jia , Xiaowei Zhou , Hujun Bao

Current text-to-3D methods excel at generating single objects but falter on compositional prompts. We argue this failure is fundamental to their optimization schedules, as simultaneous or iterative heuristics predictably collapse under a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Utkarsh Nath , Rajeev Goel , Rahul Khurana , Kyle Min , Mark Ollila , Pavan Turaga , Varun Jampani , Tejaswi Gowda

Real-world data is high-dimensional: a book, image, or musical performance can easily contain hundreds of thousands of elements even after compression. However, the most commonly used autoregressive models, Transformers, are prohibitively…

Autoregressive transformers have recently shown impressive image generation quality and efficiency on par with state-of-the-art diffusion models. Unlike diffusion architectures, autoregressive models can naturally incorporate arbitrary…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yixiao Chen , Zhiyuan Ma , Guoli Jia , Che Jiang , Jianjun Li , Bowen Zhou

We present a GAN-based Transformer for general action-conditioned 3D human motion generation, including not only single-person actions but also multi-person interactive actions. Our approach consists of a powerful Action-conditioned motion…

Computer Vision and Pattern Recognition · Computer Science 2022-11-24 Liang Xu , Ziyang Song , Dongliang Wang , Jing Su , Zhicheng Fang , Chenjing Ding , Weihao Gan , Yichao Yan , Xin Jin , Xiaokang Yang , Wenjun Zeng , Wei Wu

While generative artificial intelligence has advanced significantly across text, image, audio, and video domains, 3D generation remains comparatively underdeveloped due to fundamental challenges such as data scarcity, algorithmic…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Weiyu Li , Xuanyang Zhang , Zheng Sun , Di Qi , Hao Li , Wei Cheng , Weiwei Cai , Shihao Wu , Jiarui Liu , Zihao Wang , Xiao Chen , Feipeng Tian , Jianxiong Pan , Zeming Li , Gang Yu , Xiangyu Zhang , Daxin Jiang , Ping Tan

Generative modeling of high-dimensional data is a key problem in machine learning. Successful approaches include latent variable models and autoregressive models. The complementary strengths of these approaches, to model global and local…

Computer Vision and Pattern Recognition · Computer Science 2019-04-19 Thomas Lucas , Jakob Verbeek

Recently, generating 3D assets with the control of condition images has achieved impressive quality. However, existing 3D generation methods are limited to handling a single control objective and lack the ability to utilize multiple images…

Graphics · Computer Science 2026-02-20 Xuancheng Jin , Rengan Xie , Wenting Zheng , Rui Wang , Hujun Bao , Yuchi Huo

3D asset generation is getting massive amounts of attention, inspired by the recent success of text-guided 2D content creation. Existing text-to-3D methods use pretrained text-to-image diffusion models in an optimization problem or…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Lukas Höllein , Aljaž Božič , Norman Müller , David Novotny , Hung-Yu Tseng , Christian Richardt , Michael Zollhöfer , Matthias Nießner

Removing modeling constraints and unifying architectures across domains has been a key driver of the recent progress in training large multimodal models. However, most of these models still rely on many separately trained components such as…

Machine Learning · Computer Science 2025-05-20 Michael Tschannen , André Susano Pinto , Alexander Kolesnikov

Developing embodied AI agents requires scalable training environments that balance content diversity with physics accuracy. World simulators provide such environments but face distinct limitations: video-based methods generate diverse…

Recent progress in image-to-3D has opened up immense possibilities for design, AR/VR, and robotics. However, to use AI-generated 3D assets in real applications, a critical requirement is the capability to edit them easily. We present a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Ziqi Ma , Hongqiao Chen , Yisong Yue , Georgia Gkioxari

3D modeling is moving from virtual to physical. Existing 3D generation primarily emphasizes geometries and textures while neglecting physical-grounded modeling. Consequently, despite the rapid development of 3D generative models, the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Ziang Cao , Zhaoxi Chen , Liang Pan , Ziwei Liu

While recent advancements have shown remarkable progress in general 3D shape generation models, the challenge of leveraging these approaches to automatically generate wearable 3D assets remains unexplored. To this end, we present BAG, a…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Zhongjin Luo , Yang Li , Mingrui Zhang , Senbo Wang , Han Yan , Xibin Song , Taizhang Shang , Wei Mao , Hongdong Li , Xiaoguang Han , Pan Ji

We present Meta 3D AssetGen (AssetGen), a significant advancement in text-to-3D generation which produces faithful, high-quality meshes with texture and material control. Compared to works that bake shading in the 3D object's appearance,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Yawar Siddiqui , Tom Monnier , Filippos Kokkinos , Mahendra Kariya , Yanir Kleiman , Emilien Garreau , Oran Gafni , Natalia Neverova , Andrea Vedaldi , Roman Shapovalov , David Novotny

Recent advances in diffusion models have improved controllable streetscape generation and supported downstream perception and planning tasks. However, challenges remain in accurately modeling driving scenes and generating long videos. To…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Jianbiao Mei , Tao Hu , Xuemeng Yang , Licheng Wen , Yu Yang , Tiantian Wei , Yukai Ma , Min Dou , Botian Shi , Yong Liu

Traditional image-to-3D models often struggle with scenes containing multiple objects due to biases and occlusion complexities. To address this challenge, we present REPARO, a novel approach for compositional 3D asset generation from single…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Haonan Han , Rui Yang , Huan Liao , Jiankai Xing , Zunnan Xu , Xiaoming Yu , Junwei Zha , Xiu Li , Wanhua Li