English
Related papers

Related papers: AnimateAnyMesh++: A Flexible 4D Foundation Model f…

200 papers

We present SDXL, a latent diffusion model for text-to-image synthesis. Compared to previous versions of Stable Diffusion, SDXL leverages a three times larger UNet backbone: The increase of model parameters is mainly due to more attention…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Dustin Podell , Zion English , Kyle Lacey , Andreas Blattmann , Tim Dockhorn , Jonas Müller , Joe Penna , Robin Rombach

Urban scene generation has been developing rapidly recently. However, existing methods primarily focus on generating static and single-frame scenes, overlooking the inherently dynamic nature of real-world driving environments. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Hengwei Bian , Lingdong Kong , Haozhe Xie , Liang Pan , Yu Qiao , Ziwei Liu

Text-driven motion generation offers a powerful and intuitive way to create human movements directly from natural language. By removing the need for predefined motion inputs, it provides a flexible and accessible approach to controlling…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Ali Rida Sahili , Najett Neji , Hedi Tabia

Diffusion and flow matching models have unlocked unprecedented capabilities for creative content creation, such as interactive image and streaming video generation. The growing demand for higher resolutions, frame rates, and context…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Brian Chao , Lior Yariv , Howard Xiao , Gordon Wetzstein

Mesh is a fundamental representation of 3D assets in various industrial applications, and is widely supported by professional softwares. However, due to its irregular structure, mesh creation and manipulation is often time-consuming and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Zhaoyang Lyu , Ben Fei , Jinyi Wang , Xudong Xu , Ya Zhang , Weidong Yang , Bo Dai

Text-based motion generation models are drawing a surge of interest for their potential for automating the motion-making process in the game, animation, or robot industries. In this paper, we propose a diffusion-based motion synthesis and…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Jihoon Kim , Jiseob Kim , Sungjoon Choi

We present DreamHuman, a method to generate realistic animatable 3D human avatar models solely from textual descriptions. Recent text-to-3D methods have made considerable strides in generation, but are still lacking in important aspects.…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Nikos Kolotouros , Thiemo Alldieck , Andrei Zanfir , Eduard Gabriel Bazavan , Mihai Fieraru , Cristian Sminchisescu

3D modeling is shifting from static visual representations toward physical, articulated assets that can be directly used in simulation and interaction. However, most existing 3D generation methods overlook key physical and articulation…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Ziang Cao , Fangzhou Hong , Zhaoxi Chen , Liang Pan , Ziwei Liu

We introduce TADA, a simple-yet-effective approach that takes textual descriptions and produces expressive 3D avatars with high-quality geometry and lifelike textures, that can be animated and rendered with traditional graphics pipelines.…

Artificial Intelligence · Computer Science 2023-08-22 Tingting Liao , Hongwei Yi , Yuliang Xiu , Jiaxaing Tang , Yangyi Huang , Justus Thies , Michael J. Black

Existing methods for image-to-3D avatar generation struggle to produce highly detailed, animation-ready avatars suitable for real-world applications. We introduce AdaHuman, a novel framework that generates high-fidelity animatable 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Yangyi Huang , Ye Yuan , Xueting Li , Jan Kautz , Umar Iqbal

We present DreamWaltz, a novel framework for generating and animating complex 3D avatars given text guidance and parametric human body prior. While recent methods have shown encouraging results for text-to-3D generation of common objects,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Yukun Huang , Jianan Wang , Ailing Zeng , He Cao , Xianbiao Qi , Yukai Shi , Zheng-Jun Zha , Lei Zhang

This paper presents DriVerse, a generative model for simulating navigation-driven driving scenes from a single image and a future trajectory. Previous autonomous driving world models either directly feed the trajectory or discrete control…

Robotics · Computer Science 2026-04-28 Xiaofan Li , Chenming Wu , Zhao Yang , Zhihao Xu , Dingkang Liang , Yumeng Zhang , Ji Wan , Jun Wang

Vision-Language Models (VLMs) are increasingly tasked with ultra-long multimodal understanding. While linear architectures offer constant computation and memory footprints, they often struggle with high-frequency visual perception compared…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Hongyuan Tao , Bencheng Liao , Shaoyu Chen , Haoran Yin , Qian Zhang , Wenyu Liu , Xinggang Wang

Deep neural networks have enabled improved image quality and fast inference times for various inverse problems, including accelerated magnetic resonance imaging (MRI) reconstruction. However, such models require a large number of…

Image and Video Processing · Electrical Eng. & Systems 2022-06-20 Arjun D Desai , Beliz Gunel , Batu M Ozturkler , Harris Beg , Shreyas Vasanawala , Brian A Hargreaves , Christopher Ré , John M Pauly , Akshay S Chaudhari

In this paper, we introduce MeshGen, an advanced image-to-3D pipeline that generates high-quality 3D meshes with detailed geometry and physically based rendering (PBR) textures. Addressing the challenges faced by existing 3D native…

Graphics · Computer Science 2025-05-09 Zilong Chen , Yikai Wang , Wenqiang Sun , Feng Wang , Yiwen Chen , Huaping Liu

Impressive progress has been made in audio-driven 3D facial animation recently, but synthesizing 3D talking-head with rich emotion is still unsolved. This is due to the lack of 3D generative models and available 3D emotional dataset with…

Computer Vision and Pattern Recognition · Computer Science 2021-04-27 Qianyun Wang , Zhenfeng Fan , Shihong Xia

Advancing robotic manipulation of deformable objects can enable automation of repetitive tasks across multiple industries, from food processing to textiles and healthcare. Yet robots struggle with the high dimensionality of deformable…

Robotics · Computer Science 2024-09-26 Jan Obrist , Miguel Zamora , Hehui Zheng , Juan Zarate , Robert K. Katzschmann , Stelian Coros

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

In this work, we introduce FlexGen, a flexible framework designed to generate controllable and consistent multi-view images, conditioned on a single-view image, or a text prompt, or both. FlexGen tackles the challenges of controllable…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Xinli Xu , Wenhang Ge , Jiantao Lin , Jiawei Feng , Lie Xu , HanFeng Zhao , Shunsi Zhang , Ying-Cong Chen

Diffusion transformers have recently delivered strong text-to-image generation around 1K resolution, but we show that extending them to native 4K across diverse aspect ratios exposes a tightly coupled failure mode spanning positional…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Tian Ye , Song Fei , Lei Zhu