English
Related papers

Related papers: StructLDM: Structured Latent Diffusion for 3D Huma…

200 papers

Computed Tomography (CT) is a widely utilized imaging modality in clinical settings. Using densely acquired rotational X-ray arrays, CT can capture 3D spatial features. However, it is confronted with challenged such as significant time…

Image and Video Processing · Electrical Eng. & Systems 2025-07-16 Duoyou Chen , Yunqing Chen , Can Zhang , Zhou Wang , Cheng Chen , Ruoxiu Xiao

The rapid evolution of the fashion industry increasingly intersects with technological advancements, particularly through the integration of generative AI. This study introduces a novel generative pipeline designed to transform the fashion…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Abhishek Kumar Singh , Ioannis Patras

Diffusion models are capable of impressive feats of image generation with uncommon juxtapositions such as astronauts riding horses on the moon with properly placed shadows. These outputs indicate the ability to perform compositional…

Machine Learning · Computer Science 2024-05-01 Qiyao Liang , Ziming Liu , Ila Fiete

Humans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (LMMs), however, lack of this capability of 3D spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Wufei Ma , Luoxin Ye , Celso M de Melo , Jieneng Chen , Alan Yuille

Articulated 3D objects are central to many applications in robotics, AR/VR, and animation. Recent approaches to modeling such objects either rely on optimization-based reconstruction pipelines that require dense-view supervision or on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Chuhao Chen , Isabella Liu , Xinyue Wei , Hao Su , Minghua Liu

We present En3D, an enhanced generative scheme for sculpting high-quality 3D human avatars. Unlike previous works that rely on scarce 3D datasets or limited 2D collections with imbalanced viewing angles and imprecise pose priors, our…

Computer Vision and Pattern Recognition · Computer Science 2024-01-03 Yifang Men , Biwen Lei , Yuan Yao , Miaomiao Cui , Zhouhui Lian , Xuansong Xie

Recent text-to-scene generation approaches largely reduced the manual efforts required to create 3D scenes. However, their focus is either to generate a scene layout or to generate objects, and few generate both. The generated scene layout…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Zhenggang Tang , Yuehao Wang , Yuchen Fan , Jun-Kun Chen , Yu-Ying Yeh , Kihyuk Sohn , Zhangyang Wang , Qixing Huang , Alexander Schwing , Rakesh Ranjan , Dilin Wang , Zhicheng Yan

We introduce a novel 3D generation method for versatile and high-quality 3D asset creation. The cornerstone is a unified Structured LATent (SLAT) representation which allows decoding to different output formats, such as Radiance Fields, 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Jianfeng Xiang , Zelong Lv , Sicheng Xu , Yu Deng , Ruicheng Wang , Bowen Zhang , Dong Chen , Xin Tong , Jiaolong Yang

Recently, it has become progressively more evident that classic diagnostic labels are unable to reliably describe the complexity and variability of several clinical phenotypes. This is particularly true for a broad range of neuropsychiatric…

Machine Learning · Computer Science 2024-02-28 Giovanna Maria Dimitri , Simeon Spasov , Andrea Duggento , Luca Passamonti , Pietro Li`o , Nicola Toschi

Understanding and generating the fine-grained structure of objects -- such as birds with species-specific beaks, wings, and tails -- is a long-standing challenge in computer vision. We propose Chirpy3D, a part-aware multi-view diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Kam Woh Ng , Jing Yang , Jia Wei Sii , Chee Seng Chan , Jiankang Deng , Yi-Zhe Song , Tao Xiang , Xiatian Zhu

Diffusion-based generative models have achieved promising results recently, but raise an array of open questions in terms of conceptual understanding, theoretical analysis, algorithm improvement and extensions to discrete, structured,…

Machine Learning · Computer Science 2022-09-01 Xingchao Liu , Lemeng Wu , Mao Ye , Qiang Liu

Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step toward a single architecture for both generation (visual synthesis) and understanding…

Computation and Language · Computer Science 2026-05-11 Jiaxiu Jiang , Jingjing Ren , Wenbo Li , Bo Wang , Haoze Sun , Yijun Yang , Jianhui Liu , Yanbing Zhang , Shenghe Zheng , Yuan Zhang , Haoyang Huang , Nan Duan , Wangmeng Zuo

Previous efforts have managed to generate production-ready 3D assets from text or images. However, these methods primarily employ NeRF or 3D Gaussian representations, which are not adept at producing smooth, high-quality geometries required…

Graphics · Computer Science 2024-10-15 Rengan Xie , Wenting Zheng , Kai Huang , Yizheng Chen , Qi Wang , Qi Ye , Wei Chen , Yuchi Huo

Text-to-audio (TTA) system has recently gained attention for its ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study,…

Sound · Computer Science 2023-09-12 Haohe Liu , Zehua Chen , Yi Yuan , Xinhao Mei , Xubo Liu , Danilo Mandic , Wenwu Wang , Mark D. Plumbley

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation by combining LLM and diffusion models, the state-of-the-art in each task, respectively. Existing approaches rely on spatial visual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Kaihang Pan , Wang Lin , Zhongqi Yue , Tenglong Ao , Liyu Jia , Wei Zhao , Juncheng Li , Siliang Tang , Hanwang Zhang

This paper introduces an innovative methodology for producing high-quality 3D lung CT images guided by textual information. While diffusion-based generative models are increasingly used in medical imaging, current state-of-the-art…

Image and Video Processing · Electrical Eng. & Systems 2024-10-16 Yanwu Xu , Li Sun , Wei Peng , Shuyue Jia , Katelyn Morrison , Adam Perer , Afrooz Zandifar , Shyam Visweswaran , Motahhare Eslami , Kayhan Batmanghelich

Despite the remarkable progress in deep generative models, synthesizing high-resolution and temporally coherent videos still remains a challenge due to their high-dimensionality and complex temporal dynamics along with large spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Sihyun Yu , Kihyuk Sohn , Subin Kim , Jinwoo Shin

A complete representation of 3D objects requires characterizing the space of deformations in an interpretable manner, from articulations of a single instance to changes in shape across categories. In this work, we improve on a prior…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Tristan Aumentado-Armstrong , Stavros Tsogkas , Sven Dickinson , Allan Jepson

Precipitation nowcasting is a critical spatio-temporal prediction task for society to prevent severe damage owing to extreme weather events. Despite the advances in this field, the complex and stochastic nature of this task still poses…

Machine Learning · Computer Science 2025-12-25 Shi Quan Foo , Chi-Ho Wong , Zhihan Gao , Dit-Yan Yeung , Ka-Hing Wong , Wai-Kin Wong

Automatic layout generation that can synthesize high-quality layouts is an important tool for graphic design in many applications. Though existing methods based on generative models such as Generative Adversarial Networks (GANs) and…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Shang Chai , Liansheng Zhuang , Fengying Yan