English
Related papers

Related papers: LTM3D: Bridging Token Spaces for Conditional 3D Ge…

200 papers

Recently, multi-view diffusion-based 3D generation methods have gained significant attention. However, these methods often suffer from shape and texture misalignment across generated multi-view images, leading to low-quality 3D generation…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Zhuojiang Cai , Yiheng Zhang , Meitong Guo , Mingdao Wang , Yuwang Wang

The discovery of inorganic crystal structures with targeted properties is a significant challenge in materials science. Generative models, especially state-of-the-art diffusion models, offer the promise of modeling complex data…

Latent diffusion models (LDMs) enable high-fidelity synthesis by operating in learned latent spaces. However, training state-of-the-art LDMs requires complex staging: a tokenizer must be trained first, before the diffusion model can be…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Shivam Duggal , Xingjian Bai , Zongze Wu , Richard Zhang , Eli Shechtman , Antonio Torralba , Phillip Isola , William T. Freeman

We introduce an approach for 3D head avatar generation and editing with multi-modal conditioning based on a 3D Generative Adversarial Network (GAN) and a Latent Diffusion Model (LDM). 3D GANs can generate high-quality head avatars given a…

Computer Vision and Pattern Recognition · Computer Science 2024-02-09 Wamiq Reyaz Para , Abdelrahman Eldesokey , Zhenyu Li , Pradyumna Reddy , Jiankang Deng , Peter Wonka

Tracking of dynamic people in cluttered and crowded human-centered environments is a challenging robotics problem due to the presence of intraclass variations including occlusions, pose deformations, and lighting variations. This paper…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Angus Fung , Beno Benhabib , Goldie Nejat

Diffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text. Token-level diffusion doesn't model word-order dependencies explicitly and operates on short, fixed…

Computation and Language · Computer Science 2025-05-27 Xiaochen Zhu , Georgi Karadzhov , Chenxi Whitehouse , Andreas Vlachos

Generating visual layouts is an essential ingredient of graphic design. The ability to condition layout generation on a partial subset of component attributes is critical to real-world applications that involve user interaction. Recently,…

Computer Vision and Pattern Recognition · Computer Science 2023-03-08 Elad Levi , Eli Brosh , Mykola Mykhailych , Meir Perez

By enabling capturing of 3D point clouds that reflect the geometry of the immediate environment, LiDAR has emerged as a primary sensor for autonomous systems. If a LiDAR scan is too sparse, occluded by obstacles, or too small in range,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Ryan Faulkner , Luke Haub , Simon Ratcliffe , Anh-Dzung Doan , Ian Reid , Tat-Jun Chin

Fast and accurate 3D shape generation from point clouds is essential for applications in robotics, AR/VR, and digital content creation. We introduce ConTiCoM-3D, a continuous-time consistency model that synthesizes 3D shapes directly in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Sebastian Eilermann , René Heesch , Oliver Niggemann

Large-scale pre-training tasks like image classification, captioning, or self-supervised techniques do not incentivize learning the semantic boundaries of objects. However, recent generative foundation models built using text-based latent…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Koutilya Pnvr , Bharat Singh , Pallabi Ghosh , Behjat Siddiquie , David Jacobs

Diffusion models have recently become the de-facto approach for generative modeling in the 2D domain. However, extending diffusion models to 3D is challenging due to the difficulties in acquiring 3D ground truth data for training. On the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Jiatao Gu , Qingzhe Gao , Shuangfei Zhai , Baoquan Chen , Lingjie Liu , Josh Susskind

We propose Latent-Shift -- an efficient text-to-video generation method based on a pretrained text-to-image generation model that consists of an autoencoder and a U-Net diffusion model. Learning a video diffusion model in the latent space…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Jie An , Songyang Zhang , Harry Yang , Sonal Gupta , Jia-Bin Huang , Jiebo Luo , Xi Yin

Existing single image-to-3D creation methods typically involve a two-stage process, first generating multi-view images, and then using these images for 3D reconstruction. However, training these two stages separately leads to significant…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Hao Wen , Zehuan Huang , Yaohui Wang , Xinyuan Chen , Lu Sheng

In spite of the remarkable potential of Latent Diffusion Models (LDMs) in image generation, the desired properties and optimal design of the autoencoders have been underexplored. In this work, we analyze the role of autoencoders in LDMs and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Junho Lee , Jeongwoo Shin , Hyungwook Choi , Joonseok Lee

We present significant extensions to diffusion-based sequence generation models, blurring the line with autoregressive language models. We introduce hyperschedules, which assign distinct noise schedules to individual token positions,…

Machine Learning · Computer Science 2025-10-08 Nima Fathi , Torsten Scholak , Pierre-André Noël

The recent advancements in image-text diffusion models have stimulated research interest in large-scale 3D generative models. Nevertheless, the limited availability of diverse 3D resources presents significant challenges to learning. In…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Chi Zhang , Yiwen Chen , Yijun Fu , Zhenglin Zhou , Gang YU , Billzb Wang , Bin Fu , Tao Chen , Guosheng Lin , Chunhua Shen

Recent advances in generative models have yielded impressive progress on motion in-betweening, allowing for more complex, varied, and realistic motion transitions. However, recent methods still exhibit noticeable limitations in preserving…

Graphics · Computer Science 2026-05-14 Shiyu Fan , Paul Henderson , Edmond S. L. Ho

We present DIRECT-3D, a diffusion-based 3D generative model for creating high-quality 3D assets (represented by Neural Radiance Fields) from text prompts. Unlike recent 3D generative models that rely on clean and well-aligned 3D data,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Qihao Liu , Yi Zhang , Song Bai , Adam Kortylewski , Alan Yuille

3D shape generation aims to produce innovative 3D content adhering to specific conditions and constraints. Existing methods often decompose 3D shapes into a sequence of localized components, treating each element in isolation without…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Ruikai Cui , Weizhe Liu , Weixuan Sun , Senbo Wang , Taizhang Shang , Yang Li , Xibin Song , Han Yan , Zhennan Wu , Shenzhou Chen , Hongdong Li , Pan Ji

We present Diff3F as a simple, robust, and class-agnostic feature descriptor that can be computed for untextured input shapes (meshes or point clouds). Our method distills diffusion features from image foundational models onto input shapes.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Niladri Shekhar Dutt , Sanjeev Muralikrishnan , Niloy J. Mitra
‹ Prev 1 8 9 10 Next ›