English
Related papers

Related papers: GAOT: Generating Articulated Objects Through Text-…

200 papers

While text-to-video diffusion models have made significant strides, many still face challenges in generating videos with temporal consistency. Within diffusion frameworks, guidance techniques have proven effective in enhancing output…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyelin Nam , Jaemin Kim , Dohun Lee , Jong Chul Ye

Diffusion models have achieved remarkable results in generating high-quality, diverse, and creative images. However, when it comes to text-based image generation, they often fail to capture the intended meaning presented in the text. For…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Kota Sueyoshi , Takashi Matsubara

Diffusion models have gained tremendous success in text-to-image generation, yet still lag behind with visual understanding tasks, an area dominated by autoregressive vision-language models. We propose a large-scale and fully end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Zijie Li , Henry Li , Yichun Shi , Amir Barati Farimani , Yuval Kluger , Linjie Yang , Peng Wang

The recent surge in popularity of diffusion models for image generation has brought new attention to the potential of these models in other areas of media generation. One area that has yet to be fully explored is the application of…

Sound · Computer Science 2023-02-01 Flavio Schneider

Recent text-to-image generative models can generate high-fidelity images from text prompts. However, these models struggle to consistently generate the same objects in different contexts with the same appearance. Consistent object…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Alec Helbling , Evan Montoya , Duen Horng Chau

Generating 3D CT volumes from descriptive free-text inputs presents a transformative opportunity in diagnostics and research. In this paper, we introduce Text2CT, a novel approach for synthesizing 3D CT volumes from textual descriptions…

Image and Video Processing · Electrical Eng. & Systems 2025-05-09 Pengfei Guo , Can Zhao , Dong Yang , Yufan He , Vishwesh Nath , Ziyue Xu , Pedro R. A. S. Bassi , Zongwei Zhou , Benjamin D. Simon , Stephanie Anne Harmon , Baris Turkbey , Daguang Xu

Graph generation is a critical yet challenging task, as empirical analyses require a deep understanding of complex, non-Euclidean structures. Diffusion models have recently made significant advances in graph generation, but these models are…

Machine Learning · Computer Science 2026-03-13 Yiming Huang , Tolga Birdal

Diffusion generative models transform noise into data by inverting a process that progressively adds noise to data samples. Inspired by concepts from the renormalization group in physics, which analyzes systems across different scales, we…

Machine Learning · Computer Science 2024-10-04 Mathis Gerdes , Max Welling , Miranda C. N. Cheng

While generative artificial intelligence has advanced significantly across text, image, audio, and video domains, 3D generation remains comparatively underdeveloped due to fundamental challenges such as data scarcity, algorithmic…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Weiyu Li , Xuanyang Zhang , Zheng Sun , Di Qi , Hao Li , Wei Cheng , Weiwei Cai , Shihao Wu , Jiarui Liu , Zihao Wang , Xiao Chen , Feipeng Tian , Jianxiong Pan , Zeming Li , Gang Yu , Xiangyu Zhang , Daxin Jiang , Ping Tan

We present Text2Tex, a novel method for generating high-quality textures for 3D meshes from the given text prompts. Our method incorporates inpainting into a pre-trained depth-aware image diffusion model to progressively synthesize high…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Dave Zhenyu Chen , Yawar Siddiqui , Hsin-Ying Lee , Sergey Tulyakov , Matthias Nießner

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

Inspired by Geoffrey Hinton emphasis on generative modeling, To recognize shapes, first learn to generate them, we explore the use of 3D diffusion models for object classification. Leveraging the density estimates from these models, our…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Nursena Koprucu , Meher Shashwat Nigam , Shicheng Xu , Biruk Abere , Gabriele Dominici , Andrew Rodriguez , Sharvaree Vadgama , Berfin Inal , Alberto Tono

The efficiency of multi-agent systems driven by large language models (LLMs) largely hinges on their communication topology. However, designing an optimal topology is a non-trivial challenge, as it requires balancing competing objectives…

Computation and Language · Computer Science 2026-05-19 Eric Hanchen Jiang , Mengting Li , Guancheng Wan , Sophia Yin , Yuchen Wu , Xiao Liang , Xinfeng Li , Yizhou Sun , Wei Wang , Kai-Wei Chang , Ying Nian Wu

Generating high-quality whole-body human object interaction motion sequences is becoming increasingly important in various fields such as animation, VR/AR, and robotics. The main challenge of this task lies in determining the level of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Yonghao Zhang , Qiang He , Yanguang Wan , Yinda Zhang , Xiaoming Deng , Cuixia Ma , Hongan Wang

Recent generative models can synthesize high-quality images, but they often fail to generate humans interacting with objects using their hands. This arises mostly from the model's misunderstanding of such interactions and the hardships of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Patrick Kwon , Chen Chen , Hanbyul Joo

Probabilistic diffusion models have achieved state-of-the-art results for image synthesis, inpainting, and text-to-image tasks. However, they are still in the early stages of generating complex 3D shapes. This work proposes Diffusion-SDF, a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Gene Chou , Yuval Bahat , Felix Heide

Articulated objects (e.g., doors and drawers) exist everywhere in our life. Different from rigid objects, articulated objects have higher degrees of freedom and are rich in geometries, semantics, and part functions. Modeling different kinds…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 Yushi Du , Ruihai Wu , Yan Shen , Hao Dong

The way humans interact with each other, including interpersonal distances, spatial configuration, and motion, varies significantly across different situations. To enable machines to understand such complex, context-dependent behaviors, it…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Jeonghyeon Na , Sangwon Baik , Inhee Lee , Junyoung Lee , Hanbyul Joo

We propose CG-HOI, the first method to address the task of generating dynamic 3D human-object interactions (HOIs) from text. We model the motion of both human and object in an interdependent fashion, as semantically rich human motion rarely…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Christian Diller , Angela Dai

In this work, we propose a novel method for generating 3D point clouds that leverage properties of hyper networks. Contrary to the existing methods that learn only the representation of a 3D object, our approach simultaneously finds a…

Computer Vision and Pattern Recognition · Computer Science 2020-12-04 Przemysław Spurek , Sebastian Winczowski , Jacek Tabor , Maciej Zamorski , Maciej Zięba , Tomasz Trzciński