中文
相关论文

相关论文: StackGen: Generating Stable Structures from Silhou…

200 篇论文

Recent advances in Artificial Intelligence Generated Content (AIGC) have garnered significant interest, accompanied by an increasing need to transmit and compress the vast number of AI-generated images (AIGIs). However, there is a…

图像与视频处理 · 电气工程与系统科学 2024-12-18 Ruijie Chen , Qi Mao , Zhengxue Cheng

Generative AI models, such as score-based diffusion models, have recently advanced the field of computational materials science by enabling the generation of new materials with desired properties. In addition, these models could also be…

材料科学 · 物理学 2026-01-06 Timo Reents , Arianna Cantarella , Marnik Bercx , Pietro Bonfà , Giovanni Pizzi

Precise perception of contact interactions is essential for fine-grained manipulation skills for robots. In this paper, we present the design of feedback skills for robots that must learn to stack complex-shaped objects on top of each other…

机器人学 · 计算机科学 2024-03-26 Kei Ota , Devesh K. Jha , Krishna Murthy Jatavallabhula , Asako Kanezaki , Joshua B. Tenenbaum

We introduce ClutterGen, a physically compliant simulation scene generator capable of producing highly diverse, cluttered, and stable scenes for robot learning. Generating such scenes is challenging as each object must adhere to physical…

机器人学 · 计算机科学 2024-10-08 Yinsen Jia , Boyuan Chen

We tackle the problem of text-driven 3D generation from a geometry alignment perspective. Given a set of text prompts, we aim to generate a collection of objects with semantically corresponding parts aligned across them. Recent methods…

We generate synthetic images with the "Stable Diffusion" image generation model using the Wordnet taxonomy and the definitions of concepts it contains. This synthetic image database can be used as training data for data augmentation in…

计算机视觉与模式识别 · 计算机科学 2022-11-07 Andreas Stöckl

Image generation models trained on large datasets can synthesize high-quality images but often produce spatially inconsistent and distorted images due to limited information about the underlying structures and spatial layouts. In this work,…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Hyundo Lee , Suhyung Choi , Inwoo Hwang , Byoung-Tak Zhang

Learning disentangled representations is a key step towards effectively discovering and modelling the underlying structure of environments. In the natural sciences, physics has found great success by describing the universe in terms of…

机器学习 · 计算机科学 2020-10-27 Robin Quessard , Thomas D. Barrett , William R. Clements

While diffusion models have revolutionized text-to-image generation with their ability to synthesize realistic and diverse scenes, they continue to struggle to generate consistent and legible text within images. This shortcoming is commonly…

机器学习 · 计算机科学 2025-09-16 Tianyu Zhang , Xinyu Wang , Lu Li , Zhenghan Tai , Jijun Chi , Jingrui Tian , Hailin He , Suyuchen Wang

Intelligent agents, such as robots and virtual agents, must understand the dynamics of complex social interactions to interact with humans. Effectively representing social dynamics is challenging because we require multi-modal, synchronized…

机器学习 · 计算机科学 2025-01-22 Antonio Lech Martin-Ozimek , Isuru Jayarathne , Su Larb Mon , Jouh Yeong Chew

Diffusion plays an important role in a wide variety of phenomena, from bacterial quorum sensing to the dynamics of traffic flow. While it generally tends to level out gradients and inhomogeneities, diffusion has nonetheless been shown to…

斑图形成与孤子 · 物理学 2024-07-03 Alexandre Champagne-Ruel , Sascha Zakaib-Bernier , Paul Charbonneau

Animating virtual avatars to make co-speech gestures facilitates various applications in human-machine interaction. The existing methods mainly rely on generative adversarial networks (GANs), which typically suffer from notorious mode…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Lingting Zhu , Xian Liu , Xuanyu Liu , Rui Qian , Ziwei Liu , Lequan Yu

Nature evolves creatures with a high complexity of morphological and behavioral intelligence, meanwhile computational methods lag in approaching that diversity and efficacy. Co-optimization of artificial creatures' morphology and control in…

Text-to-image models are showcasing the impressive ability to create high-quality and diverse generative images. Nevertheless, the transition from freehand sketches to complex scene images remains challenging using diffusion models. In this…

计算机视觉与模式识别 · 计算机科学 2024-07-10 Tianyu Zhang , Xiaoxuan Xie , Xusheng Du , Haoran Xie

From serving a cup of coffee to positioning mechanical parts during assembly, stable object placement is a crucial skill for future robots. It becomes particularly challenging under geometric uncertainties, e.g., when the object pose or…

机器人学 · 计算机科学 2025-12-02 Linfeng Li , Gang Yang , Lin Shao , David Hsu

While diffusion models can successfully generate data and make predictions, they are predominantly designed for static images. We propose an approach for efficiently training diffusion models for probabilistic spatiotemporal forecasting,…

机器学习 · 计算机科学 2023-10-12 Salva Rühling Cachay , Bo Zhao , Hailey Joren , Rose Yu

Diffusion-based generative models' impressive ability to create convincing images has garnered global attention. However, their complex structures and operations often pose challenges for non-experts to grasp. We present Diffusion…

The computational intensity of detector simulation and event reconstruction poses a significant difficulty for data analysis in collider experiments. This challenge inspires the continued development of machine learning techniques to serve…

高能物理 - 实验 · 物理学 2024-11-22 Dmitrii Kobylianskii , Nathalie Soybelman , Nilotpal Kakati , Etienne Dreyer , Benjamin Nachman , Eilam Gross

We fine-tuned a foundational stable diffusion model using X-ray scattering images and their corresponding descriptions to generate new scientific images from given prompts. However, some of the generated images exhibit significant…

图像与视频处理 · 电气工程与系统科学 2024-08-26 Zhuowen Zhao , Xiaoya Chong , Tanny Chavez , Alexander Hexemer

Generative foundation models like Stable Diffusion comprise a diverse spectrum of knowledge in computer vision with the potential for transfer learning, e.g., via generating data to train student models for downstream tasks. This could…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Leonhard Hennicke , Christian Medeiros Adriano , Holger Giese , Jan Mathias Koehler , Lukas Schott