English
Related papers

Related papers: Structural Energy-Guided Sampling for View-Consist…

200 papers

Open-vocabulary 3D scene understanding is crucial for applications requiring natural language-driven spatial interpretation, such as robotics and augmented reality. While 3D Gaussian Splatting (3DGS) offers a powerful representation for…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Wei Sun , Yanzhao Zhou , Jianbin Jiao , Yuan Li

Over the last two decades, Electron Energy Loss Spectroscopy (EELS) imaging with a scanning transmission electron microscope (STEM) has emerged as a technique of choice for visualizing complex chemical, electronic, plasmonic, and phononic…

3D Gaussian Splatting (3DGS) has emerged as a real-time, differentiable representation for neural scene understanding. However, existing 3DGS-based methods struggle to represent hierarchical 3D semantic structures and capture whole-part…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Jingbin You , Zehao Li , Hao Jiang , Xinzhu Ma , Shuqin Gao , Honglong Zhao , Congcong Zheng , Tianlu Mao , Feng Dai , Yucheng Zhang , Zhaoqi Wang

Auto-regressive frameworks for next-scale prediction of 2D images have demonstrated strong potential for producing diverse and sophisticated content by progressively refining a coarse input. However, extending this paradigm to 3D object…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Quanyuan Ruan , Kewei Shi , Jiabao Lei , Xifeng Gao , Xiaoguang Han

Recent progress in text-to-3D object generation enables the synthesis of detailed geometry from text input by leveraging 2D diffusion models and differentiable 3D representations. However, the approaches often suffer from limited…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Ming He , Zhixiang Chen , Steve Maddock

Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. In this work, we enhance the T2I diversity through a geometric lens. Unlike most existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Ye Zhu , Kaleb S. Newman , Johannes F. Lutzeyer , Adriana Romero-Soriano , Michal Drozdzal , Olga Russakovsky

Generating high-quality 3D assets from textual descriptions remains a pivotal challenge in computer graphics and vision research. Due to the scarcity of 3D data, state-of-the-art approaches utilize pre-trained 2D diffusion priors, optimized…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Ling Yang , Zixiang Zhang , Junlin Han , Bohan Zeng , Runjia Li , Philip Torr , Wentao Zhang

Text-to-3D generation aims to create 3D assets from text-to-image diffusion models. However, existing methods face an inherent bottleneck in generation quality because the widely-used objectives such as Score Distillation Sampling (SDS)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-24 Zixuan Chen , Ruijie Su , Jiahao Zhu , Lingxiao Yang , Jian-Huang Lai , Xiaohua Xie

Recent works on text-to-3d generation show that using only 2D diffusion supervision for 3D generation tends to produce results with inconsistent appearances (e.g., faces on the back view) and inaccurate shapes (e.g., animals with extra…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Cheng Chen , Xiaofeng Yang , Fan Yang , Chengzeng Feng , Zhoujie Fu , Chuan-Sheng Foo , Guosheng Lin , Fayao Liu

The performance of existing point cloud-based 3D object detection methods heavily relies on large-scale high-quality 3D annotations. However, such annotations are often tedious and expensive to collect. Semi-supervised learning is a good…

Computer Vision and Pattern Recognition · Computer Science 2021-03-18 Na Zhao , Tat-Seng Chua , Gim Hee Lee

Token-based masked generative models are gaining popularity for their fast inference time with parallel decoding. While recent token-based approaches achieve competitive performance to diffusion-based models, their generation performance is…

Machine Learning · Computer Science 2023-04-05 Jaewoong Lee , Sangwon Jang , Jaehyeong Jo , Jaehong Yoon , Yunji Kim , Jin-Hwa Kim , Jung-Woo Ha , Sung Ju Hwang

Generating 3D content from a single image remains a fundamentally challenging and ill-posed problem due to the inherent absence of geometric and textural information in occluded regions. While state-of-the-art generative models can…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Zehua Ma , Hanhui Li , Zhenyu Xie , Xiaonan Luo , Michael Kampffmeyer , Feng Gao , Xiaodan Liang

Optimization-based text-to-3D methods distill guidance from 2D generative models via Score Distillation Sampling (SDS), but implicitly treat this guidance as static. This work shows that ignoring source dynamics yields inconsistent…

Machine Learning · Computer Science 2025-11-18 Jiayin Zhu , Linlin Yang , Yicong Li , Angela Yao

The increased demand for 3D data in AR/VR, robotics and gaming applications, gave rise to powerful generative pipelines capable of synthesizing high-quality 3D objects. Most of these models rely on the Score Distillation Sampling (SDS)…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Ziyu Wan , Despoina Paschalidou , Ian Huang , Hongyu Liu , Bokui Shen , Xiaoyu Xiang , Jing Liao , Leonidas Guibas

We present SPAD, a novel approach for creating consistent multi-view images from text prompts or single images. To enable multi-view generation, we repurpose a pretrained 2D diffusion model by extending its self-attention layers with…

Computer Vision and Pattern Recognition · Computer Science 2024-02-09 Yash Kant , Ziyi Wu , Michael Vasilkovsky , Guocheng Qian , Jian Ren , Riza Alp Guler , Bernard Ghanem , Sergey Tulyakov , Igor Gilitschenski , Aliaksandr Siarohin

3D Gaussian Splatting (3DGS) has recently emerged as a promising approach for 3D reconstruction, providing explicit, point-based representations and enabling high-quality real time rendering. However, when trained with sparse input views,…

Image and Video Processing · Electrical Eng. & Systems 2026-02-09 Chaeyoung Jeong , Kwangsu Kim

Current text-to-3D generation methods based on score distillation often suffer from geometric inconsistencies, leading to repeated patterns across different poses of 3D assets. This issue, known as the Multi-Face Janus problem, arises…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Chenxi Zheng , Yihong Lin , Bangzhen Liu , Xuemiao Xu , Yongwei Nie , Shengfeng He

State-of-the-art text-to-image models excel at photorealistic rendering but often struggle to capture the layout and object relationships implied by complex prompts. Scene graphs provide a natural structural prior, yet previous graph-guided…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Thanh-Nhan Vo , Trong-Thuan Nguyen , Tam V. Nguyen , Minh-Triet Tran

Structured Prediction Energy Networks (SPENs) are a simple, yet expressive family of structured prediction models (Belanger and McCallum, 2016). An energy function over candidate structured outputs is given by a deep network, and…

Machine Learning · Statistics 2017-07-18 David Belanger , Bishan Yang , Andrew McCallum

Video-based bug reports are increasingly being used to document bugs for programs centered around a graphical user interface (GUI). However, developing automated techniques to manage video-based reports is challenging as it requires…

Software Engineering · Computer Science 2024-07-12 Yanfu Yan , Nathan Cooper , Oscar Chaparro , Kevin Moran , Denys Poshyvanyk