中文
相关论文

相关论文: Structural Energy-Guided Sampling for View-Consist…

200 篇论文

Open-vocabulary 3D scene understanding is crucial for applications requiring natural language-driven spatial interpretation, such as robotics and augmented reality. While 3D Gaussian Splatting (3DGS) offers a powerful representation for…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Wei Sun , Yanzhao Zhou , Jianbin Jiao , Yuan Li

Over the last two decades, Electron Energy Loss Spectroscopy (EELS) imaging with a scanning transmission electron microscope (STEM) has emerged as a technique of choice for visualizing complex chemical, electronic, plasmonic, and phononic…

3D Gaussian Splatting (3DGS) has emerged as a real-time, differentiable representation for neural scene understanding. However, existing 3DGS-based methods struggle to represent hierarchical 3D semantic structures and capture whole-part…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Jingbin You , Zehao Li , Hao Jiang , Xinzhu Ma , Shuqin Gao , Honglong Zhao , Congcong Zheng , Tianlu Mao , Feng Dai , Yucheng Zhang , Zhaoqi Wang

Auto-regressive frameworks for next-scale prediction of 2D images have demonstrated strong potential for producing diverse and sophisticated content by progressively refining a coarse input. However, extending this paradigm to 3D object…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Quanyuan Ruan , Kewei Shi , Jiabao Lei , Xifeng Gao , Xiaoguang Han

Recent progress in text-to-3D object generation enables the synthesis of detailed geometry from text input by leveraging 2D diffusion models and differentiable 3D representations. However, the approaches often suffer from limited…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Ming He , Zhixiang Chen , Steve Maddock

Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. In this work, we enhance the T2I diversity through a geometric lens. Unlike most existing methods…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Ye Zhu , Kaleb S. Newman , Johannes F. Lutzeyer , Adriana Romero-Soriano , Michal Drozdzal , Olga Russakovsky

Generating high-quality 3D assets from textual descriptions remains a pivotal challenge in computer graphics and vision research. Due to the scarcity of 3D data, state-of-the-art approaches utilize pre-trained 2D diffusion priors, optimized…

计算机视觉与模式识别 · 计算机科学 2024-10-14 Ling Yang , Zixiang Zhang , Junlin Han , Bohan Zeng , Runjia Li , Philip Torr , Wentao Zhang

Text-to-3D generation aims to create 3D assets from text-to-image diffusion models. However, existing methods face an inherent bottleneck in generation quality because the widely-used objectives such as Score Distillation Sampling (SDS)…

计算机视觉与模式识别 · 计算机科学 2024-06-24 Zixuan Chen , Ruijie Su , Jiahao Zhu , Lingxiao Yang , Jian-Huang Lai , Xiaohua Xie

Recent works on text-to-3d generation show that using only 2D diffusion supervision for 3D generation tends to produce results with inconsistent appearances (e.g., faces on the back view) and inaccurate shapes (e.g., animals with extra…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Cheng Chen , Xiaofeng Yang , Fan Yang , Chengzeng Feng , Zhoujie Fu , Chuan-Sheng Foo , Guosheng Lin , Fayao Liu

The performance of existing point cloud-based 3D object detection methods heavily relies on large-scale high-quality 3D annotations. However, such annotations are often tedious and expensive to collect. Semi-supervised learning is a good…

计算机视觉与模式识别 · 计算机科学 2021-03-18 Na Zhao , Tat-Seng Chua , Gim Hee Lee

Token-based masked generative models are gaining popularity for their fast inference time with parallel decoding. While recent token-based approaches achieve competitive performance to diffusion-based models, their generation performance is…

机器学习 · 计算机科学 2023-04-05 Jaewoong Lee , Sangwon Jang , Jaehyeong Jo , Jaehong Yoon , Yunji Kim , Jin-Hwa Kim , Jung-Woo Ha , Sung Ju Hwang

Generating 3D content from a single image remains a fundamentally challenging and ill-posed problem due to the inherent absence of geometric and textural information in occluded regions. While state-of-the-art generative models can…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Zehua Ma , Hanhui Li , Zhenyu Xie , Xiaonan Luo , Michael Kampffmeyer , Feng Gao , Xiaodan Liang

Optimization-based text-to-3D methods distill guidance from 2D generative models via Score Distillation Sampling (SDS), but implicitly treat this guidance as static. This work shows that ignoring source dynamics yields inconsistent…

机器学习 · 计算机科学 2025-11-18 Jiayin Zhu , Linlin Yang , Yicong Li , Angela Yao

The increased demand for 3D data in AR/VR, robotics and gaming applications, gave rise to powerful generative pipelines capable of synthesizing high-quality 3D objects. Most of these models rely on the Score Distillation Sampling (SDS)…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Ziyu Wan , Despoina Paschalidou , Ian Huang , Hongyu Liu , Bokui Shen , Xiaoyu Xiang , Jing Liao , Leonidas Guibas

We present SPAD, a novel approach for creating consistent multi-view images from text prompts or single images. To enable multi-view generation, we repurpose a pretrained 2D diffusion model by extending its self-attention layers with…

3D Gaussian Splatting (3DGS) has recently emerged as a promising approach for 3D reconstruction, providing explicit, point-based representations and enabling high-quality real time rendering. However, when trained with sparse input views,…

图像与视频处理 · 电气工程与系统科学 2026-02-09 Chaeyoung Jeong , Kwangsu Kim

Current text-to-3D generation methods based on score distillation often suffer from geometric inconsistencies, leading to repeated patterns across different poses of 3D assets. This issue, known as the Multi-Face Janus problem, arises…

计算机视觉与模式识别 · 计算机科学 2025-02-19 Chenxi Zheng , Yihong Lin , Bangzhen Liu , Xuemiao Xu , Yongwei Nie , Shengfeng He

State-of-the-art text-to-image models excel at photorealistic rendering but often struggle to capture the layout and object relationships implied by complex prompts. Scene graphs provide a natural structural prior, yet previous graph-guided…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Thanh-Nhan Vo , Trong-Thuan Nguyen , Tam V. Nguyen , Minh-Triet Tran

Structured Prediction Energy Networks (SPENs) are a simple, yet expressive family of structured prediction models (Belanger and McCallum, 2016). An energy function over candidate structured outputs is given by a deep network, and…

机器学习 · 统计学 2017-07-18 David Belanger , Bishan Yang , Andrew McCallum

Video-based bug reports are increasingly being used to document bugs for programs centered around a graphical user interface (GUI). However, developing automated techniques to manage video-based reports is challenging as it requires…

软件工程 · 计算机科学 2024-07-12 Yanfu Yan , Nathan Cooper , Oscar Chaparro , Kevin Moran , Denys Poshyvanyk