中文
相关论文

相关论文: Geometry-Aware Attention Guidance for Diffusion Mo…

200 篇论文

Models such as VGGT and $\pi^3$ have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-attention variants offer partial speedups, yet lack a…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Xianbing Sun , Zhikai Zhu , Zhengyu Lou , Bo Yang , Jinyang Tang , Liqing Zhang , He Wang , Jianfu Zhang

Classifier-Free Guidance (CFG) enhances the quality and condition adherence of text-to-image diffusion models. It operates by combining the conditional and unconditional predictions using a fixed weight. However, recent works vary the…

计算机视觉与模式识别 · 计算机科学 2025-01-23 Xi Wang , Nicolas Dufour , Nefeli Andreou , Marie-Paule Cani , Victoria Fernandez Abrevaya , David Picard , Vicky Kalogeiton

We propose a novel attention gate (AG) model for medical imaging that automatically learns to focus on target structures of varying shapes and sizes. Models trained with AGs implicitly learn to suppress irrelevant regions in an input image…

Discrete diffusion models have achieved strong empirical performance in text and other symbolic domains, but, especially for uniform-rate models, they often require many steps to generate a single sample. Existing acceleration methods…

机器学习 · 计算机科学 2026-05-27 Yuchen Liang , Ness Shroff , Yingbin Liang

Training-free guidance enables controlled generation in diffusion and flow models, but most methods rely on gradients and assume differentiable objectives. This work focuses on training-free guidance addressing challenges from…

机器学习 · 计算机科学 2025-06-12 Yingqing Guo , Yukang Yang , Hui Yuan , Mengdi Wang

Analytical diffusion models offer a mathematically transparent path to generative modeling by formulating the denoising score as an empirical-Bayes posterior mean. However, this interpretability comes at a prohibitive cost: the standard…

机器学习 · 计算机科学 2026-02-19 Xinyi Shang , Peng Sun , Jingyu Lin , Zhiqiang Shen

Conditional diffusion models can create unseen images in various settings, aiding image interpolation. Interpolation in latent spaces is well-studied, but interpolation with specific conditions like text or poses is less understood. Simple…

计算机视觉与模式识别 · 计算机科学 2024-10-07 Qiyuan He , Jinghao Wang , Ziwei Liu , Angela Yao

Despite recent advancements in latent diffusion models that generate high-dimensional image data and perform various downstream tasks, there has been little exploration into perceptual consistency within these models on the task of…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Shreshth Saini , Ru-Ling Liao , Yan Ye , Alan C. Bovik

Diffusion-based Transformers have demonstrated impressive generative capabilities, but their high computational costs hinder practical deployment, for example, generating an $8192\times 8192$ image can take over an hour on an A100 GPU. In…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Sucheng Ren , Qihang Yu , Ju He , Alan Yuille , Liang-Chieh Chen

3D Gaussian Splatting (3DGS) is a powerful alternative to Neural Radiance Fields (NeRF), excelling in complex scene reconstruction and efficient rendering. However, it relies on high-quality point clouds from Structure-from-Motion (SfM),…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Ziao Liu , Zhenjia Li , Yifeng Shi , Xiangang Li

Guidance is a widely used technique for diffusion models to enhance sample quality. Technically, guidance is realised by using an auxiliary model that generalises more broadly than the primary model. Using a 2D toy example, we first show…

计算机视觉与模式识别 · 计算机科学 2025-09-01 Nikolas Adaloglou , Tim Kaiser , Damir Iagudin , Markus Kollmann

Attention is a general reasoning mechanism than can flexibly deal with image information, but its memory requirements had made it so far impractical for high resolution image generation. We present Grid Partitioned Attention (GPA), a new…

计算机视觉与模式识别 · 计算机科学 2021-07-09 Nikolay Jetchev , Gökhan Yildirim , Christian Bracher , Roland Vollgraf

Ground Penetrating Radar (GPR) has emerged as a pivotal tool for non-destructive evaluation of subsurface road defects. However, conventional GPR image interpretation remains heavily reliant on subjective expertise, introducing…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Haotian Lv , Yuhui Zhang , Jiangbo Dai , Hanli Wu , Jiaji Wang , Dawei Wang

Few-shot object detection (FSOD) localizes and classifies objects in an image given only a few data samples. Recent trends in FSOD research show the adoption of metric and meta-learning techniques, which are prone to catastrophic forgetting…

计算机视觉与模式识别 · 计算机科学 2021-11-15 Ashutosh Agarwal , Anay Majee , Anbumani Subramanian , Chetan Arora

Recently, image editing based on Diffusion-in-Transformer models has undergone rapid development. However, existing editing methods often lack effective control over the degree of editing, limiting their ability to achieve more customized…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Xuanpu Zhang , Xuesong Niu , Ruidong Chen , Dan Song , Jianhao Zeng , Penghui Du , Haoxiang Cao , Kai Wu , An-an Liu

We propose the joint graph attention neural network (GAT), clustering with adaptive neighbors (CAN) and probabilistic graphical model for dynamic power flow analysis and fault characteristics. In fact, computational efficiency is the main…

机器学习 · 计算机科学 2025-03-25 Tan Le , Van Le

2D irregular packing is a classic combinatorial optimization problem with various applications, such as material utilization and texture atlas generation. This NP-hard problem requires efficient algorithms to optimize space utilization.…

人工智能 · 计算机科学 2024-06-13 Tianyang Xue , Lin Lu , Yang Liu , Mingdong Wu , Hao Dong , Yanbin Zhang , Renmin Han , Baoquan Chen

The generation of high-quality, diverse, and prompt-aligned images is a central goal in image-generating diffusion models. The popular classifier-free guidance (CFG) approach improves quality and alignment at the cost of reduced variation,…

机器学习 · 计算机科学 2025-10-21 Enhao Gu , Haolin Hou

Many generative models attempt to replicate the density of their input data. However, this approach is often undesirable, since data density is highly affected by sampling biases, noise, and artifacts. We propose a method called SUGAR…

机器学习 · 计算机科学 2018-09-10 Ofir Lindenbaum , Jay S. Stanley , Guy Wolf , Smita Krishnaswamy

Classifier-free guidance (CFG) is a key technique for improving conditional generation in diffusion models, enabling more accurate control while enhancing sample quality. It is natural to extend this technique to video diffusion, which…

机器学习 · 计算机科学 2025-07-25 Kiwhan Song , Boyuan Chen , Max Simchowitz , Yilun Du , Russ Tedrake , Vincent Sitzmann