中文
相关论文

相关论文: GASS: Geometry-Aware Spherical Sampling for Disent…

200 篇论文

Medical image synthesis is crucial for alleviating data scarcity and privacy constraints. However, fine-tuning general text-to-image (T2I) models remains challenging, mainly due to the significant modality gap between complex visual details…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Xin Huang , Junjie Liang , Qingshan Hou , Peng Cao , Jinzhu Yang , Xiaoli Liu , Osmar R. Zaiane

Reconstructing Dynamic 3D Gaussian Splatting (3DGS) from low-framerate RGB videos is challenging. This is because large inter-frame motions will increase the uncertainty of the solution space. For example, one pixel in the first frame might…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Junhao He , Jiaxu Wang , Jia Li , Mingyuan Sun , Qiang Zhang , Jiahang Cao , Ziyi Zhang , Yi Gu , Jingkai Sun , Renjing Xu

Gaussian Splatting (GS), a recent technique for converting discrete points into continuous spatial representations, has shown promising results in 3D scene modeling and 2D image super-resolution. In this paper, we explore its untapped…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Hongyu Li , Chaofeng Chen , Xiaoming Li , Guangming Lu

Sampling-based motion planning algorithms are widely used for motion planning of robotic manipulators, but they often struggle with sample inefficiency in high-dimensional configuration spaces due to their reliance on uniform or…

机器人学 · 计算机科学 2026-03-06 Davood Soleymanzadeh , Xiao Liang , Minghui Zheng

Due to the current lack of large-scale datasets at the million-scale level, tasks involving panoramic images predominantly rely on existing two-dimensional pre-trained image benchmark models as backbone networks. However, these networks are…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Jingguo Liu , Han Yu , Shigang Li , Jianfeng Li

Latent flow matching for image generation usually transports Gaussian noise to variational autoencoder latents along linear paths. Both endpoints, however, concentrate in thin spherical shells, and a Euclidean chord leaves those shells even…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Tuna Han Salih Meral , Kaan Oktay , Hidir Yesiltepe , Adil Kaan Akan , Pinar Yanardag

A significant ``modality gap" exists between the abundance of text-only data and the increasing power of multimodal models. This work systematically investigates whether images generated on-the-fly by Text-to-Image (T2I) models can serve as…

多媒体 · 计算机科学 2026-03-04 Yuesheng Huang , Peng Zhang , Xiaoxin Wu , Riliang Liu , Jiaqi Liang

Flow-based text-to-image models follow deterministic trajectories, making it costly to explore diverse modes under limited sampling budgets. Existing approaches to improving diversity often rely on retraining or degrade image fidelity. To…

人工智能 · 计算机科学 2026-05-21 Jingxuan Wu , Zhenglin Wan , Xingrui Yu , Yuzhe Yang , Bo An , Ivor Tsang , Yang You

Text-to-image (T2I) customization empowers users to adapt the T2I diffusion model to new concepts absent in the pre-training dataset. On this basis, capturing multiple new concepts from a single image has emerged as a new task, allowing the…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Junjie Shentu , Matthew Watson , Noura Al Moubayed

General object composition (GOC) aims to seamlessly integrate a target object into a background scene with desired geometric properties, while simultaneously preserving its fine-grained appearance details. Recent approaches derive semantic…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Jianman Lin , Haojie Li , Chunmei Qing , Zhijing Yang , Liang Lin , Tianshui Chen

Learning interpretable multimodal representations inherently relies on uncovering the conditional dependencies between heterogeneous features. However, sparse graph estimation techniques, such as Graphical Lasso (GLasso), to…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Fei Wang , Yutong Zhang , Xiong Wang

Text-to-image generative models excel in creating images from text but struggle with ensuring alignment and consistency between outputs and prompts. This paper introduces TextMatch, a novel framework that leverages multimodal optimization…

计算机视觉与模式识别 · 计算机科学 2025-01-28 Yucong Luo , Mingyue Cheng , Jie Ouyang , Xiaoyu Tao , Qi Liu

We present SPAD, a novel approach for creating consistent multi-view images from text prompts or single images. To enable multi-view generation, we repurpose a pretrained 2D diffusion model by extending its self-attention layers with…

Text-to-3D synthesis has recently emerged as a new approach to sampling 3D models by adopting pretrained text-to-image models as guiding visual priors. An intriguing but underexplored problem with existing text-to-3D methods is that 3D…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Uy Dieu Tran , Minh Luu , Phong Ha Nguyen , Khoi Nguyen , Binh-Son Hua

This paper addresses the limitations of existing 3D Gaussian Splatting (3DGS) methods, particularly their reliance on adaptive density control, which can lead to floating artifacts and inefficient resource usage. We propose a novel densify…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Phurtivilai Patt , Leyang Huang , Yinqiang Zhang , Yang Lei

Precise Text-to-Image (T2I) generation has achieved great success but is hindered by the limited relational reasoning of static text encoders and the error accumulation in open-loop sampling. Without real-time feedback, initial semantic…

人工智能 · 计算机科学 2026-03-20 Ping Chen , Daoxuan Zhang , Xiangming Wang , Yungeng Liu , Haijin Zeng , Yongyong Chen

Text-to-image (T2I) diffusion models have revolutionized generative modeling by producing high-fidelity, diverse, and visually realistic images from textual prompts. Despite these advances, existing models struggle with complex prompts…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Eric Hanchen Jiang , Yasi Zhang , Zhi Zhang , Yixin Wan , Andrew Lizarraga , Shufan Li , Ying Nian Wu

Despite advancements in text-to-image generation (T2I), prior methods often face text-image misalignment problems such as relation confusion in generated images. Existing solutions involve cross-attention manipulation for better…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Leigang Qu , Wenjie Wang , Yongqi Li , Hanwang Zhang , Liqiang Nie , Tat-Seng Chua

Automated axon tracing via fully supervised learning requires large amounts of 3D brain imagery, which is time consuming and laborious to obtain. It also requires expertise. Thus, there is a need for more efficient segmentation and…

图像与视频处理 · 电气工程与系统科学 2023-11-08 Nina I. Shamsi , Alex S. Xu , Lars A. Gjesteby , Laura J. Brattain

Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in remote sensing (RS) images based on textual descriptions. Although Segment Anything Model 2 (SAM2) has shown remarkable performance in various…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Fu Rong , Meng Lan , Qian Zhang , Lefei Zhang