中文
相关论文

相关论文: Why Settle for Mid: A Probabilistic Viewpoint to S…

200 篇论文

Although recent text-to-image generative models have achieved impressive performance, they still often struggle with capturing the compositional complexities of prompts including attribute binding, and spatial relationships between…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Seyed Mohammad Hadi Hosseini , Amir Mohammad Izadi , Ali Abdollahi , Armin Saghafian , Mahdieh Soleymani Baghshah

Text-to-image diffusion models have demonstrated remarkable capability in generating realistic images from arbitrary text prompts. However, they often produce inconsistent results for compositional prompts such as "two dogs" or "a penguin…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Shuangqi Li , Hieu Le , Jingyi Xu , Mathieu Salzmann

Accurate localization on autonomous driving cars is essential for autonomy and driving safety, especially for complex urban streets and search-and-rescue subterranean environments where high-accurate GPS is not available. However current…

计算机视觉与模式识别 · 计算机科学 2021-08-31 Peng Yin , Lingyun Xu , Ziyue Feng , Anton Egorov , Bing Li

State-of-the-art T2I models are capable of generating high-resolution images given textual prompts. However, they still struggle with accurately depicting compositional scenes that specify multiple objects, attributes, and spatial…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Yixin Wan , Kai-Wei Chang

In this paper we contribute a novel algorithm family, which generalizes many unsupervised techniques including unnormalized and energy models, and allows us to infer different statistical modalities (e.g. data likelihood and ratio between…

机器学习 · 计算机科学 2021-01-14 Dmitry Kopitkov , Vadim Indelman

Optimal design for linear regression is a fundamental task in statistics. For finite design spaces, recent progress has shown that random designs drawn using proportional volume sampling (PVS) lead to approximation guarantees for A-optimal…

统计计算 · 统计学 2021-02-02 Arnaud Poinas , Rémi Bardenet

We address the problem of estimating the alignment pose between two models using structure-specific local descriptors. Our descriptors are generated using a combination of 2D image data and 3D contextual shape data, resulting in a set of…

计算机视觉与模式识别 · 计算机科学 2017-08-24 Anders Glent Buch , Dirk Kraft , Joni-Kristian Kamarainen , Henrik Gordon Petersen , Norbert Krüger

Recent works favored dense signals (e.g., depth, DensePose), as an alternative to sparse signals (e.g., OpenPose), to provide detailed spatial guidance for pose-guided text-to-image generation. However, dense representations raised new…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Wenjie Xuan , Jing Zhang , Juhua Liu , Bo Du , Dacheng Tao

Spatial perception aims to estimate camera motion and scene structure from visual observations, a problem traditionally addressed through geometric modeling and physical consistency constraints. Recent learning-based methods have…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Haichao Zhu , Zhaorui Yang , Qian Zhang

Generating images from text involving complex and novel object arrangements remains a significant challenge for current text-to-image (T2I) models. Although prior layout-based methods improve object arrangements using spatial constraints…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Zeeshan Khan , Shizhe Chen , Cordelia Schmid

While computer vision and machine learning have made great progress, their robustness is still challenged by two key issues: data distribution shift and label noise. When domain generalization (DG) encounters noise, noisy labels further…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Wang Lu , Jindong Wang

Despite remarkable advances in video generative models, they still struggle to generate physically realistic videos, frequently exhibiting appearance drift, implausible motion, and temporal inconsistencies. In this work, we address this…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Manjin Kim , Suha Kwak , Minsu Cho

Understanding spatial relations is a crucial cognitive ability for both humans and AI. While current research has predominantly focused on the benchmarking of text-to-image (T2I) models, we propose a more comprehensive evaluation that…

计算机视觉与模式识别 · 计算机科学 2024-11-13 Shang Hong Sim , Clarence Lee , Alvin Tan , Cheston Tan

We study a systematic bias in modern image generation models: the mention order of entities in text spuriously determines spatial layout and entity--role binding. We term this phenomenon Order-to-Space Bias (OTS) and show that it arises in…

计算与语言 · 计算机科学 2026-03-05 Yongkang Zhang , Zonglin Zhao , Yuechen Zhang , Fei Ding , Pei Li , Wenxuan Wang

Content safety is a fundamental challenge for text-to-image (T2I) models, yet prevailing methods enforce a debilitating trade-off between safety and generation quality. We argue that mitigating this trade-off hinges on addressing systemic…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Shouwei Ruan , Zhenyu Wu , Yao Huang , Ruochen Zhang , Yitong Sun , Caixin Kang , Shiji Zhao , Xingxing Wei

Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Gaoyang Zhang , Bingtao Fu , Qingnan Fan , Qi Zhang , Runxing Liu , Hong Gu , Huaqi Zhang , Xinguo Liu

Pose Machines provide a sequential prediction framework for learning rich implicit spatial models. In this work we show a systematic design for how convolutional networks can be incorporated into the pose machine framework for learning…

计算机视觉与模式识别 · 计算机科学 2016-04-13 Shih-En Wei , Varun Ramakrishna , Takeo Kanade , Yaser Sheikh

Estimating position and orientation change of a mobile platform from two consecutive point clouds provided by a high-resolution sensor is a key problem in autonomous navigation. In particular, scan matching algorithms aim to find the…

信号处理 · 电气工程与系统科学 2021-06-09 Rico Mendrzik , Florian Meyer

Deep networks are increasingly being applied to problems involving image synthesis, e.g., generating images from textual descriptions and reconstructing an input image from a compact representation. Supervised training of image-synthesis…

机器学习 · 计算机科学 2017-01-25 Jake Snell , Karl Ridgeway , Renjie Liao , Brett D. Roads , Michael C. Mozer , Richard S. Zemel

State-of-the-art text-to-image models produce visually impressive results but often struggle with precise alignment to text prompts, leading to missing critical elements or unintended blending of distinct concepts. We propose a novel…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Paul Grimal , Michaël Soumm , Hervé Le Borgne , Olivier Ferret , Akihiro Sugimoto