中文
相关论文

相关论文: Getting it Right: Improving Spatial Consistency in…

200 篇论文

In recent years, Text-to-Image (T2I) models have been extensively studied, especially with the emergence of diffusion models that achieve state-of-the-art results on T2I synthesis tasks. However, existing benchmarks heavily rely on…

计算机视觉与模式识别 · 计算机科学 2023-11-27 Eslam Mohamed Bakr , Pengzhan Sun , Xiaoqian Shen , Faizan Farooq Khan , Li Erran Li , Mohamed Elhoseiny

Spatial confounding poses a significant challenge in scientific studies involving spatial data, where unobserved spatial variables can influence both treatment and outcome, possibly leading to spurious associations. To address this problem,…

Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single-view perception and reason consistently about objects, visibility, geometry, and interactions across multiple viewpoints. However, progress in…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Wei Wang , Yuqian Yuan , Tianwei Lin , Wenqiao Zhang , Siliang Tang , Jun Xiao , Yueting Zhuang

The text to medical image (T2MedI) with latent diffusion model has great potential to alleviate the scarcity of medical imaging data and explore the underlying appearance distribution of lesions in a specific patient status description.…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Xu Han , Fangfang Fan , Jingzhao Rong , Zhen Li , Georges El Fakhri , Qingyu Chen , Xiaofeng Liu

The integration of language and 3D perception is critical for embodied AI and robotic systems to perceive, understand, and interact with the physical world. Spatial reasoning, a key capability for understanding spatial relationships between…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Jiaxin Huang , Ziwen Li , Hanlve Zhang , Runnan Chen , Xiao He , Yandong Guo , Wenping Wang , Tongliang Liu , Mingming Gong

Zero-shot Image Captioning (ZIC) increasingly utilizes synthetic datasets generated by text-to-image (T2I) models to mitigate the need for costly manual annotation. However, these T2I models often produce images that exhibit semantic…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Si-Woo Kim , MinJu Jeon , Ye-Chan Kim , Soeun Lee , Taewhan Kim , Dong-Jin Kim

The issue of generative pretraining for vision models has persisted as a long-standing conundrum. At present, the text-to-image (T2I) diffusion model demonstrates remarkable proficiency in generating high-definition images matching textual…

计算机视觉与模式识别 · 计算机科学 2023-12-25 Qiang Wan , Zilong Huang , Bingyi Kang , Jiashi Feng , Li Zhang

Traditionally, sparse retrieval systems relied on lexical representations to retrieve documents, such as BM25, dominated information retrieval tasks. With the onset of pre-trained transformer models such as BERT, neural sparse retrieval has…

信息检索 · 计算机科学 2023-07-21 Nandan Thakur , Kexin Wang , Iryna Gurevych , Jimmy Lin

Data augmentation has been recently leveraged as an effective regularizer in various vision-language deep neural networks. However, in text-to-image synthesis (T2Isyn), current augmentation wisdom still suffers from the semantic mismatch…

计算机视觉与模式识别 · 计算机科学 2023-12-14 Zhaorui Tan , Xi Yang , Kaizhu Huang

Image-to-text tasks, such as open-ended image captioning and controllable image description, have received extensive attention for decades. Here, we further advance this line of work by presenting Visual Spatial Description (VSD), a new…

计算机视觉与模式识别 · 计算机科学 2022-10-27 Yu Zhao , Jianguo Wei , Zhichao Lin , Yueheng Sun , Meishan Zhang , Min Zhang

Frame of Reference (FoR) is a fundamental concept in spatial reasoning that humans utilize to comprehend and describe space. With the rapid progress in Multimodal Language models, the moment has come to integrate this long-overlooked…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Tanawan Premsri , Parisa Kordjamshidi

Many deep learning tasks require annotations that are too time consuming for human operators, resulting in small dataset sizes. This is especially true for dense regression problems such as crowd counting which requires the location of…

计算机视觉与模式识别 · 计算机科学 2023-02-01 Arian Bakhtiarnia , Qi Zhang , Alexandros Iosifidis

Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralistic alignment, where an AI understands and is steerable towards diverse, and often conflicting,…

Text-to-image (T2I) models have made substantial progress in generating images from textual prompts. However, they frequently fail to produce images consistent with physical commonsense, a vital capability for applications in world…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Fanqing Meng , Wenqi Shao , Lixin Luo , Yahong Wang , Yiran Chen , Quanfeng Lu , Yue Yang , Tianshuo Yang , Kaipeng Zhang , Yu Qiao , Ping Luo

Image inpainting aims to fill missing pixels in damaged images and has achieved significant progress with cut-edging learning techniques. Nevertheless, state-of-the-art inpainting methods are mainly designed for nature images and cannot…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Liang Zhao , Qing Guo , Xiaoguang Li , Song Wang

Personalizing text-to-image diffusion models involves integrating novel visual concepts from a small set of reference images while retaining the model's original generative capabilities. However, this process often leads to overfitting,…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Gihoon Kim , Hyungjin Park , Taesup Kim

Generative foundation models have advanced large-scale text-driven natural image generation, becoming a prominent research trend across various vertical domains. However, in the remote sensing field, there is still a lack of research on…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Chenyang Liu , Keyan Chen , Rui Zhao , Zhengxia Zou , Zhenwei Shi

Unified multimodal generation architectures that jointly produce text and images have recently emerged as a promising direction for text-to-image (T2I) synthesis. However, many existing systems rely on explicit modality switching,…

Reconstructing 3D scenes with high fidelity and efficiency remains a central pursuit in computer vision and graphics. Recent advances in 3D Gaussian Splatting (3DGS) enable photorealistic rendering with Gaussian primitives, yet the modeling…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Guangchi Fang , Bing Wang

Recent action recognition models have achieved impressive results by integrating objects, their locations and interactions. However, obtaining dense structured annotations for each frame is tedious and time-consuming, making these methods…

计算机视觉与模式识别 · 计算机科学 2022-11-30 Elad Ben-Avraham , Roei Herzig , Karttikeya Mangalam , Amir Bar , Anna Rohrbach , Leonid Karlinsky , Trevor Darrell , Amir Globerson
‹ 上一页 1 8 9 10 下一页 ›