中文
相关论文

相关论文: MegaSR: Mining Customized Semantics and Expressive…

200 篇论文

Recent advances in text-to-image (T2I) generation have achieved impressive results, yet existing models often struggle with simple or underspecified prompts, leading to suboptimal image-text alignment, aesthetics, and quality. We propose a…

计算与语言 · 计算机科学 2025-10-16 Ruibo Chen , Jiacheng Pan , Heng Huang , Zhenheng Yang

Spatial understanding is a fundamental aspect of computer vision and integral for human-level reasoning about images, making it an important component for grounded language understanding. While recent text-to-image synthesis (T2I) models…

计算机视觉与模式识别 · 计算机科学 2023-10-30 Tejas Gokhale , Hamid Palangi , Besmira Nushi , Vibhav Vineet , Eric Horvitz , Ece Kamar , Chitta Baral , Yezhou Yang

The objective of image super-resolution is to generate clean and high-resolution images from degraded versions. Recent advancements in diffusion modeling have led to the emergence of various image super-resolution techniques that leverage…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Haolan Chen , Jinhua Hao , Kai Zhao , Kun Yuan , Ming Sun , Chao Zhou , Wei Hu

Advancements in text-to-image generative AI with large multimodal models are spreading into the field of image compression, creating high-quality representation of images at extremely low bit rates. This work introduces novel components to…

图像与视频处理 · 电气工程与系统科学 2025-06-02 Cheng-Lin Wu , Hyomin Choi , Ivan V. Bajić

Single image super-resolution (SISR) is the task of inferring a high-resolution image from a single low-resolution image. Recent research on super-resolution has achieved great progress due to the development of deep convolutional neural…

图像与视频处理 · 电气工程与系统科学 2019-11-22 Zhengyang Lu , Ying Chen

Text-to-image (T2I) generation has achieved remarkable progress in instruction following and aesthetics. However, a persistent challenge is the prevalence of physical artifacts, such as anatomical and structural flaws, which severely…

计算机视觉与模式识别 · 计算机科学 2025-09-15 Jia Wang , Jie Hu , Xiaoqi Ma , Hanghang Ma , Yanbing Zeng , Xiaoming Wei

Medical image super-resolution (MedSR) is essential for improving diagnostic precision across diverse imaging modalities such as MRI, CT, X-ray, Ultrasound, and Fundus imaging. Despite rapid advances in deep learning, challenges remain in…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Subhash Gurappa , Trivikram Satharasi , Yashas Hariprasad , Sundararaj Sitharama Iyengar

Semantic communications have gained significant attention as a promising approach to address the transmission bottleneck, especially with the continuous development of 6G techniques. Distinct from the well investigated physical channel…

信号处理 · 电气工程与系统科学 2024-03-15 Xiang Peng , Zhijin Qin , Xiaoming Tao , Jianhua Lu , Khaled B. Letaief

Text-to-image (T2I) generation models have made significant strides but still struggle with prompt sensitivity: even minor changes in prompt wording can yield inconsistent or inaccurate outputs. To address this challenge, we introduce a…

机器学习 · 计算机科学 2025-07-31 Mohammad Abdul Hafeez Khan , Yash Jain , Siddhartha Bhattacharyya , Vibhav Vineet

Recent advances in text-to-image (T2I) diffusion models have significantly improved the quality of generated images. However, providing efficient control over individual subjects, particularly the attributes characterizing them, remains a…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Stefan Andreas Baumann , Felix Krause , Michael Neumayr , Nick Stracke , Melvin Sevi , Vincent Tao Hu , Björn Ommer

Recent methods exploit the powerful text-to-image (T2I) diffusion models for real-world image super-resolution (Real-ISR) and achieve impressive results compared to previous models. However, we observe two kinds of inconsistencies in…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Junhao Gu , Peng-Tao Jiang , Hao Zhang , Mi Zhou , Jinwei Chen , Wenming Yang , Bo Li

Although deep learning has advanced remote sensing change detection (RSCD), most methods rely solely on image modality, limiting feature representation, change pattern modeling, and generalization especially under illumination and noise…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Yijun Zhou , Yikui Zhai , Zilu Ying , Tingfeng Xian , Wenlve Zhou , Zhiheng Zhou , Xiaolin Tian , Xudong Jia , Hongsheng Zhang , C. L. Philip Chen

Text-to-Image (T2I) and multimodal large language models (MLLMs) have been adopted in solutions for several computer vision and multimodal learning tasks. However, it has been found that such vision-language models lack the ability to…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Agneet Chatterjee , Yiran Luo , Tejas Gokhale , Yezhou Yang , Chitta Baral

Existing diffusion-based super-resolution approaches often exhibit semantic ambiguities due to inaccuracies and incompleteness in their text conditioning, coupled with the inherent tendency for cross-attention to divert towards irrelevant…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Chen Chen , Majid Abdolshah , Violetta Shevchenko , Hongdong Li , Chang Xu , Pulak Purkait

Rapid advances in text-to-image (T2I) generation have raised higher requirements for evaluation methodologies. Existing benchmarks center on objective capabilities and dimensions, but lack an application-scenario perspective, limiting…

人工智能 · 计算机科学 2025-09-23 Xiaojing Dong , Weilin Huang , Liang Li , Yiying Li , Shu Liu , Tongtong Ou , Shuang Ouyang , Yu Tian , Fengxuan Zhao

Referring Image Segmentation (RIS) consistently requires language and appearance semantics to more understand each other. The need becomes acute especially under hard situations. To achieve, existing works tend to resort to various…

计算机视觉与模式识别 · 计算机科学 2024-05-16 Jiaxing Yang , Lihe Zhang , Jiayu Sun , Huchuan Lu

We introduce ``Idea to Image,'' a system that enables multimodal iterative self-refinement with GPT-4V(ision) for automatic image design and generation. Humans can quickly identify the characteristics of different text-to-image (T2I) models…

计算机视觉与模式识别 · 计算机科学 2024-08-15 Zhengyuan Yang , Jianfeng Wang , Linjie Li , Kevin Lin , Chung-Ching Lin , Zicheng Liu , Lijuan Wang

Real-world image super-resolution (Real-ISR) must handle complex degradations and inherent reconstruction ambiguities. While generative models have improved perceptual quality, a key trade-off remains with computational cost. One-step…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Yun Kai Zhuang

Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Shilong Zhang , He Zhang , Zhifei Zhang , Chongjian Ge , Shuchen Xue , Shaoteng Liu , Mengwei Ren , Soo Ye Kim , Yuqian Zhou , Qing Liu , Daniil Pakhomov , Kai Zhang , Zhe Lin , Ping Luo

Image super-resolution (SR) has attracted increasing attention due to its wide applications. However, current SR methods generally suffer from over-smoothing and artifacts, and most work only with fixed magnifications. This paper introduces…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Sicheng Gao , Xuhui Liu , Bohan Zeng , Sheng Xu , Yanjing Li , Xiaoyan Luo , Jianzhuang Liu , Xiantong Zhen , Baochang Zhang