中文
相关论文

相关论文: Semantic Image Synthesis with Semantically Coupled…

200 篇论文

Semantic image synthesis (SIS) refers to the problem of generating realistic imagery given a semantic segmentation mask that defines the spatial layout of object classes. Most of the approaches in the literature, other than the quality of…

计算机视觉与模式识别 · 计算机科学 2023-07-12 Tomaso Fontanini , Claudio Ferrari , Massimo Bertozzi , Andrea Prati

Despite their recent successes, GAN models for semantic image synthesis still suffer from poor image quality when trained with only adversarial supervision. Historically, additionally employing the VGG-based perceptual loss has helped to…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Vadim Sushko , Edgar Schönfeld , Dan Zhang , Juergen Gall , Bernt Schiele , Anna Khoreva

Deep neural networks with discrete latent variables offer the promise of better symbolic reasoning, and learning abstractions that are more useful to new tasks. There has been a surge in interest in discrete latent variable models, however,…

机器学习 · 计算机科学 2018-07-23 Aurko Roy , Ashish Vaswani , Arvind Neelakantan , Niki Parmar

World model-based policy evaluation is a practical proxy for testing real-world robot control by rolling out candidate actions in action-conditioned video diffusion models. As these models increasingly adopt latent diffusion modeling (LDM),…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Nilaksh , Saurav Jha , Artem Zholus , Sarath Chandar

This work presents a dual-agent \ac{llm}-based reasoning framework for automated planar mechanism synthesis that tightly couples linguistic specification with symbolic representation and simulation. From a natural-language task description,…

人工智能 · 计算机科学 2025-10-09 João Pedro Gandarela , Thiago Rios , Stefan Menzel , André Freitas

Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs remains a challenge. Conventional approaches rely on conditioning mechanisms, such as text…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Nithesh Chandher Karthikeyan , Jonas Unger , Gabriel Eilertsen

Data-free quantization (DFQ) enables model quantization without accessing real data, addressing concerns regarding data security and privacy. With the growing adoption of Vision Transformers (ViTs), DFQ for ViTs has garnered significant…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Yunshan Zhong , Yuyao Zhou , Yuxin Zhang , Wanchen Sui , Shen Li , Yong Li , Fei Chao , Rongrong Ji

We propose a method for scene-level sketch-to-photo synthesis with text guidance. Although object-level sketch-to-photo synthesis has been widely studied, whole-scene synthesis is still challenging without reference photos that adequately…

计算机视觉与模式识别 · 计算机科学 2023-02-15 AprilPyone MaungMaung , Makoto Shing , Kentaro Mitsui , Kei Sawada , Fumio Okura

Automatic image synthesis research has been rapidly growing with deep networks getting more and more expressive. In the last couple of years, we have observed images of digits, indoor scenes, birds, chairs, etc. being automatically…

计算机视觉与模式识别 · 计算机科学 2016-12-02 Levent Karacan , Zeynep Akata , Aykut Erdem , Erkut Erdem

Although autoregressive models have achieved promising results on image generation, their unidirectional generation process prevents the resultant images from fully reflecting global contexts. To address the issue, we propose an effective…

计算机视觉与模式识别 · 计算机科学 2022-06-10 Doyup Lee , Chiheon Kim , Saehoon Kim , Minsu Cho , Wook-Shin Han

Synthesizing photo-realistic images from text descriptions is a challenging problem. Previous studies have shown remarkable progresses on visual quality of the generated images. In this paper, we consider semantics from the input text…

计算机视觉与模式识别 · 计算机科学 2019-04-03 Guojun Yin , Bin Liu , Lu Sheng , Nenghai Yu , Xiaogang Wang , Jing Shao

Visual generative AI models often encounter challenges related to text-image alignment and reasoning limitations. This paper presents a novel method for selectively enhancing the signal at critical denoising steps, optimizing image…

计算机视觉与模式识别 · 计算机科学 2025-04-25 Paul Grimal , Hervé Le Borgne , Olivier Ferret

Discretization of semantic features enables interoperability between semantic and digital communication systems, showing significant potential for practical applications. The fundamental difficulty in digitizing semantic features stems from…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Jianqiao Chen , Tingting Zhu , Huishi Song , Nan Ma , Xiaodong Xu

Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Shilong Zhang , He Zhang , Zhifei Zhang , Chongjian Ge , Shuchen Xue , Shaoteng Liu , Mengwei Ren , Soo Ye Kim , Yuqian Zhou , Qing Liu , Daniil Pakhomov , Kai Zhang , Zhe Lin , Ping Luo

A latent denoising semantic communication (SemCom) framework is proposed for robust image transmission over noisy channels. By incorporating a learnable latent denoiser into the receiver, the received signals are preprocessed to effectively…

机器学习 · 计算机科学 2025-05-19 Mingkai Xu , Yongpeng Wu , Yuxuan Shi , Xiang-Gen Xia , Wenjun Zhang , Ping Zhang

This paper addresses the limitations of adverse weather image restoration approaches trained on synthetic data when applied to real-world scenarios. We formulate a semi-supervised learning framework employing vision-language models to…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Jiaqi Xu , Mengyang Wu , Xiaowei Hu , Chi-Wing Fu , Qi Dou , Pheng-Ann Heng

ControlNet has enabled detailed spatial control in text-to-image diffusion models by incorporating additional visual conditions such as depth or edge maps. However, its effectiveness heavily depends on the availability of visual conditions…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Woosung Joung , Daewon Chae , Jinkyu Kim

Deep neural networks can form high-level hierarchical representations of input data. Various researchers have demonstrated that these representations can be used to enable a variety of useful applications. However, such representations are…

计算机视觉与模式识别 · 计算机科学 2020-02-25 Burkay Donderici , Caleb New , Chenliang Xu

Auto-Regressive (AR) models have recently gained prominence in image generation, often matching or even surpassing the performance of diffusion models. However, one major limitation of AR models is their sequential nature, which processes…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Doohyuk Jang , Sihwan Park , June Yong Yang , Yeonsung Jung , Jihun Yun , Souvik Kundu , Sung-Yub Kim , Eunho Yang

Vector quantization approaches (VQ-VAE, VQ-GAN) learn discrete neural representations of images, but these representations are inherently position-dependent: codes are spatially arranged and contextually entangled, requiring autoregressive…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Jamie S. J. Stirling , Noura Al-Moubayed , Hubert P. H. Shum