中文
相关论文

相关论文: Semantically Multi-modal Image Synthesis

200 篇论文

Multimodal Sarcasm Explanation (MuSE) is a new yet challenging task, which aims to generate a natural language sentence for a multimodal social post (an image as well as its caption) to explain why it contains sarcasm. Although the existing…

计算与语言 · 计算机科学 2023-06-30 Liqiang Jing , Xuemeng Song , Kun Ouyang , Mengzhao Jia , Liqiang Nie

Cross-domain mapping has been a very active topic in recent years. Given one image, its main purpose is to translate it to the desired target domain, or multiple domains in the case of multiple labels. This problem is highly challenging due…

计算机视觉与模式识别 · 计算机科学 2019-09-06 Andrés Romero , Pablo Arbeláez , Luc Van Gool , Radu Timofte

Recent advancement in computer vision has significantly lowered the barriers to artistic creation. Exemplar-based image translation methods have attracted much attention due to flexibility and controllability. However, these methods hold…

计算机视觉与模式识别 · 计算机科学 2024-01-26 Wei Guo , Yuqi Zhang , De Ma , Qian Zheng

We introduce a segmentation-guided approach to synthesise images that integrate features from two distinct domains. Images synthesised by our dual-domain model belong to one domain within the semantic mask, and to another in the rest of the…

计算机视觉与模式识别 · 计算机科学 2022-04-20 Dena Bazazian , Andrew Calway , Dima Damen

The emergence of diverse generative vision models has recently enabled the synthesis of visually realistic images, underscoring the critical need for effectively detecting these generated images from real photos. Despite advances in this…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Huangsen Cao , Yongwei Wang , Yinfeng Liu , Sixian Zheng , Kangtao Lv , Zhimeng Zhang , Bo Zhang , Xin Ding , Fei Wu

Recent generative models produce images with a level of authenticity that makes them nearly indistinguishable from real photos and artwork. Potential harmful use cases of these models, necessitate the creation of robust synthetic image…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Delyan Boychev , Radostin Cholakov

Image compression techniques typically focus on compressing rectangular images for human consumption, however, resulting in transmitting redundant content for downstream applications. To overcome this limitation, some previous works propose…

图像与视频处理 · 电气工程与系统科学 2025-03-04 Ruoyu Feng , Yixin Gao , Xin Jin , Runsen Feng , Zhibo Chen

Multimodal summarization (MS) aims to generate a summary from multimodal input. Previous works mainly focus on textual semantic coverage metrics such as ROUGE, which considers the visual content as supplemental data. Therefore, the summary…

人工智能 · 计算机科学 2023-02-21 Litian Zhang , Xiaoming Zhang , Ziming Guo , Zhipeng Liu

In hyperspectral remote sensing field, some downstream dense prediction tasks, such as semantic segmentation (SS) and change detection (CD), rely on supervised learning to improve model performance and require a large amount of manually…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Wendi Liu , Pei Yang , Wenhui Hong , Xiaoguang Mei , Jiayi Ma

Deep learning algorithms require extensive data to achieve robust performance. However, data availability is often restricted in the medical domain due to patient privacy concerns. Synthetic data presents a possible solution to these…

Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote…

图像与视频处理 · 电气工程与系统科学 2024-10-31 Jialin Luo , Yuanzhi Wang , Ziqi Gu , Yide Qiu , Shuaizhen Yao , Fuyun Wang , Chunyan Xu , Wenhua Zhang , Dan Wang , Zhen Cui

Synthesis of face images from visual attributes is an important problem in computer vision and biometrics due to its applications in law enforcement and entertainment. Recent advances in deep generative networks have made it possible to…

计算机视觉与模式识别 · 计算机科学 2022-01-14 Xing Di , Vishal M. Patel

Synthesizing high-quality, realistic images from text-descriptions is a challenging task, and current methods synthesize images from text in a multi-stage manner, typically by first generating a rough initial image and then refining image…

计算机视觉与模式识别 · 计算机科学 2021-10-18 Amrit Diggavi Seshadri , Balaraman Ravindran

We present an end-to-end, multimodal, fully convolutional network for extracting semantic structures from document images. We consider document semantic structure extraction as a pixel-wise segmentation task, and propose a unified model…

计算机视觉与模式识别 · 计算机科学 2017-06-09 Xiao Yang , Ersin Yumer , Paul Asente , Mike Kraley , Daniel Kifer , C. Lee Giles

Cross-modal retrieval between visual data and natural language description remains a long-standing challenge in multimedia. While recent image-text retrieval methods offer great promise by learning deep representations aligned across…

Scene text image super-resolution has significantly improved the accuracy of scene text recognition. However, many existing methods emphasize performance over efficiency and ignore the practical need for lightweight solutions in deployment…

计算机视觉与模式识别 · 计算机科学 2024-03-21 LeoWu TomyEnrique , Xiangcheng Du , Kangliang Liu , Han Yuan , Zhao Zhou , Cheng Jin

Single-Image-Super-Resolution (SISR) is a classical computer vision problem that has benefited from the recent advancements in deep learning methods, especially the advancements of convolutional neural networks (CNN). Although…

计算机视觉与模式识别 · 计算机科学 2022-04-26 Mustafa Ayazoglu

Training of semantic segmentation models for material analysis requires micrographs and their corresponding masks. It is quite unlikely that perfect masks will be drawn, especially at the edges of objects, and sometimes the amount of data…

计算机视觉与模式识别 · 计算机科学 2024-08-02 Matias Oscar Volman Stern , Dominic Hohs , Andreas Jansche , Timo Bernthaler , Gerhard Schneider

Digital art synthesis is receiving increasing attention in the multimedia community because of engaging the public with art effectively. Current digital art synthesis methods usually use single-modality inputs as guidance, thereby limiting…

计算机视觉与模式识别 · 计算机科学 2022-09-29 Nisha Huang , Fan Tang , Weiming Dong , Changsheng Xu

Conditional image synthesis for generating photorealistic images serves various applications for content editing to content generation. Previous conditional image synthesis algorithms mostly rely on semantic maps, and often fail in complex…

计算机视觉与模式识别 · 计算机科学 2020-04-23 Aysegul Dundar , Karan Sapra , Guilin Liu , Andrew Tao , Bryan Catanzaro