English
Related papers

Related papers: MaGIC: Multi-modality Guided Image Completion

200 papers

Few-shot anomaly generation is a key challenge in industrial quality control. Although diffusion models are promising, existing methods struggle: global prompt-guided approaches corrupt normal regions, and existing inpainting-based methods…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 JaeHyuck Choi , MinJun Kim , Je Hyeong Hong

Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these VLMs are task-specific and assume that both video and language inputs are complete.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Xiang Fang , Wanlong Fang , Changshuo Wang , Keke Tang , Daizong Liu , Siyi Wang , Wei Ji

Depth prediction is a critical problem in robotics applications especially autonomous driving. Generally, depth prediction based on binocular stereo matching and fusion of monocular image and laser point cloud are two mainstream methods.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Guancheng Chen , Junli Lin , Huabiao Qin

Image completion with large-scale free-form missing regions is one of the most challenging tasks for the computer vision community. While researchers pursue better solutions, drawbacks such as pattern unawareness, blurry textures, and…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Xingqian Xu , Shant Navasardyan , Vahram Tadevosyan , Andranik Sargsyan , Yadong Mu , Humphrey Shi

Digital art synthesis is receiving increasing attention in the multimedia community because of engaging the public with art effectively. Current digital art synthesis methods usually use single-modality inputs as guidance, thereby limiting…

Computer Vision and Pattern Recognition · Computer Science 2022-09-29 Nisha Huang , Fan Tang , Weiming Dong , Changsheng Xu

Adapting to diverse user needs at test time is a key challenge in controllable multi-objective generation. Existing methods are insufficient: merging-based approaches provide indirect, suboptimal control at the parameter level, often…

Machine Learning · Computer Science 2025-10-17 Guofu Xie , Chen Zhang , Xiao Zhang , Yunsheng Shi , Ting Yao , Jun Xu

Text-guided image editing has been allowing users to transform and synthesize images through natural language instructions, offering considerable flexibility. However, most existing image editing models naively attempt to follow all user…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Hyunseung Kim , Chiho Choi , Srikanth Malla , Sai Prahladh Padmanabhan , Saurabh Bagchi , Joon Hee Choi

The existing image manipulation localization (IML) models mainly relies on visual cues, but ignores the semantic logical relationships between content features. In fact, the content semantics conveyed by real images often conform to human…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Songlin Li , Zhiqing Guo , Yuanman Li , Zeyu Li , Yunfeng Diao , Gaobo Yang , Liejun Wang

Multi-modal knowledge graph completion (MMKGC) aims to discover unobserved knowledge from given knowledge graphs, collaboratively leveraging structural information from the triples and multi-modal information of the entities to overcome the…

Artificial Intelligence · Computer Science 2024-12-17 Yichi Zhang , Zhuo Chen , Lingbing Guo , Yajing Xu , Binbin Hu , Ziqi Liu , Wen Zhang , Huajun Chen

Recent years have seen a surge of interest in anomaly detection for tackling industrial defect detection, event detection, etc. However, existing unsupervised anomaly detectors, particularly those for the vision modality, face significant…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Dong Chen , Kaihang Pan , Guoming Wang , Yueting Zhuang , Siliang Tang

Large-scale language-vision pre-training models, such as CLIP, have achieved remarkable text-guided image morphing results by leveraging several unconditional generative models. However, existing CLIP-guided image morphing methods encounter…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 Yeongtak Oh , Saehyung Lee , Uiwon Hwang , Sungroh Yoon

Multi-Modal Knowledge Graphs (MMKGs) benefit from visual information, yet large-scale image collection is hard to curate and often excludes ambiguous but relevant visuals (e.g., logos, symbols, abstract scenes). We present Beyond Images, an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Pengyu Zhang , Klim Zaporojets , Jie Liu , Jia-Hong Huang , Paul Groth

Depth completion aims to recover dense depth maps from sparse ones, where color images are often used to facilitate this task. Recent depth methods primarily focus on image guided learning frameworks. However, blurry guidance in the image…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Zhiqiang Yan , Xiang Li , Le Hui , Zhenyu Zhang , Jun Li , Jian Yang

Developing effective path representations has become increasingly essential across various fields within intelligent transportation. Although pre-trained path representation learning models have shown improved performance, they…

Machine Learning · Computer Science 2025-01-03 Ronghui Xu , Hanyin Cheng , Chenjuan Guo , Hongfan Gao , Jilin Hu , Sean Bin Yang , Bin Yang

Generative image compression has recently shown impressive perceptual quality, but often suffers from semantic deviations caused by generative hallucinations at ultra-low bitrate (bpp < 0.05), limiting its reliable deployment in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Kaile Wang , Lijun He , Haisheng Fu , Haixia Bi , Fan Li

For improving image composition and aesthetic quality, most existing methods modulate the captured images by striking out redundant content near the image borders. However, such image cropping methods are limited in the range of image…

Computer Vision and Pattern Recognition · Computer Science 2023-09-22 Xiaoyu Liu , Ming Liu , Junyi Li , Shuai Liu , Xiaotao Wang , Lei Lei , Wangmeng Zuo

There is growing interest in multi-label image classification due to its critical role in web-based image analytics-based applications, such as large-scale image retrieval and browsing. Matrix completion has recently been introduced as a…

Machine Learning · Statistics 2019-04-09 Yong Luo , Tongliang Liu , Dacheng Tao , Chao Xu

Vision-language foundation models, represented by Contrastive Language-Image Pre-training (CLIP), have gained increasing attention for jointly understanding both vision and textual tasks. However, existing approaches primarily focus on…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Bowen Shi , Peisen Zhao , Zichen Wang , Yuhang Zhang , Yaoming Wang , Jin Li , Wenrui Dai , Junni Zou , Hongkai Xiong , Qi Tian , Xiaopeng Zhang

Recent advances in personalized generative models have demonstrated impressive capabilities in producing identity-consistent images of the same individual across diverse scenes. However, most existing methods lack explicit viewpoint control…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Hengjia Li , Jianjin Xu , Keli Cheng , Lei Wang , Ning Bi , Boxi Wu , Fernando De la Torre , Deng Cai

Person re-identification (ReID) has recently benefited from large pretrained vision-language models such as Contrastive Language-Image Pre-Training (CLIP). However, the absence of concrete descriptions necessitates the use of implicit text…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Qianru Han , Xinwei He , Zhi Liu , Sannyuya Liu , Ying Zhang , Jinhai Xiang