English
Related papers

Related papers: See and Fix the Flaws: Enabling VLMs and Diffusion…

200 papers

Despite the remarkable capabilities of text-to-image (T2I) generation models, real-world applications often demand fine-grained, iterative image editing that existing methods struggle to provide. Key challenges include granular instruction…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Zihan Liang , Jiahao Sun , Haoran Ma

Weakly-supervised semantic segmentation (WSSS) has achieved remarkable progress using only image-level labels. However, most existing WSSS methods focus on designing new network structures and loss functions to generate more accurate dense…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Wangyu Wu , Xianglin Qiu , Siqi Song , Zhenhong Chen , Xiaowei Huang , Fei Ma , Jimin Xiao

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

Existing Image Restoration (IR) studies typically focus on task-specific or universal modes individually, relying on the mode selection of users and lacking the cooperation between multiple task-specific/universal restoration modes. This…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Bingchen Li , Xin Li , Yiting Lu , Zhibo Chen

Web agents struggle to adapt to new websites due to the scarcity of environment specific tasks and demonstrations. Recent works have explored synthetic data generation to address this challenge, however, they suffer from data quality issues…

Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Yunlong Lin , Zixu Lin , Kunjie Lin , Jinbin Bai , Panwang Pan , Chenxin Li , Haoyu Chen , Zhongdao Wang , Xinghao Ding , Wenbo Li , Shuicheng Yan

Facing scaling laws, video data from the internet becomes increasingly important. However, collecting extensive videos that meet specific needs is extremely labor-intensive and time-consuming. In this work, we study the way to expedite this…

Artificial Intelligence · Computer Science 2025-09-26 Yidan Zhang , Mutian Xu , Yiming Hao , Kun Zhou , Jiahao Chang , Xiaoqiang Liu , Pengfei Wan , Hongbo Fu , Xiaoguang Han

Nowadays, a huge number of images are available. However, retrieving a required image for an ordinary user is a challenging task in computer vision systems. During the past two decades, many types of research have been introduced to improve…

Multimedia · Computer Science 2020-01-30 Amir Vatani , Milad Taleby Ahvanooey , Mostafa Rahimi

To detect GAN generated images, conventional supervised machine learning algorithms require collection of a number of real and fake images from the targeted GAN model. However, the specific model used by the attacker is often unavailable.…

Computer Vision and Pattern Recognition · Computer Science 2019-10-17 Xu Zhang , Svebor Karaman , Shih-Fu Chang

Multimodal generative AI systems like Stable Diffusion, DALL-E, and MidJourney have fundamentally changed how synthetic images are created. These tools drive innovation but also enable the spread of misleading content, false information,…

Recent advances in Artificial Intelligence Generated Content have led to highly realistic synthetic videos, particularly in human-centric scenarios involving speech, gestures, and full-body motion, posing serious threats to information…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Zhipei Xu , Xuanyu Zhang , Qing Huang , Xing Zhou , Jian Zhang

Deep neural networks can generate images that are astonishingly realistic, so much so that it is often hard for humans to distinguish them from actual photos. These achievements have been largely made possible by Generative Adversarial…

Computer Vision and Pattern Recognition · Computer Science 2020-06-29 Joel Frank , Thorsten Eisenhofer , Lea Schönherr , Asja Fischer , Dorothea Kolossa , Thorsten Holz

With the advancement of AIGC (AI-generated content) technologies, an increasing number of generative models are revolutionizing fields such as video editing, music generation, and even film production. However, due to the limitations of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Daoan Zhang , Wenlin Yao , Xiaoyang Wang , Yebowen Hu , Jiebo Luo , Dong Yu

Stylized Text-to-Image Generation (STIG) aims to generate images from text prompts and style reference images. In this paper, we present ArtWeaver, a novel framework that leverages pretrained Stable Diffusion (SD) to address challenges such…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Chengming Xu , Kai Hu , Qilin Wang , Donghao Luo , Jiangning Zhang , Xiaobin Hu , Yanwei Fu , Chengjie Wang

Advances in deep generative networks have led to impressive results in recent years. Nevertheless, such models can often waste their capacity on the minutiae of datasets, presumably due to weak inductive biases in their decoders. This is…

Computer Vision and Pattern Recognition · Computer Science 2018-04-05 Yaroslav Ganin , Tejas Kulkarni , Igor Babuschkin , S. M. Ali Eslami , Oriol Vinyals

Generative AI models can produce high-quality images based on text prompts. The generated images often appear indistinguishable from images generated by conventional optical photography devices or created by human artists (i.e., real…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Yuying Li , Zeyan Liu , Junyi Zhao , Liangqin Ren , Fengjun Li , Jiebo Luo , Bo Luo

We introduce the concept of "Design Agents" for engineering applications, particularly focusing on the automotive design process, while emphasizing that our approach can be readily extended to other engineering and design domains. Our…

Artificial Intelligence · Computer Science 2025-12-04 Mohamed Elrefaie , Janet Qian , Raina Wu , Qian Chen , Angela Dai , Faez Ahmed

Visual reasoning -- the ability to interpret the visual world -- is crucial for embodied agents that operate within three-dimensional scenes. Progress in AI has led to vision and language models capable of answering questions from images.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Damiano Marsili , Rohun Agrawal , Yisong Yue , Georgia Gkioxari

Image Restoration (IR) agents, leveraging multimodal large language models to perceive degradation and invoke restoration tools, have shown promise in automating IR tasks. However, existing IR agents typically lack an insight summarization…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Yijian Wang , Qingsen Yan , Jiantao Zhou , Duwei Dai , Wei Dong

This paper proposes an extension to the Generative Adversarial Networks (GANs), namely as ARTGAN to synthetically generate more challenging and complex images such as artwork that have abstract characteristics. This is in contrast to most…

Computer Vision and Pattern Recognition · Computer Science 2017-04-20 Wei Ren Tan , Chee Seng Chan , Hernan Aguirre , Kiyoshi Tanaka