English
Related papers

Related papers: EditRefiner: A Human-Aligned Agentic Framework for…

200 papers

Recent advances in text-to-image generation have produced strong single-shot models, yet no individual system reliably executes the long, compositional prompts typical of creative workflows. We introduce Image-POSER, a reflective…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Hossein Mohebbi , Mohammed Abdulrahman , Yanting Miao , Pascal Poupart , Suraj Kothawade

The rapid evolution of AI-generated images poses growing challenges to information integrity and media authenticity. Existing detection approaches face limitations in robustness, interpretability, and generalization across diverse…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Mengfei Liang , Yiting Qu , Yukun Jiang , Michael Backes , Yang Zhang

In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Runze He , Yiji Cheng , Tiankai Hang , Zhimin Li , Yu Xu , Zijin Yin , Shiyi Zhang , Wenxun Dai , Penghui Du , Ao Ma , Chunyu Wang , Qinglin Lu , Jizhong Han , Jiao Dai

Person re-identification (Re-ID) is a crucial task in computer vision, aiming to recognize individuals across non-overlapping camera views. While recent advanced vision-language models (VLMs) excel in logical reasoning and multi-task…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Ke Niu , Haiyang Yu , Mengyang Zhao , Teng Fu , Siyang Yi , Wei Lu , Bin Li , Xuelin Qian , Xiangyang Xue

Recent progress in controllable image generation and editing is largely driven by diffusion-based methods. Although diffusion models perform exceptionally well in specific tasks with tailored designs, establishing a unified model is still…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Jiteng Mu , Nuno Vasconcelos , Xiaolong Wang

Recent Text-to-Image (T2I) generation models such as Stable Diffusion and Imagen have made significant progress in generating high-resolution images based on text descriptions. However, many generated images still suffer from issues such as…

Replacing the background and simultaneously adjusting foreground objects is a challenging task in image editing. Current techniques for generating such images relies heavily on user interactions with image editing softwares, which is a…

Computer Vision and Pattern Recognition · Computer Science 2019-01-15 Yunxuan Xiao , Yikai Li , Yuwei Wu , Lizhen Zhu

Existing text-to-image (T2I) evaluation metrics mainly assess whether generated images align with information explicitly stated in the prompt, but often fail to capture factual requirements that are implicit, externally grounded, or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Youngsun Lim , Cusuh Ham , Pin-Yu Chen , Deepti Ghadiyaram

Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance the reasoning-grounded planning for image editing, we propose DDA-Thinker, a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Hanqing Yang , Qiang Zhou , Yongchao Du , Sashuai Zhou , Zhibin Wang , Jun Song , Tiezheng Ge , Cheng Yu , Bo Zheng

Recent studies have demonstrated the exceptional potentials of leveraging human preference datasets to refine text-to-image generative models, enhancing the alignment between generated images and textual prompts. Despite these advances,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Xun Wu , Shaohan Huang , Furu Wei

The notable gap between user-provided and model-preferred prompts poses a significant challenge for generating high-quality images with text-to-image models, compelling the need for prompt engineering. Current studies on prompt engineering…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Shiyu Wu , Mingzhen Sun , Weining Wang , Yequan Wang , Jing Liu

Text-guided image editing model has achieved great success in general domain. However, directly applying these models to the fashion domain may encounter two issues: (1) Inaccurate localization of editing region; (2) Weak editing magnitude.…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Zechao Zhan , Dehong Gao , Jinxia Zhang , Jiale Huang , Yang Hu , Xin Wang

Non-autoregressive models greatly improve decoding speed over typical sequence-to-sequence models, but suffer from degraded performance. Infilling and iterative refinement models make up some of this gap by editing the outputs of a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-28 Ethan A. Chi , Julian Salazar , Katrin Kirchhoff

Conditional image editing aims to modify a source image according to textual prompts and optional reference guidance. Such editing is crucial in scenarios requiring strict structural control (i.e., anomaly insertion in driving scenes and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Yuhan Pu , Hao Zheng , Ziqian Mo , Hill Zhang , Tianyi Fan , Shuhong Wu , Jiaheng Wei

Diffusion models have achieved remarkable success in image synthesis. However, addressing artifacts and unrealistic regions remains a critical challenge. We propose self-refining diffusion, a novel framework that enhances image generation…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Seoyeon Lee , Gwangyeol Yu , Chaewon Kim , Jonghyuk Park

Recent visual generation models have made major progress in photorealism, typography, instruction following, and interactive editing, yet they still struggle with spatial reasoning, persistent state, long-horizon consistency, and causal…

Diffusion models have achieved remarkable success in generating realistic images but suffer from generating accurate human hands, such as incorrect finger counts or irregular shapes. This difficulty arises from the complex task of learning…

Computer Vision and Pattern Recognition · Computer Science 2024-08-19 Wenquan Lu , Yufei Xu , Jing Zhang , Chaoyue Wang , Dacheng Tao

When generating images from prompts that include specific entities, the model must retain as much entity-specific knowledge as possible. However, the number of entities is almost countless, and new entities emerge; memorizing all of them…

Presentation generation requires deep content research, coherent visual design, and iterative refinement based on observation. However, existing presentation agents often rely on predefined workflows and fixed templates. To address this, we…

Artificial Intelligence · Computer Science 2026-04-21 Hao Zheng , Guozhao Mo , Xinru Yan , Qianhao Yuan , Wenkai Zhang , Xuanang Chen , Yaojie Lu , Hongyu Lin , Xianpei Han , Le Sun

Existing Image Restoration (IR) studies typically focus on task-specific or universal modes individually, relying on the mode selection of users and lacking the cooperation between multiple task-specific/universal restoration modes. This…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Bingchen Li , Xin Li , Yiting Lu , Zhibo Chen
‹ Prev 1 4 5 6 7 8 10 Next ›