中文
相关论文

相关论文: A$^2$-Edit: Precise Reference-Guided Image Editing…

200 篇论文

While recent flow-based image editing models demonstrate general-purpose capabilities across diverse tasks, they often struggle to specialize in challenging scenarios -- particularly those involving large-scale shape transformations. When…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Zeqian Long , Mingzhe Zheng , Kunyu Feng , Xinhua Zhang , Hongyu Liu , Harry Yang , Linfeng Zhang , Qifeng Chen , Yue Ma

Corner cases are crucial for training and validating autonomous driving systems, yet collecting them from the real world is often costly and hazardous. Editing objects within captured sensor data offers an effective alternative for…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Jiusi Li , Jackson Jiang , Jinyu Miao , Miao Long , Tuopu Wen , Peijin Jia , Shengxiang Liu , Chunlei Yu , Maolin Liu , Yuzhan Cai , Kun Jiang , Mengmeng Yang , Diange Yang

Diffusion-based text-to-image (T2I) models have demonstrated remarkable results in global video editing tasks. However, their focus is primarily on global video modifications, and achieving desired attribute-specific changes remains a…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Haoyu Zheng , Wenqiao Zhang , Zheqi Lv , Yu Zhong , Yang Dai , Jianxiang An , Yongliang Shen , Juncheng Li , Dongping Zhang , Siliang Tang , Yueting Zhuang

Editing images via instruction provides a natural way to generate interactive content, but it is a big challenge due to the higher requirement of scene understanding and generation. Prior work utilizes a chain of large language models,…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Liya Ji , Chenyang Qi , Qifeng Chen

Image editing has achieved impressive results with the development of large-scale generative models. However, existing models mainly focus on the editing effects of intended objects and regions, often leading to unwanted changes in…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Yuhui Wu , Chenxi Xie , Ruibin Li , Liyi Chen , Qiaosi Yi , Lei Zhang

Advanced image editing software enables easy creation of highly convincing image manipulations, which has been made even more accessible in recent years due to advances in generative AI. Manipulated images, while often harmless, could…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Keanu Nichols , Divya Appapogu , Giscard Biamby , Dina Bashkirova , Anna Rohrbach , Bryan A. Plummer

We present UniRef-Image-Edit, a high-performance multi-modal generation system that unifies single-image editing and multi-image composition within a single framework. Existing diffusion-based editing methods often struggle to maintain…

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Ye Liu , Zongyang Ma , Junfu Pu , Zhongang Qi , Yang Wu , Ying Shan , Chang Wen Chen

Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still struggle with handling diverse and complex image editing tasks that require implicit reasoning,…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Feng Han , Yibin Wang , Chenglin Li , Zheming Liang , Dianyi Wang , Yang Jiao , Zhipeng Wei , Chao Gong , Cheng Jin , Jingjing Chen , Jiaqi Wang

Deep convolutional neural network (DCNN) is the state-of-the-art method for image segmentation, which is one of key challenging computer vision tasks. However, DCNN requires a lot of training images with corresponding image masks to get a…

计算机视觉与模式识别 · 计算机科学 2018-09-19 Chuanhai Zhang , Kurt Loken , Zhiyu Chen , Zhiyong Xiao , Gary Kunkel

Image compositing is one of the most fundamental steps in creative workflows. It involves taking objects/parts of several images to create a new image, called a composite. Currently, this process is done manually by creating accurate masks…

计算机视觉与模式识别 · 计算机科学 2022-05-03 Kerem Turgutlu , Sanat Sharma , Jayant Kumar

Cartoon editing, appreciated by both professional illustrators and hobbyists, allows extensive creative freedom and the development of original narratives within the cartoon domain. However, the existing literature on cartoon editing is…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Jian Lin , Chengze Li , Xueting Liu , Zhongping Ge

Recent advances in unsupervised learning for object detection, segmentation, and tracking hold significant promise for applications in robotics. A common approach is to frame these tasks as inference in probabilistic latent-variable models.…

机器人学 · 计算机科学 2021-09-14 Yizhe Wu , Oiwi Parker Jones , Martin Engelcke , Ingmar Posner

This paper extends the popular task of multi-object tracking to multi-object tracking and segmentation (MOTS). Towards this goal, we create dense pixel-level annotations for two existing tracking datasets using a semi-automatic annotation…

计算机视觉与模式识别 · 计算机科学 2019-04-09 Paul Voigtlaender , Michael Krause , Aljosa Osep , Jonathon Luiten , Berin Balachandar Gnana Sekar , Andreas Geiger , Bastian Leibe

Stereo image matching is a fundamental task in computer vision, photogrammetry and remote sensing, but there is an almost unexplored field, i.e., polygon matching, which faces the following challenges: disparity discontinuity, scale…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Chang Li , Xingtao Peng

Masked (or absorbing) diffusion is actively explored as an alternative to autoregressive models for generative modeling of discrete data. However, existing work in this area has been hindered by unnecessarily complex model formulations and…

机器学习 · 计算机科学 2025-01-17 Jiaxin Shi , Kehang Han , Zhe Wang , Arnaud Doucet , Michalis K. Titsias

We introduce Alterbute, a diffusion-based method for editing an object's intrinsic attributes in an image. We allow changing color, texture, material, and even the shape of an object, while preserving its perceived identity and scene…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Tal Reiss , Daniel Winter , Matan Cohen , Alex Rav-Acha , Yael Pritch , Ariel Shamir , Yedid Hoshen

State-of-the-art image inpainting approaches can suffer from generating distorted structures and blurry textures in high-resolution images (e.g., 512x512). The challenges mainly drive from (1) image content reasoning from distant contexts,…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Yanhong Zeng , Jianlong Fu , Hongyang Chao , Baining Guo

Diffusion-based Image Editing (DIE) is an emerging research hot-spot, which often applies a semantic mask to control the target area for diffusion-based editing. However, most existing solutions obtain these masks via manual operations or…

计算机视觉与模式识别 · 计算机科学 2024-01-24 Siyu Zou , Jiji Tang , Yiyi Zhou , Jing He , Chaoyi Zhao , Rongsheng Zhang , Zhipeng Hu , Xiaoshuai Sun

The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Zhihong Chen , Xuehai Bai , Yang Shi , Chaoyou Fu , Huanyu Zhang , Haotian Wang , Xiaoyan Sun , Zhang Zhang , Liang Wang , Yuanxing Zhang , Pengfei Wan , Yi-Fan Zhang
‹ 上一页 1 8 9 10 下一页 ›