中文
相关论文

相关论文: Group Editing: Edit Multiple Images in One Go

200 篇论文

Scene graphs offer a structured, hierarchical representation of images, with nodes and edges symbolizing objects and the relationships among them. It can serve as a natural interface for image editing, dramatically improving precision and…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Zhiyuan Zhang , DongDong Chen , Jing Liao

We present a method for finding cross-modal space-time correspondences. Given two images from different visual modalities, such as an RGB image and a depth map, our model identifies which pairs of pixels correspond to the same physical…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Ayush Shrivastava , Andrew Owens

The ability of Generative Adversarial Networks to encode rich semantics within their latent space has been widely adopted for facial image editing. However, replicating their success with videos has proven challenging. Sets of high-quality…

计算机视觉与模式识别 · 计算机科学 2022-01-24 Rotem Tzaban , Ron Mokady , Rinon Gal , Amit H. Bermano , Daniel Cohen-Or

Despite remarkable advances in image synthesis research, existing works often fail in manipulating images under the context of large geometric transformations. Synthesizing person images conditioned on arbitrary poses is one of the most…

计算机视觉与模式识别 · 计算机科学 2019-01-14 Haoye Dong , Xiaodan Liang , Ke Gong , Hanjiang Lai , Jia Zhu , Jian Yin

With the rapid advancement of commercial multi-modal models, image editing has garnered significant attention due to its widespread applicability in daily life. Despite impressive progress, existing image editing systems, particularly…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Yiran Zhao , Yaoqi Ye , Xiang Liu , Michael Qizhe Shieh , Trung Bui

Text-driven image editing has achieved remarkable success in following single instructions. However, real-world scenarios often involve complex, multi-step instructions, particularly ``chain'' instructions where operations are…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Chenglin Wang , Yucheng Zhou , Qianning Wang , Zhe Wang , Kai Zhang

The visual dialog task attempts to train an agent to answer multi-turn questions given an image, which requires the deep understanding of interactions between the image and dialog history. Existing researches tend to employ the…

计算与语言 · 计算机科学 2022-02-23 Tong Ye , Shijing Si , Jianzong Wang , Rui Wang , Ning Cheng , Jing Xiao

The accelerated advancement of generative AI significantly enhance the viability and effectiveness of generative regional editing methods. This evolution render the image manipulation more accessible, thereby intensifying the risk of…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Zhihao Sun , Haipeng Fang , Xinying Zhao , Danding Wang , Juan Cao

Recently, large pretrained models (e.g., BERT, StyleGAN, CLIP) have shown great knowledge transfer and generalization capability on various downstream tasks within their domains. Inspired by these efforts, in this paper we propose a unified…

计算机视觉与模式识别 · 计算机科学 2021-12-02 Jing Shi , Ning Xu , Haitian Zheng , Alex Smith , Jiebo Luo , Chenliang Xu

Learning inter-image similarity is crucial for 3D medical images self-supervised pre-training, due to their sharing of numerous same semantic regions. However, the lack of the semantic prior in metrics and the semantic-independent variation…

计算机视觉与模式识别 · 计算机科学 2023-03-03 Yuting He , Guanyu Yang , Rongjun Ge , Yang Chen , Jean-Louis Coatrieux , Boyu Wang , Shuo Li

Recent progress in generative models has significantly advanced image editing capabilities, yet precise and intuitive user control remains difficult. Specifically, users often struggle to communicate both exact spatial layouts and specific…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Anya Ji , George Ma , Téa Wright , Yiming Zhang , David M. Chan , Alane Suhr , Somayeh Sojoudi

Image captioning attempts to generate a sentence composed of several linguistic words, which are used to describe objects, attributes, and interactions in an image, denoted as visual semantic units in this paper. Based on this view, we…

计算机视觉与模式识别 · 计算机科学 2019-08-07 Longteng Guo , Jing Liu , Jinhui Tang , Jiangwei Li , Wei Luo , Hanqing Lu

Given an original image, image editing aims to generate an image that align with the provided instruction. The challenges are to accept multimodal inputs as instructions and a scarcity of high-quality training data, including crucial…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Zhen Han , Chaojie Mao , Zeyinzi Jiang , Yulin Pan , Jingfeng Zhang

The ability to synthesize personalized group photos and specify the positions of each identity offers immense creative potential. While such imagery can be visually appealing, it presents significant challenges for existing technologies. A…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Yimeng Zhang , Tiancheng Zhi , Jing Liu , Shen Sang , Liming Jiang , Qing Yan , Sijia Liu , Linjie Luo

We study the 3D-aware image attribute editing problem in this paper, which has wide applications in practice. Recent methods solved the problem by training a shared encoder to map images into a 3D generator's latent space or by per-image…

计算机视觉与模式识别 · 计算机科学 2023-04-21 Jianhui Li , Jianmin Li , Haoji Zhang , Shilong Liu , Zhengyi Wang , Zihao Xiao , Kaiwen Zheng , Jun Zhu

Utilizing a shared embedding space, emerging multimodal models exhibit unprecedented zero-shot capabilities. However, the shared embedding space could lead to new vulnerabilities if different modalities can be misaligned. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Shaeke Salman , Md Montasir Bin Shams , Xiuwen Liu

Text-based 2D diffusion models have demonstrated impressive capabilities in image generation and editing. Meanwhile, the 2D diffusion models also exhibit substantial potentials for 3D editing tasks. However, how to achieve consistent edits…

计算机视觉与模式识别 · 计算机科学 2024-06-26 Ruihuang Li , Liyi Chen , Zhengqiang Zhang , Varun Jampani , Vishal M. Patel , Lei Zhang

Scalable Vector Graphics are a standard representation for editable visual design, yet they are usually authored as single view two dimensional illustrations. This limits their use in applications that require object level assets to remain…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Mengnan Jiang , Zhaolin Sun , Christian Franke , Michele Franco Adesso , Antonio Haas , Grace Li Zhang

Instance-level object segmentation across disparate egocentric and exocentric views is a fundamental challenge in visual understanding, critical for applications in embodied AI and remote collaboration. This task is exceptionally difficult…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Yulu Gao , Bohao Zhang , Zongheng Tang , Jitong Liao , Wenjun Wu , Si Liu

We propose A^2-Edit, a unified inpainting framework for arbitrary object categories, which allows users to replace any target region with a reference object using only a coarse mask. To address the issues of severe homogenization and…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Huayu Zheng , Guangzhao Li , Baixuan Zhao , Siqi Luo , Hantao Jiang , Guangtao Zhai , Xiaohong Liu