中文
相关论文

相关论文: RealDrag: The First Dragging Benchmark with Real T…

200 篇论文

The rapid expansion of mobile internet has resulted in a substantial increase in user-generated content (UGC) images, thereby making the thorough assessment of UGC images both urgent and essential. Recently, multimodal large language models…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Mingxing Li , Rui Wang , Lei Sun , Yancheng Bai , Xiangxiang Chu

Solving the camera-to-robot pose is a fundamental requirement for vision-based robot control, and is a process that takes considerable effort and cares to make accurate. Traditional approaches require modification of the robot via markers,…

计算机视觉与模式识别 · 计算机科学 2023-03-22 Jingpei Lu , Florian Richter , Michael C. Yip

Single image deraining (SID) in real scenarios attracts increasing attention in recent years. Due to the difficulty in obtaining real-world rainy/clean image pairs, previous real datasets suffer from low-resolution images, homogeneous rain…

计算机视觉与模式识别 · 计算机科学 2022-11-22 Wei Li , Qiming Zhang , Jing Zhang , Zhen Huang , Xinmei Tian , Dacheng Tao

The transformative potential of 3D content creation has been progressively unlocked through advancements in generative models. Recently, intuitive drag editing with geometric changes has attracted significant attention in 2D editing yet…

计算机视觉与模式识别 · 计算机科学 2026-01-14 Jiahua Dong , Yu-Xiong Wang

Over the past years, significant progress has been made in creating photorealistic and drivable 3D avatars solely from videos of real humans. However, a core remaining challenge is the fine-grained and user-friendly editing of clothing…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Basavaraj Sunagad , Heming Zhu , Mohit Mendiratta , Adam Kortylewski , Christian Theobalt , Marc Habermann

One of the key shortcomings in current text-to-image (T2I) models is their inability to consistently generate images which faithfully follow the spatial relationships specified in the text prompt. In this paper, we offer a comprehensive…

Drag-based image editing has long suffered from distortions in the target region, largely because the priors of earlier base models, Stable Diffusion, are insufficient to project optimized latents back onto the natural image manifold. With…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Zihan Zhou , Shilin Lu , Shuli Leng , Shaocong Zhang , Zhuming Lian , Xinlei Yu , Adams Wai-Kin Kong

The remarkable ease of use of diffusion models for image generation has led to a proliferation of synthetic content online. While these models are often employed for legitimate purposes, they are also used to generate fake images that…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Giulia Bertazzini , Daniele Baracchi , Dasara Shullani , Isao Echizen , Alessandro Piva

While modern visual generation models excel at creating aesthetically pleasing natural images, they struggle with producing or editing structured visuals like charts, diagrams, and mathematical figures, which demand composition planning,…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Le Zhuo , Songhao Han , Yuandong Pu , Boxiang Qiu , Sayak Paul , Yue Liao , Yihao Liu , Jie Shao , Xi Chen , Si Liu , Hongsheng Li

Recent advancements in diffusion-based image editing pose a significant threat to the authenticity of digital visual content. Traditional embedding-based watermarking methods often introduce perceptible perturbations to maintain robustness,…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Pengzhen Chen , Yanwei Liu , Xiaoyan Gu , Xiaojun Chen , Wu Liu , Weiping Wang

Graphical user interface (GUI) grounding, the process of mapping human instructions to GUI actions, serves as a fundamental basis to autonomous GUI agents. While existing grounding models achieve promising performance to simulate the mouse…

人机交互 · 计算机科学 2026-01-13 Zeyi Liao , Yadong Lu , Boyu Gou , Huan Sun , Ahmed Awadallah

Image editing models are advancing rapidly, yet comprehensive evaluation remains a significant challenge. Existing image editing benchmarks generally suffer from limited task scopes, insufficient evaluation dimensions, and heavy reliance on…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Juntong Wang , Jiarui Wang , Huiyu Duan , Jiaxiang Kang , Guangtao Zhai , Xiongkuo Min

Diffusion-based editing enables realistic modification of local image regions, making AI-generated content harder to detect. Existing AIGC detection benchmarks focus on classifying entire images, overlooking the localization of…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Hai Ci , Ziheng Peng , Pei Yang , Yingxin Xuan , Mike Zheng Shou

Scene text recognition (STR) is a challenging task in computer vision due to the large number of possible text appearances in natural scenes. Most STR models rely on synthetic datasets for training since there are no sufficiently big and…

计算机视觉与模式识别 · 计算机科学 2021-08-17 Rowel Atienza

We introduce $\texttt{ReMOVE}$, a novel reference-free metric for assessing object erasure efficacy in diffusion-based image editing models post-generation. Unlike existing measures such as LPIPS and CLIPScore, $\texttt{ReMOVE}$ addresses…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Aditya Chandrasekar , Goirik Chakrabarty , Jai Bardhan , Ramya Hebbalaguppe , Prathosh AP

All current benchmarks for multimodal deepfake detection manipulate entire frames using various generation techniques, resulting in oversaturated detection accuracies exceeding 94% at the video-level classification. However, these…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Juho Jung , Sangyoun Lee , Jooeon Kang , Yunjin Na

We present Magic Insert, a method for dragging-and-dropping subjects from a user-provided image into a target image of a different style in a physically plausible manner while matching the style of the target image. This work formalizes the…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Nataniel Ruiz , Yuanzhen Li , Neal Wadhwa , Yael Pritch , Michael Rubinstein , David E. Jacobs , Shlomi Fruchter

Recent generative models produce images with a level of authenticity that makes them nearly indistinguishable from real photos and artwork. Potential harmful use cases of these models, necessitate the creation of robust synthetic image…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Delyan Boychev , Radostin Cholakov

We propose a novel meta-learning framework for real-time object tracking with efficient model adaptation and channel pruning. Given an object tracker, our framework learns to fine-tune its model parameters in only a few iterations of…

计算机视觉与模式识别 · 计算机科学 2019-12-05 Ilchae Jung , Kihyun You , Hyeonwoo Noh , Minsu Cho , Bohyung Han

We propose a large-scale dataset of real-world rainy and clean image pairs and a method to remove degradations, induced by rain streaks and rain accumulation, from the image. As there exists no real-world dataset for deraining, current…