中文
相关论文

相关论文: EdiVal-Agent: An Object-Centric Framework for Auto…

200 篇论文

Object detection (OD) in computer vision has made significant progress in recent years, transitioning from closed-set labels to open-vocabulary detection (OVD) based on large-scale vision-language pre-training (VLP). However, current…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Yiyang Yao , Peng Liu , Tiancheng Zhao , Qianqian Zhang , Jiajia Liao , Chunxin Fang , Kyusong Lee , Qing Wang

Vision-language pre-training (VLP) has shown impressive performance on a wide range of cross-modal tasks, where VLP models without reliance on object detectors are becoming the mainstream due to their superior computation efficiency and…

计算机视觉与模式识别 · 计算机科学 2022-11-23 Yuan Yao , Qianyu Chen , Ao Zhang , Wei Ji , Zhiyuan Liu , Tat-Seng Chua , Maosong Sun

While diffusion models have achieved remarkable success in text-to-image generation, they encounter significant challenges with instruction-driven image editing. Our research highlights a key challenge: these models particularly struggle…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Yujia Hu , Songhua Liu , Zhenxiong Tan , Xingyi Yang , Xinchao Wang

The rapid development of large language models (LLMs) and large vision models (LVMs) have propelled the evolution of multi-modal AI systems, which have demonstrated the remarkable potential for industrial applications by emulating…

计算机视觉与模式识别 · 计算机科学 2025-01-10 Di Jin , Xing Liu , Yu Liu , Jia Qing Yap , Andrea Wong , Adriana Crespo , Qi Lin , Zhiyuan Yin , Qiang Yan , Ryan Ye

We introduce a new task called Defeasible Visual Entailment (DVE), where the goal is to allow the modification of the entailment relationship between an image premise and a text hypothesis based on an additional update. While this concept…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Yue Zhang , Liqiang Jing , Vibhav Gogate

Unified vision-language models (VLMs) promise to streamline computer vision pipelines by reframing multiple visual tasks such as classification, detection, and keypoint localization within a single language-driven interface. This…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Conor Wallace , Isaac Corley , Jonathan Lwowski

Natural language instructions are a powerful interface for editing the outputs of text-to-image diffusion models. However, several challenges need to be addressed: 1) underspecification (the need to model the implicit meaning of…

计算与语言 · 计算机科学 2023-10-31 Tuhin Chakrabarty , Kanishk Singh , Arkadiy Saakyan , Smaranda Muresan

Recent advances in text-to-image (T2I) diffusion models have significantly improved semantic image editing, yet most methods fall short in performing 3D-aware object manipulation. In this work, we present FFSE, a 3D-aware autoregressive…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Xincheng Shuai , Zhenyuan Qin , Henghui Ding , Dacheng Tao

Video editing and synthesis often introduce object inconsistencies, such as frame flicker and identity drift that degrade perceptual quality. To address these issues, we introduce ObjectAlign, a novel framework that seamlessly blends…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Mustafa Munir , Harsh Goel , Xiwen Wei , Minkyu Choi , Sahil Shah , Kartikeya Bhardwaj , Paul Whatmough , Sandeep Chinchali , Radu Marculescu

With the rapid advancement of commercial multi-modal models, image editing has garnered significant attention due to its widespread applicability in daily life. Despite impressive progress, existing image editing systems, particularly…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Yiran Zhao , Yaoqi Ye , Xiang Liu , Michael Qizhe Shieh , Trung Bui

Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities. However, evaluating their capacity for human-like understanding in One-Image Guides remains insufficiently explored. One-Image Guides are…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Jiancong Xie , Wenjin Wang , Zhuomeng Zhang , Zihan Liu , Qi Liu , Ke Feng , Zixun Sun , Yuedong Yang

Instruction-based image editing, which aims to modify the image faithfully according to the instruction while preserving irrelevant content unchanged, has made significant progress. However, there still lacks a comprehensive metric for…

图形学 · 计算机科学 2025-06-18 Zhuoying Li , Zhu Xu , Yuxin Peng , Yang Liu

Vision-language instruction-tuning models have recently achieved significant performance improvements. In this work, we discover that large-scale 3D parallel training on those models leads to an imbalanced computation load across different…

人工智能 · 计算机科学 2025-10-14 Yongqiang Yao , Jingru Tan , Feizhao Zhang , Jiahao Hu , Yazhe Niu , Xin Jin , Bo Li , Pengfei Liu , Ruihao Gong , Dahua Lin , Ningyi Xu

An image editing model should be able to perform diverse edits, ranging from object replacement, changing attributes or style, to performing actions or movement, which require many forms of reasoning. Current general instruction-guided…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Benno Krojer , Dheeraj Vattikonda , Luis Lara , Varun Jampani , Eva Portelance , Christopher Pal , Siva Reddy

Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different from the assessment of text-driven image generation,…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Bowen Qu , Shangkun Sun , Xiaoyu Liang , Wei Gao

We introduce V-Agent, a novel multi-agent platform designed for advanced video search and interactive user-system conversations. By fine-tuning a vision-language model (VLM) with a small video preference dataset and enhancing it with a…

计算机视觉与模式识别 · 计算机科学 2026-01-08 SunYoung Park , Jong-Hyeon Lee , Youngjune Kim , Daegyu Sung , Younghyun Yu , Young-rok Cha , Jeongho Ju

We investigate fine-tuning Vision-Language Models (VLMs) for multi-task medical image understanding, focusing on detection, localization, and counting of findings in medical images. Our objective is to evaluate whether instruction-tuned…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Sushant Gautam , Michael A. Riegler , Pål Halvorsen

Blind deconvolution aims to recover an original image from a blurred version in the case where the blurring kernel is unknown. It has wide applications in diverse fields such as astronomy, microscopy, and medical imaging. Blind…

数值分析 · 数学 2024-02-06 Markus Haltmeier , Gyeongha Hwang

Evaluating text-guided image editing (TIE) methods remains a challenging problem, as reliable assessment should simultaneously consider perceptual quality, alignment with textual instructions, and preservation of original image content.…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Shiqi Gao , Zitong Xu , Kang Fu , Huiyu Duan , Xiongkuo Min , Jia wang

With the rapid advancements in Artificial Intelligence Generated Image (AGI) technology, the accurate assessment of their quality has become an increasingly vital requirement. Prevailing methods typically rely on cross-modal models like…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Qiang Li , Qingsen Yan , Haojian Huang , Peng Wu , Haokui Zhang , Yanning Zhang