中文
相关论文

相关论文: X-Edit: Exact, Explicit, and Explainable Null-Spac…

200 篇论文

We introduce MATEX (Multi-scale Attention and Text-guided Explainability), a novel framework that advances interpretability in medical vision-language models by incorporating anatomically informed spatial reasoning. MATEX synergistically…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Muhammad Imran , Chi Lee , Yugyung Lee

Recent advances in diffusion models have enabled high-quality image generation, leading to increasing demand for post-generation editing that modifies local regions while preserving global structure. Achieving such flexible and precise…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Hanyi Wang , Han Fang , Zheng Wang , Shilin Wang , Ee-Chien Chang

Instruction-based image editing enables precise modifications via natural language prompts, but existing methods face a precision-efficiency tradeoff: fine-tuning demands massive datasets (>10M) and computational resources, while…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Zechuan Zhang , Ji Xie , Yu Lu , Zongxin Yang , Yi Yang

Recently vision transformer models have become prominent models for a range of tasks. These models, however, usually suffer from intensive computational costs and heavy memory requirements, making them impractical for deployment on edge…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Lu Yu , Wei Xiang

Recent text-guided diffusion models provide powerful image generation capabilities. Currently, a massive effort is given to enable the modification of these images using text only as means to offer intuitive and versatile editing. To edit a…

计算机视觉与模式识别 · 计算机科学 2022-11-18 Ron Mokady , Amir Hertz , Kfir Aberman , Yael Pritch , Daniel Cohen-Or

Model editing aims to correct outdated or erroneous knowledge in large models without costly retraining. Recent research discovered that the mid-layer representation of the subject's final token in a prompt has a strong influence on factual…

计算机视觉与模式识别 · 计算机科学 2025-01-24 Qizhou Chen , Taolin Zhang , Chengyu Wang , Xiaofeng He , Dakan Wang , Tingting Liu

Text-driven 3D stylization is a complex and crucial task in the fields of computer vision (CV) and computer graphics (CG), aimed at transforming a bare mesh to fit a target text. Prior methods adopt text-independent multilayer perceptrons…

计算机视觉与模式识别 · 计算机科学 2023-08-07 Yiwei Ma , Xiaioqing Zhang , Xiaoshuai Sun , Jiayi Ji , Haowei Wang , Guannan Jiang , Weilin Zhuang , Rongrong Ji

Combining Vision Large Language Models (VLLMs) with diffusion models offers a powerful method for executing image editing tasks based on human language instructions. However, language instructions alone often fall short in accurately…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Tianshuo Yuan , Yuxiang Lin , Jue Wang , Zhi-Qi Cheng , Xiaolong Wang , Jiao GH , Wei Chen , Xiaojiang Peng

Chest computed tomography (CT) is central to the detection and management of thoracic disease, yet the growing scale and complexity of volumetric imaging increasingly exceed what can be addressed by scan-level prediction alone. Clinically…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Xuguang Bai , Mingxuan Liu , Tongxi Song , Yifei Chen , Hongjia Yang , Kasidit Anmahapong , Zihan Li , Ying Zhou , Qiyuan Tian

Image-driven video editing aims to propagate edit contents from the modified first frame to the remaining frames. Existing methods usually invert the source video to noise using a pre-trained image-to-video (I2V) model and then guide the…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Maomao Li , Yunfei Liu , Yu Li

Vision-language models (VLMs) have recently shown remarkable zero-shot performance in medical image understanding, yet their grounding ability, the extent to which textual concepts align with visual evidence, remains underexplored. In the…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Haozhe Luo , Shelley Zixin Shu , Ziyu Zhou , Sebastian Otalora , Mauricio Reyes

Flow matching models have emerged as a strong alternative to diffusion models, but existing inversion and editing methods designed for diffusion are often ineffective or inapplicable to them. The straight-line, non-crossing trajectories of…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Guanlong Jiao , Biqing Huang , Kuan-Chieh Wang , Renjie Liao

Text-based 2D image editing models have recently reached an impressive level of maturity, motivating a growing body of work that heavily depends on these models to drive 3D edits. While effective for appearance-based modifications, such…

图形学 · 计算机科学 2026-04-30 Etai Sella , Hao Phung , Nitay Amiel , Or Litany , Or Patashnik , Hadar Averbuch-Elor

Vision Transformers (ViTs) have demonstrated superior performance over Convolutional Neural Networks (CNNs) in various vision-related tasks such as classification, object detection, and segmentation due to their use of self-attention…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Fereshteh Baradaran , Mohsen Raji , Azadeh Baradaran , Arezoo Baradaran , Reihaneh Akbarifard

Introducing user-specified visual concepts in image editing is highly practical as these concepts convey the user's intent more precisely than text-based descriptions. We propose FreeEdit, a novel approach for achieving such reference-based…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Runze He , Kai Ma , Linjiang Huang , Shaofei Huang , Jialin Gao , Xiaoming Wei , Jiao Dai , Jizhong Han , Si Liu

Editing complex visual content from ambiguous or partially specified instructions remains a core challenge in vision-language modeling. Existing models can contextualize content but often fail to infer the underlying intent within a…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Umar Khalid , Kashif Munir , Hasan Iqbal , Azib Farooq , Jing Hua , Nazanin Rahnavard , Chen Chen , Victor Zhu , Zhengping Ji

Instruction-based image editing enables natural-language control over visual modifications, yet existing models falter under Instruction-Visual Complexity (IV-Complexity), where intricate instructions meet cluttered or ambiguous scenes. We…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Tianyuan Qu , Lei Ke , Xiaohang Zhan , Longxiang Tang , Yuqi Liu , Bohao Peng , Bei Yu , Dong Yu , Jiaya Jia

Medical image analysis often faces significant challenges due to limited expert-annotated data, hindering both model generalization and clinical adoption. We propose an expert-guided explainable few-shot learning framework that integrates…

图像与视频处理 · 电气工程与系统科学 2025-09-12 Ifrat Ikhtear Uddin , Longwei Wang , KC Santosh

Implicit surface representations are valued for their compactness and continuity, but they pose significant challenges for editing. Despite recent advancements, existing methods often fail to preserve identity and maintain geometric…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Nail Ibrahimli , Julian F. P. Kooij , Liangliang Nan

In recent years, vision transformers (ViTs) have emerged as powerful and promising techniques for computer vision tasks such as image classification, object detection, and segmentation. Unlike convolutional neural networks (CNNs), which…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Shaibal Saha , Lanyu Xu