English
Related papers

Related papers: VTEdit-Bench: A Comprehensive Benchmark for Multi-…

200 papers

Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Juntong Wang , Jiarui Wang , Huiyu Duan , Guangtao Zhai , Xiongkuo Min

We introduce $\texttt{Complex-Edit}$, a comprehensive benchmark designed to systematically evaluate instruction-based image editing models across instructions of varying complexity. To develop this benchmark, we harness GPT-4o to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Siwei Yang , Mude Hui , Bingchen Zhao , Yuyin Zhou , Nataniel Ruiz , Cihang Xie

The recent introduction of Unified Virtual Memory (UVM) in GPUs offers a new programming model that allows GPUs and CPUs to share the same virtual memory space, which shifts the complex memory management from programmers to GPU driver/…

Hardware Architecture · Computer Science 2020-10-22 Yongbin Gu , Wenxuan Wu , Yunfan Li , Lizhong Chen

Understanding how effectively large vision language models (VLMs) compare visual inputs is crucial across numerous applications, yet this fundamental capability remains insufficiently assessed. While VLMs are increasingly deployed for tasks…

Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks which couple powerful vision-language model (VLM) with…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Yuxin Song , Wenkai Dong , Shizun Wang , Qi Zhang , Song Xue , Tao Yuan , Hu Yang , Haocheng Feng , Hang Zhou , Xinyan Xiao , Jingdong Wang

Recent efforts to accelerate inference in Multimodal Large Language Models (MLLMs) have largely focused on visual token compression. The effectiveness of these methods is commonly evaluated by measuring the accuracy drop on existing MLLM…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Chenfei Liao , Wensong Wang , Zichen Wen , Xu Zheng , Yiyu Wang , Haocong He , Yuanhuiyi Lyu , Lutao Jiang , Xin Zou , Yuqian Fu , Bin Ren , Linfeng Zhang , Xuming Hu

Recent advancements in Unified Multimodal Models (UMMs) have enabled remarkable image understanding and generation capabilities. However, while models like Gemini-2.5-Flash-Image show emerging abilities to reason over multiple related…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Mingrui Wu , Hang Liu , Jiayi Ji , Xiaoshuai Sun , Rongrong Ji

Text-guided image editing has seen significant progress in natural image domains, but its application in medical imaging remains limited and lacks standardized evaluation frameworks. Such editing could revolutionize clinical practices by…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Minghao Liu , Zhitao He , Zhiyuan Fan , Qingyun Wang , Yi R. Fung

Multipanel images, commonly seen as web screenshots, posters, etc., pervade our daily lives. These images, characterized by their composition of multiple subfigures in distinct layouts, effectively convey information to people. Toward…

Computer Vision and Pattern Recognition · Computer Science 2024-06-28 Yue Fan , Jing Gu , Kaiwen Zhou , Qianqi Yan , Shan Jiang , Ching-Chen Kuo , Xinze Guan , Xin Eric Wang

Multimodal large language models (MLLMs) have significantly advanced the integration of visual and textual understanding. However, their ability to generate code from multimodal inputs remains limited. In this work, we introduce VisCodex, a…

Computation and Language · Computer Science 2025-08-14 Lingjie Jiang , Shaohan Huang , Xun Wu , Yixia Li , Dongdong Zhang , Furu Wei

Despite the promising performance of existing visual models on public benchmarks, the critical assessment of their robustness for real-world applications remains an ongoing challenge. To bridge this gap, we propose an explainable visual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Qiang Li , Dan Zhang , Shengzhao Lei , Xun Zhao , Porawit Kamnoedboon , WeiWei Li , Junhao Dong , Shuyan Li

End-to-end In-Image Machine Translation (IIMT) aims to convert text embedded within an image into a target language while preserving the original visual context, layout, and rendering style. However, existing IIMT benchmarks are largely…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Jiahao Lyu , Pei Fu , Zhenhang Li , Weichao Zeng , Shaojie Zhang , Jiahui Yang , Can Ma , Yu Zhou , Zhenbo Luo , Jian Luan

Large Vision-Language Models (LVLMs) have become essential for advancing the integration of visual and linguistic information. However, the evaluation of LVLMs presents significant challenges as the evaluation benchmark always demands lots…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Han Bao , Yue Huang , Yanbo Wang , Jiayi Ye , Xiangqi Wang , Xiuying Chen , Yue Zhao , Tianyi Zhou , Mohamed Elhoseiny , Xiangliang Zhang

Virtual Try-ON (VTON) aims to synthesis specific person images dressed in given garments, which recently receives numerous attention in online shopping scenarios. Currently, the core challenges of the VTON task mainly lie in the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Jiabao Wei , Zhiyuan Ma

With the rapid advancement of Generative AI technology, Multimodal Large Language Models(MLLMs) have the potential to act as AI software engineers capable of executing complex web application development. Considering that the model requires…

Computation and Language · Computer Science 2025-06-10 Zhiyu Lin , Zhengda Zhou , Zhiyuan Zhao , Tianrui Wan , Yilun Ma , Junyu Gao , Xuelong Li

Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, especially for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Xuehai Bai , Yang Shi , Yi-Fan Zhang , Xuanyu Zhu , Yuran Wang , Yifan Dai , Xinyu Liu , Yiyan Ji , Xiaoling Gu , Yuanxing Zhang

Vision-language models (VLMs) have made significant progress in recent visual-question-answering (VQA) benchmarks that evaluate complex visio-linguistic reasoning. However, are these models truly effective? In this work, we show that VLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Baiqi Li , Zhiqiu Lin , Wenxuan Peng , Jean de Dieu Nyandwi , Daniel Jiang , Zixian Ma , Simran Khanuja , Ranjay Krishna , Graham Neubig , Deva Ramanan

Recently, knowledge editing on large language models (LLMs) has received considerable attention. Compared to this, editing Large Vision-Language Models (LVLMs) faces extra challenges from diverse data modalities and complicated model…

Computation and Language · Computer Science 2024-10-30 Han Huang , Haitian Zhong , Tao Yu , Qiang Liu , Shu Wu , Liang Wang , Tieniu Tan

Current Visual Language Models (VLMs) show impressive image understanding but struggle with visual illusions, especially in real-world scenarios. Existing benchmarks focus on classical cognitive illusions, which have been learned by…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Yiming Zhang , Zicheng Zhang , Xinyi Wei , Xiaohong Liu , Guangtao Zhai , Xiongkuo Min

Person re-identification (ReID) aims to retrieve images of a target person from the gallery set, with wide applications in medical rehabilitation and public security. However, traditional person ReID models are typically uni-modal,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Jinhao Li , Zijian Chen , Lirong Deng , Guangtao Zhai , Changbo Wang