English
Related papers

Related papers: VTEdit-Bench: A Comprehensive Benchmark for Multi-…

200 papers

Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing benchmarks predominantly focus on pure-text queries. In real-world scenarios, language also…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Qing'an Liu , Juntong Feng , Yuhao Wang , Xinzhe Han , Yujie Cheng , Yue Zhu , Haiwen Diao , Yunzhi Zhuge , Huchuan Lu

Understanding multi-image, multi-turn scenarios is a critical yet underexplored capability for Large Vision-Language Models (LVLMs). Existing benchmarks predominantly focus on static or horizontal comparisons -- e.g., spotting visual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Wenbo Lyu , Yingjun Du , Jinglin Zhao , Xianton Zhen , Ling Shao

As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured footage to meet professional requirements. Yet the field still lacks both a large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xiangbo Gao , Sicong Jiang , Bangya Liu , Xinghao Chen , Minglai Yang , Siyuan Yang , Mingyang Wu , Jiongze Yu , Qi Zheng , Haozhi Wang , Jiayi Zhang , Jie Yang , Zihan Wang , Qing Yin , Zhengzhong Tu

Visualization, a domain-specific yet widely used form of imagery, is an effective way to turn complex datasets into intuitive insights, and its value depends on whether data are faithfully represented, clearly communicated, and…

Computation and Language · Computer Science 2026-03-03 Yupeng Xie , Zhiyang Zhang , Yifan Wu , Sirong Lu , Jiayi Zhang , Zhaoyang Yu , Jinlin Wang , Sirui Hong , Bang Liu , Chenglin Wu , Yuyu Luo

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of…

Text rendering has recently emerged as one of the most challenging frontiers in visual generation, drawing significant attention from large-scale diffusion and multimodal models. However, text editing within images remains largely…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Rui Gui , Yang Wan , Haochen Han , Dongxing Mao , Fangming Liu , Min Li , Alex Jinpeng Wang

Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information, achieving near-human proficiency in tasks like object recognition, captioning, and visual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Zhikai Wang , Jiashuo Sun , Wenqi Zhang , Zhiqiang Hu , Xin Li , Fan Wang , Deli Zhao

On the way towards general Visual Question Answering (VQA) systems that are able to answer arbitrary questions, the need arises for evaluation beyond single-metric leaderboards for specific datasets. To this end, we propose a browser-based…

Computer Vision and Pattern Recognition · Computer Science 2021-10-12 Dirk Väth , Pascal Tilli , Ngoc Thang Vu

Knowledge editing techniques have emerged as essential tools for updating the factual knowledge of large language models (LLMs) and multimodal models (LMMs), allowing them to correct outdated or inaccurate information without retraining…

Computation and Language · Computer Science 2025-03-04 Yuntao Du , Kailin Jiang , Zhi Gao , Chenrui Shi , Zilong Zheng , Siyuan Qi , Qing Li

We introduce VisIT-Bench (Visual InsTruction Benchmark), a benchmark for evaluation of instruction-following vision-language models for real-world use. Our starting point is curating 70 'instruction families' that we envision instruction…

Computation and Language · Computer Science 2023-12-27 Yonatan Bitton , Hritik Bansal , Jack Hessel , Rulin Shao , Wanrong Zhu , Anas Awadalla , Josh Gardner , Rohan Taori , Ludwig Schmidt

Latest advances have achieved realistic virtual try-on (VTON) through localized garment inpainting using latent diffusion models, significantly enhancing consumers' online shopping experience. However, existing VTON technologies neglect the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Fei Shen , Xin Jiang , Xin He , Hu Ye , Cong Wang , Xiaoyu Du , Zechao Li , Jinhui Tang

Image generation has witnessed significant advancements in the past few years. However, evaluating the performance of image generation models remains a formidable challenge. In this paper, we propose ICE-Bench, a unified and comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Yulin Pan , Xiangteng He , Chaojie Mao , Zhen Han , Zeyinzi Jiang , Jingfeng Zhang , Yu Liu

Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for professional workflows. Existing methods predominantly rely on…

Computation and Language · Computer Science 2026-05-18 Xiaoyan Su , Peijie Dong , Zhenheng Tang , Song Tang , Yuyao Zhai , Kaitao Lin , Liang Chen , Gai Yuhang , Yuyu Luo , Qiang Wang , Xiaowen Chu

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Ziqi Huang , Fan Zhang , Xiaojie Xu , Yinan He , Jiashuo Yu , Ziyue Dong , Qianli Ma , Nattapol Chanpaisit , Chenyang Si , Yuming Jiang , Yaohui Wang , Xinyuan Chen , Ying-Cong Chen , Limin Wang , Dahua Lin , Yu Qiao , Ziwei Liu

Multimodal embedding models have been crucial in enabling various downstream tasks such as semantic similarity, information retrieval, and clustering over different modalities. However, existing multimodal embeddings like VLM2Vec, E5-V, GME…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Rui Meng , Ziyan Jiang , Ye Liu , Mingyi Su , Xinyi Yang , Yuepeng Fu , Can Qin , Zeyuan Chen , Ran Xu , Caiming Xiong , Yingbo Zhou , Wenhu Chen , Semih Yavuz

While Multimodal Large Language Models (MLLMs) perform strongly on single-turn chart generation, their ability to support real-world exploratory data analysis remains underexplored. In practice, users iteratively refine visualizations…

Computation and Language · Computer Science 2026-02-18 Manav Nitin Kapadnis , Lawanya Baghel , Atharva Naik , Carolyn Rosé

Virtual try-on (VTON) has been widely explored for rendering garments onto person images, while its inverse task, virtual try-off (VTOFF), remains largely overlooked. VTOFF aims to recover standardized product images of garments directly…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Davide Lobba , Fulvio Sanguigni , Bin Ren , Marcella Cornia , Rita Cucchiara , Nicu Sebe

While real-world applications increasingly demand intricate scene manipulation, existing instruction-guided image editing benchmarks often oversimplify task complexity and lack comprehensive, fine-grained instructions. To bridge this gap,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Bohan Jia , Wenxuan Huang , Yuntian Tang , Junbo Qiao , Jincheng Liao , Shaosheng Cao , Fei Zhao , Zhaopeng Feng , Zhouhong Gu , Zhenfei Yin , Lei Bai , Wanli Ouyang , Lin Chen , Fei Zhao , Yao Hu , Zihan Wang , Yuan Xie , Shaohui Lin

Text-to-3D (T23D) generation has emerged as a crucial visual generation task, aiming at synthesizing 3D content from textual descriptions. Studies of this task are currently shifting from per-scene T23D, which requires optimization of the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Xiao Cai , Sitong Su , Jingkuan Song , Pengpeng Zeng , Ji Zhang , Qinhong Du , Mengqi Li , Heng Tao Shen , Lianli Gao

Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesion recognition.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Abdurrahim Yilmaz , Ozan Erdem , Ece Gokyayla , Ayda Acar , Burc Bugra Dagtas , Dilara Ilhan Erdil , Gulsum Gencoglan , Burak Temelkuran