English
Related papers

Related papers: AIGVE-MACS: Unified Multi-Aspect Commenting and Sc…

200 papers

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in general video understanding, yet they often struggle with the fine-grained comprehension crucial for real-world applications requiring nuanced interpretation of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Gueter Josmy Faure , Min-Hung Chen , Jia-Fong Yeh , Hung-Ting Su , Winston H. Hsu

Human matting is a foundation task in image and video processing, where human foreground pixels are extracted from the input. Prior works either improve the accuracy by additional guidance or improve the temporal consistency of a single…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Chuong Huynh , Seoung Wug Oh , Abhinav Shrivastava , Joon-Young Lee

Affective Image Manipulation (AIM) aims to evoke specific emotions through targeted editing. Current image editing benchmarks primarily focus on object-level modifications in general scenarios, lacking the fine-grained granularity to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Shi Chen , Xuecheng Wu , Heli Sun , Yunyun Shi , Xinyi Yin , Fengjian Xue , Jinheng Xie , Dingkang Yang , Hao Wang , Junxiao Xue , Liang He

Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these generated outputs presents ongoing challenges. Although numerous…

Computation and Language · Computer Science 2025-06-13 Tian Lan , Yang-Hao Zhou , Zi-Ao Ma , Fanshu Sun , Rui-Qing Sun , Junyu Luo , Rong-Cheng Tu , Heyan Huang , Chen Xu , Zhijing Wu , Xian-Ling Mao

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the absence of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Yuansen Liu , Haiming Tang , Jinlong Peng , Jiangning Zhang , Xiaozhong Ji , Qingdong He , Wenbin Wu , Donghao Luo , Zhenye Gan , Junwei Zhu , Yunhang Shen , Chaoyou Fu , Chengjie Wang , Xiaobin Hu , Shuicheng Yan

Numerous text-to-video (T2V) editing methods have emerged recently, but the lack of a standardized benchmark for fair evaluation has led to inconsistent claims and an inability to assess model sensitivity to hyperparameters. Fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Minghan Li , Chenxi Xie , Yichen Wu , Lei Zhang , Mengyu Wang

This work addresses the lack of multimodal generative models capable of producing high-quality videos with spatially aligned audio. While recent advancements in generative models have been successful in video generation, they often overlook…

Sound · Computer Science 2026-02-05 Kazuki Shimada , Christian Simon , Takashi Shibuya , Shusuke Takahashi , Yuki Mitsufuji

We provide a new multi-task benchmark for evaluating text-to-image models. We perform a human evaluation comparing the most common open-source (Stable Diffusion) and commercial (DALL-E 2) models. Twenty computer science AI graduate students…

Generating controllable videos conforming to user intentions is an appealing yet challenging topic in computer vision. To enable maneuverable control in line with user intentions, a novel video generation task, named Text-Image-to-Video…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Yaosi Hu , Chong Luo , Zhenzhong Chen

Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primarily on single-turn interactions, failing to capture the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Wei Chow , Jiachun Pan , Yongyuan Liang , Mingze Zhou , Xue Song , Liyu Jia , Saining Zhang , Siliang Tang , Juncheng Li , Fengda Zhang , Weijia Wu , Hanwang Zhang , Tat-Seng Chua

Understanding real-world videos such as movies requires integrating visual and dialogue cues. Yet existing VideoQA benchmarks struggle to capture this multimodal reasoning and, given the difficulty of evaluating free-form answers, largely…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Shaden Shaar , Bradon Thymes , Sirawut Chaixanien , Claire Cardie , Bharath Hariharan

Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insufficient. To address this, we introduce InstructionBench, an…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Haiwan Wei , Yitian Yuan , Xiaohan Lan , Wei Ke , Lin Ma

Video generation models, as one form of world models, have emerged as one of the most exciting frontiers in AI, promising agents the ability to imagine the future by modeling the temporal evolution of complex scenes. In autonomous driving,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yang Zhou , Hao Shao , Letian Wang , Zhuofan Zong , Hongsheng Li , Steven L. Waslander

Recent breakthroughs in large multimodal models (LMMs) have significantly advanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, many generated images still suffer from issues related to perceptual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Jiarui Wang , Huiyu Duan , Yu Zhao , Juntong Wang , Guangtao Zhai , Xiongkuo Min

As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured footage to meet professional requirements. Yet the field still lacks both a large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xiangbo Gao , Sicong Jiang , Bangya Liu , Xinghao Chen , Minglai Yang , Siyuan Yang , Mingyang Wu , Jiongze Yu , Qi Zheng , Haozhi Wang , Jiayi Zhang , Jie Yang , Zihan Wang , Qing Yin , Zhengzhong Tu

Recent advancements in video generation have significantly improved the ability to synthesize videos from text instructions. However, existing models still struggle with key challenges such as instruction misalignment, content…

Computer Vision and Pattern Recognition · Computer Science 2025-02-10 Haibo Tong , Zhaoyang Wang , Zhaorun Chen , Haonian Ji , Shi Qiu , Siwei Han , Kexin Geng , Zhongkai Xue , Yiyang Zhou , Peng Xia , Mingyu Ding , Rafael Rafailov , Chelsea Finn , Huaxiu Yao

While text-to-visual models now produce photo-realistic images and videos, they struggle with compositional text prompts involving attributes, relationships, and higher-order reasoning such as logic and comparison. In this work, we conduct…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Baiqi Li , Zhiqiu Lin , Deepak Pathak , Jiayao Li , Yixin Fei , Kewen Wu , Tiffany Ling , Xide Xia , Pengchuan Zhang , Graham Neubig , Deva Ramanan

The development of Large Language Models (LLM) and Diffusion Models brings the boom of Artificial Intelligence Generated Content (AIGC). It is essential to build an effective quality assessment framework to provide a quantifiable evaluation…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Xi Fang , Weigang Wang , Xiaoxin Lv , Jun Yan

In recent years, image generation technology has rapidly advanced, resulting in the creation of a vast array of AI-generated images (AIGIs). However, the quality of these AIGIs is highly inconsistent, with low-quality AIGIs severely…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Jiquan Yuan , Fanyi Yang , Jihe Li , Xinyan Cao , Jinming Che , Jinlong Lin , Xixin Cao

With the rapid development of generative technologies, AI-Generated Images (AIGIs) have been widely applied in various aspects of daily life. However, due to the immaturity of the technology, the quality of the generated images varies, so…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Zhenchen Tang , Zichuan Wang , Bo Peng , Jing Dong