English
Related papers

Related papers: UniVBench: Towards Unified Evaluation for Video Fo…

200 papers

Text-to-video generative models have made significant strides in recent years, producing high-quality videos that excel in both aesthetic appeal and accurate instruction following, and have become central to digital art creation and user…

Machine Learning · Computer Science 2025-05-02 Xuyang Guo , Jiayan Huo , Zhenmei Shi , Zhao Song , Jiahao Zhang , Jiale Zhao

The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on image-level pretraining and adpation, which are limited for…

Computer Vision and Pattern Recognition · Computer Science 2022-12-08 Yi Wang , Kunchang Li , Yizhuo Li , Yinan He , Bingkun Huang , Zhiyu Zhao , Hongjie Zhang , Jilan Xu , Yi Liu , Zun Wang , Sen Xing , Guo Chen , Junting Pan , Jiashuo Yu , Yali Wang , Limin Wang , Yu Qiao

Large multimodal models (LMMs) are processing increasingly longer and richer inputs. Albeit the progress, few public benchmark is available to measure such development. To mitigate this gap, we introduce LongVideoBench, a question-answering…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Haoning Wu , Dongxu Li , Bei Chen , Junnan Li

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and…

Given the enormous number of instructional videos available online, learning a diverse array of multi-step task models from videos is an appealing goal. We introduce a new pre-trained video model, VideoTaskformer, focused on representing…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Medhini Narasimhan , Licheng Yu , Sean Bell , Ning Zhang , Trevor Darrell

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image features remains…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Shuo Cao , Jiayang Li , Xiaohui Li , Yuandong Pu , Kaiwen Zhu , Yuanting Gao , Siqi Luo , Yi Xin , Qi Qin , Yu Zhou , Xiangyu Chen , Wenlong Zhang , Bin Fu , Yu Qiao , Yihao Liu

Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and locations across shots remains a challenge over long sequences. Existing evaluations typically use…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Ruozhen He , Meng Wei , Ziyan Yang , Vicente Ordonez

Video is a promising source of knowledge for embodied agents to learn models of the world's dynamics. Large deep networks have become increasingly effective at modeling complex video data in a self-supervised manner, as evaluated by metrics…

Computer Vision and Pattern Recognition · Computer Science 2023-04-27 Stephen Tian , Chelsea Finn , Jiajun Wu

Recent advances in 3D vision have led to specialized models for either 3D understanding (e.g., shape classification, segmentation, reconstruction) or 3D generation (e.g., synthesis, completion, and editing). However, these tasks are often…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Peng Huang , Yifeng Chen , Zeyu Zhang , Hao Tang

While image editing has advanced rapidly, video editing remains less explored, facing challenges in consistency, control, and generalization. We study the design space of data, architecture, and control, and introduce \emph{EasyV2V}, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jinjie Mai , Chaoyang Wang , Guocheng Gordon Qian , Willi Menapace , Sergey Tulyakov , Bernard Ghanem , Peter Wonka , Ashkan Mirzaei

Cinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Xinran Wang , Songyu Xu , Xiangxuan Shan , Yuxuan Zhang , Muxi Diao , Xueyan Duan , Yanhua Huang , Kongming Liang , Zhanyu Ma

Video generation models have developed rapidly in recent years, where generating natural human motion plays a pivotal role. However, accurately evaluating the quality of generated human motion video remains a significant challenge. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Bingzi Zhang , Kaisi Guan , Ruihua Song

Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lack a systematic exploration of the specific tasks where…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Zimo Wen , Boxiu Li , Wanbo Zhang , Junxiang Lei , Xiaoyu Chen , Yijia Fan , Qi Zhang , Yujiang Wang , Lili Qiu , Bo Li , Ziwei Liu , Caihua Shan , Yifan Yang , Yifei Shen

Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of using generated videos as scalable supervision for robot learning. However, for embodied manipulation,…

We propose a new "Unbiased through Textual Description (UTD)" video benchmark based on unbiased subsets of existing video classification and retrieval datasets to enable a more robust assessment of video understanding capabilities. Namely,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Nina Shvetsova , Arsha Nagrani , Bernt Schiele , Hilde Kuehne , Christian Rupprecht

Existing AI-generated video quality assessment (AIGVQA) methods mainly focus on global perceptual realism and coarse text-video alignment, while overlooking a critical requirement in educational scenarios: concept correctness. In early…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Baoliang Chen , Xinlong Bu , Hanwei Zhu , Lingyu Zhu , Jieyu Zhan

A critical yet frequently overlooked challenge in the field of deepfake detection is the lack of a standardized, unified, comprehensive benchmark. This issue leads to unfair performance comparisons and potentially misleading results.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zhiyuan Yan , Yong Zhang , Xinhang Yuan , Siwei Lyu , Baoyuan Wu

Document understanding is a critical capability in financial credit review, onboarding, and remote verification, where both decision accuracy and evidence traceability matter. Compared with static document images, document videos present a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Runze Cui , Fangxin Shang , Yehui Yang , Qing Yang , Yanwu Xu , Tao Chen

Instruction-based image editing has advanced rapidly, yet reliable and interpretable evaluation remains a bottleneck. Current protocols either (i) depend on paired reference images, resulting in limited coverage and inheriting biases from…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Tianyu Chen , Yasi Zhang , Zhi Zhang , Peiyu Yu , Shu Wang , Zhendong Wang , Kevin Lin , Xiaofei Wang , Zhengyuan Yang , Linjie Li , Chung-Ching Lin , Jianwen Xie , Oscar Leong , Lijuan Wang , Ying Nian Wu , Mingyuan Zhou

Visually-grounded dialog systems, which integrate multiple modes of communication such as text and visual inputs, have become an increasingly popular area of investigation. However, the absence of a standardized evaluation framework poses a…

Computation and Language · Computer Science 2023-09-15 Yunshui Li , Binyuan Hui , Zhaochao Yin , Wanwei He , Run Luo , Yuxing Long , Min Yang , Fei Huang , Yongbin Li