English
Related papers

Related papers: VMDT: Decoding the Trustworthiness of Video Founda…

200 papers

We introduce MMVU, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. MMVU includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science,…

While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under-explored frontier. To bridge this gap, we present RISE-Video, a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Mingxin Liu , Shuran Ma , Shibei Meng , Xiangyu Zhao , Zicheng Zhang , Shaofeng Zhang , Zhihang Zhong , Peixian Chen , Haoyu Cao , Xing Sun , Haodong Duan , Xue Yang

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Jianhui Wei , Xiaotian Zhang , Yichen Li , Yuan Wang , Yan Zhang , Ziyi Chen , Zhihang Tang , Wei Xu , Zuozhu Liu

Evaluating model robustness is critical when developing trustworthy models not only to gain deeper understanding of model behavior, strengths, and weaknesses, but also to develop future models that are generalizable and robust across…

Computation and Language · Computer Science 2021-04-27 Maria Glenski , Ellyn Ayton , Robin Cosbey , Dustin Arendt , Svitlana Volkova

Rapid advancements in video diffusion models have enabled the creation of realistic videos, raising concerns about unauthorized use and driving the demand for techniques to protect model ownership. Existing watermarking methods, while…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 MinHyuk Jang , Youngdong Jang , JaeHyeok Lee , Feng Yang , Gyeongrok Oh , Jongheon Jeong , Sangpil Kim

Modern video-text retrieval (VTR) models excel on in-distribution benchmarks but are highly vulnerable to real-world query shifts, where the distribution of query data deviates from the training domain, leading to a sharp performance drop.…

Information Retrieval · Computer Science 2026-04-24 Bingqing Zhang , Zhuo Cao , Heming Du , Yang Li , Xue Li , Jiajun Liu , Sen Wang

Comprehensive and constructive evaluation protocols play an important role in the development of sophisticated text-to-video (T2V) generation models. Existing evaluation protocols primarily focus on temporal consistency and content…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Mingxiang Liao , Hannan Lu , Xinyu Zhang , Fang Wan , Tianyu Wang , Yuzhong Zhao , Wangmeng Zuo , Qixiang Ye , Jingdong Wang

Vision-Language Models (VLMs) are increasingly used as perceptual modules for visual content reasoning, including through captioning and DeepFake detection. In this work, we expose a critical vulnerability of VLMs when exposed to subtle,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Jordan Vice , Naveed Akhtar , Yansong Gao , Richard Hartley , Ajmal Mian

YouTube is a major platform for information and entertainment, but its wide accessibility also makes it attractive for scammers to upload deceptive or malicious content. Prior detection approaches rely largely on textual or statistical…

Cryptography and Security · Computer Science 2026-04-02 Ummay Kulsum , Aafaq Sabir , Abhinaya S. B. , Anupam Das

In text-to-image (T2I) generation, a prevalent training technique involves utilizing Vision Language Models (VLMs) for image re-captioning. Even though VLMs are known to exhibit hallucination, generating descriptive content that deviates…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Weichen Yu , Ziyan Yang , Shanchuan Lin , Qi Zhao , Jianyi Wang , Liangke Gui , Matt Fredrikson , Lu Jiang

Diffusion-based text-to-video generation has witnessed impressive progress in the past year yet still falls behind text-to-image generation. One of the key reasons is the limited scale of publicly available data (e.g., 10M video-text pairs…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Xiang Wang , Shiwei Zhang , Hangjie Yuan , Zhiwu Qing , Biao Gong , Yingya Zhang , Yujun Shen , Changxin Gao , Nong Sang

Existing codecs are designed to eliminate intrinsic redundancies to create a compact representation for compression. However, strong external priors from Multimodal Large Language Models (MLLMs) have not been explicitly explored in video…

Computer Vision and Pattern Recognition · Computer Science 2025-02-17 Pingping Zhang , Jinlong Li , Kecheng Chen , Meng Wang , Long Xu , Haoliang Li , Nicu Sebe , Sam Kwong , Shiqi Wang

As Video Large Multimodal Models (VLMMs) rapidly advance, their inherent complexity introduces significant safety challenges, particularly the issue of mismatched generalization where static safety alignments fail to transfer to dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yixu Wang , Jiaxin Song , Yifeng Gao , Xin Wang , Yang Yao , Yan Teng , Xingjun Ma , Yingchun Wang , Yu-Gang Jiang

Human intelligence requires correctness and robustness, with the former being foundational for the latter. In video understanding, correctness ensures the accurate interpretation of visual content, and robustness maintains consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yuanhan Zhang , Yunice Chew , Yuhao Dong , Aria Leo , Bo Hu , Ziwei Liu

Large-scale video generation models have demonstrated high visual realism in diverse contexts, spurring interest in their potential as general-purpose world simulators. Existing benchmarks focus on individual subjects rather than scenes…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Aaron Appelle , Jerome P. Lynch

We introduce a new benchmark, COVID-VTS, for fact-checking multi-modal information involving short-duration videos with COVID19- focused information from both the real world and machine generation. We propose, TwtrDetective, an effective…

Computer Vision and Pattern Recognition · Computer Science 2023-02-17 Fuxiao Liu , Yaser Yacoob , Abhinav Shrivastava

Recent text-to-video (T2V) models can synthesize complex videos from lightweight natural language prompts, raising urgent concerns about safety alignment in the event of misuse in the real world. Prior jailbreak attacks typically rewrite…

Cryptography and Security · Computer Science 2026-03-10 Moyang Chen , Zonghao Ying , Wenzhuo Xu , Quancheng Zou , Deyue Zhang , Dongdong Yang , Xiangzheng Zhang

Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models" -- interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal medium, as they…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Xuehai He , Weixi Feng , Kaizhi Zheng , Yujie Lu , Wanrong Zhu , Jiachen Li , Yue Fan , Jianfeng Wang , Linjie Li , Zhengyuan Yang , Kevin Lin , William Yang Wang , Lijuan Wang , Xin Eric Wang

Text-to-image (T2I) models are capable of generating visually impressive images, yet they often fail to accurately capture specific attributes in user prompts, such as the correct number of objects with the specified colors. The diversity…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Kevin David Hayes , Micah Goldblum , Vikash Sehwag , Gowthami Somepalli , Ashwinee Panda , Tom Goldstein