English
Related papers

Related papers: VACT: A Video Automatic Causal Testing System and …

200 papers

Vision Large Language Models (VLLMs) are widely acknowledged to be prone to hallucinations. Existing research addressing this problem has primarily been confined to image inputs, with limited exploration of video-based hallucinations.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Wey Yeh Choong , Yangyang Guo , Mohan Kankanhalli

The integration of visual encoders and large language models (LLMs) has driven recent progress in multimodal large language models (MLLMs). However, the scarcity of high-quality instruction-tuning data for vision-language tasks remains a…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Bin Wang , Fan Wu , Xiao Han , Jiahui Peng , Huaping Zhong , Pan Zhang , Xiaoyi Dong , Weijia Li , Wei Li , Jiaqi Wang , Conghui He

Humans are able to perceive, understand and reason about causal events. Developing models with similar physical and causal understanding capabilities is a long-standing goal of artificial intelligence. As a step towards this direction, we…

Artificial Intelligence · Computer Science 2022-03-02 Tayfun Ates , M. Samil Atesoglu , Cagatay Yigit , Ilker Kesen , Mert Kobas , Erkut Erdem , Aykut Erdem , Tilbe Goksun , Deniz Yuret

Large language models (LLMs) have shown great promise in machine translation, but they still struggle with contextually dependent terms, such as new or domain-specific words. This leads to inconsistencies and errors that are difficult to…

Computation and Language · Computer Science 2024-10-29 Meiqi Chen , Fandong Meng , Yingxue Zhang , Yan Zhang , Jie Zhou

Video quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: \textit{poor…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Linhan Cao , Wei Sun , Weixia Zhang , Xiangyang Zhu , Jun Jia , Kaiwei Zhang , Dandan Zhu , Guangtao Zhai , Xiongkuo Min

Real-world reasoning often requires combining information across modalities, connecting textual context with visual cues in a multi-hop process. Yet, most multimodal benchmarks fail to capture this ability: they typically rely on single…

Machine Learning · Computer Science 2026-04-03 Junyoung Sung , Seungwoo Lyu , Minjun Kim , Sumin An , Arsha Nagrani , Paul Hongsuck Seo

Video-generative world models are increasingly used as neural simulators for embodied planning and policy learning, yet their ability to predict physical risk and severe consequences is rarely evaluated.We find that these models often…

Robotics · Computer Science 2026-04-21 Zhenglin Lai , Sirui Huang , Yuteng Li , Changxin Huang , Jianqiang Li , Bingzhe Wu

Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily generalizable to diverse tasks, domains, and datasets. To…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Linjie Li , Jie Lei , Zhe Gan , Licheng Yu , Yen-Chun Chen , Rohit Pillai , Yu Cheng , Luowei Zhou , Xin Eric Wang , William Yang Wang , Tamara Lee Berg , Mohit Bansal , Jingjing Liu , Lijuan Wang , Zicheng Liu

Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Keming Wu , Yijing Cui , Wenhan Xue , Qijie Wang , Xuan Luo , Zhiyuan Feng , Zuhao Yang , Sudong Wang , Sicong Jiang , Haowei Zhu , Zihan Wang , Ping Nie , Wenhu Chen , Bin Wang

The automatic generation of visualizations is an old task that, through the years, has shown more and more interest from the research and practitioner communities. Recently, large language models (LLM) have become an interesting option for…

Human-Computer Interaction · Computer Science 2024-02-06 Luca Podo , Muhammad Ishmal , Marco Angelini

AI-driven video analytics has become increasingly important across diverse domains. However, existing systems are often constrained to specific, predefined tasks, limiting their adaptability in open-ended analytical scenarios. The recent…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Yuxuan Yan , Shiqi Jiang , Ting Cao , Yifan Yang , Qianqian Yang , Yuanchao Shu , Yuqing Yang , Lili Qiu

The quality of video-text pairs fundamentally determines the upper bound of text-to-video models. Currently, the datasets used for training these models suffer from significant shortcomings, including low temporal consistency, poor-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Zhiyu Tan , Xiaomeng Yang , Luozheng Qin , Hao Li

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual-text processing. However, existing static image-text benchmarks are insufficient for evaluating their dynamic perception and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Xiangxi Zheng , Linjie Li , Zhengyuan Yang , Ping Yu , Alex Jinpeng Wang , Rui Yan , Yuan Yao , Lijuan Wang

Given the accelerating progress of vision and language modeling, accurate evaluation of machine-generated image captions remains critical. In order to evaluate captions more closely to human preferences, metrics need to discriminate between…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Koki Maeda , Shuhei Kurita , Taiki Miyanishi , Naoaki Okazaki

Beneath the stunning visual fidelity of modern AIGC models lies a "logical desert", where systems fail tasks that require physical, causal, or complex spatial reasoning. Current evaluations largely rely on superficial metrics or fragmented…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Haonan Han , Jiancheng Huang , Xiaopeng Sun , Junyan He , Rui Yang , Jie Hu , Xiaojiang Peng , Lin Ma , Xiaoming Wei , Xiu Li

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

In the realm of vision models, the primary mode of representation is using pixels to rasterize the visual world. Yet this is not always the best or unique way to represent visual content, especially for designers and artists who depict the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Bocheng Zou , Mu Cai , Jianrui Zhang , Yong Jae Lee

Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing benchmarks predominantly focus on pure-text queries. In real-world scenarios, language also…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Qing'an Liu , Juntong Feng , Yuhao Wang , Xinzhe Han , Yujie Cheng , Yue Zhu , Haiwen Diao , Yunzhi Zhuge , Huchuan Lu

Video causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily executed in a question-answering paradigm and focusing on…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Tieyuan Chen , Huabin Liu , Tianyao He , Yihang Chen , Chaofan Gan , Xiao Ma , Cheng Zhong , Yang Zhang , Yingxue Wang , Hui Lin , Weiyao Lin

Existing methods for video question answering (VideoQA) often suffer from spurious correlations between different modalities, leading to a failure in identifying the dominant visual evidence and the intended question. Moreover, these…

Computer Vision and Pattern Recognition · Computer Science 2023-08-02 Yushen Wei , Yang Liu , Hong Yan , Guanbin Li , Liang Lin