中文
相关论文

相关论文: GenVideoLens: Where LVLMs Fall Short in AI-Generat…

200 篇论文

This paper does not introduce a novel method but instead establishes a straightforward, incremental, yet essential baseline for video temporal grounding (VTG), a core capability in video understanding. While multimodal large language models…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Jun Zhang , Teng Wang , Yuying Ge , Yixiao Ge , Xinhao Li , Ying Shan , Limin Wang

Recently, there is a surge in interest surrounding video large language models (Video LLMs). However, existing benchmarks fail to provide a comprehensive feedback on the temporal perception ability of Video LLMs. On the one hand, most of…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Yuanxin Liu , Shicheng Li , Yi Liu , Yuxiang Wang , Shuhuai Ren , Lei Li , Sishuo Chen , Xu Sun , Lu Hou

Recently, Vision Large Language Models (VLLMs) integrated with vision encoders have shown promising performance in vision understanding. The key of VLLMs is to encode visual content into sequences of visual tokens, enabling VLLMs to…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Zhuqiang Lu , Zhenfei Yin , Mengwei He , Zhihui Wang , Zicheng Liu , Zhiyong Wang , Kun Hu

The rapid advancement of large multimodal models (LMMs) has led to the rapid expansion of artificial intelligence generated videos (AIGVs), which highlights the pressing need for effective video quality assessment (VQA) models designed…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Jiarui Wang , Huiyu Duan , Guangtao Zhai , Juntong Wang , Xiongkuo Min

Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and challenging frontier. Current video benchmarks predominantly…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Andong Deng , Taojiannan Yang , Shoubin Yu , Lincoln Spencer , Mohit Bansal , Chen Chen , Serena Yeung-Levy , Xiaohan Wang

Advertisement videos serve as a rich and valuable source of purpose-driven information, encompassing high-quality visual, textual, and contextual cues designed to engage viewers. They are often more complex than general videos of similar…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Zheyuan Zhang , Monica Dou , Linkai Peng , Hongyi Pan , Ulas Bagci , Boqing Gong

The emergence of Large Vision-Language Models (LVLMs) marks significant strides towards achieving general artificial intelligence. However, these advancements are accompanied by concerns about biased outputs, a challenge that has yet to be…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Sibo Wang , Xiangkui Cao , Jie Zhang , Zheng Yuan , Shiguang Shan , Xilin Chen , Wen Gao

Large Video Models (LVMs) build on the semantic capabilities of Large Language Models (LLMs) and vision modules by integrating temporal information to better understand dynamic video content. Despite their progress, LVMs are prone to…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Garry Yang , Zizhe Chen , Man Hon Wong , Haoyu Lei , Yongqiang Chen , Zhenguo Li , Kaiwen Zhou , James Cheng

Video-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user inquiries. In pursuit of the ultimate goal of achieving…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Munan Ning , Bin Zhu , Yujia Xie , Bin Lin , Jiaxi Cui , Lu Yuan , Dongdong Chen , Li Yuan

In this paper, we propose XGC-AVis, a multi-agent framework that enhances the audio-video temporal alignment capabilities of multimodal large models (MLLMs) and improves the efficiency of retrieving key video segments through 4 stages:…

多媒体 · 计算机科学 2025-09-30 Yuqin Cao , Xiongkuo Min , Yixuan Gao , Wei Sun , Zicheng Zhang , Jinliang Han , Guangtao Zhai

The generative model has made significant advancements in the creation of realistic videos, which causes security issues. However, this emerging risk has not been adequately addressed due to the absence of a benchmark dataset for…

计算机视觉与模式识别 · 计算机科学 2024-05-08 Peisong He , Leyao Zhu , Jiaxing Li , Shiqi Wang , Haoliang Li

Humans perform visual perception at multiple levels, including low-level object recognition and high-level semantic interpretation such as behavior understanding. Subtle differences in low-level details can lead to substantial changes in…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Guanzhen Li , Yuxi Xie , Min-Yen Kan

Evaluating generative video models remains an open problem. Reference-based metrics such as Structural Similarity Index Measure (SSIM) and Peak Signal to Noise Ratio (PSNR) reward pixel fidelity over semantic correctness, while Frechet…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Karthik Inbasekar , Guy Rom , Omer Shlomovits

Video-based quality assurance (QA) for long-form gameplay video is labor-intensive and error-prone, yet valuable for assessing game stability and visual correctness over extended play sessions. Vision language models (VLMs) promise…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Wentao Lu , Alexander Senchenko , Alan Sayle , Abram Hindle , Cor-Paul Bezemer

Despite recent advances in Video Large Language Models (VideoLLMs), effectively understanding long-form videos remains a significant challenge. Perceiving lengthy videos containing thousands of frames poses substantial computational burden.…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Linli Yao , Haoning Wu , Kun Ouyang , Yuanxing Zhang , Caiming Xiong , Bei Chen , Xu Sun , Junnan Li

Recent advances in deep generative models have led to significant progress in video generation, yet the fidelity of AI-generated videos remains limited. Synthesized content often exhibits visual artifacts such as temporally inconsistent…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Jiahao Lin , Weixuan Peng , Bojia Zi , Yifeng Gao , Xianbiao Qi , Xingjun Ma , Yu-Gang Jiang

Large vision-language models (LVLMs) have been regarded as a breakthrough advance in an astoundingly variety of tasks, from content generation to virtual assistants and multimodal search or retrieval. However, for many of these…

计算机视觉与模式识别 · 计算机科学 2025-02-03 Kailash Hambarde , Pranita Samale , Hugo Proença

Recently, Large Vision-Language Models (LVLMs) have made significant strides across diverse multimodal tasks and benchmarks. This paper reveals a largely under-explored problem from existing video-involved LVLMs - language bias, where…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Yiming Yang , Yangyang Guo , Hui Lu , Yan Wang

Large Vision Language Models (LVLMs) have shown remarkable capabilities in multimodal tasks like visual question answering or image captioning. However, inconsistencies between the visual information and the generated text, a phenomenon…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Laura Fieback , Jakob Spiegelberg , Hanno Gottschalk

Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Lei Li , Yuancheng Wei , Zhihui Xie , Xuqing Yang , Yifan Song , Peiyi Wang , Chenxin An , Tianyu Liu , Sujian Li , Bill Yuchen Lin , Lingpeng Kong , Qi Liu