中文
相关论文

相关论文: LongViTU: Instruction Tuning for Long-Form Video U…

200 篇论文

Accurate evaluation of financial question answering (QA) systems necessitates a comprehensive dataset encompassing diverse question types and contexts. However, current financial QA datasets lack scope diversity and question complexity.…

计算与语言 · 计算机科学 2025-03-04 Jian Chen , Peilin Zhou , Yining Hua , Yingxin Loh , Kehui Chen , Ziyuan Li , Bing Zhu , Junwei Liang

Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily focus on short audio and video clips ranging from 10 seconds…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Keda Tao , Yuhua Zheng , Jia Xu , Wenjie Du , Kele Shao , Hesong Wang , Xueyi Chen , Xin Jin , Junhan Zhu , Bohan Yu , Weiqiang Wang , Jian Liu , Can Qin , Yulun Zhang , Ming-Hsuan Yang , Huan Wang

Humans naturally share information with those they are connected to, and video has become one of the dominant mediums for communication and expression on the Internet. To support the creation of high-quality large-scale video content, a…

The rapid advancement of large multimodal models (LMMs) has led to the rapid expansion of artificial intelligence generated videos (AIGVs), which highlights the pressing need for effective video quality assessment (VQA) models designed…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Jiarui Wang , Huiyu Duan , Guangtao Zhai , Juntong Wang , Xiongkuo Min

Generating high-quality and temporally synchronized audio from video content is essential for video editing and post-production tasks, enabling the creation of semantically aligned audio for silent videos. However, most existing approaches…

声音 · 计算机科学 2025-08-18 Haomin Zhang , Kristin Qi , Shuxin Yang , Zihao Chen , Chaofan Ding , Xinhan Di

Long video understanding has become a critical task in computer vision, driving advancements across numerous applications from surveillance to content retrieval. Existing video understanding methods suffer from two challenges when dealing…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Zeng You , Zhiquan Wen , Yaofo Chen , Xin Li , Runhao Zeng , Yaowei Wang , Mingkui Tan

Document Visual Question Answering (Document VQA) faces significant challenges when processing long documents in low-resource environments due to context limitations and insufficient training data. This paper presents AdaDocVQA, a unified…

计算与语言 · 计算机科学 2025-08-20 Haoxuan Li , Wei Song , Aofan Liu , Peiwu Qin

Video quality assessment (VQA) is an important processing task, aiming at predicting the quality of videos in a manner highly consistent with human judgments of perceived quality. Traditional VQA models based on natural image and/or video…

图像与视频处理 · 电气工程与系统科学 2024-12-12 Qi Zheng , Yibo Fan , Leilei Huang , Tianyu Zhu , Jiaming Liu , Zhijian Hao , Shuo Xing , Chia-Ju Chen , Xiongkuo Min , Alan C. Bovik , Zhengzhong Tu

Online surgical phase recognition plays a significant role towards building contextual tools that could quantify performance and oversee the execution of surgical workflows. Current approaches are limited since they train spatial feature…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Yang Liu , Maxence Boels , Luis C. Garcia-Peraza-Herrera , Tom Vercauteren , Prokar Dasgupta , Alejandro Granados , Sebastien Ourselin

Despite the growing development of long-context large language models (LLMs), data-centric approaches relying on synthetic data have been hindered by issues related to faithfulness, which limit their effectiveness in enhancing model…

计算与语言 · 计算机科学 2025-05-30 Cehao Yang , Xueyuan Lin , Chengjin Xu , Xuhui Jiang , Shengjie Ma , Aofan Liu , Hui Xiong , Jian Guo

We propose a novel video understanding task by fusing knowledge-based and video question answering. First, we introduce KnowIT VQA, a video dataset with 24,282 human-generated question-answer pairs about a popular sitcom. The dataset…

计算机视觉与模式识别 · 计算机科学 2020-04-21 Noa Garcia , Mayu Otani , Chenhui Chu , Yuta Nakashima

The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond short, isolated events to encompass the continuous,…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Aniket Rege , Arka Sadhu , Yuliang Li , Kejie Li , Ramya Korlakai Vinayak , Yuning Chai , Yong Jae Lee , Hyo Jin Kim

State-of-the-art large language models (LLMs) are now claiming remarkable supported context lengths of 256k or even more. In contrast, the average context lengths of mainstream benchmarks are insufficient (5k-21k), and they suffer from…

计算与语言 · 计算机科学 2025-10-23 Tao Yuan , Xuefei Ning , Dong Zhou , Zhijie Yang , Shiyao Li , Minghui Zhuang , Zheyue Tan , Zhuyu Yao , Dahua Lin , Boxun Li , Guohao Dai , Shengen Yan , Yu Wang

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activity analysis, and…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Wenqi Zhou , Kai Cao , Hao Zheng , Yunze Liu , Xinyi Zheng , Miao Liu , Per Ola Kristensson , Walterio Mayol-Cuevas , Fan Zhang , Weizhe Lin , Junxiao Shen

Recent advancements in language-model-based video understanding have been progressing at a remarkable pace, spurred by the introduction of Large Language Models (LLMs). However, the focus of prior research has been predominantly on devising…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Yizhou Wang , Ruiyi Zhang , Haoliang Wang , Uttaran Bhattacharya , Yun Fu , Gang Wu

Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over omni-modal (visual, audio, and textual) signals in long and…

Short-form video poses new challenges to the quality assessment of user-generated content (UGC) due to its complex generation pipeline, rapid content variation, and mixed distortions. To address this challenge, we propose an end-to-end…

图像与视频处理 · 电气工程与系统科学 2026-05-20 Xinyi Wang , Angeliki Katsenou , Junxiao Shen , David Bull

We present Q-ViD, a simple approach for video question answering (video QA), that unlike prior methods, which are based on complex architectures, computationally expensive pipelines or use closed models like GPTs, Q-ViD relies on a single…

计算机视觉与模式识别 · 计算机科学 2024-07-23 David Romero , Thamar Solorio

Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able…

We propose the LEHA-CVQAD (Large-scale Enriched Human-Annotated Compressed Video Quality Assessment) dataset, which comprises 6,240 clips for compression-oriented video quality assessment. 59 source videos are encoded with 186 codec-preset…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Aleksandr Gushchin , Maksim Smirnov , Dmitriy Vatolin , Anastasia Antsiferova