English
Related papers

Related papers: TimeChat: A Time-sensitive Multimodal Large Langua…

200 papers

Video generation is a critical pathway toward world models, with efficient long video inference as a key capability. Toward this end, we introduce LongCat-Video, a foundational video generation model with 13.6B parameters, delivering strong…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Meituan LongCat Team , Xunliang Cai , Qilong Huang , Zhuoliang Kang , Hongyu Li , Shijun Liang , Liya Ma , Siyu Ren , Xiaoming Wei , Rixu Xie , Tong Zhang

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Keshigeyan Chandrasegaran , Agrim Gupta , Lea M. Hadzic , Taran Kota , Jimming He , Cristóbal Eyzaguirre , Zane Durante , Manling Li , Jiajun Wu , Li Fei-Fei

This paper presents VideoLoom, a unified Video Large Language Model (Video LLM) for joint spatial-temporal understanding. To facilitate the development of fine-grained spatial and temporal localization capabilities, we curate LoomData-8.7k,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jiapeng Shi , Junke Wang , Zuyao You , Bo He , Zuxuan Wu

Long-context video modeling is critical for multimodal large language models (MLLMs), enabling them to process movies, online video streams, and so on. Despite its advances, handling long videos remains challenging due to the difficulty in…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Xinhao Li , Yi Wang , Jiashuo Yu , Xiangyu Zeng , Yuhan Zhu , Haian Huang , Jianfei Gao , Kunchang Li , Yinan He , Chenting Wang , Yu Qiao , Yali Wang , Limin Wang

With the success of large language models (LLMs), integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However, existing LLM-based large multimodal models (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Bo He , Hengduo Li , Young Kyun Jang , Menglin Jia , Xuefei Cao , Ashish Shah , Abhinav Shrivastava , Ser-Nam Lim

Recent advancements in large-scale video-language models have shown significant potential for real-time planning and detailed interactions. However, their high computational demands and the scarcity of annotated datasets limit their…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Yuxuan Wang , Yiqi Song , Cihang Xie , Yang Liu , Zilong Zheng

Empowered by Large Language Models (LLMs), recent advancements in Video-based LLMs (VideoLLMs) have driven progress in various video understanding tasks. These models encode video representations through pooling or query aggregation over a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Yuetian Weng , Mingfei Han , Haoyu He , Xiaojun Chang , Bohan Zhuang

Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hours) -- existing benchmarks either neglect this multi-scale…

Long video understanding remains a fundamental challenge for multimodal large language models (MLLMs), particularly in tasks requiring precise temporal reasoning and event localization. Existing approaches typically adopt uniform frame…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Chao Yuan , Yang Yang , Yehui Yang , Zach Cheng

Human experts typically integrate numerical and textual multimodal information to analyze time series. However, most traditional deep learning predictors rely solely on unimodal numerical data, using a fixed-length window for training and…

Computation and Language · Computer Science 2024-12-17 Chengsen Wang , Qi Qi , Jingyu Wang , Haifeng Sun , Zirui Zhuang , Jinming Wu , Lei Zhang , Jianxin Liao

Large Multimodal Models (LMMs) for video-audio understanding have traditionally been evaluated only on shorter videos of a few minutes long. In this paper, we introduce QMAVIS (Q Team-Multimodal Audio Video Intelligent Sensemaking), a novel…

Artificial Intelligence · Computer Science 2026-01-13 Zixing Lin , Jiale Wang , Gee Wah Ng , Lee Onn Mak , Chan Zhi Yang Jeriel , Jun Yang Lee , Yaohao Li

The ability to perceive how objects change over time is a crucial ingredient in human intelligence. However, current benchmarks cannot faithfully reflect the temporal understanding abilities of video-language models (VidLMs) due to the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Shicheng Li , Lei Li , Shuhuai Ren , Yuanxin Liu , Yi Liu , Rundong Gao , Xu Sun , Lu Hou

We introduce Long-VITA, a simple yet effective large multi-modal model for long-context visual-language understanding tasks. It is adept at concurrently processing and analyzing modalities of image, video, and text over 4K frames or 1M…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Yunhang Shen , Chaoyou Fu , Shaoqi Dong , Xiong Wang , Yi-Fan Zhang , Peixian Chen , Mengdan Zhang , Haoyu Cao , Ke Li , Shaohui Lin , Xiawu Zheng , Yan Zhang , Yiyi Zhou , Ran He , Caifeng Shan , Rongrong Ji , Xing Sun

Long-term video understanding requires interpreting complex temporal events and reasoning over procedural activities. While instructional video corpora, like HowTo100M, offer rich resources for model training, they present significant…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Mingji Ge , Qirui Chen , Zeqian Li , Weidi Xie

This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation. The InternVid dataset contains over 7…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 Yi Wang , Yinan He , Yizhuo Li , Kunchang Li , Jiashuo Yu , Xin Ma , Xinhao Li , Guo Chen , Xinyuan Chen , Yaohui Wang , Conghui He , Ping Luo , Ziwei Liu , Yali Wang , Limin Wang , Yu Qiao

Despite recent progress on the short-video Text-Visual Question Answering (ViteVQA) task - largely driven by benchmarks such as M4-ViteVQA - existing datasets still suffer from limited video duration and narrow evaluation scopes, making it…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Yangyang Zhong , Ji Qi , Yuan Yao , Pengxin Luo , Yunfeng Yan , Donglian Qi , Zhiyuan Liu , Tat-Seng Chua

Recent advances in video large language models have demonstrated strong capabilities in understanding short clips. However, scaling them to hours- or days-long videos remains highly challenging due to limited context capacity and the loss…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Woongyeong Yeo , Kangsan Kim , Jaehong Yoon , Sung Ju Hwang

Long video question answering is a challenging task that involves recognizing short-term activities and reasoning about their fine-grained relationships. State-of-the-art video Large Language Models (vLLMs) hold promise as a viable solution…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Reuben Tan , Ximeng Sun , Ping Hu , Jui-hsien Wang , Hanieh Deilamsalehy , Bryan A. Plummer , Bryan Russell , Kate Saenko

Despite advancements in multimodal large language models (MLLMs), current approaches struggle in medium-to-long video understanding due to frame and context length limitations. As a result, these models often depend on frame sampling, which…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Shehreen Azad , Vibhav Vineet , Yogesh Singh Rawat

Large video-language models (LVLMs) have shown remarkable performance across various video-language tasks. However, they encounter significant challenges when processing long videos because of the large number of video frames involved.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Junwen Pan , Rui Zhang , Xin Wan , Yuan Zhang , Ming Lu , Qi She