English
Related papers

Related papers: MEVA: A Large-Scale Multiview, Multimodal Video Da…

200 papers

Despite significant breakthroughs in video analysis driven by the rapid development of large multimodal models (LMMs), there remains a lack of a versatile evaluation benchmark to comprehensively assess these models' performance in video…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Yunxin Li , Xinyu Chen , Baotian Hu , Longyue Wang , Haoyuan Shi , Min Zhang

Nowadays, short-form videos (SVs) are essential to web information acquisition and sharing in our daily life. The prevailing use of SVs to spread emotions leads to the necessity of conducting video emotion analysis (VEA) towards SVs.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Xuecheng Wu , Heli Sun , Junxiao Xue , Jiayu Nie , Xiangyan Kong , Ruofan Zhai , Liang He

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Keshigeyan Chandrasegaran , Agrim Gupta , Lea M. Hadzic , Taran Kota , Jimming He , Cristóbal Eyzaguirre , Zane Durante , Manling Li , Jiajun Wu , Li Fei-Fei

There are substantial instructional videos on the Internet, which enables us to acquire knowledge for completing various tasks. However, most existing datasets for instructional video analysis have the limitations in diversity and…

Computer Vision and Pattern Recognition · Computer Science 2019-03-08 Yansong Tang , Dajun Ding , Yongming Rao , Yu Zheng , Danyang Zhang , Lili Zhao , Jiwen Lu , Jie Zhou

Idle animations are essential for virtual characters, as they convey realistic behaviour during inactive states. While automatic animation generation has been widely studied, limited attention has been given to idle motion due to the…

Graphics · Computer Science 2026-05-14 Eneko Atxa Landa , Igor Rodriguez , Elena Lazkano , Taras Kucherenko

We introduce Human-like Video Models (HVM-1), large-scale video models pretrained with nearly 5000 hours of curated human-like video data (mostly egocentric, temporally extended, continuous video recordings), using the spatiotemporal masked…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 A. Emin Orhan

In this era, the success of large language models and text-to-image models can be attributed to the driving force of large-scale datasets. However, in the realm of 3D vision, while remarkable progress has been made with models trained on…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Zhangyang Xiong , Chenghong Li , Kenkun Liu , Hongjie Liao , Jianqiao Hu , Junyi Zhu , Shuliang Ning , Lingteng Qiu , Chongjie Wang , Shijie Wang , Shuguang Cui , Xiaoguang Han

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

Computer Vision and Pattern Recognition · Computer Science 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

Spatio-temporal action detection is an important and challenging problem in video understanding. However, the application of the existing large-scale spatio-temporal action datasets in specific fields is limited, and there is currently no…

Computer Vision and Pattern Recognition · Computer Science 2022-04-27 Fan Yang

Automatic detection of natural disasters and incidents has become more important as a tool for fast response. There have been many studies to detect incidents using still images and text. However, the number of approaches that exploit…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Duygu Sesver , Alp Eren Gençoğlu , Çağrı Emre Yıldız , Zehra Günindi , Faeze Habibi , Ziya Ata Yazıcı , Hazım Kemal Ekenel

Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Baoyu Liang , Qile Su , Shoutai Zhu , Yuchen Liang , Chao Tong

Video understanding with multimodal large language models (MLLMs) remains challenging due to the long token sequences of videos, which contain extensive temporal dependencies and redundant frames. Existing approaches typically treat MLLMs…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Yaolun Zhang , Ruohui Wang , Jiahao Wang , Yepeng Tang , Xuanyu Zheng , Haonan Duan , Hao Lu , Hanming Deng , Lewei Lu

Event cameras, with their high temporal and dynamic range and minimal memory usage, have found applications in various fields. However, their potential in static traffic monitoring remains largely unexplored. To facilitate this exploration,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Aayush Atul Verma , Bharatesh Chakravarthi , Arpitsinh Vaghela , Hua Wei , Yezhou Yang

In this era, the success of large language models and text-to-image models can be attributed to the driving force of large-scale datasets. However, in the realm of 3D vision, while significant progress has been achieved in object-centric…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Chenghong Li , Hongjie Liao , Yihao Zhi , Xihe Yang , Zhengwentai Sun , Jiahao Chang , Shuguang Cui , Xiaoguang Han

Current developments in temporal event or action localization usually target actions captured by a single camera. However, extensive events or actions in the wild may be captured as a sequence of shots by multiple cameras at different…

Computer Vision and Pattern Recognition · Computer Science 2021-04-16 Xiaolong Liu , Yao Hu , Song Bai , Fei Ding , Xiang Bai , Philip H. S. Torr

Unmanned Aerial Vehicles (UAVs) offer wide-ranging applications but also pose significant safety and privacy violation risks in areas like airport and infrastructure inspection, spurring the rapid development of Anti-UAV technologies in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Chunhui Zhang , Li Liu , Zhipeng Zhang , Yong Wang , Hao Wen , Xi Zhou , Shiming Ge , Yanfeng Wang

Understanding events in texts is a core objective of natural language understanding, which requires detecting event occurrences, extracting event arguments, and analyzing inter-event relationships. However, due to the annotation challenges…

Computation and Language · Computer Science 2024-06-21 Xiaozhi Wang , Hao Peng , Yong Guan , Kaisheng Zeng , Jianhui Chen , Lei Hou , Xu Han , Yankai Lin , Zhiyuan Liu , Ruobing Xie , Jie Zhou , Juanzi Li

Recognizing the motion of Micro Aerial Vehicles (MAVs) is crucial for enabling cooperative perception and control in autonomous aerial swarms. Yet, vision-based recognition models relying only on RGB data often fail to capture the complex…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Nengbo Zhang , Hann Woei Ho

Activity detection in security videos is a difficult problem due to multiple factors such as large field of view, presence of multiple activities, varying scales and viewpoints, and its untrimmed nature. The existing research in activity…

Computer Vision and Pattern Recognition · Computer Science 2020-05-20 Mamshad Nayeem Rizve , Ugur Demir , Praveen Tirupattur , Aayush Jung Rana , Kevin Duarte , Ishan Dave , Yogesh Singh Rawat , Mubarak Shah

Autonomous driving, particularly navigating complex and unanticipated scenarios, demands sophisticated reasoning and planning capabilities. While Multi-modal Large Language Models (MLLMs) offer a promising avenue for this, their use has…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Hidehisa Arai , Keita Miwa , Kento Sasaki , Yu Yamaguchi , Kohei Watanabe , Shunsuke Aoki , Issei Yamamoto