English
Related papers

Related papers: MEVA: A Large-Scale Multiview, Multimodal Video Da…

200 papers

Existing multimodal machine translation (MMT) datasets consist of images and video captions or instructional video subtitles, which rarely contain linguistic ambiguity, making visual information ineffective in generating appropriate…

Computation and Language · Computer Science 2023-11-01 Yihang Li , Shuichiro Shimizu , Chenhui Chu , Sadao Kurohashi , Wei Li

Referring Atomic Video Action Recognition (RAVAR) aims to recognize fine-grained, atomic-level actions of a specific person of interest conditioned on natural language descriptions. Distinct from conventional action recognition and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Kunyu Peng , Di Wen , Jia Fu , Jiamin Wu , Kailun Yang , Junwei Zheng , Ruiping Liu , Yufan Chen , Yuqian Fu , Danda Pani Paudel , Luc Van Gool , Rainer Stiefelhagen

We introduce Replay, a collection of multi-view, multi-modal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with several static cameras, as well as wearable action cameras,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Roman Shapovalov , Yanir Kleiman , Ignacio Rocco , David Novotny , Andrea Vedaldi , Changan Chen , Filippos Kokkinos , Ben Graham , Natalia Neverova

This notebook paper presents an overview and comparative analysis of our system designed for activity detection in extended videos (ActEV-PC) in ActivityNet Challenge 2019. Specifically, we exploit person/vehicle detections in spatial level…

Computer Vision and Pattern Recognition · Computer Science 2019-06-21 Fuchen Long , Qi Cai , Zhaofan Qiu , Zhijian Hou , Yingwei Pan , Ting Yao , Chong-Wah Ngo

Video Individual Counting (VIC) has received increasing attention for its importance in intelligent video surveillance. Existing works are limited in two aspects, i.e., dataset and method. Previous datasets are captured with fixed or rarely…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Yaowu Fan , Jia Wan , Tao Han , Antoni B. Chan , Andy J. Ma

Interaction data is widely used in multiple domains such as cognitive science, visualization, human computer interaction, and cybersecurity, among others. Applications range from cognitive analyses over user/behavior modeling, adaptation,…

The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Yuanhan Zhang , Jinming Wu , Wei Li , Bo Li , Zejun Ma , Ziwei Liu , Chunyuan Li

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera…

Understanding how humans interact with each other is key to building realistic multi-human virtual reality systems. This area remains relatively unexplored due to the lack of large-scale datasets. Recent datasets focusing on this issue…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Rawal Khirodkar , Jyun-Ting Song , Jinkun Cao , Zhengyi Luo , Kris Kitani

The application of methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS) have steadily gained popularity in the field of 3D object segmentation in static scenes. These approaches demonstrate efficacy in a range of…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Bangning Wei , Joshua Maraval , Meriem Outtas , Kidiyo Kpalma , Nicolas Ramin , Lu Zhang

Unmanned Aerial Vehicles (UAVs) have quickly become common in various airspaces, representing a wide range of applications from recreation flying to commercial photography and package delivery. With the increasing prevalence of UAVs, it…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Knut Peterson , Zaid Mayers , Azmain Yousuf , Priontu Chowdhury , Asher Zaczepinski , Solmaz Arezoomandan , Reihaneh Maarefdoust , David Han

Human motion is created by, and constrained by, our muscles. We take a first step at building computer vision methods that represent the internal muscle activity that causes motion. We present a new dataset, Muscles in Action (MIA), to…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Mia Chiquier , Carl Vondrick

Thanks to the substantial and explosively inscreased instructional videos on the Internet, novices are able to acquire knowledge for completing various tasks. Over the past decade, growing efforts have been devoted to investigating the…

Computer Vision and Pattern Recognition · Computer Science 2020-03-23 Yansong Tang , Jiwen Lu , Jie Zhou

We present our three branch solutions for International Challenge on Activity Recognition at CVPR2019. This model seeks to fuse richer information of global video clip, short human attention and long-term human activity into a unified…

Computer Vision and Pattern Recognition · Computer Science 2019-08-14 Jin Xia , Jiajun Tang , Cewu Lu

The advancement of safety-critical research in driving behavior in ADAS-equipped vehicles require real-world datasets that not only include diverse traffic scenarios but also capture high-risk edge cases such as near-miss events and system…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Shaoyan Zhai , Mohamed Abdel-Aty , Chenzhu Wang , Rodrigo Vena Garcia

Researchers have used machine learning approaches to identify motion sickness in VR experience. These approaches demand an accurately-labeled, real-world, and diverse dataset for high accuracy and generalizability. As a starting point to…

Artificial Intelligence · Computer Science 2023-06-07 Elliott Wen , Chitralekha Gupta , Prasanth Sasikumar , Mark Billinghurst , James Wilmott , Emily Skow , Arindam Dey , Suranga Nanayakkara

Despite the fact that many 3D human activity benchmarks being proposed, most existing action datasets focus on the action recognition tasks for the segmented videos. There is a lack of standard large-scale benchmarks, especially for current…

Computer Vision and Pattern Recognition · Computer Science 2017-03-29 Chunhui Liu , Yueyu Hu , Yanghao Li , Sijie Song , Jiaying Liu

Semi-supervised video anomaly detection (VAD) is a critical task in the intelligent surveillance system. However, an essential type of anomaly in VAD named scene-dependent anomaly has not received the attention of researchers. Moreover,…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Congqi Cao , Yue Lu , Peng Wang , Yanning Zhang

A comprehensive understanding of interested human-to-human interactions in video streams, such as queuing, handshaking, fighting and chasing, is of immense importance to the surveillance of public security in regions like campuses, squares…

Computer Vision and Pattern Recognition · Computer Science 2023-08-14 Zhenhua Wang , Kaining Ying , Jiajun Meng , Jifeng Ning

Short-form videos (SVs) have become a vital part of our online routine for acquiring and sharing information. Their multimodal complexity poses new challenges for video analysis, highlighting the need for video emotion analysis (VEA) within…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xuecheng Wu , Dingkang Yang , Danlei Huang , Xinyi Yin , Yifan Wang , Jia Zhang , Jiayu Nie , Liangyu Fu , Yang Liu , Junxiao Xue , Hadi Amirpour , Wei Zhou
‹ Prev 1 4 5 6 7 8 10 Next ›