中文
相关论文

相关论文: AVECL-UMONS database for audio-visual event classi…

200 篇论文

We introduce the first very large detection dataset for event cameras. The dataset is composed of more than 39 hours of automotive recordings acquired with a 304x240 ATIS sensor. It contains open roads and very diverse driving scenarios,…

计算机视觉与模式识别 · 计算机科学 2020-02-03 Pierre de Tournemire , Davide Nitti , Etienne Perot , Davide Migliore , Amos Sironi

Multiple-object tracking and behavior analysis have been the essential parts of surveillance video analysis for public security and urban management. With billions of surveillance video captured all over the world, multiple-object tracking…

计算机视觉与模式识别 · 计算机科学 2020-07-29 Guojun Yin , Bin Liu , Huihui Zhu , Tao Gong , Nenghai Yu

We collected a new dataset that includes approximately eight hours of audiovisual recordings of a group of students and their self-evaluation scores for classroom engagement. The dataset and data analysis scripts are available on our…

人机交互 · 计算机科学 2023-04-19 Alpay Sabuncuoglu , T. Metin Sezgin

Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations…

计算机视觉与模式识别 · 计算机科学 2022-05-03 Fadime Sener , Dibyadip Chatterjee , Daniel Shelepov , Kun He , Dipika Singhania , Robert Wang , Angela Yao

Current developments in temporal event or action localization usually target actions captured by a single camera. However, extensive events or actions in the wild may be captured as a sequence of shots by multiple cameras at different…

计算机视觉与模式识别 · 计算机科学 2021-04-16 Xiaolong Liu , Yao Hu , Song Bai , Fei Ding , Xiang Bai , Philip H. S. Torr

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene understanding for…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Ziheng Zhou , Jinxing Zhou , Wei Qian , Shengeng Tang , Xiaojun Chang , Dan Guo

Visual crowd counting has been recently studied as a way to enable people counting in crowd scenes from images. Albeit successful, vision-based crowd counting approaches could fail to capture informative features in extreme conditions,…

计算机视觉与模式识别 · 计算机科学 2020-05-19 Di Hu , Lichao Mou , Qingzhong Wang , Junyu Gao , Yuansheng Hua , Dejing Dou , Xiao Xiang Zhu

In this work, we present the development of a new database, namely Sound Localization and Classification (SLoClas) corpus, for studying and analyzing sound localization and classification. The corpus contains a total of 23.27 hours of data…

声音 · 计算机科学 2021-08-06 Xinyuan Qian , Bidisha Sharma , Amine El Abridi , Haizhou Li

When humans perceive the world, they naturally integrate multiple audio-visual tasks within dynamic, real-world scenes. However, current works such as event localization, parsing, segmentation and question answering are mostly explored…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Guangyao Li , Xin Wang , Wenwu Zhu

Audio-Visual Event Localization (AVEL) is the task of temporally localizing and classifying \emph{audio-visual events}, i.e., events simultaneously visible and audible in a video. In this paper, we solve AVEL in a weakly-supervised setting,…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Kalyan Ramakrishnan

This paper introduces a new challenge and datasets to foster research toward designing systems that can understand medical videos and provide visual answers to natural language questions. We believe medical videos may provide the best…

计算机视觉与模式识别 · 计算机科学 2022-02-01 Deepak Gupta , Kush Attal , Dina Demner-Fushman

This paper describes the AVA-Kinetics localized human actions video dataset. The dataset is collected by annotating videos from the Kinetics-700 dataset using the AVA annotation protocol, and extending the original AVA dataset with these…

计算机视觉与模式识别 · 计算机科学 2020-05-21 Ang Li , Meghana Thotakuri , David A. Ross , João Carreira , Alexander Vostrikov , Andrew Zisserman

We present a dataset with models of 14 articulated objects commonly found in human environments and with RGB-D video sequences and wrenches recorded of human interactions with them. The 358 interaction sequences total 67 minutes of human…

机器人学 · 计算机科学 2018-06-19 Roberto Martín-Martín , Clemens Eppner , Oliver Brock

Along with the increasing use of unmanned aerial vehicles (UAVs), large volumes of aerial videos have been produced. It is unrealistic for humans to screen such big data and understand their contents. Hence methodological research on the…

计算机视觉与模式识别 · 计算机科学 2020-06-26 Lichao Mou , Yuansheng Hua , Pu Jin , Xiao Xiang Zhu

We present an approach to labeling short video clips with English verbs as event descriptions. A key distinguishing aspect of this work is that it labels videos with verbs that describe the spatiotemporal interaction between event…

We live in a world filled with never-ending streams of multimodal information. As a more natural recording of the real scenario, long form audio-visual videos are expected as an important bridge for better exploring and understanding the…

多媒体 · 计算机科学 2023-06-19 Wenxuan Hou , Guangyao Li , Yapeng Tian , Di Hu

Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to differentiate sounds…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Haytham M. Fayek , Anurag Kumar

As audio-visual systems increasingly bring immersive and interactive capabilities into our work and leisure activities, so the need for naturalistic test material grows. New volumetric datasets have captured high-quality 3D video, but…

多媒体 · 计算机科学 2021-05-04 Hanne Stenzel , Davide Berghi , Marco Volino , Philip J. B. Jackson

Understanding the semantics of human movement -- the what, how and why of the movement -- is an important problem that requires datasets of human actions with semantic labels. Existing datasets take one of two approaches. Large-scale video…

计算机视觉与模式识别 · 计算机科学 2021-06-25 Abhinanda R. Punnakkal , Arjun Chandrasekaran , Nikos Athanasiou , Alejandra Quiros-Ramirez , Michael J. Black

Can our video understanding systems perceive objects when a heavy occlusion exists in a scene? To answer this question, we collect a large-scale dataset called OVIS for occluded video instance segmentation, that is, to simultaneously…

计算机视觉与模式识别 · 计算机科学 2022-05-18 Jiyang Qi , Yan Gao , Yao Hu , Xinggang Wang , Xiaoyu Liu , Xiang Bai , Serge Belongie , Alan Yuille , Philip H. S. Torr , Song Bai