English
Related papers

Related papers: The SkatingVerse Workshop & Challenge: Methods and…

200 papers

We collected a new dataset that includes approximately eight hours of audiovisual recordings of a group of students and their self-evaluation scores for classroom engagement. The dataset and data analysis scripts are available on our…

Human-Computer Interaction · Computer Science 2023-04-19 Alpay Sabuncuoglu , T. Metin Sezgin

The TREC Video Retrieval Evaluation (TRECVID) 2019 was a TREC-style video analysis and retrieval evaluation, the goal of which remains to promote progress in research and development of content-based exploitation and retrieval of…

This work studies the problem of predicting the sequence of future actions for surround vehicles in real-world driving scenarios. To this aim, we make three main contributions. The first contribution is an automatic method to convert the…

Computer Vision and Pattern Recognition · Computer Science 2020-04-30 Jan-Nico Zaech , Dengxin Dai , Alexander Liniger , Luc Van Gool

Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insufficient. To address this, we introduce InstructionBench, an…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Haiwan Wei , Yitian Yuan , Xiaohan Lan , Wei Ke , Lin Ma

Part-level Action Parsing aims at part state parsing for boosting action recognition in videos. Despite of dramatic progresses in the area of video classification research, a severe problem faced by the community is that the detailed…

Computer Vision and Pattern Recognition · Computer Science 2021-11-08 Xuanhan Wang , Xiaojia Chen , Lianli Gao , Lechao Chen , Jingkuan Song

The goal of this work is to generate step-by-step visual instructions in the form of a sequence of images, given an input image that provides the scene context and the sequence of textual instructions. This is a challenging problem as it…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Tomáš Souček , Prajwal Gatti , Michael Wray , Ivan Laptev , Dima Damen , Josef Sivic

This article describes a dataset collected in a set of experiments that involves human participants and a robot. The set of experiments was conducted in the computing science robotics lab in Simon Fraser University, Burnaby, BC, Canada, and…

Computer Vision and Pattern Recognition · Computer Science 2020-10-29 Zhitian Zhang , Jimin Rhim , Taher Ahmadi , Kefan Yang , Angelica Lim , Mo Chen

Supervised Learning is a way of developing Artificial Intelligence systems in which a computer algorithm is trained on labeled data inputs. Effectiveness of a Supervised Learning algorithm is determined by its performance on a given dataset…

Computers and Society · Computer Science 2024-10-29 Shubhi Bansal , Atharva Tendulkar , Nagendra Kumar

Video games are a compelling source of annotated data as they can readily provide fine-grained groundtruth for diverse tasks. However, it is not clear whether the synthetically generated data has enough resemblance to the real-world images…

Computer Vision and Pattern Recognition · Computer Science 2016-08-16 Alireza Shafaei , James J. Little , Mark Schmidt

Vertebral labelling and segmentation are two fundamental tasks in an automated spine processing pipeline. Reliable and accurate processing of spine images is expected to benefit clinical decision-support systems for diagnosis, surgery…

Human Activity Recognition in RGB-D videos has been an active research topic during the last decade. However, no efforts have been found in the literature, for recognizing human activity in RGB-D videos where several performers are…

Computer Vision and Pattern Recognition · Computer Science 2018-07-10 Snehasis Mukherjee , Leburu Anvitha , T. Mohana Lahari

With the recent progress in sports analytics, deep learning approaches have demonstrated the effectiveness of mining insights into players' tactics for improving performance quality and fan engagement. This is attributed to the availability…

Machine Learning · Computer Science 2023-06-09 Wei-Yao Wang , Yung-Chang Huang , Tsi-Ui Ik , Wen-Chih Peng

Despite the recent emergence of video captioning models, how to generate the text description with specific entity names and fine-grained actions is far from being solved, which however has great applications such as basketball live text…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Zeyu Xi , Ge Shi , Xuefen Li , Junchi Yan , Zun Li , Lifang Wu , Zilin Liu , Liang Wang

This paper presents the details of the Audio-Visual Scene Classification task in the DCASE 2021 Challenge (Task 1 Subtask B). The task is concerned with classification using audio and video modalities, using a dataset of synchronized…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-21 Shanshan Wang , Toni Heittola , Annamaria Mesaros , Tuomas Virtanen

Despite the number of currently available datasets on video question answering, there still remains a need for a dataset involving multi-step and non-factoid answers. Moreover, relying on video transcripts remains an under-explored topic.…

Computation and Language · Computer Science 2020-06-02 Anthony Colas , Seokhwan Kim , Franck Dernoncourt , Siddhesh Gupte , Daisy Zhe Wang , Doo Soon Kim

This paper provides an overview of the Fake-EmoReact 2021 Challenge, held at the 9th SocialNLP Workshop, in conjunction with NAACL 2021. The challenge requires predicting the authenticity of tweets using reply context and augmented GIF…

Computation and Language · Computer Science 2024-06-10 Chien-Kun Huang , Yi-Ting Chang , Lun-Wei Ku , Cheng-Te Li , Hong-Han Shuai

The proliferation of generative video technologies has intensified the need for reliable methods to detect and characterize synthetic media. To address this challenge, we organized the \href{https://safe-video-2025.dsri.org}{SAFE: Synthetic…

On public benchmarks, current action recognition techniques have achieved great success. However, when used in real-world applications, e.g. sport analysis, which requires the capability of parsing an activity into phases and…

Computer Vision and Pattern Recognition · Computer Science 2020-04-15 Dian Shao , Yue Zhao , Bo Dai , Dahua Lin

This paper presents the results of a study conducted on the perceptual acceptability of audio-video desynchronization for sports videos. The study was conducted with 45 videos generated by applying 8 audio-video offsets on 5 source…

Image and Video Processing · Electrical Eng. & Systems 2022-12-06 Joshua Peter Ebenezer

Most existing cross-modal language-to-video retrieval (VR) research focuses on single-modal input from video, i.e., visual representation, while the text is omnipresent in human environments and frequently critical to understand video. To…

Computer Vision and Pattern Recognition · Computer Science 2023-05-08 Weijia Wu , Yuzhong Zhao , Zhuang Li , Jiahong Li , Hong Zhou , Mike Zheng Shou , Xiang Bai
‹ Prev 1 3 4 5 6 7 10 Next ›