English
Related papers

Related papers: Technical Report for SoccerNet Challenge 2022 -- R…

200 papers

In the task of temporal action localization of ActivityNet-1.3 datasets, we propose to locate the temporal boundaries of each action and predict action class in untrimmed videos. We first apply VideoSwinTransformer as feature extractor to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Shimin Chen , Wei Li , Jianyang Gu , Chen Chen , Yandong Guo

The SoccerNet 2022 challenges were the second annual video understanding challenges organized by the SoccerNet team. In 2022, the challenges were composed of 6 vision-based tasks: (1) action spotting, focusing on retrieving action…

The task of video grounding, which temporally localizes a natural language description in a video, plays an important role in understanding videos. Existing studies have adopted strategies of sliding window over the entire video or…

Computer Vision and Pattern Recognition · Computer Science 2019-01-23 Dongliang He , Xiang Zhao , Jizhou Huang , Fu Li , Xiao Liu , Shilei Wen

Temporal action detection (TAD) aims to detect the semantic labels and boundaries of action instances in untrimmed videos. Current mainstream approaches are multi-step solutions, which fall short in efficiency and flexibility. In this…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Shimin Chen , Chen Chen , Wei Li , Xunqiang Tao , Yandong Guo

The SoccerNet 2024 challenges represent the fourth annual video understanding challenges organized by the SoccerNet team. These challenges aim to advance research across multiple themes in football, including broadcast video understanding,…

The SoccerNet 2025 Challenges mark the fifth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in football video understanding. This year's challenges span four vision-based tasks: (1)…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Silvio Giancola , Anthony Cioppa , Marc Gutiérrez-Pérez , Jan Held , Carlos Hinojosa , Victor Joos , Arnaud Leduc , Floriane Magera , Karen Sanchez , Vladimir Somers , Artur Xarles , Antonio Agudo , Alexandre Alahi , Olivier Barnich , Albert Clapés , Christophe De Vleeschouwer , Sergio Escalera , Bernard Ghanem , Thomas B. Moeslund , Marc Van Droogenbroeck , Tomoki Abe , Saad Alotaibi , Faisal Altawijri , Steven Araujo , Xiang Bai , Xiaoyang Bi , Jiawang Cao , Vanyi Chao , Kamil Czarnogórski , Fabian Deuser , Mingyang Du , Tianrui Feng , Patrick Frenzel , Mirco Fuchs , Jorge García , Konrad Habel , Takaya Hashiguchi , Sadao Hirose , Xinting Hu , Yewon Hwang , Ririko Inoue , Riku Itsuji , Kazuto Iwai , Hongwei Ji , Yangguang Ji , Licheng Jiao , Yuto Kageyama , Yuta Kamikawa , Yuuki Kanasugi , Hyungjung Kim , Jinwook Kim , Takuya Kurihara , Bozheng Li , Lingling Li , Xian Li , Youxing Lian , Dingkang Liang , Hongkai Lin , Jiadong Lin , Jian Liu , Liang Liu , Shuaikun Liu , Zhaohong Liu , Yi Lu , Federico Méndez , Huadong Ma , Wenping Ma , Jacek Maksymiuk , Henry Mantilla , Ismail Mathkour , Daniel Matthes , Ayaha Motomochi , Amrulloh Robbani Muhammad , Haruto Nakayama , Joohyung Oh , Yin May Oo , Marcelo Ortega , Norbert Oswald , Rintaro Otsubo , Fabian Perez , Mengshi Qi , Cristian Rey , Abel Reyes-Angulo , Oliver Rose , Hoover Rueda-Chacón , Hideo Saito , Jose Sarmiento , Kanta Sawafuji , Atom Scott , Xi Shen , Pragyan Shrestha , Jae-Young Sim , Long Sun , Yuyang Sun , Tomohiro Suzuki , Licheng Tang , Masato Tonouchi , Ikuma Uchida , Henry O. Velesaca , Tiancheng Wang , Rio Watanabe , Jay Wu , Yongliang Wu , Shunzo Yamagishi , Di Yang , Xu Yang , Yuxin Yang , Hao Ye , Xinyu Ye , Calvin Yeung , Xuanlong Yu , Chao Zhang , Dingyuan Zhang , Kexing Zhang , Zhe Zhao , Xin Zhou , Wenbo Zhu , Julian Ziegler

We propose TAL-Net, an improved approach to temporal action localization in video that is inspired by the Faster R-CNN object detection framework. TAL-Net addresses three key shortcomings of existing approaches: (1) we improve receptive…

Computer Vision and Pattern Recognition · Computer Science 2018-04-23 Yu-Wei Chao , Sudheendra Vijayanarasimhan , Bryan Seybold , David A. Ross , Jia Deng , Rahul Sukthankar

This technical report analyzes a temporal action localization method we used in the HACS competition which is hosted in Activitynet Challenge 2020.The goal of our task is to locate the start time and end time of the action in the untrimmed…

Computer Vision and Pattern Recognition · Computer Science 2020-06-16 Zhiwu Qing , Xiang Wang , Yongpeng Sang , Changxin Gao , Shiwei Zhang , Nong Sang

Video action detection (spatio-temporal action localization) is usually the starting point for human-centric intelligent analysis of videos nowadays. It has high practical impacts for many applications across robotics, security, healthcare,…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Xin Hu , Zhenyu Wu , Hao-Yu Miao , Siqi Fan , Taiyu Long , Zhenyu Hu , Pengcheng Pi , Yi Wu , Zhou Ren , Zhangyang Wang , Gang Hua

With rapidly evolving internet technologies and emerging tools, sports related videos generated online are increasing at an unprecedentedly fast pace. To automate sports video editing/highlight generation process, a key task is to precisely…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Xin Zhou , Le Kang , Zhiyu Cheng , Bo He , Jingyu Xin

Video temporal grounding aims to localize relevant temporal boundaries in a video given a textual prompt. Recent work has focused on enabling Video LLMs to perform video temporal grounding via next-token prediction of temporal timestamps.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Xizi Wang , Feng Cheng , Ziyang Wang , Huiyu Wang , Md Mohaiminul Islam , Lorenzo Torresani , Mohit Bansal , Gedas Bertasius , David Crandall

Temporal action detection (TAD) aims to detect all action boundaries and their corresponding categories in an untrimmed video. The unclear boundaries of actions in videos often result in imprecise predictions of action boundaries by…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Dingfeng Shi , Qiong Cao , Yujie Zhong , Shan An , Jian Cheng , Haogang Zhu , Dacheng Tao

The SoccerNet 2023 challenges were the third annual video understanding challenges organized by the SoccerNet team. For this third edition, the challenges were composed of seven vision-based tasks split into three main themes. The first…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Anthony Cioppa , Silvio Giancola , Vladimir Somers , Floriane Magera , Xin Zhou , Hassan Mkhallati , Adrien Deliège , Jan Held , Carlos Hinojosa , Amir M. Mansourian , Pierre Miralles , Olivier Barnich , Christophe De Vleeschouwer , Alexandre Alahi , Bernard Ghanem , Marc Van Droogenbroeck , Abdullah Kamal , Adrien Maglo , Albert Clapés , Amr Abdelaziz , Artur Xarles , Astrid Orcesi , Atom Scott , Bin Liu , Byoungkwon Lim , Chen Chen , Fabian Deuser , Feng Yan , Fufu Yu , Gal Shitrit , Guanshuo Wang , Gyusik Choi , Hankyul Kim , Hao Guo , Hasby Fahrudin , Hidenari Koguchi , Håkan Ardö , Ibrahim Salah , Ido Yerushalmy , Iftikar Muhammad , Ikuma Uchida , Ishay Be'ery , Jaonary Rabarisoa , Jeongae Lee , Jiajun Fu , Jianqin Yin , Jinghang Xu , Jongho Nang , Julien Denize , Junjie Li , Junpei Zhang , Juntae Kim , Kamil Synowiec , Kenji Kobayashi , Kexin Zhang , Konrad Habel , Kota Nakajima , Licheng Jiao , Lin Ma , Lizhi Wang , Luping Wang , Menglong Li , Mengying Zhou , Mohamed Nasr , Mohamed Abdelwahed , Mykola Liashuha , Nikolay Falaleev , Norbert Oswald , Qiong Jia , Quoc-Cuong Pham , Ran Song , Romain Hérault , Rui Peng , Ruilong Chen , Ruixuan Liu , Ruslan Baikulov , Ryuto Fukushima , Sergio Escalera , Seungcheon Lee , Shimin Chen , Shouhong Ding , Taiga Someya , Thomas B. Moeslund , Tianjiao Li , Wei Shen , Wei Zhang , Wei Li , Wei Dai , Weixin Luo , Wending Zhao , Wenjie Zhang , Xinquan Yang , Yanbiao Ma , Yeeun Joo , Yingsen Zeng , Yiyang Gan , Yongqiang Zhu , Yujie Zhong , Zheng Ruan , Zhiheng Li , Zhijian Huang , Ziyu Meng

We address the problem of temporal localization of repetitive activities in a video, i.e., the problem of identifying all segments of a video that contain some sort of repetitive or periodic motion. To do so, the proposed method represents…

Computer Vision and Pattern Recognition · Computer Science 2019-10-15 Giorgos Karvounas , Iason Oikonomidis , Antonis Argyros

The main challenge of Temporal Action Localization is to retrieve subtle human actions from various co-occurring ingredients, e.g., context and background, in an untrimmed video. While prior approaches have achieved substantial progress…

Computer Vision and Pattern Recognition · Computer Science 2022-06-24 Kun Xia , Le Wang , Sanping Zhou , Nanning Zheng , Wei Tang

Temporal action detection (TAD) is a fundamental video understanding task that aims to identify human actions and localize their temporal boundaries in videos. Although this field has achieved remarkable progress in recent years, further…

Many methods have been developed to help people find the video contents they want efficiently. However, there are still some unsolved problems in this area. For example, given a query video and a reference video, how to accurately localize…

Computer Vision and Pattern Recognition · Computer Science 2018-08-07 Yang Feng , Lin Ma , Wei Liu , Tong Zhang , Jiebo Luo

The recently proposed action spotting task consists in finding the exact timestamp in which an event occurs. This task fits particularly well for soccer videos, where events correspond to salient actions strictly defined by soccer rules (a…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Matteo Tomei , Lorenzo Baraldi , Simone Calderara , Simone Bronzin , Rita Cucchiara

This technical report presents an overview of our solution used in the submission to 2021 HACS Temporal Action Localization Challenge on both Supervised Learning Track and Weakly-Supervised Learning Track. Temporal Action Localization (TAL)…

Computer Vision and Pattern Recognition · Computer Science 2021-07-28 Haisheng Su , Peiqin Zhuang , Yukun Li , Dongliang Wang , Weihao Gan , Wei Wu , Yu Qiao

Temporal language grounding in videos aims to localize the temporal span relevant to the given query sentence. Previous methods treat it either as a boundary regression task or a span extraction task. This paper will formulate temporal…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Jialin Gao , Xin Sun , Mengmeng Xu , Xi Zhou , Bernard Ghanem
‹ Prev 1 2 3 10 Next ›