English
Related papers

Related papers: Technical Report for Soccernet 2023 -- Dense Video…

200 papers

Despite the recent emergence of video captioning models, how to generate vivid, fine-grained video descriptions based on the background knowledge (i.e., long and informative commentary about the domain-specific scenes with appropriate…

Computer Vision and Pattern Recognition · Computer Science 2023-10-06 Ji Qi , Jifan Yu , Teng Tu , Kunyu Gao , Yifan Xu , Xinyu Guan , Xiaozhi Wang , Yuxiao Dong , Bin Xu , Lei Hou , Juanzi Li , Jie Tang , Weidong Guo , Hui Liu , Yu Xu

The SoccerNet 2023 challenges were the third annual video understanding challenges organized by the SoccerNet team. For this third edition, the challenges were composed of seven vision-based tasks split into three main themes. The first…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Anthony Cioppa , Silvio Giancola , Vladimir Somers , Floriane Magera , Xin Zhou , Hassan Mkhallati , Adrien Deliège , Jan Held , Carlos Hinojosa , Amir M. Mansourian , Pierre Miralles , Olivier Barnich , Christophe De Vleeschouwer , Alexandre Alahi , Bernard Ghanem , Marc Van Droogenbroeck , Abdullah Kamal , Adrien Maglo , Albert Clapés , Amr Abdelaziz , Artur Xarles , Astrid Orcesi , Atom Scott , Bin Liu , Byoungkwon Lim , Chen Chen , Fabian Deuser , Feng Yan , Fufu Yu , Gal Shitrit , Guanshuo Wang , Gyusik Choi , Hankyul Kim , Hao Guo , Hasby Fahrudin , Hidenari Koguchi , Håkan Ardö , Ibrahim Salah , Ido Yerushalmy , Iftikar Muhammad , Ikuma Uchida , Ishay Be'ery , Jaonary Rabarisoa , Jeongae Lee , Jiajun Fu , Jianqin Yin , Jinghang Xu , Jongho Nang , Julien Denize , Junjie Li , Junpei Zhang , Juntae Kim , Kamil Synowiec , Kenji Kobayashi , Kexin Zhang , Konrad Habel , Kota Nakajima , Licheng Jiao , Lin Ma , Lizhi Wang , Luping Wang , Menglong Li , Mengying Zhou , Mohamed Nasr , Mohamed Abdelwahed , Mykola Liashuha , Nikolay Falaleev , Norbert Oswald , Qiong Jia , Quoc-Cuong Pham , Ran Song , Romain Hérault , Rui Peng , Ruilong Chen , Ruixuan Liu , Ruslan Baikulov , Ryuto Fukushima , Sergio Escalera , Seungcheon Lee , Shimin Chen , Shouhong Ding , Taiga Someya , Thomas B. Moeslund , Tianjiao Li , Wei Shen , Wei Zhang , Wei Li , Wei Dai , Weixin Luo , Wending Zhao , Wenjie Zhang , Xinquan Yang , Yanbiao Ma , Yeeun Joo , Yingsen Zeng , Yiyang Gan , Yongqiang Zhu , Yujie Zhong , Zheng Ruan , Zhiheng Li , Zhijian Huang , Ziyu Meng

The recently proposed action spotting task consists in finding the exact timestamp in which an event occurs. This task fits particularly well for soccer videos, where events correspond to salient actions strictly defined by soccer rules (a…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Matteo Tomei , Lorenzo Baraldi , Simone Calderara , Simone Bronzin , Rita Cucchiara

In video understanding, action spotting consists in temporally localizing human-induced events annotated with single timestamps. In this paper, we propose a novel loss function that specifically considers the temporal context naturally…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Anthony Cioppa , Adrien Deliège , Silvio Giancola , Bernard Ghanem , Marc Van Droogenbroeck , Rikke Gade , Thomas B. Moeslund

The SoccerNet 2025 Challenges mark the fifth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in football video understanding. This year's challenges span four vision-based tasks: (1)…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Silvio Giancola , Anthony Cioppa , Marc Gutiérrez-Pérez , Jan Held , Carlos Hinojosa , Victor Joos , Arnaud Leduc , Floriane Magera , Karen Sanchez , Vladimir Somers , Artur Xarles , Antonio Agudo , Alexandre Alahi , Olivier Barnich , Albert Clapés , Christophe De Vleeschouwer , Sergio Escalera , Bernard Ghanem , Thomas B. Moeslund , Marc Van Droogenbroeck , Tomoki Abe , Saad Alotaibi , Faisal Altawijri , Steven Araujo , Xiang Bai , Xiaoyang Bi , Jiawang Cao , Vanyi Chao , Kamil Czarnogórski , Fabian Deuser , Mingyang Du , Tianrui Feng , Patrick Frenzel , Mirco Fuchs , Jorge García , Konrad Habel , Takaya Hashiguchi , Sadao Hirose , Xinting Hu , Yewon Hwang , Ririko Inoue , Riku Itsuji , Kazuto Iwai , Hongwei Ji , Yangguang Ji , Licheng Jiao , Yuto Kageyama , Yuta Kamikawa , Yuuki Kanasugi , Hyungjung Kim , Jinwook Kim , Takuya Kurihara , Bozheng Li , Lingling Li , Xian Li , Youxing Lian , Dingkang Liang , Hongkai Lin , Jiadong Lin , Jian Liu , Liang Liu , Shuaikun Liu , Zhaohong Liu , Yi Lu , Federico Méndez , Huadong Ma , Wenping Ma , Jacek Maksymiuk , Henry Mantilla , Ismail Mathkour , Daniel Matthes , Ayaha Motomochi , Amrulloh Robbani Muhammad , Haruto Nakayama , Joohyung Oh , Yin May Oo , Marcelo Ortega , Norbert Oswald , Rintaro Otsubo , Fabian Perez , Mengshi Qi , Cristian Rey , Abel Reyes-Angulo , Oliver Rose , Hoover Rueda-Chacón , Hideo Saito , Jose Sarmiento , Kanta Sawafuji , Atom Scott , Xi Shen , Pragyan Shrestha , Jae-Young Sim , Long Sun , Yuyang Sun , Tomohiro Suzuki , Licheng Tang , Masato Tonouchi , Ikuma Uchida , Henry O. Velesaca , Tiancheng Wang , Rio Watanabe , Jay Wu , Yongliang Wu , Shunzo Yamagishi , Di Yang , Xu Yang , Yuxin Yang , Hao Ye , Xinyu Ye , Calvin Yeung , Xuanlong Yu , Chao Zhang , Dingyuan Zhang , Kexing Zhang , Zhe Zhao , Xin Zhou , Wenbo Zhu , Julian Ziegler

In this paper, we propose a new approach for retrieval of video segments using natural language queries. Unlike most previous approaches such as concept-based methods or rule-based structured models, the proposed method uses image…

Computer Vision and Pattern Recognition · Computer Science 2017-07-04 Sangkuk Lee , Daesik Kim , Myunggi Lee , Jihye Hwang , Nojun Kwak

The SoccerNet 2024 challenges represent the fourth annual video understanding challenges organized by the SoccerNet team. These challenges aim to advance research across multiple themes in football, including broadcast video understanding,…

In this paper, we propose a study on multi-modal (audio and video) action spotting and classification in soccer videos. Action spotting and classification are the tasks that consist in finding the temporal anchors of events in a video and…

Computer Vision and Pattern Recognition · Computer Science 2020-11-10 Bastien Vanderplaetse , Stéphane Dupont

Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the {sparsity dilemma} in video annotations, which fails to provide the context information between…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Hongxiang Li , Meng Cao , Xuxin Cheng , Zhihong Zhu , Yaowei Li , Yuexian Zou

Dense video captioning, a task of localizing meaningful moments and generating relevant captions for videos, often requires a large, expensive corpus of annotated video segments paired with text. In an effort to minimize the annotation…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Yongrae Jo , Seongyun Lee , Aiden SJ Lee , Hyunji Lee , Hanseok Oh , Minjoon Seo

Recent advances in 3D human motion and language integration have primarily focused on text-to-motion generation, leaving the task of motion understanding relatively unexplored. We introduce Dense Motion Captioning, a novel task that aims to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Shiyao Xu , Benedetta Liberatori , Gül Varol , Paolo Rota

Soccer broadcast video understanding has been drawing a lot of attention in recent years within data scientists and industrial companies. This is mainly due to the lucrative potential unlocked by effective deep learning techniques developed…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Anthony Cioppa , Adrien Deliège , Floriane Magera , Silvio Giancola , Olivier Barnich , Bernard Ghanem , Marc Van Droogenbroeck

Existing dense or paragraph video captioning approaches rely on holistic representations of videos, possibly coupled with learned object/action representations, to condition hierarchical language decoders. However, they fundamentally lack…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Shih-Han Chou , James J. Little , Leonid Sigal

There has been significant attention to the research on dense video captioning, which aims to automatically localize and caption all events within untrimmed video. Several studies introduce methods by designing dense video captioning as a…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Minkuk Kim , Hyeon Bae Kim , Jinyoung Moon , Jinwoo Choi , Seong Tae Kim

We propose a new task and model for dense video object captioning -- detecting, tracking and captioning trajectories of objects in a video. This task unifies spatial and temporal localization in video, whilst also requiring fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Xingyi Zhou , Anurag Arnab , Chen Sun , Cordelia Schmid

The task of action spotting consists in both identifying actions and precisely localizing them in time with a single timestamp in long, untrimmed video streams. Automatically extracting those actions is crucial for many sports applications,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Silvio Giancola , Anthony Cioppa , Bernard Ghanem , Marc Van Droogenbroeck

We introduce a method to learn unsupervised semantic visual information based on the premise that complex events can be decomposed into simpler events and that these simple events are shared across several complex events. We first employ a…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Valter Estevam , Rayson Laroca , Helio Pedrini , David Menotti

This paper focuses on a novel and challenging vision task, dense video captioning, which aims to automatically describe a video clip with multiple informative and diverse caption sentences. The proposed method is trained without explicit…

Computer Vision and Pattern Recognition · Computer Science 2017-04-06 Zhiqiang Shen , Jianguo Li , Zhou Su , Minjun Li , Yurong Chen , Yu-Gang Jiang , Xiangyang Xue

Video captioning, i.e. the task of generating captions from video sequences creates a bridge between the Natural Language Processing and Computer Vision domains of computer science. The task of generating a semantically accurate description…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Md. Mushfiqur Rahman , Thasin Abedin , Khondokar S. S. Prottoy , Ayana Moshruba , Fazlul Hasan Siddiqui

Dense video captioning is a challenging video understanding task which aims to simultaneously segment the video into a sequence of meaningful consecutive events and to generate detailed captions to accurately describe each event. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 AJ Piergiovanni , Ganesh Satish Mallya , Dahun Kim , Anelia Angelova