中文
相关论文

相关论文: Towards Universal Soccer Video Understanding

200 篇论文

Comprehensive understanding of key players and actions in multiplayer sports broadcast videos is a challenging problem. Unlike in news or finance videos, sports videos have limited text. While both action recognition for multiplayer sports…

多媒体 · 计算机科学 2021-11-02 Avijit Shah , Topojoy Biswas , Sathish Ramadoss , Deven Santosh Shah

The Long-form Video Question-Answering task requires the comprehension and analysis of extended video content to respond accurately to questions by utilizing both temporal and contextual information. In this paper, we present…

计算机视觉与模式识别 · 计算机科学 2024-06-26 Yongliang Wu , Bozheng Li , Jiawang Cao , Wenbo Zhu , Yi Lu , Weiheng Chi , Chuyun Xie , Haolin Zheng , Ziyue Su , Jay Wu , Xu Yang

Sports broadcasters inject drama into play-by-play commentary by building team and player narratives through subjective analyses and anecdotes. Prior studies based on small datasets and manual coding show that such theatrics evince…

计算与语言 · 计算机科学 2019-10-22 Jack Merullo , Luke Yeh , Abram Handler , Alvin Grissom , Brendan O'Connor , Mohit Iyyer

Humans are routinely asked to evaluate the performance of other individuals, separating success from failure and affecting outcomes from science to education and sports. Yet, in many contexts, the metrics driving the human evaluation…

物理与社会 · 物理学 2017-12-07 Luca Pappalardo , Paolo Cintia , Dino Pedreschi , Fosca Giannotti , Albert-Laszlo Barabasi

We present a novel framework for predicting next actions in soccer possessions by leveraging path signatures to encode their complex spatio-temporal structure. Unlike existing approaches, we do not rely on fixed historical windows and…

机器学习 · 统计学 2025-08-19 David Hirnschall , Robert Bajons

Event detection is an important step in extracting knowledge from the video. In this paper, we propose a deep learning approach to detect events in a soccer match emphasizing the distinction between images of red and yellow cards and the…

计算机视觉与模式识别 · 计算机科学 2021-02-09 Ali Karimi , Ramin Toosi , Mohammad Ali Akhaee

In video understanding, action spotting consists in temporally localizing human-induced events annotated with single timestamps. In this paper, we propose a novel loss function that specifically considers the temporal context naturally…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Anthony Cioppa , Adrien Deliège , Silvio Giancola , Bernard Ghanem , Marc Van Droogenbroeck , Rikke Gade , Thomas B. Moeslund

In many real-world complex systems, the behavior can be observed as a collection of discrete events generated by multiple interacting agents. Analyzing the dynamics of these multi-agent systems, especially team sports, often relies on…

人工智能 · 计算机科学 2025-05-23 Rikuhei Umemoto , Keisuke Fujii

True understanding of videos comes from a joint analysis of all its modalities: the video frames, the audio track, and any accompanying text such as closed captions. We present a way to learn a compact multimodal feature representation that…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Vivek Sharma , Makarand Tapaswi , Rainer Stiefelhagen

Computer Vision developments are enabling significant advances in many fields, including sports. Many applications built on top of Computer Vision technologies, such as tracking data, are nowadays essential for every top-level analyst,…

计算机视觉与模式识别 · 计算机科学 2023-01-19 Tiago Mendes-Neves , Luís Meireles , João Mendes-Moreira

We present a large-scale video subtitle translation dataset, BigVideo, to facilitate the study of multi-modality machine translation. Compared with the widely used How2 and VaTeX datasets, BigVideo is more than 10 times larger, consisting…

计算机视觉与模式识别 · 计算机科学 2023-07-04 Liyan Kang , Luyang Huang , Ningxin Peng , Peihao Zhu , Zewei Sun , Shanbo Cheng , Mingxuan Wang , Degen Huang , Jinsong Su

We present an open, sport-agnostic platform that turns tracking into comparable spatial measures across professional Ultimate, basketball, and soccer. Coaches in all three sports ask the same question: where is the usable space, and when…

Summarization of multimedia data becomes increasingly significant as it is the basis for many real-world applications, such as question answering, Web search, and so forth. Most existing multi-modal summarization works however have used…

计算与语言 · 计算机科学 2020-09-18 Xiyan Fu , Jun Wang , Zhenglu Yang

Given the explosive growth of online videos, it is becoming increasingly important to relieve the tedious work of browsing and managing the video content of interest. Video summarization aims at providing such a technique by transforming…

计算机视觉与模式识别 · 计算机科学 2017-07-14 Zhong Ji , Yaru Ma , Yanwei Pang , Xuelong Li

In soccer, scoring goals is a fundamental objective which depends on many conditions and constraints. Considering the RoboCup soccer 2D-simulator, this paper presents a data mining-based decision system to identify the best time and…

人工智能 · 计算机科学 2013-06-28 Renato Oliveira , Paulo Adeodato , Arthur Carvalho , Icamaan Viegas , Christian Diego , Tsang Ing-Ren

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Tanzila Rahman , Renjie Liao , Leonid Sigal

Foundation models are used for many real-world applications involving language generation from temporally-ordered multimodal events. In this work, we study the ability of models to identify the most important sub-events in a video, which is…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Aditya K Surikuchi , Raquel Fernández , Sandro Pezzelle

Video action recognition is one of the representative tasks for video understanding. Over the last decade, we have witnessed great advancements in video action recognition thanks to the emergence of deep learning. But we also encountered…

计算机视觉与模式识别 · 计算机科学 2020-12-14 Yi Zhu , Xinyu Li , Chunhui Liu , Mohammadreza Zolfaghari , Yuanjun Xiong , Chongruo Wu , Zhi Zhang , Joseph Tighe , R. Manmatha , Mu Li

Game State Reconstruction (GSR), a critical task in Sports Video Understanding, involves precise tracking and localization of all individuals on the football field-players, goalkeepers, referees, and others - in real-world coordinates. This…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Vladimir Golovkin , Nikolay Nemtsev , Vasyl Shandyba , Oleg Udin , Nikita Kasatkin , Pavel Kononov , Anton Afanasiev , Sergey Ulasen , Andrei Boiarov

The popularity of racket sports (e.g., tennis and table tennis) leads to high demands for data analysis, such as notational analysis, on player performance. While sports videos offer many benefits for such analysis, retrieving accurate…

人机交互 · 计算机科学 2021-05-21 Dazhen Deng , Jiang Wu , Jiachen Wang , Yihong Wu , Xiao Xie , Zheng Zhou , Hui Zhang , Xiaolong Zhang , Yingcai Wu