中文
相关论文

相关论文: TennisExpert: Towards Expert-Level Analytical Spor…

200 篇论文

Sports tracking data are the high-resolution spatiotemporal observations of a competitive event. The growing collection of these data in professional sport allows us to address a fundamental problem of modern sport: how to attribute value…

应用统计 · 统计学 2020-05-27 Stephanie Kovalchik , Martin Ingram , Kokum Weeratunga , Cagatay Goncu

Large multimodal models (LMMs) are processing increasingly longer and richer inputs. Albeit the progress, few public benchmark is available to measure such development. To mitigate this gap, we introduce LongVideoBench, a question-answering…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Haoning Wu , Dongxu Li , Bei Chen , Junnan Li

While Multimodal Large Language Models (MLLMs) excel at generic video understanding, their ability to support specialized, rule-grounded decision-making remains insufficiently explored. In this paper, we introduce RefereeBench, the first…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Yichen Xu , Yuanhang Liu , Chuhan Wang , Zihan Zhao , jinghan luo , Jianzhe Ma , Wenxuan Wang , Qin Jin

Video Temporal Grounding (VTG) aims to precisely identify video event segments in response to textual queries. The outputs of VTG tasks manifest as sequences of events, each defined by precise timestamps, saliency scores, and textual…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Zuhao Yang , Yingchen Yu , Yunqing Zhao , Shijian Lu , Song Bai

Sports have long attracted broad attention as they push the limits of human physical and cognitive capabilities. Amid growing interest in spatial intelligence for vision-language models (VLMs), sports provide a natural testbed for…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Yuchen Yang , Yuqing Shao , Duxiu Huang , Linfeng Dong , Yifei Liu , Suixin Tang , Xiang Zhou , Yuanyuan Gao , Wei Wang , Yue Zhou , Xue Yang , Yanfeng Wang , Xiao Sun , Zhihang Zhong

Fine-grained analysis of complex and high-speed sports like badminton presents a significant challenge for Multimodal Large Language Models (MLLMs), despite their notable advancements in general video understanding. This difficulty arises…

多媒体 · 计算机科学 2025-08-12 Xusheng He , Wei Liu , Shanshan Ma , Qian Liu , Chenghao Ma , Jianlong Wu

We present a comprehensive video-based analytics framework for tennis doubles that addresses the lack of automated analysis tools for this strategically complex sport. Our approach introduces a standardised annotation methodology…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Jia Wei Chen

Despite significant breakthroughs in video analysis driven by the rapid development of large multimodal models (LMMs), there remains a lack of a versatile evaluation benchmark to comprehensively assess these models' performance in video…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Yunxin Li , Xinyu Chen , Baotian Hu , Longyue Wang , Haoyuan Shi , Min Zhang

This paper addresses the challenge of automated sports video analysis, which has traditionally been limited by computationally intensive models requiring server-side processing and lacking fine-grained understanding of athletic movements.…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Sai Varun Kodathala , Yashwanth Reddy Vutukoori , Rakesh Vunnam

Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benchmarks mostly resemble static image benchmarks and are…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Mu Cai , Reuben Tan , Jianrui Zhang , Bocheng Zou , Kai Zhang , Feng Yao , Fangrui Zhu , Jing Gu , Yiwu Zhong , Yuzhang Shang , Yao Dou , Jaden Park , Jianfeng Gao , Yong Jae Lee , Jianwei Yang

The immense popularity of racket sports has fueled substantial demand in tactical analysis with broadcast videos. However, existing manual methods require laborious annotation, and recent attempts leveraging video perception models are…

计算机视觉与模式识别 · 计算机科学 2024-02-27 Yuchen He , Zeqing Yuan , Yihong Wu , Liqi Cheng , Dazhen Deng , Yingcai Wu

Evaluating the nuanced human-centric video understanding capabilities of Multimodal Large Language Models (MLLMs) remains a great challenge, as existing benchmarks often overlook the intricacies of emotion, behavior, and cross-modal…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Ting Zhou , Daoyuan Chen , Qirui Jiao , Bolin Ding , Yaliang Li , Ying Shen

Despite recent progress on the short-video Text-Visual Question Answering (ViteVQA) task - largely driven by benchmarks such as M4-ViteVQA - existing datasets still suffer from limited video duration and narrow evaluation scopes, making it…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Yangyang Zhong , Ji Qi , Yuan Yao , Pengxin Luo , Yunfeng Yan , Donglian Qi , Zhiyuan Liu , Tat-Seng Chua

Sports video data is recorded for nearly every major tournament but remains archived and inaccessible to large scale data mining and analytics. It can only be viewed sequentially or manually tagged with higher-level labels which is time…

计算机视觉与模式识别 · 计算机科学 2017-12-27 Anurag Ghosh , Suriya Singh , C. V. Jawahar

Deeply understanding sports requires an intricate blend of fine-grained visual perception and rule-based reasoning - a challenge that pushes the limits of current multimodal models. To succeed, models must master three critical…

We introduce a novel method for collecting table tennis video data and perform stroke detection and classification. A diverse dataset containing video data of 11 basic strokes obtained from 14 professional table tennis players, summing up…

计算机视觉与模式识别 · 计算机科学 2021-06-02 Kaustubh Milind Kulkarni , Sucheth Shenoy

With the recent development of Deep Learning applied to Computer Vision, sport video understanding has gained a lot of attention, providing much richer information for both sport consumers and leagues. This paper introduces…

计算机视觉与模式识别 · 计算机科学 2022-08-18 Gabriel Van Zandycke , Vladimir Somers , Maxime Istasse , Carlo Del Don , Davide Zambrano

Recent advances of deep learning makes it possible to identify specific events in videos with greater precision. This has great relevance in sports like tennis in order to e.g., automatically collect game statistics, or replay actions of…

计算机视觉与模式识别 · 计算机科学 2024-02-06 Emil Hovad , Therese Hougaard-Jensen , Line Katrine Harder Clemmensen

We introduce MMVU, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. MMVU includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science,…

Long-form video understanding poses a significant challenge for video large language models (VideoLLMs) due to prohibitively high computational and memory demands. In this paper, we propose FlexSelect, a flexible and efficient token…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Yunzhu Zhang , Yu Lu , Tianyi Wang , Fengyun Rao , Yi Yang , Linchao Zhu