中文
相关论文

相关论文: BoxComm: Benchmarking Category-Aware Commentary Ge…

200 篇论文

Recently, significant advances have been made in Video Large Language Models (Video LLMs) in both academia and industry. However, methods to evaluate and benchmark the performance of different Video LLMs, especially their fine-grained,…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Kuangzhi Ge , Lingjun Chen , Kevin Zhang , Yulin Luo , Tianyu Shi , Liaoyuan Fan , Xiang Li , Guanqun Wang , Shanghang Zhang

In competitive combat sports like boxing, analyzing a boxers's performance statics is crucial for evaluating the quantity and variety of punches delivered during bouts. These statistics provide valuable data and feedback, which are…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Shashikanta Sahoo

Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and livestreaming. Commentary generation involves two essential…

Soccer is a globally popular sport with a vast audience, in this paper, we consider constructing an automatic soccer game commentary model to improve the audiences' viewing experience. In general, we make the following contributions: First,…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Jiayuan Rao , Haoning Wu , Chang Liu , Yanfeng Wang , Weidi Xie

Soccer is a globally popular sporting event, typically characterized by long matches and distinctive highlight moments. Recent advances in Multimodal Large Language Models (MLLMs) offer promising capabilities in temporal grounding and video…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Ling You , Wenxuan Huang , Xinni Xie , Xiangyi Wei , Bangyan Li , Shaohui Lin , Yang Li , Changbo Wang

The advent of large (visual) language models (LLM / LVLM) have led to a deluge of automated human-like systems in several domains including social media content generation, search and recommendation, healthcare prognosis, AI assistants for…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Sauptik Dhar , Nicholas Buoncristiani , Joe Anakata , Haoyu Zhang , Michelle Munson

While there is overall agreement that future technology for organizing, browsing and searching videos hinges on the development of methods for high-level semantic understanding of video, so far no consensus has been reached on the best way…

计算机视觉与模式识别 · 计算机科学 2017-06-20 Du Tran , Maksim Bolonkin , Manohar Paluri , Lorenzo Torresani

Accurate analysis of combat sports using computer vision has gained traction in recent years, yet the development of robust datasets remains a major bottleneck due to the dynamic, unstructured nature of actions and variations in recording…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Rahul Kumar , Vipul Baghel , Sudhanshu Singh , Bikash Kumar Badatya , Shivam Yadav , Babji Srinivasan , Ravi Hegde

Understanding the world and explaining it with scientific theories is a central aspiration of artificial intelligence research. Proposing theories, designing experiments to test them, and then revising them based on data are fundamental to…

In the pursuit of natural language understanding, there has been a long standing interest in tracking state changes throughout narratives. Impressive progress has been made in modeling the state of transaction-centric dialogues and…

计算与语言 · 计算机科学 2021-06-04 Ruochen Zhang , Carsten Eickhoff

Despite the recent emergence of video captioning models, how to generate vivid, fine-grained video descriptions based on the background knowledge (i.e., long and informative commentary about the domain-specific scenes with appropriate…

计算机视觉与模式识别 · 计算机科学 2023-10-06 Ji Qi , Jifan Yu , Teng Tu , Kunyu Gao , Yifan Xu , Xinyu Guan , Xiaozhi Wang , Yuxiao Dong , Bin Xu , Lei Hou , Juanzi Li , Jie Tang , Weidong Guo , Hui Liu , Yu Xu

The advent of artificial intelligence has propelled AI-Generated Game Commentary (AI-GGC) into a rapidly expanding field, offering benefits such as unlimited availability and personalized narration. However, current researches in this area…

计算与语言 · 计算机科学 2025-10-21 Qirui Zheng , Xingbo Wang , Keyuan Cheng , Muhammad Asif Ali , Yunlong Lu , Wenxin Li

In this paper, we tackle the problem of how to build and benchmark a large motion model (LMM). The ultimate goal of LMM is to serve as a foundation model for versatile motion-related tasks, e.g., human motion generation, with…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Liang Xu , Shaoyang Hua , Zili Lin , Yifan Liu , Feipeng Ma , Yichao Yan , Xin Jin , Xiaokang Yang , Wenjun Zeng

Sports videos are a challenging domain for multimodal understanding because they involve complex and dynamic human activities. Despite rapid progress in Multimodal Large Language Models (MLLMs), long-horizon reasoning in sports videos…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Siyu Cao , Lu Zhang , Ruizhe Zeng , Zhi-yong Liu

Learning commonsense reasoning from visual contexts and scenes in real-world is a crucial step toward advanced artificial intelligence. However, existing video reasoning benchmarks are still inadequate since they were mainly designed for…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Andong Wang , Bo Wu , Sunli Chen , Zhenfang Chen , Haotian Guan , Wei-Ning Lee , Li Erran Li , Chuang Gan

Deeply understanding sports requires an intricate blend of fine-grained visual perception and rule-based reasoning - a challenge that pushes the limits of current multimodal models. To succeed, models must master three critical…

Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories:…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Hui Han , Siyuan Li , Jiaqi Chen , Yiwen Yuan , Yuling Wu , Chak Tou Leong , Hanwen Du , Junchen Fu , Youhua Li , Jie Zhang , Chi Zhang , Li-jia Li , Yongxin Ni

Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. To address this gap, we introduce \textbf{\textit{CultSportQA}}, a benchmark designed to assess LMs'…

The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, track object state…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Yang Dai , Dian Jiao , Tianwei Lin , Wenqiao Zhang

While Multimodal Large Language Models (MLLMs) excel at generic video understanding, their ability to support specialized, rule-grounded decision-making remains insufficiently explored. In this paper, we introduce RefereeBench, the first…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Yichen Xu , Yuanhang Liu , Chuhan Wang , Zihan Zhao , jinghan luo , Jianzhe Ma , Wenxuan Wang , Qin Jin
‹ 上一页 1 2 3 10 下一页 ›