中文
相关论文

相关论文: Game-MUG: Multimodal Oriented Game Situation Under…

200 篇论文

Human team tactics emerge from each player's individual perspective and their ability to anticipate, interpret, and adapt to teammates' intentions. While advances in video understanding have improved the modeling of team interactions in…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Yunzhe Wang , Soham Hans , Volkan Ustun

Social media in present times has a significant and growing influence. Fake news being spread on these platforms have a disruptive and damaging impact on our lives. Furthermore, as multimedia content improves the visibility of posts more…

多媒体 · 计算机科学 2024-06-13 Mudit Dhawan , Shakshi Sharma , Aditya Kadam , Rajesh Sharma , Ponnurangam Kumaraguru

Sports video understanding requires perceiving high-speed dynamics, complex rules, and long temporal contexts. Yet, current Multimodal Large Language Models (MLLMs) remain narrowly focused on single sports, specific tasks, or training-free…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Junbo Zou , Haotian Xia , Zhen Ye , Shengjie Zhang , Christopher Lai , Vicente Ordonez , Weining Shen , Hanjie Chen

Sports analysis and viewing play a pivotal role in the current sports domain, offering significant value not only to coaches and athletes but also to fans and the media. In recent years, the rapid development of virtual reality (VR) and…

计算机视觉与模式识别 · 计算机科学 2024-05-03 Wenxuan Guo , Zhiyu Pan , Ziheng Xi , Alapati Tuerxun , Jianjiang Feng , Jie Zhou

In the Massive Open Online Courses (MOOC) learning scenario, the semantic information of instructional videos has a crucial impact on learners' emotional state. Learners mainly acquire knowledge by watching instructional videos, and the…

多媒体 · 计算机科学 2024-04-12 Yuan Zhang , Xiaomei Tao , Hanxu Ai , Tao Chen , Yanling Gan

This paper introduces the schemes of Team LingJing's experiments in NLPCC-2022-Shared-Task-4 Multi-modal Dialogue Understanding and Generation (MDUG). The MDUG task can be divided into two phases: multi-modal context understanding and…

计算与语言 · 计算机科学 2022-07-06 Bin Li , Yixuan Weng , Ziyu Ma , Bin Sun , Shutao Li

Foundation models are used for many real-world applications involving language generation from temporally-ordered multimodal events. In this work, we study the ability of models to identify the most important sub-events in a video, which is…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Aditya K Surikuchi , Raquel Fernández , Sandro Pezzelle

In most team-based esports, voice communications are prominent in the team efficiency and synergy. In fact it has been observed that not only the skill aspect of the team but also the team effective voice communication comes into play when…

声音 · 计算机科学 2024-12-02 Aymeric Vinot , Nicolas Perez

Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long…

计算机视觉与模式识别 · 计算机科学 2025-01-24 Haomiao Xiong , Zongxin Yang , Jiazuo Yu , Yunzhi Zhuge , Lu Zhang , Jiawen Zhu , Huchuan Lu

The development of video game streaming has grown rapidly, with major platforms such as YouTube and Twitch using different codecs. To support quality assessment models that work consistently across any codec, it is necessary to have access…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Rajesh Sureddi , Shreshth Saini , Avinab Saha , Alan C. Bovik

Contemporary news reporting increasingly features multimedia content, motivating research on multimedia event extraction. However, the task lacks annotated multimodal training data and artificially generated training data suffer from…

多媒体 · 计算机科学 2023-08-14 Zilin Du , Yunxin Li , Xu Guo , Yidan Sun , Boyang Li

Quantitative analysis of Game User eXperience (GUX) is important to the game industry. Different from the typical questionnaire analysis, this paper focuses on the computational analysis of GUX. We aim to analyze the relationship between…

人机交互 · 计算机科学 2021-12-23 Zhitao Liu , Ning Xie , Guobiao Yang , Jiale Dou , Lanxiao Huang , Guang Yang , Lin Yuan

The rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despite their potential,…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Yuxuan Wang , Yueqian Wang , Bo Chen , Tong Wu , Dongyan Zhao , Zilong Zheng

Traditional esports scouting workflows rely heavily on manual video review and aggregate performance metrics, which often fail to capture the nuanced decision-making patterns necessary to determine if a prospect fits a specific tactical…

机器学习 · 计算机科学 2026-04-17 Qing Yan , Wenyu Yang , Yufei Wang , Wenhao Ma , Linchong Hu , Yifei Jin , Anton Dahbura

eSports is a developing multidisciplinary research area. At present, there is a lack of relevant data collected from real eSports athletes and lack of platforms which could be used for the data collection and further analysis. In this…

人机交互 · 计算机科学 2019-08-20 Alexander Korotin , Nikita Khromov , Anton Stepanov , Andrey Lange , Evgeny Burnaev , Andrey Somov

The rapid evolution of digital sports media necessitates sophisticated information retrieval systems that can efficiently parse extensive multimodal datasets. This paper introduces SoccerRAG, an innovative framework designed to harness the…

信息检索 · 计算机科学 2025-08-26 Aleksander Theo Strand , Sushant Gautam , Cise Midoglu , Pål Halvorsen

The integration of artificial intelligence in sports analytics has transformed soccer video understanding, enabling real-time, automated insights into complex game dynamics. Traditional approaches rely on isolated data streams, limiting…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Sushant Gautam , Cise Midoglu , Vajira Thambawita , Michael A. Riegler , Pål Halvorsen , Mubarak Shah

Dark humor in online memes poses unique challenges due to its reliance on implicit, sensitive, and culturally contextual cues. To address the lack of resources and methods for detecting dark humor in multimodal content, we introduce a novel…

Understanding videos is an important research topic for multimodal learning. Leveraging large-scale datasets of web-crawled video-text pairs as weak supervision has become a pre-training paradigm for learning joint representations and…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Gengyuan Zhang , Jinhe Bi , Jindong Gu , Yanyu Chen , Volker Tresp

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual-text processing. However, existing static image-text benchmarks are insufficient for evaluating their dynamic perception and…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Xiangxi Zheng , Linjie Li , Zhengyuan Yang , Ping Yu , Alex Jinpeng Wang , Rui Yan , Yuan Yao , Lijuan Wang