English
Related papers

Related papers: Game-MUG: Multimodal Oriented Game Situation Under…

200 papers

Sports game summarization aims at generating sports news from live commentaries. However, existing datasets are all constructed through automated collection and cleaning processes, resulting in a lot of noise. Besides, current works neglect…

Computation and Language · Computer Science 2021-11-25 Jiaan Wang , Zhixu Li , Tingyi Zhang , Duo Zheng , Jianfeng Qu , An Liu , Lei Zhao , Zhigang Chen

Sports channel video portals offer an exciting domain for research on multimodal, multilingual analysis. We present methods addressing the problem of automatic video highlight prediction based on joint visual features and textual analysis…

Computation and Language · Computer Science 2017-07-27 Cheng-Yang Fu , Joon Lee , Mohit Bansal , Alexander C. Berg

Real-time understanding of long video streams remains challenging for multimodal large language models (VLMs) due to redundant frame processing and rapid forgetting of past context. Existing streaming systems rely on fixed-interval decoding…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Zhenghui Guo , Yuanbin Man , Junyuan Sheng , Bowen Lin , Ahmed Ahmed , Bo Jiang , Boyuan Zhang , Miao Yin , Sian Jin , Omprakash Gnawal , Chengming Zhang

The increasing complexity of Industry 4.0 systems brings new challenges regarding predictive maintenance tasks such as fault detection and diagnosis. A corresponding and realistic setting includes multi-source data streams from different…

Machine Learning · Computer Science 2024-02-23 Victor Pellegrain , Myriam Tami , Michel Batteux , Céline Hudelot

Is it possible to predict moment-to-moment gameplay engagement based solely on game telemetry? Can we reveal engaging moments of gameplay by observing the way the viewers of the game behave? To address these questions in this paper, we…

Human-Computer Interaction · Computer Science 2020-08-18 David Melhart , Daniele Gravina , Georgios N. Yannakakis

News videos are carefully edited multimodal narratives that combine narration, visuals, and external quotations into coherent storylines. In recent years, there have been significant advances in evaluating multimodal large language models…

Machine Learning · Computer Science 2026-01-08 Zibo Liu , Muyang Li , Zhe Jiang , Shigang Chen

Multimodal video-audio-text understanding and generation can benefit from datasets that are narrow but rich. The narrowness allows bite-sized challenges that the research community can make progress on. The richness ensures we are making…

Computer Vision and Pattern Recognition · Computer Science 2022-04-29 Thomas Hayes , Songyang Zhang , Xi Yin , Guan Pang , Sasha Sheng , Harry Yang , Songwei Ge , Qiyuan Hu , Devi Parikh

Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over omni-modal (visual, audio, and textual) signals in long and…

Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Weihan Wang , Zehai He , Wenyi Hong , Yean Cheng , Xiaohan Zhang , Ji Qi , Xiaotao Gu , Shiyu Huang , Bin Xu , Yuxiao Dong , Ming Ding , Jie Tang

Recent video+language datasets cover domains where the interaction is highly structured, such as instructional videos, or where the interaction is scripted, such as TV shows. Both of these properties can lead to spurious cues to be…

Computer Vision and Pattern Recognition · Computer Science 2022-11-10 Alessandro Suglia , José Lopes , Emanuele Bastianelli , Andrea Vanzo , Shubham Agarwal , Malvina Nikandrou , Lu Yu , Ioannis Konstas , Verena Rieser

Esports has emerged as a popular genre for players as well as spectators, supporting a global entertainment industry. Esports analytics has evolved to address the requirement for data-driven feedback, and is focused on cyber-athlete…

Artificial Intelligence · Computer Science 2017-11-20 Victoria Hodge , Sam Devlin , Nick Sephton , Florian Block , Anders Drachen , Peter Cowling

The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Tianhao Peng , Haochen Wang , Yuanxing Zhang , Zekun Wang , Zili Wang , Gavin Chang , Jian Yang , Shihao Li , Yanghai Wang , Xintao Wang , Houyi Li , Wei Ji , Pengfei Wan , Steven Huang , Zhaoxiang Zhang , Jiaheng Liu

This paper is the basis paper for the accepted IJCNN challenge One-Minute Gradual-Emotion Recognition (OMG-Emotion) by which we hope to foster long-emotion classification using neural models for the benefit of the IJCNN community. The…

Human-Computer Interaction · Computer Science 2018-05-23 Pablo Barros , Nikhil Churamani , Egor Lakomkin , Henrique Siqueira , Alexander Sutherland , Stefan Wermter

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Jielin Qiu , Jiacheng Zhu , William Han , Aditesh Kumar , Karthik Mittal , Claire Jin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Ding Zhao , Bo Li , Lijuan Wang

In the context of today's high-pressure, aging society, the demand for large-scale emotional models capable of providing empathetic support is more critical than ever. However, existing benchmarks fail to simultaneously achieve ecological…

Computation and Language · Computer Science 2026-05-12 Pengze Guo , Jingxi Liang , Zhiwen Xie , Qifeng Wang , Derek F. Wong

How are we able to learn about complex current events just from short snippets of video? While natural language enables straightforward ways to represent under-specified, partially observable events, visual data does not facilitate…

Computation and Language · Computer Science 2024-10-08 Kate Sanders , Reno Kriz , David Etter , Hannah Recknor , Alexander Martin , Cameron Carpenter , Jingyang Lin , Benjamin Van Durme

Competitive games pose steep learning curves and strong social pressures, often discouraging novice players and limiting sustained engagement. To address these challenges, this study introduces LeagueBot, a large language model-based voice…

Human-Computer Interaction · Computer Science 2026-02-03 Jungmin Lee , Inhee Cho , Youngjae Yoo

Proper training and analytics in eSports require accurately collected and annotated data. Most eSports research focuses exclusively on in-game data analysis, and there is a lack of prior work involving eSports athletes' psychophysiological…

Human-Computer Interaction · Computer Science 2021-08-24 Anton Smerdov , Bo Zhou , Paul Lukowicz , Andrey Somov

People are sharing their opinions, stories and reviews through online video sharing websites every day. Studying sentiment and subjectivity in these opinion videos is experiencing a growing attention from academia and industry. While…

Computation and Language · Computer Science 2016-11-18 Amir Zadeh , Rowan Zellers , Eli Pincus , Louis-Philippe Morency

When people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information describing the important high-level details (what, where,…

Computer Vision and Pattern Recognition · Computer Science 2021-05-11 Mathew Monfort , SouYoung Jin , Alexander Liu , David Harwath , Rogerio Feris , James Glass , Aude Oliva