中文
相关论文

相关论文: Game-MUG: Multimodal Oriented Game Situation Under…

200 篇论文

Video temporal understanding is crucial for multimodal large language models (MLLMs) to reason over events in videos. Despite recent advances in general video understanding, current MLLMs still struggle with fine-grained temporal reasoning.…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Fuwen Luo , Shengfeng Lou , Chi Chen , Ziyue Wang , Chenliang Li , Weizhou Shen , Jiyue Guo , Peng Li , Ming Yan , Ji Zhang , Fei Huang , Yang Liu

Recent multimodal large language models (MLLMs) have shown strong capabilities in general video understanding, driving growing interest in automatic sports commentary generation. However, existing benchmarks for this task focus exclusively…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Kaiwen Wang , Kaili Zheng , Rongrong Deng , Yiming Shi , Chenyi Guo , Ji Wu

What if emotion could be captured in a general and subject-agnostic fashion? Is it possible, for instance, to design general-purpose representations that detect affect solely from the pixels and audio of a human-computer interaction video?…

人机交互 · 计算机科学 2021-03-01 Konstantinos Makantasis , Antonios Liapis , Georgios N. Yannakakis

Given a video with aligned dialogue, people can often infer what is more likely to happen next. Making such predictions requires not only a deep understanding of the rich dynamics underlying the video and dialogue, but also a significant…

计算与语言 · 计算机科学 2020-10-19 Jie Lei , Licheng Yu , Tamara L. Berg , Mohit Bansal

Multimodal large language models (MLLMs) have shown strong performance on offline video understanding, but most are limited to offline inference or have weak online reasoning, making multi-turn interaction over continuously arriving video…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Lu Wang , Zhuoran Jin , Yupu Hao , Yubo Chen , Kang Liu , Yulong Ao , Jun Zhao

Multi-modal object tracking (MMOT) is an emerging field that combines data from various modalities, \eg vision (RGB), depth, thermal infrared, event, language and audio, to estimate the state of an arbitrary object in a video sequence. It…

计算机视觉与模式识别 · 计算机科学 2024-06-03 Chunhui Zhang , Li Liu , Hao Wen , Xi Zhou , Yanfeng Wang

Chatbots via large language models (LLMs) generate fluent responses but often struggle with when to speak, especially for brief, timely listener reactions during ongoing dialogue. We present a multimodal strategy for LLMs, which leverages…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Zikai Liao , Yi Ouyang , Yi-Lun Lee , Chen-Ping Yu , Yi-Hsuan Tsai , Zhaozheng Yin

GuessWhich is an engaging visual dialogue game that involves interaction between a Questioner Bot (QBot) and an Answer Bot (ABot) in the context of image-guessing. In this game, QBot's objective is to locate a concealed image solely through…

人工智能 · 计算机科学 2024-08-19 Wei Pang , Ruixue Duan , Jinfu Yang , Ning Li

Sports game summarization aims to generate sports news based on real-time commentaries. The task has attracted wide research attention but is still under-explored probably due to the lack of corresponding English datasets. Therefore, in…

计算与语言 · 计算机科学 2022-07-19 Jiaan Wang , Tingyi Zhang , Haoxiang Shi

Humans express feelings or emotions via different channels. Take language as an example, it entails different sentiments under different visual-acoustic contexts. To precisely understand human intentions as well as reduce the…

人工智能 · 计算机科学 2021-11-17 Ting Wu , Junjie Peng , Wenqiang Zhang , Huiran Zhang , Chuanshuai Ma , Yansong Huang

As humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future. We introduce MERLOT, a model that learns multimodal script knowledge by…

计算机视觉与模式识别 · 计算机科学 2021-10-25 Rowan Zellers , Ximing Lu , Jack Hessel , Youngjae Yu , Jae Sung Park , Jize Cao , Ali Farhadi , Yejin Choi

Accurate emotion understanding in videos necessitates effectively recognizing and interpreting emotional states by integrating visual, textual, auditory, and contextual cues. Although recent Large Multimodal Models (LMMs) have exhibited…

Online education platforms have experienced explosive growth over the past decade, generating massive volumes of user-generated content in the form of reviews, ratings, and behavioral logs. These heterogeneous signals provide unprecedented…

图形学 · 计算机科学 2026-04-14 Arman Bekov , Azamat Nurgali

Videoconferencing is now a frequent mode of communication in both professional and informal settings, yet it often lacks the fluidity and enjoyment of in-person conversation. This study leverages multimodal machine learning to predict…

机器学习 · 计算机科学 2025-03-11 Andrew Chang , Viswadruth Akkaraju , Ray McFadden Cogliano , David Poeppel , Dustin Freeman

Large Language Models (\textbf{LLMs}), e.g. ChatGPT, have been widely adopted in real-world dialogue applications. However, LLMs' robustness, especially in handling long complex dialogue sessions, including frequent motivation transfer,…

计算与语言 · 计算机科学 2025-09-16 Chenghao Yang , Yinbo Luo , Zhoufutu Wen , Qi Chu , Tao Gong , Longxiang Liu , Kaiyuan Zhang , Jianpeng Jiao , Ge Zhang , Wenhao Huang , Nenghai Yu

Streaming video understanding demands more than watching longer videos: assistants must decide when to speak in real time, balancing responsiveness against verbosity. Yet most video-language models (VideoLLMs) are trained for offline…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Zichen Wen , Boxue Yang , Junlong Ke , Jiajie Huang , Chenfei Liao , Junxi Wang , Xuyang Liu , Linfeng Zhang

The rapid growth of live-streaming platforms such as Twitch has introduced complex challenges in moderating toxic behavior. Traditional moderation approaches, such as human annotation and keyword-based filtering, have demonstrated utility,…

计算与语言 · 计算机科学 2026-02-05 Baktash Ansari , Elias Martin , Afra Mashhadi

Compared to traditional sentiment analysis, which only considers text, multimodal sentiment analysis needs to consider emotional signals from multimodal sources simultaneously and is therefore more consistent with the way how humans process…

计算与语言 · 计算机科学 2024-08-19 Hao Yang , Yanyan Zhao , Yang Wu , Shilong Wang , Tian Zheng , Hongbo Zhang , Zongyang Ma , Wanxiang Che , Bing Qin

Recent large vision-language models have achieved strong performance on short- and medium-length video understanding, yet they remain inadequate for ultra-long or even infinite video reasoning, where models must preserve coherent memory…

人工智能 · 计算机科学 2026-05-08 Peizheng Yan , Yu Zhao , Liang Xie , Juntong Qi , Mingming Wang , Erwei Yin

Soccer video understanding has motivated the creation of datasets for tasks such as temporal action localization, spatiotemporal action detection (STAD), or multiobject tracking (MOT). The annotation of structured sequences of events (who…

人工智能 · 计算机科学 2025-11-21 Jeremie Ochin , Raphael Chekroun , Bogdan Stanciulescu , Sotiris Manitsaris