English
Related papers

Related papers: MTP: A Dataset for Multi-Modal Turning Points in C…

200 papers

Traffic accident prediction in driving videos aims to provide an early warning of the accident occurrence, and supports the decision making of safe driving systems. Previous works usually concentrate on the spatial-temporal correlation of…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Jianwu Fang , Lei-Lei Li , Kuan Yang , Zhedong Zheng , Jianru Xue , Tat-Seng Chua

Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of multi-turn…

Computation and Language · Computer Science 2025-07-18 Yiming Lei , Zhizheng Yang , Zeming Liu , Haitao Leng , Shaoguo Liu , Tingting Gao , Qingjie Liu , Yunhong Wang

Table pretrain-then-finetune paradigm has been proposed and employed at a rapid pace after the success of pre-training in the natural language domain. Despite the promising findings in tabular pre-trained language models (TPLMs), there is…

Computation and Language · Computer Science 2023-02-21 Nuo Chen , Linjun Shou , Ming Gong , Jian Pei , Chenyu You , Jianhui Chang , Daxin Jiang , Jia Li

Tipping points are moments of change that characterise crucial turning points in a piece of music. This study presents a first step towards quantitatively and systematically describing the musical properties of tipping points. Timing…

Sound · Computer Science 2024-12-17 Canishk Naik , Elaine Chew

Large language models (LLMs) excel at solving problems with clear and complete statements, but often struggle with nuanced environments or interactive tasks which are common in most real-world scenarios. This highlights the critical need…

As multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: What essential capabilities are still missing? A critical aspect of human learning is continuous interaction with the environment --…

Current vision-language multimodal models are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Dewen Zhang , Wangpeng An , Hayaru Shouno

Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but extending this paradigm to video remains challenging due to the difficulty of extracting…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Patrick Knab , Sascha Marton , Philipp J. Schubert , Drago Guggiana , Christian Bartelt

Recent advances in text-to-music (TTM) generation have enabled controllable and expressive music creation using natural language prompts. However, the emotional fidelity of TTM systems remains largely underexplored compared to human…

Sound · Computer Science 2025-09-05 Gyehun Go , Satbyul Han , Ahyeon Choi , Eunjin Choi , Juhan Nam , Jeong Mi Park

Accurate recognition of human emotions is critical for adaptive human-computer interaction, yet remains challenging in dynamic, conversation-like settings. This work presents a personality-aware multimodal framework that integrates…

Language confusion -- where large language models (LLMs) generate unintended languages against the user's need -- remains a critical challenge, especially for English-centric models. We present the first mechanistic interpretability (MI)…

Computation and Language · Computer Science 2025-09-19 Ercong Nie , Helmut Schmid , Hinrich Schütze

Deception detection has garnered increasing attention in recent years due to the significant growth of digital media and heightened ethical and security concerns. It has been extensively studied using multimodal methods, including video,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Cong Cai , Shan Liang , Xuefei Liu , Kang Zhu , Zhengqi Wen , Jianhua Tao , Heng Xie , Jizhou Cui , Yiming Ma , Zhenhua Cheng , Hanzhe Xu , Ruibo Fu , Bin Liu , Yongwei Li

This paper presents ConvBench, a novel multi-turn conversation evaluation benchmark tailored for Large Vision-Language Models (LVLMs). Unlike existing benchmarks that assess individual capabilities in single-turn dialogues, ConvBench adopts…

Multimedia · Computer Science 2024-04-26 Shuo Liu , Kaining Ying , Hao Zhang , Yue Yang , Yuqi Lin , Tianle Zhang , Chuanhao Li , Yu Qiao , Ping Luo , Wenqi Shao , Kaipeng Zhang

Understanding sentiment in multimodal conversations is a complex yet crucial challenge toward building emotionally intelligent AI systems. The Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA) Challenge invited participants…

Computation and Language · Computer Science 2025-12-30 Zhiqiang Gao , Shihao Gao , Zixing Zhang , Yihao Guo , Hongyu Chen , Jing Han

Large Language Models (LLMs) have demonstrated impressive capabilities as intelligent agents capable of solving complex problems. However, effective planning in scenarios involving dependencies between API or tool calls-particularly in…

Emotion detection is a central problem in NLP, with recent progress driven by transformer-based models trained on established datasets. However, little is known about the linguistic regularities that characterize how emotions are expressed…

Computation and Language · Computer Science 2026-03-24 Florian Lecourt , Madalina Croitoru , Konstantin Todorov

In this paper, we introduce the action language C-MT (Mind Transition Language). It is built on top of answer set programming (ASP) and transition systems to represent how human mental states evolve in response to sequences of observable…

Artificial Intelligence · Computer Science 2026-05-13 Andreas Brännström , Juan Carlos Nieves

Autonomous Vehicle decisions rely on multimodal prediction models that account for multiple route options and the inherent uncertainty in human behavior. However, models can suffer from mode collapse, where only the most likely mode is…

Robotics · Computer Science 2025-07-01 Maarten Hugenholtz , Anna Meszaros , Jens Kober , Zlatan Ajanovic

Sentiment analysis, also referred to as opinion mining, primarily tries to extract opinion from any text-based data. In the context of movie reviews and critics, sentimental analysis can be a helpful tool to predict whether a movie review…

Computation and Language · Computer Science 2026-05-22 Dip Biswas Shanto , Mitali Yadav , Prajwal Panth , Suresh Chandra Satapathy

Emotion recognition in conversation (ERC) has been attracting attention by methods for modeling multi-turn contexts. The multi-turn input to a pretraining model implicitly assumes that the current turn and other turns are distinguished…

Computation and Language · Computer Science 2025-01-03 Junya Ono , Hiromi Wakaki