English
Related papers

Related papers: MISID: A Multimodal Multi-turn Dataset for Complex…

200 papers

Traditional recommendation models trained on observational interaction data have generated large impacts in a wide range of applications, it faces bias problems that cover users' true intent and thus deteriorate the recommendation…

Information Retrieval · Computer Science 2022-02-08 Xiangmeng Wang , Qian Li , Dianer Yu , Peng Cui , Zhichao Wang , Guandong Xu

While Large Language Model (LLM) capabilities have scaled, safety guardrails remain largely stateless, treating multi-turn dialogues as a series of disconnected events. This lack of temporal awareness facilitates a "Safety Gap" where…

Artificial Intelligence · Computer Science 2026-02-20 Justin Albrethsen , Yash Datta , Kunal Kumar , Sharath Rajasekar

The emergence of multimodal large language models has redefined the agent paradigm by integrating language and vision modalities with external data sources, enabling agents to better interpret human instructions and execute increasingly…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Le Wang , Zonghao Ying , Tianyuan Zhang , Siyuan Liang , Shengshan Hu , Mingchuan Zhang , Aishan Liu , Xianglong Liu

In high-conflict mixed-traffic scenarios involving human-driven and autonomous vehicles, most existing autonomous driving systems default to overly conservative behaviors, lack proactive interaction, and consequently suffer from limited…

Robotics · Computer Science 2026-04-28 Xinwei Dong , Jiyang Li , Jiabin Xie , Yang Yi , Tianshang Jia , Shiyu Fang , Ye Tian , Peng Hang

Theory of Mind (ToM) in Large Language Models (LLMs) refers to the model's ability to infer the mental states of others, with failures in this ability often manifesting as systemic implicit biases. Assessing this challenge is difficult, as…

Computation and Language · Computer Science 2026-01-19 Yanlin Li , Hao Liu , Huimin Liu , Kun Wang , Yinwei Wei , Yupeng Hu

In healthcare intelligence, the ability to fuse heterogeneous, multi-intent information from diverse clinical sources is fundamental to building reliable decision-making systems. Large Language Model (LLM)-driven information interaction…

Computation and Language · Computer Science 2025-07-04 Dingkang Yang , Jinjie Wei , Mingcheng Li , Jiyao Liu , Lihao Liu , Ming Hu , Junjun He , Yakun Ju , Wei Zhou , Yang Liu , Lihua Zhang

Multimodal fusion leverages information across modalities to learn better feature representations with the goal of improving performance in fusion-based tasks. However, multimodal datasets, especially in medical settings, are typically…

Machine Learning · Computer Science 2025-02-05 Alejandro Guerra-Manzanares , Farah E. Shamout

Identifying intents from dialogue utterances forms an integral component of task-oriented dialogue systems. Intent-related tasks are typically formulated either as a classification task, where the utterances are classified into predefined…

Computation and Language · Computer Science 2023-10-26 Bhavuk Singhal , Ashim Gupta , Shivasankaran V P , Amrith Krishna

Identifying user intents from natural language utterances is a crucial step in conversational systems that has been extensively studied as a supervised classification problem. However, in practice, new intents emerge after deploying an…

Computation and Language · Computer Science 2021-02-08 A. B. Siddique , Fuad Jamour , Luxun Xu , Vagelis Hristidis

Identifying the unknown (novel) user intents that have never appeared in the training set is a challenging task in the dialogue system. In this paper, we present a two-stage method for detecting unknown intents. We use bidirectional long…

Computation and Language · Computer Science 2019-06-04 Ting-En Lin , Hua Xu

Hallucination continues to pose a major obstacle in the reasoning capabilities of large language models (LLMs). Although the Multi-Agent Debate (MAD) paradigm offers a promising solution by promoting consensus among multiple agents to…

Artificial Intelligence · Computer Science 2025-11-17 Dayong Liang , Xiao-Yong Wei , Changmeng Zheng

Active Real-time interaction with video LLMs introduces a new paradigm for human-computer interaction, where the model not only understands user intent but also responds while continuously processing streaming video on the fly. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Rui Qian , Shuangrui Ding , Xiaoyi Dong , Pan Zhang , Yuhang Zang , Yuhang Cao , Dahua Lin , Jiaqi Wang

Identifying user intent from mobile UI operation trajectories is critical for advancing UI understanding and enabling task automation agents. While Multimodal Large Language Models (MLLMs) excel at video understanding tasks, their real-time…

Artificial Intelligence · Computer Science 2025-12-23 Zhe Yang , Xiaoshuang Sheng , Zhengnan Zhang , Jidong Wu , Zexing Wang , Xin He , Shenghua Xu , Guanjing Xiong

Stance detection, which aims to identify public opinion towards specific targets using social media data, is an important yet challenging task. With the proliferation of diverse multimodal social media content including text, and images…

Multimedia · Computer Science 2024-09-04 Fuqiang Niu , Zebang Cheng , Xianghua Fu , Xiaojiang Peng , Genan Dai , Yin Chen , Hu Huang , Bowen Zhang

Despite the remarkable advancements of Large Vision-Language Models (LVLMs), the mechanistic interpretability remains underexplored. Existing analyses are insufficiently comprehensive and lack examination covering visual and textual tokens,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Qiming Li , Zekai Ye , Xiaocheng Feng , Weihong Zhong , Weitao Ma , Xiachong Feng

Employing skin-like tactile sensors on robots enhances both the safety and usability of collaborative robots by adding the capability to detect human contact. Unfortunately, simple binary tactile sensors alone cannot determine the context…

Robotics · Computer Science 2023-04-20 Christopher Yee Wong , Lucas Vergez , Wael Suleiman

Recently, multimodal large language models (MLLMs) have been widely applied to reasoning tasks. However, they suffer from limited multi-rationale semantic modeling, insufficient logical robustness, and are susceptible to misleading…

Artificial Intelligence · Computer Science 2025-12-08 Chuang Yu , Jinmiao Zhao , Mingxuan Zhao , Yunpeng Liu , Xiujun Shu , Yuanhao Feng , Bo Wang , Xiangyu Yue

Multimodal Misinformation Detection (MMD) refers to the task of detecting social media posts involving misinformation, where the post often contains text and image modalities. However, by observing the MMD posts, we hold that the text…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Bing Wang , Ximing Li , Yanjun Wang , Changchun Li , Lin Yuanbo Wu , Buyu Wang , Shengsheng Wang

Existing sequential recommendation models, even advanced diffusion-based approaches, often struggle to capture the rich semantic intent underlying user behavior, especially for new users or long-tail items. This limitation stems from their…

Information Retrieval · Computer Science 2026-01-08 Bo-Chian Chen , Manel Slokom

In conversational AI systems, a critical challenge in training effective multi-turn intent classification models lies in the generation of large-scale, domain-specific, multilingual dialogue datasets. In this paper, we introduce…

Computation and Language · Computer Science 2025-09-03 Junhua Liu , Yong Keat Tan , Bin Fu , Kwan Hui Lim
‹ Prev 1 3 4 5 6 7 10 Next ›