English
Related papers

Related papers: Plug-and-Play Clarifier: A Zero-Shot Multimodal Fr…

200 papers

The ultimate goal of embodied agents is to create collaborators that can interact with humans, not mere executors that passively follow instructions. This requires agents to communicate, coordinate, and adapt their actions based on human…

Robotics · Computer Science 2025-09-22 Xingyao Lin , Xinghao Zhu , Tianyi Lu , Sicheng Xie , Hui Zhang , Xipeng Qiu , Zuxuan Wu , Yu-Gang Jiang

Resolving ambiguities through interaction is a hallmark of natural language, and modeling this behavior is a core challenge in crafting AI assistants. In this work, we study such behavior in LMs by proposing a task-agnostic framework for…

Computation and Language · Computer Science 2023-11-17 Michael J. Q. Zhang , Eunsol Choi

Large language models (LLMs) have shown remarkable progress in understanding and generating natural language across various applications. However, they often struggle with resolving ambiguities in real-world, enterprise-level interactions,…

Computation and Language · Computer Science 2025-03-28 John Murzaku , Zifan Liu , Vaishnavi Muppala , Md Mehrab Tanjim , Xiang Chen , Yunyao Li

The Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Mingliang Zhai , Cheng Li , Zengyuan Guo , Ningrui Yang , Xiameng Qin , Sanyuan Zhao , Junyu Han , Ji Tao , Yuwei Wu , Yunde Jia

Multimodal Large Language Models (MLLMs) have achieved impressive progress in vision-language alignment, yet they remain limited in visual-spatial reasoning. We first identify that this limitation arises from the attention mechanism: visual…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Zhaozhi Wang , Tong Zhang , Mingyue Guo , Yaowei Wang , Qixiang Ye

Large Language Models are rapidly emerging as web-native interfaces to social platforms. On the social web, users frequently have ambiguous and dynamic goals, making complex intent understanding-rather than single-turn execution-the…

Artificial Intelligence · Computer Science 2026-01-27 Zenghua Liao , Jinzhi Liao , Xiang Zhao

Low-light image enhancement is a crucial visual task, and many unsupervised methods tend to overlook the degradation of visible information in low-light scenes, which adversely affects the fusion of complementary information and hinders the…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Xiaofeng Zhang , Zishan Xu , Hao Tang , Chaochen Gu , Wei Chen , Shanying Zhu , Xinping Guan

Egocentric AI agents, such as smart glasses, rely on pointing gestures to resolve referential ambiguities in natural language commands. However, despite advancements in Multimodal Large Language Models (MLLMs), current systems often fail to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Chentao Li , Zirui Gao , Mingze Gao , Yinglian Ren , Jianjiang Feng , Jie Zhou

In-context learning (ICL) allows large models to adapt to tasks using a few examples, yet its extension to vision-language models (VLMs) remains fragile. Our analysis reveals that the fundamental limitation lies in an inductive gap, models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Haoyu Wang , Haonan Wang , Yuyan Chen , Jun Chen , Gang Liu , Qian Wang , Jiahong Yan , Yanghua Xiao

In visual question answering (VQA) context, users often pose ambiguous questions to visual language models (VLMs) due to varying expression habits. Existing research addresses such ambiguities primarily by rephrasing questions. These…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Pu Jian , Donglei Yu , Wen Yang , Shuo Ren , Jiajun Zhang

Explanations are crucial for building trustworthy AI systems, but a gap often exists between the explanations provided by models and those needed by users. To address this gap, we introduce MetaExplainer, a neuro-symbolic framework designed…

Human-Computer Interaction · Computer Science 2025-09-11 Shruthi Chari , Oshani Seneviratne , Prithwish Chakraborty , Pablo Meyer , Deborah L. McGuinness

Intelligent virtual assistants are currently designed to perform tasks or services explicitly mentioned by users, so multiple related domains or tasks need to be performed one by one through a long conversation with many explicit intents.…

Computation and Language · Computer Science 2023-06-07 Hui-Chi Kuo , Yun-Nung Chen

Conversational AI has a fundamental flaw as a knowledge interface: sycophantic chatbots induce epistemic entrenchment and delusional belief spirals even in rational agents. We propose the problem does not stem from the AI model, rooted…

Artificial Intelligence · Computer Science 2026-05-12 Will Beaumaster , Paul Schrater

Explainable Artificial Intelligence (XAI) is critical for ensuring trust and accountability, yet its development remains predominantly visual. For blind and low-vision (BLV) users, the lack of accessible explanations creates a fundamental…

Human-Computer Interaction · Computer Science 2026-05-05 Abu Noman Md Sakib , Protik Dey , Zijie Zhang , Taslima Akter

Accurate visual understanding is imperative for advancing autonomous systems and intelligent robots. Despite the powerful capabilities of vision-language models (VLMs) in processing complex visual scenes, precisely recognizing obscured or…

Computer Vision and Pattern Recognition · Computer Science 2024-06-03 Huaxiang Zhang , Yaojia Mu , Guo-Niu Zhu , Zhongxue Gan

Recent advances in multimodal learning have significantly enhanced the reasoning capabilities of vision-language models (VLMs). However, state-of-the-art approaches rely heavily on large-scale human-annotated datasets, which are costly and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Han Wang , Yi Yang , Jingyuan Hu , Minfeng Zhu , Wei Chen

Conversational agents often encounter ambiguous user requests, requiring an effective clarification to successfully complete tasks. While recent advancements in real-world applications favor multi-agent architectures to manage complex…

Artificial Intelligence · Computer Science 2025-12-16 Emre Can Acikgoz , Jinoh Oh , Joo Hyuk Jeon , Jie Hao , Heng Ji , Dilek Hakkani-Tür , Gokhan Tur , Xiang Li , Chengyuan Ma , Xing Fan

Multimodal large language models (MLLMs) such as GPT-4o, Gemini Pro, and Claude 3.5 have enabled unified reasoning over text and visual inputs, yet they often hallucinate in real world scenarios especially when small objects or fine spatial…

When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facilitate an iterative reasoning and elicitation process, where the human's role is to guide the AI…

Artificial Intelligence · Computer Science 2026-04-10 Guilhem Fouilhé , Rebecca Eifler , Antonin Poché , Sylvie Thiébaux , Nicholas Asher

Vision language models (VLMs) have achieved remarkable success in broad visual understanding, yet they remain challenged by object-centric reasoning on rare objects due to the scarcity of such instances in pretraining data. While prior…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Xin Hu , Haomiao Ni , Yunbei Zhang , Jihun Hamm , Zechen Li , Zhengming Ding
‹ Prev 1 2 3 10 Next ›