English
Related papers

Related papers: ProactiveVideoQA: A Comprehensive Benchmark Evalua…

200 papers

Proactive streaming video understanding requires models to continuously process video streams and decide when to respond, rather than merely what to respond. This naturally introduces a decision-making problem under partial observations,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Ao Li , Zihan Xiao , Zihao Yue , Boshen Xu , Linli Yao , Jiaze Li , Pei Fu , Jianzhong Ju , Jian Luan , Qin Jin

Dialogue models are inherently reactive, responding to the current user turn without anticipating upcoming intents, which leads to redundant interactions in multi-intent settings. We address this limitation by introducing a lightweight…

Computation and Language · Computer Science 2026-05-01 Yang Luo

Conversational systems based on Large Language Models (LLMs), such as ChatGPT, show exceptional proficiency in context understanding and response generation. However, despite their impressive capabilities, they still possess limitations,…

Computation and Language · Computer Science 2023-10-17 Yang Deng , Lizi Liao , Liang Chen , Hongru Wang , Wenqiang Lei , Tat-Seng Chua

Conversational systems have made significant progress in generating natural language responses. However, their potential as conversational search systems is currently limited due to their passive role in the information-seeking process. One…

Computation and Language · Computer Science 2024-02-27 Pierre Erbacher , Jian-Yun Nie , Philippe Preux , Laure Soulier

Predicting and planning interactive behaviors in complex traffic situations presents a challenging task. Especially in scenarios involving multiple traffic participants that interact densely, autonomous vehicles still struggle to interpret…

Multiagent Systems · Computer Science 2021-02-12 Julian Bernhard , Klemens Esterle , Patrick Hart , Tobias Kessler

Envision an AI capable of functioning in human-like settings, moving beyond mere observation to actively understand, anticipate, and proactively respond to unfolding events. Towards this vision, we focus on the innovative task where, given…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yulin Zhang , Cheng Shi , Yang Wang , Sibei Yang

User engagement is a critical metric for evaluating the quality of open-domain dialogue systems. Prior work has focused on conversation-level engagement by using heuristically constructed features such as the number of turns and the total…

Computation and Language · Computer Science 2020-01-27 Sarik Ghazarian , Ralph Weischedel , Aram Galstyan , Nanyun Peng

On the way towards general Visual Question Answering (VQA) systems that are able to answer arbitrary questions, the need arises for evaluation beyond single-metric leaderboards for specific datasets. To this end, we propose a browser-based…

Computer Vision and Pattern Recognition · Computer Science 2021-10-12 Dirk Väth , Pascal Tilli , Ngoc Thang Vu

Perceiving multi-modal information and fulfilling dialogues with humans is a long-term goal of artificial intelligence. Pre-training is commonly regarded as an effective approach for multi-modal dialogue. However, due to the limited…

Computation and Language · Computer Science 2023-06-14 Yunshui Li , Binyuan Hui , ZhiChao Yin , Min Yang , Fei Huang , Yongbin Li

An intelligent dialogue system in a multi-turn setting should not only generate the responses which are of good quality, but it should also generate the responses which can lead to long-term success of the dialogue. Although, the current…

Computation and Language · Computer Science 2023-01-12 Anant Khandelwal

Physical AI aims to develop models that can perceive and predict real-world dynamics; yet, the extent to which current multi-modal large language models and video generative models support these abilities is insufficiently understood. We…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Fengzhe Zhou , Jiannan Huang , Jialuo Li , Deva Ramanan , Humphrey Shi

In this work we propose a blackbox intervention method for visual dialog models, with the aim of assessing the contribution of individual linguistic or visual components. Concretely, we conduct structured or randomized interventions that…

Computer Vision and Pattern Recognition · Computer Science 2017-12-06 Mircea Mironenco , Dana Kianfar , Ke Tran , Evangelos Kanoulas , Efstratios Gavves

Full-duplex spoken dialogue systems promise to transform human-machine interaction from a rigid, turn-based protocol into a fluid, natural conversation. However, the central challenge to realizing this vision, managing overlapping speech,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-28 Guan-Ting Lin , Shih-Yun Shan Kuan , Qirui Wang , Jiachen Lian , Tingle Li , Shinji Watanabe , Hung-yi Lee

With the deep integration of artificial intelligence and interactive technology, Graphical User Interface (GUI) Agent, as the carrier connecting goal-oriented natural language and real-world devices, has received widespread attention from…

Artificial Intelligence · Computer Science 2025-11-13 Leyang Yang , Ziwei Wang , Xiaoxuan Tang , Sheng Zhou , Dajun Chen , Wei Jiang , Yong Li

Interactive video generation models such as Genie, YUME, HY-World, and Matrix-Game are advancing rapidly, yet every model is evaluated on its own benchmark with private scenes and trajectories, making fair cross-model comparison impossible.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Xiaojie Xu , Zhengyuan Lin , Kang He , Yukang Feng , Xiaofeng Mao , Yuanyang Yin , Kaipeng Zhang , Yongtao Ge

Real-time duplex interaction is essential for multimodal AI systems operating in real-world scenarios, where models must continuously process streaming inputs and respond at appropriate moments. However, most existing multimodal large…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Chaoqun He , Mingyang Xiang , Yingjing Xu , Bokai Xu , Junbo Cui , Jie Zhou , Yuan Yao , Lijie Wen

Assessing drivers' interaction capabilities is crucial for understanding human driving behavior and enhancing the interactive abilities of autonomous vehicles. In scenarios involving strong interaction, existing metrics focused on…

Robotics · Computer Science 2024-05-07 Jiaqi Liu , Peng Hang , Xiangwang Hu , Jian Sun

Visual events are a composition of temporal actions involving actors spatially interacting with objects. When developing computer vision models that can reason about compositional spatio-temporal events, we need benchmarks that can analyze…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Madeleine Grunde-McLaughlin , Ranjay Krishna , Maneesh Agrawala

Despite decades of work, surveillance still struggles to find specific targets across long, multi-camera video. Prior methods -- tracking pipelines, CLIP based models, and VideoRAG -- require heavy manual filtering, capture only shallow…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Hyojin Park , Yi Li , Janghoon Cho , Sungha Choi , Jungsoo Lee , Taotao Jing , Shuai Zhang , Munawar Hayat , Dashan Gao , Ning Bi , Fatih Porikli

The next step for In-vehicle Conversational Assistants (IVCAs) will be their capability to initiate and automate proactive system interactions throughout journeys. However, diverse drivers make it challenging to design voice interventions…

Human-Computer Interaction · Computer Science 2026-01-28 Josh Susak , Yifu Liu , Pascal Jansen , Mark Colley
‹ Prev 1 3 4 5 6 7 10 Next ›