中文
相关论文

相关论文: ProactiveVideoQA: A Comprehensive Benchmark Evalua…

200 篇论文

Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smiles, and gestures. To support successful human-agent…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Amrita Mazumdar , Seonwook Park , Rajarshi Roy , Nikhil Srihari , Shengze Wang , Yuhao Zhou , Julia Wang , Koki Nagano , Shalini De Mello

LLaVA-Interactive is a research prototype for multimodal human-AI interaction. The system can have multi-turn dialogues with human users by taking multimodal user inputs and generating multimodal responses. Importantly, LLaVA-Interactive…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Wei-Ge Chen , Irina Spiridonova , Jianwei Yang , Jianfeng Gao , Chunyuan Li

Video Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and sophisticated cross-modal interaction for accurate inference.…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Ting Yu , Kunhao Fu , Shuhui Wang , Qingming Huang , Jun Yu

There have been significant innovations in media technologies in the recent years. While these developments have improved experiences for individual users, design of multi-user interfaces still remains a challenge. A relatively unexplored…

人机交互 · 计算机科学 2018-11-20 Sumit Shekhar , Aditya Siddhant , Anindya Shankar Bhandari , Nishant Yadav

Conversational agents often encounter ambiguous user requests, requiring an effective clarification to successfully complete tasks. While recent advancements in real-world applications favor multi-agent architectures to manage complex…

人工智能 · 计算机科学 2025-12-16 Emre Can Acikgoz , Jinoh Oh , Joo Hyuk Jeon , Jie Hao , Heng Ji , Dilek Hakkani-Tür , Gokhan Tur , Xiang Li , Chengyuan Ma , Xing Fan

Real-time conversational assistants for procedural tasks often depend on video input, which can be computationally expensive and compromise user privacy. For the first time, we propose a real-time conversational assistant that provides…

多媒体 · 计算机科学 2026-02-18 Rehana Mahfuz , Yinyi Guo , Erik Visser , Phanidhar Chinchili

One of the long-standing aspirations in conversational AI is to allow them to autonomously take initiatives in conversations, i.e., being proactive. This is especially challenging for multi-party conversations. Prior NLP research focused…

人机交互 · 计算机科学 2025-02-19 Xingyu Bruce Liu , Shitao Fang , Weiyan Shi , Chien-Sheng Wu , Takeo Igarashi , Xiang Anthony Chen

The field of eXplainable Artificial Intelligence (XAI) is increasingly recognizing the need to personalize and/or interactively adapt the explanation to better reflect users' explanation needs. While dialogue-based approaches to XAI have…

机器学习 · 计算机科学 2024-08-15 Dimitry Mindlin , Amelie Sophie Robrecht , Michael Morasch , Philipp Cimiano

Interactive communication (IC), i.e., the reciprocal exchange of information between two or more interactive partners, is a fundamental part of human nature. As such, it has been studied across multiple scientific disciplines with different…

Empowering large language models with long-term memory is crucial for building agents that adapt to users' evolving needs. Existing evaluations of this capability typically interleave preference-related dialogues with irrelevant…

In recent years, large language models have shown exceptional performance in fulfilling diverse human needs. However, their training data can introduce harmful content, underscoring the necessity for robust value alignment. Mainstream…

人工智能 · 计算机科学 2024-12-19 Rui Zou , Mengqi Wei , Jintian Feng , Qian Wan , Jianwen Sun , Sannyuya Liu

Individuals often align their speaking patterns with their interlocutors, a phenomenon linked to engagement and rapport. While well documented in task-oriented dialogues, less is known about entrainment in naturalistic, non-task and virtual…

人机交互 · 计算机科学 2026-04-20 Thanushi Withanage , Elizabeth Redcay , Carol Espy-Wilson

Temporal logical understanding, a core facet of human cognition, plays a pivotal role in capturing complex sequential events and their temporal relationships within videos. This capability is particularly crucial in tasks like Video…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Sirnam Swetha , Hilde Kuehne , Mubarak Shah

Querying generative AI models, e.g., large language models (LLMs), has become a prevalent method for information acquisition. However, existing query-answer datasets primarily focus on textual responses, making it challenging to address…

人工智能 · 计算机科学 2025-06-03 Shuting Wang , Yunqi Liu , Zixin Yang , Ning Hu , Zhicheng Dou , Chenyan Xiong

Chat-oriented dialogue systems designed to provide tangible benefits, such as sharing the latest news or preventing frailty in senior citizens, often require Proactive acquisition of specific user Information via chats on user-faVOred…

计算与语言 · 计算机科学 2025-04-11 Shiki Sato , Jun Baba , Asahi Hentona , Shinji Iwata , Akifumi Yoshimoto , Koichiro Yoshino

User Satisfaction Estimation (USE) is an important yet challenging task in goal-oriented conversational systems. Whether the user is satisfied with the system largely depends on the fulfillment of the user's needs, which can be implicitly…

计算与语言 · 计算机科学 2022-02-08 Yang Deng , Wenxuan Zhang , Wai Lam , Hong Cheng , Helen Meng

Establishing evaluation schemes for spoken dialogue systems is important, but it can also be challenging. While subjective evaluations are commonly used in user experiments, objective evaluations are necessary for research comparison and…

计算与语言 · 计算机科学 2024-01-24 Koji Inoue , Divesh Lala , Keiko Ochi , Tatsuya Kawahara , Gabriel Skantze

Recent research has shown that multi-task pre-training greatly improves the model's robustness and transfer ability, which is crucial for building a high-quality dialog system. However, most previous works on multi-task pre-training rely…

计算与语言 · 计算机科学 2023-09-21 Yucheng Cai , Wentao Ma , Yuchuan Wu , Shuzheng Si , Yuan Shao , Zhijian Ou , Yongbin Li

Building a reliable and automated evaluation metric is a necessary but challenging problem for open-domain dialogue systems. Recent studies proposed evaluation metrics that assess generated responses by considering their relevance to…

计算与语言 · 计算机科学 2024-07-19 ChaeHun Park , Minseok Choi , Dohyun Lee , Jaegul Choo

LLM-based agents can complete tasks correctly yet still frustrate users through poor interaction patterns, such as excessive confirmations, opaque reasoning, or misaligned pacing. Current benchmarks evaluate task accuracy but overlook how…

人机交互 · 计算机科学 2026-02-09 Jialin Li , Zhenhao Chen , Hanjun Luo , Hanan Salam