中文
相关论文

相关论文: CAVEN: An Embodied Conversational Agent for Effici…

200 篇论文

Recent dense audio-visual (AV) models achieve impressive retrieval and emergent localization, but almost all evidence comes from English-centric, caption-rich web video. It is unclear whether these objectives survive in low-resource,…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Sajay Raj

Vision-and-Language Navigation (VLN) task aims to enable AI agents to accurately understand and follow natural language instructions to navigate through real-world environments, ultimately reaching specific target locations. We recognise a…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Qi Chen , Dileepa Pitawela , Chongyang Zhao , Gengze Zhou , Hsiang-Ting Chen , Qi Wu

In visual semantic navigation, the robot navigates to a target object with egocentric visual observations and the class label of the target is given. It is a meaningful task inspiring a surge of relevant research. However, most of the…

人工智能 · 计算机科学 2021-09-21 Xinzhu Liu , Di Guo , Huaping Liu , Fuchun Sun

Due to the complexity of the natural world, a programmer cannot foresee all possible situations, a connected and autonomous vehicle (CAV) will face during its operation, and hence, CAVs will need to learn to make decisions autonomously. Due…

多智能体系统 · 计算机科学 2018-08-24 Varuna De Silva , Xiongzhao Wang , Deniz Aladagli , Ahmet Kondoz , Erhan Ekmekcioglu

Building an end-to-end conversational agent for multi-domain task-oriented dialogues has been an open challenge for two main reasons. First, tracking dialogue states of multiple domains is non-trivial as the dialogue agent must obtain…

计算与语言 · 计算机科学 2020-11-17 Hung Le , Doyen Sahoo , Chenghao Liu , Nancy F. Chen , Steven C. H. Hoi

We propose MAViD, a novel Multimodal framework for Audio-Visual Dialogue understanding and generation. Existing approaches primarily focus on non-interactive systems and are limited to producing constrained and unnatural human speech. The…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Youxin Pang , Jiajun Liu , Lingfeng Tan , Yong Zhang , Feng Gao , Xiang Deng , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

Vision language navigation is the task that requires an agent to navigate through a 3D environment based on natural language instructions. One key challenge in this task is to ground instructions with the current visual information that the…

计算与语言 · 计算机科学 2021-04-21 Jialu Li , Hao Tan , Mohit Bansal

We propose a novel formulation of the "effectiveness problem" in communications, put forth by Shannon and Weaver in their seminal work [2], by considering multiple agents communicating over a noisy channel in order to achieve better…

信号处理 · 电气工程与系统科学 2021-04-02 Tze-Yang Tung , Szymon Kobus , Joan Roig Pujol , Deniz Gunduz

Understanding and following natural language instructions while navigating through complex, real-world environments poses a significant challenge for general-purpose robots. These environments often include obstacles and pedestrians, making…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Xiwen Liang , Liang Ma , Shanshan Guo , Jianhua Han , Hang Xu , Shikui Ma , Xiaodan Liang

Safe, agile, and socially compliant multi-robot navigation in cluttered and constrained environments remains a critical challenge. This is especially difficult with self-interested agents with unique, unknown priorities in decentralized…

机器人学 · 计算机科学 2026-05-12 Vagul Mahadevan , Shangtong Zhang , Rohan Chandra

Current vision and language tasks usually take complete visual data (e.g., raw images or videos) as input, however, practical scenarios may often consist the situations where part of the visual information becomes inaccessible due to…

计算机视觉与模式识别 · 计算机科学 2021-06-29 Ye Zhu , Yu Wu , Yi Yang , Yan Yan

The problem of building a coherent and non-monotonous conversational agent with proper discourse and coverage is still an area of open research. Current architectures only take care of semantic and contextual information for a given query…

计算与语言 · 计算机科学 2025-04-22 Gaurav Kumar , Rishabh Joshi , Jaspreet Singh , Promod Yenigalla

A core challenge in AI-guided autonomy is enabling agents to navigate realistically and effectively in previously unseen environments based on natural language commands. We propose UAV-VLN, a novel end-to-end Vision-Language Navigation…

机器人学 · 计算机科学 2025-10-01 Pranav Saxena , Nishant Raghuvanshi , Neena Goveas

Language-guided Embodied AI benchmarks requiring an agent to navigate an environment and manipulate objects typically allow one-way communication: the human user gives a natural language command to the agent, and the agent can only follow…

人工智能 · 计算机科学 2022-08-17 Xiaofeng Gao , Qiaozi Gao , Ran Gong , Kaixiang Lin , Govind Thattai , Gaurav S. Sukhatme

We present Vision-based Navigation with Language-based Assistance (VNLA), a grounded vision-language task where an agent with visual perception is guided via language to find objects in photorealistic indoor environments. The task emulates…

机器学习 · 计算机科学 2019-04-09 Khanh Nguyen , Debadeepta Dey , Chris Brockett , Bill Dolan

Object Goal Navigation (ObjectNav) refers to an agent navigating to an object in an unseen environment, which is an ability often required in the accomplishment of complex tasks. While existing methods demonstrate proficiency in isolated…

机器人学 · 计算机科学 2026-04-15 Jiahua Pei , Yi Liu , Guoping Pan , Yuanhao Jiang , Houde Liu , Xueqian Wang

There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represent knowledge using…

音频与语音处理 · 电气工程与系统科学 2024-02-22 Xianghu Yue , Xiaohai Tian , Lu Lu , Malu Zhang , Zhizheng Wu , Haizhou Li

Visual content and accompanied audio signals naturally formulate a joint representation to improve audio-visual (AV) related applications. While studies develop various AV representation learning frameworks, the importance of AV data…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Shentong Mo , Yibing Song

Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an {\em interactive object-aware audio generation} model that grounds…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Tingle Li , Baihe Huang , Xiaobin Zhuang , Dongya Jia , Jiawei Chen , Yuping Wang , Zhuo Chen , Gopala Anumanchipalli , Yuxuan Wang

Aerial Vision-and-Language Navigation (VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and navigate complex urban environments using onboard visual observation. This task holds promise for…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Huilin Xu , Zhuoyang Liu , Yixiang Luomei , Feng Xu