中文
相关论文

相关论文: Multimodal Datasets and Benchmarks for Reasoning a…

200 篇论文

The EmbodiedQA is a task of training an embodied agent by intelligently navigating in a simulated environment and gathering visual information to answer questions. Existing approaches fail to explicitly model the mental imagery function of…

计算机视觉与模式识别 · 计算机科学 2019-08-06 Juncheng Li , Siliang Tang , Fei Wu , Yueting Zhuang

Situated embodied conversation requires robots to interleave real-time dialogue with active perception: deciding what to look at, when to look, and what to say under tight latency constraints. We present a simple, minimal system recipe that…

机器人学 · 计算机科学 2026-02-05 Dong Won Lee , Sarah Gillet , Louis-Philippe Morency , Cynthia Breazeal , Hae Won Park

Multi-modal AI systems will likely become a ubiquitous presence in our everyday lives. A promising approach to making these systems more interactive is to embody them as agents within physical and virtual environments. At present, systems…

Embodiment is an important characteristic for all intelligent agents (creatures and robots), while existing scene description tasks mainly focus on analyzing images passively and the semantic understanding of the scenario is separated from…

机器人学 · 计算机科学 2020-05-08 Sinan Tan , Huaping Liu , Di Guo , Xinyu Zhang , Fuchun Sun

We present Habitat, a platform for research in embodied artificial intelligence (AI). Habitat enables training embodied agents (virtual robots) in highly efficient photorealistic 3D simulation. Specifically, Habitat consists of: (i)…

Human intelligence can remarkably adapt quickly to new tasks and environments. Starting from a very young age, humans acquire new skills and learn how to solve new tasks either by imitating the behavior of others or by following provided…

Most existing robotic datasets capture static scene data and thus are limited in evaluating robots' dynamic performance. To address this, we present a mobile robot oriented large-scale indoor dataset, denoted as THUD (Tsinghua University…

机器人学 · 计算机科学 2024-07-02 Yifan Tang , Cong Tai , Fangxing Chen , Wanting Zhang , Tao Zhang , Xueping Liu , Yongjin Liu , Long Zeng

We report on our effort to create a corpus dataset of different social context situations in an office setting for further disciplinary and interdisciplinary research in computer vision, psychology, and human-robot-interaction. For social…

机器人学 · 计算机科学 2023-11-14 Stefan Schiffer , Astrid Rosenthal-von der Pütten , Bastian Leibe

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

We propose a new task to benchmark scene understanding of embodied agents: Situated Question Answering in 3D Scenes (SQA3D). Given a scene context (e.g., 3D scan), SQA3D requires the tested agent to first understand its situation (position,…

计算机视觉与模式识别 · 计算机科学 2023-04-14 Xiaojian Ma , Silong Yong , Zilong Zheng , Qing Li , Yitao Liang , Song-Chun Zhu , Siyuan Huang

Automatic Emotion Detection (ED) aims to build systems to identify users' emotions automatically. This field has the potential to enhance HCI, creating an individualised experience for the user. However, ED systems tend to perform poorly on…

人机交互 · 计算机科学 2023-07-27 Annanda Sousa , Karen Young , Mathieu D'aquin , Manel Zarrouk , Jennifer Holloway

Unmanned surface vehicles can encounter a number of varied visual circumstances during operation, some of which can be very difficult to interpret. While most cases can be solved only using color camera images, some weather and lighting…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Jon Muhovič , Janez Perš

Embodied foundation models are increasingly performant in real-world domains such as robotics or autonomous driving. These models are often deployed in interactive or assistive settings, where it is important that these assistive models…

机器人学 · 计算机科学 2026-03-06 Pradyumna Tambwekar , Andrew Silva , Deepak Gopinath , Jonathan DeCastro , Xiongyi Cui , Guy Rosman

The impressive advances and applications of large language and joint language-and-visual understanding models has led to an increased need for methods of probing their potential reasoning capabilities. However, the difficulty of gather…

机器学习 · 计算机科学 2023-06-05 Nathan Vaska , Victoria Helus

Embodied AI is a prominent research topic in both academia and industry. Current research centers on completing tasks based on explicit user instructions. However, for robots to integrate into human society, they must understand which…

The development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step…

机器人学 · 计算机科学 2024-04-16 Roberto Bigazzi , Marcella Cornia , Silvia Cascianelli , Lorenzo Baraldi , Rita Cucchiara

Understanding road scenes is essential for autonomous driving, as it enables systems to interpret visual surroundings to aid in effective decision-making. We present Roadscapes, a multitask multimodal dataset consisting of upto 9,000 images…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Vijayasri Iyer , Maahin Rathinagiriswaran , Jyothikamalesh S

With the surge in the development of large language models, embodied intelligence has attracted increasing attention. Nevertheless, prior works on embodied intelligence typically encode scene or historical memory in an unimodal manner,…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Yang Liu , Xinshuai Song , Kaixuan Jiang , Weixing Chen , Jingzhou Luo , Guanbin Li , Liang Lin

This paper presents a computational model of the processing of dynamic spatial relations occurring in an embodied robotic interaction setup. A complete system is introduced that allows autonomous robots to produce and interpret dynamic…

计算与语言 · 计算机科学 2016-07-27 Michael Spranger , Jakob Suchan , Mehul Bhatt , Manfred Eppe

Multimodal learning, a rapidly evolving field in artificial intelligence, seeks to construct more versatile and robust systems by integrating and analyzing diverse types of data, including text, images, audio, and video. Inspired by the…