中文
相关论文

相关论文: Towards Unified Interactive Visual Grounding in Th…

200 篇论文

This paper describes the development of a real-time Human-Robot Interaction (HRI) system for a service robot based on 3D human activity recognition and human-like decision mechanism. The Human-Robot Interactive (HRI) system, which allows…

人机交互 · 计算机科学 2019-01-14 Kang Li , Jinting Wu , Xiaoguang Zhao , Min Tan

Human-robot collaboration (HRC) has become increasingly relevant in industrial, household, and commercial settings. However, the effectiveness of such collaborations is highly dependent on the human and robots' situational awareness of the…

机器人学 · 计算机科学 2023-05-09 Chelsea Zou , Kishan Chandan , Yan Ding , Shiqi Zhang

Human-robot interaction (HRI) research is progressively addressing multi-party scenarios, where a robot interacts with more than one human user at the same time. Conversely, research is still at an early stage for human-robot collaboration.…

机器学习 · 计算机科学 2023-11-16 Francesco Semeraro , Jon Carberry , Angelo Cangelosi

Dialogue systems can leverage large pre-trained language models and knowledge to generate fluent and informative responses. However, these models are still prone to produce hallucinated responses not supported by the input source, which…

计算与语言 · 计算机科学 2023-05-15 Ziwei Ji , Zihan Liu , Nayeon Lee , Tiezheng Yu , Bryan Wilie , Min Zeng , Pascale Fung

Human-robot teaming (HRT) systems often rely on large-scale datasets of human and robot interactions, especially for close-proximity collaboration tasks such as human-robot handovers. Learning robot manipulation policies from raw,…

机器人学 · 计算机科学 2025-08-14 Yuekun Wu , Yik Lung Pang , Andrea Cavallaro , Changjae Oh

In recent years, the demand for social robots has grown, requiring them to adapt their behaviors based on users' states. Accurately assessing user experience (UX) in human-robot interaction (HRI) is crucial for achieving this adaptability.…

机器人学 · 计算机科学 2025-08-01 Ryo Miyoshi , Yuki Okafuji , Takuya Iwamoto , Junya Nakanishi , Jun Baba

Machine Interpreting systems are currently implemented as unimodal, real-time speech-to-speech architectures, processing translation exclusively on the basis of the linguistic signal. Such reliance on a single modality, however, constrains…

计算与语言 · 计算机科学 2025-09-30 Claudio Fantinuoli

Digital human motion synthesis is a vibrant research field with applications in movies, AR/VR, and video games. Whereas methods were proposed to generate natural and realistic human motions, most only focus on modeling humans and largely…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Quanzhou Li , Jingbo Wang , Chen Change Loy , Bo Dai

Understanding human instructions is essential for enabling smooth human-robot interaction. In this work, we focus on object grounding, i.e., localizing an object of interest in a visual scene (e.g., an image) based on verbal human…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Joel Alberto Santos , Zongwei Wu , Xavier Alameda-Pineda , Radu Timofte

Physical Human-Humanoid Interaction (pHHI) is a rapidly advancing field with significant implications for deploying robots in unstructured, human-centric environments. In this review, we examine the current state of the art in pHHI through…

机器人学 · 计算机科学 2026-05-19 Gustavo A. Cardona , Shubham S. Kumbhar , Panagiotis Artemiadis

Interpreting human intent accurately is a central challenge in human-robot interaction (HRI) and a key requirement for achieving more natural and intuitive collaboration between humans and machines. This work presents a novel multimodal HRI…

机器人学 · 计算机科学 2026-02-25 Guanting Shen , Zi Tian

Referring expressions are commonly used when referring to a specific target in people's daily dialogue. In this paper, we develop a novel task of audio-visual grounding referring expression for robotic manipulation. The robot leverages both…

机器人学 · 计算机科学 2021-09-23 Yefei Wang , Kaili Wang , Yi Wang , Di Guo , Huaping Liu , Fuchun Sun

Trust in human-robot interactions (HRI) is measured in two main ways: through subjective questionnaires and through behavioral tasks. To optimize measurements of trust through questionnaires, the field of HRI faces two challenges: the…

人机交互 · 计算机科学 2021-04-26 Meia Chita-Tegmark , Theresa Law , Nicholas Rabb , Matthias Scheutz

We present Human to Humanoid (H2O), a reinforcement learning (RL) based framework that enables real-time whole-body teleoperation of a full-sized humanoid robot with only an RGB camera. To create a large-scale retargeted motion dataset of…

机器人学 · 计算机科学 2024-03-08 Tairan He , Zhengyi Luo , Wenli Xiao , Chong Zhang , Kris Kitani , Changliu Liu , Guanya Shi

Creating an intelligent conversational system that understands vision and language is one of the ultimate goals in Artificial Intelligence (AI)~\cite{winograd1972understanding}. Extensive research has focused on vision-to-language…

计算与语言 · 计算机科学 2018-05-10 Jiaping Zhang , Tiancheng Zhao , Zhou Yu

Visual grounding, which aims to ground a visual region via natural language, is a task that heavily relies on cross-modal alignment. Existing works utilized uni-modal pre-trained models to transfer visual or linguistic knowledge separately…

计算机视觉与模式识别 · 计算机科学 2024-09-06 Linhui Xiao , Xiaoshan Yang , Fang Peng , Yaowei Wang , Changsheng Xu

There are many examples of cases where access to improved models of human behavior and cognition has allowed creation of robots which can better interact with humans, and not least in road vehicle automation this is a rapidly growing area…

机器人学 · 计算机科学 2022-08-25 Gustav Markkula , Mehmet Dogar

A robot's ability to understand or ground natural language instructions is fundamentally tied to its knowledge about the surrounding world. We present an approach to grounding natural language utterances in the context of factual…

机器人学 · 计算机科学 2018-11-19 Rohan Paul , Andrei Barbu , Sue Felshin , Boris Katz , Nicholas Roy

We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first step toward this goal,…

Situationally Induced Impairments and Disabilities (SIIDs) can significantly hinder user experience in contexts such as poor lighting, noise, and multi-tasking. While prior research has introduced algorithms and systems to address these…

人机交互 · 计算机科学 2025-02-19 Xingyu Bruce Liu , Jiahao Nick Li , David Kim , Xiang 'Anthony' Chen , Ruofei Du