中文
相关论文

相关论文: Understanding Embodied Reference with Touch-Line T…

200 篇论文

Embodied reference understanding is crucial for intelligent agents to predict referents based on human intention through gesture signals and language descriptions. This paper introduces the Attention-Dynamic DINO, a novel framework designed…

计算机视觉与模式识别 · 计算机科学 2024-11-14 Hao Guo , Wei Fan , Baichun Wei , Jianfei Zhu , Jin Tian , Chunzhi Yi , Feng Jiang

We study the understanding of embodied reference: One agent uses both language and gesture to refer to an object to another agent in a shared physical environment. Of note, this new visual task requires understanding multimodal cues with…

计算机视觉与模式识别 · 计算机科学 2021-09-16 Yixin Chen , Qing Li , Deqian Kong , Yik Lun Kei , Song-Chun Zhu , Tao Gao , Yixin Zhu , Siyuan Huang

One of the main goals of robotics and intelligent agent research is to enable natural communication with humans in physically situated settings. While recent work has focused on verbal modes such as language and speech, non-verbal…

机器人学 · 计算机科学 2025-09-17 Anna Deichler , Siyang Wang , Simon Alexanderson , Jonas Beskow

In face-to-face interaction, we use multiple modalities, including speech and gestures, to communicate information and resolve references to objects. However, how representational co-speech gestures refer to objects remains understudied…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Esam Ghaleb , Bulat Khaertdinov , Aslı Özyürek , Raquel Fernández

Embodied Reference Understanding studies the reference understanding in an embodied fashion, where a receiver is required to locate a target object referred to by both language and gesture of the sender in a shared physical environment. Its…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Cheng Shi , Sibei Yang

Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While prior works have shown progress in open-vocabulary object detection, they often fail in…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Fevziye Irem Eyiokur , Dogucan Yaman , Hazım Kemal Ekenel , Alexander Waibel

Generalization in embodied AI is hindered by the "seeing-to-doing gap," which stems from data scarcity and embodiment heterogeneity. To address this, we pioneer "pointing" as a unified, embodiment-agnostic intermediate representation,…

机器人学 · 计算机科学 2026-04-07 Yifu Yuan , Haiqin Cui , Yaoting Huang , Yibin Chen , Fei Ni , Zibin Dong , Pengyi Li , Yan Zheng , Hongyao Tang , Jianye Hao

Hand gesture is one of the most important means of touchless communication between human and machines. There is a great interest for commanding electronic equipment in surgery rooms by hand gesture for reducing the time of surgery and the…

计算机视觉与模式识别 · 计算机科学 2017-10-24 Ebrahim Nasr-Esfahani , Nader Karimi , S. M. Reza Soroushmehr , M. Hossein Jafari , M. Amin Khorsandi , Shadrokh Samavi , Kayvan Najarian

In measurement, a reference frame is needed to compare the measured object to something already known. This raises the neuroscientific question of which reference frame is used by humans when exploring the environment. Previous studies…

神经元与认知 · 定量生物学 2024-01-24 François Le Jeune , Marco D'Alonzo , Valeria Piombino , Alessia Noccaro , Domenico Formica , Giovanni Di Pino

Embodied AI models often employ off the shelf vision backbones like CLIP to encode their visual observations. Although such general purpose representations encode rich syntactic and semantic information about the scene, much of this…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Ainaz Eftekhar , Kuo-Hao Zeng , Jiafei Duan , Ali Farhadi , Ani Kembhavi , Ranjay Krishna

3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pure language-based 3D…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Atharv Mahesh Mane , Dulanga Weerakoon , Vigneshwaran Subbaraju , Sougata Sen , Sanjay E. Sarma , Archan Misra

As robots enter human workspaces, there is a crucial need for them to comprehend embodied human instructions, enabling intuitive and fluent human-robot interaction (HRI). However, accurate comprehension is challenging due to a lack of…

机器人学 · 计算机科学 2025-12-09 Md Mofijul Islam , Alexi Gladstone , Sujan Sarker , Ganesh Nanduru , Md Fahim , Keyan Du , Aman Chadha , Tariq Iqbal

Referring expression grounding aims at locating certain objects or persons in an image with a referring expression, where the key challenge is to comprehend and align various types of information from visual and textual domain, such as…

计算机视觉与模式识别 · 计算机科学 2019-04-03 Xihui Liu , Zihao Wang , Jing Shao , Xiaogang Wang , Hongsheng Li

Imagine sitting at your desk, looking at objects on it. You do not know their exact distances from your eye in meters, but you can immediately reach out and touch them. Instead of an externally defined unit, your sense of distance is tied…

机器人学 · 计算机科学 2025-09-16 Levi Burner , Cornelia Fermüller , Yiannis Aloimonos

An important goal of computer vision is to build systems that learn visual representations over time that can be applied to many tasks. In this paper, we investigate a vision-language embedding as a core representation and show that it…

计算机视觉与模式识别 · 计算机科学 2017-10-17 Tanmay Gupta , Kevin Shih , Saurabh Singh , Derek Hoiem

In recent years, as robotics has advanced, human-robot collaboration has gained increasing importance. However, current robots struggle to fully and accurately interpret human intentions from voice commands alone. Traditional gripper and…

机器人学 · 计算机科学 2024-12-17 Junliang Li , Kai Ye , Haolan Kang , Mingxuan Liang , Yuhang Wu , Zhenhua Liu , Huiping Zhuang , Rui Huang , Yongquan Chen

Whenever we are addressing a specific object or refer to a certain spatial location, we are using referential or deictic gestures usually accompanied by some verbal description. Especially pointing gestures are necessary to dissolve…

计算机视觉与模式识别 · 计算机科学 2019-12-16 Doreen Jirak , David Biertimpel , Matthias Kerzel , Stefan Wermter

Observing touch on another's body can elicit corresponding tactile sensations in the observer, a phenomenon termed mirror touch that supports empathy and social perception. This visuo-tactile resonance is thought to rely on structural…

机器人学 · 计算机科学 2026-05-15 Tianfang Zhu , Ning An , Rui Wang , Jiasi Gao , Qingming Luo , Anan Li , Guyue Zhou

Palpation, the use of touch in medical examination, is almost exclusively performed by humans. We investigate a proof of concept for an artificial palpation method based on self-supervised learning. Our key idea is that an encoder-decoder…

机器学习 · 计算机科学 2025-11-21 Zohar Rimon , Elisei Shafer , Tal Tepper , Efrat Shimron , Aviv Tamar

In human interaction, gestures serve various functions such as marking speech rhythm, highlighting key elements, and supplementing information. These gestures are also observed in explanatory contexts. However, the impact of gestures on…

人机交互 · 计算机科学 2024-08-15 Amelie Sophie Robrecht , Hendric Voss , Lisa Gottschalk , Stefan Kopp
‹ 上一页 1 2 3 10 下一页 ›