中文
相关论文

相关论文: TAVIS: A Benchmark for Egocentric Active Vision an…

200 篇论文

Human gaze is known to be an intention-revealing signal in human demonstrations of tasks. In this work, we use gaze cues from human demonstrators to enhance the performance of agents trained via three popular imitation learning methods --…

机器学习 · 计算机科学 2021-04-23 Akanksha Saran , Ruohan Zhang , Elaine Schaertl Short , Scott Niekum

Humans often resolve visual uncertainty by comparing an image with relevant examples, but ViTs lack the ability to identify which examples would improve their predictions. We present Task-Aligned Context Selection (TACS), a framework that…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Jingyu Guo , Emir Konuk , Fredrik Strand , Christos Matsoukas , Kevin Smith

Grasping the intricacies of human motion, which involve perceiving spatio-temporal dependence and multi-scale effects, is essential for predicting human motion. While humans inherently possess the requisite skills to navigate this issue, it…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Kedi Lyu , Haipeng Chen , Zhenguang Liu , Yifang Yin , Yukang Lin , Yingying Jiao

In robot-assisted minimally invasive surgery (RMIS), reduced haptic feedback and depth cues increase reliance on expert visual perception, motivating gaze-guided training and learning-based surgical perception models. However, operative…

机器人学 · 计算机科学 2026-05-20 Yizhou Li , Shuyuan Yang , Jiaji Su , Zonghe Chua

Equipping multi-fingered robots with tactile sensing is crucial for achieving the precise, contact-rich, and dexterous manipulation that humans excel at. However, relying solely on tactile sensing fails to provide adequate cues for…

机器人学 · 计算机科学 2023-09-22 Irmak Guzey , Yinlong Dai , Ben Evans , Soumith Chintala , Lerrel Pinto

Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we identify a critical Perception Gap in distillation: student models frequently…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Linquan Wu , Tianxiang Jiang , Yifei Dong , Haoyu Yang , Fengji Zhang , Shichaang Meng , Ai Xuan , Linqi Song , Jacky Keung

Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing. In contrast, robot learning systems typically rely on passive, uniform…

机器人学 · 计算机科学 2025-09-23 Ian Chuang , Jinyu Zou , Andrew Lee , Dechen Gao , Iman Soltani

Masked video modeling~(MVM) has emerged as a highly effective pre-training strategy for visual foundation models, whereby the model reconstructs masked spatiotemporal tokens using information from visible tokens. However, a key challenge in…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Ayush K. Rai , Kyle Min , Tarun Krishna , Feiyan Hu , Alan F. Smeaton , Noel E. O'Connor

Human-object interaction is one of the most important visual cues and we propose a novel way to represent human-object interactions for egocentric action anticipation. We propose a novel transformer variant to model interactions by…

计算机视觉与模式识别 · 计算机科学 2024-01-12 Debaditya Roy , Ramanathan Rajendiran , Basura Fernando

Understanding human actions from videos plays a critical role across various domains, including sports analytics. In figure skating, accurately recognizing the type and timing of jumps a skater performs is essential for objective…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Ryota Tanaka , Tomohiro Suzuki , Keisuke Fujii

Previous studies have illustrated the potential of analysing gaze behaviours in collaborative learning to provide educationally meaningful information for students to reflect on their learning. Over the past decades, machine learning…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Junyuan Liang , Qi Zhou , Sahan Bulathwela , Mutlu Cukurova

Egocentric gaze anticipation serves as a key building block for the emerging capability of Augmented Reality. Notably, gaze behavior is driven by both visual cues and audio signals during daily activities. Motivated by this observation, we…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Bolin Lai , Fiona Ryan , Wenqi Jia , Miao Liu , James M. Rehg

Vision-Language Navigation requires the agent to follow natural language instructions to reach a specific target. The large discrepancy between seen and unseen environments makes it challenging for the agent to generalize well. Previous…

计算机视觉与模式识别 · 计算机科学 2022-09-13 Yujie Lu , Huiliang Zhang , Ping Nie , Weixi Feng , Wenda Xu , Xin Eric Wang , William Yang Wang

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans actively coordinate…

机器人学 · 计算机科学 2026-03-11 Justin Yu , Yide Shentu , Di Wu , Pieter Abbeel , Ken Goldberg , Philipp Wu

We present a simple and intuitive approach for interactive control of physically simulated characters. Our work builds upon generative adversarial networks (GAN) and reinforcement learning, and introduces an imitation learning framework…

图形学 · 计算机科学 2022-01-03 Pei Xu , Ioannis Karamouzas

Gaze estimation has become a subject of growing interest in recent research. Most of the current methods rely on single-view facial images as input. Yet, it is hard for these approaches to handle large head angles, leading to potential…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Ruicong Liu , Feng Lu

When intelligent agents learn visuomotor behaviors from human demonstrations, they may benefit from knowing where the human is allocating visual attention, which can be inferred from their gaze. A wealth of information regarding intelligent…

计算机视觉与模式识别 · 计算机科学 2018-06-12 Ruohan Zhang , Zhuode Liu , Luxin Zhang , Jake A. Whritner , Karl S. Muller , Mary M. Hayhoe , Dana H. Ballard

Imitation learning for acquiring generalizable policies often requires a large volume of demonstration data, making the process significantly costly. One promising strategy to address this challenge is to leverage the cognitive and…

机器人学 · 计算机科学 2025-06-09 Yutaro Ishida , Takamitsu Matsubara , Takayuki Kanai , Kazuhiro Shintani , Hiroshi Bito

Gaze is a crucial social cue in any interacting scenario and drives many mechanisms of social cognition (joint and shared attention, predicting human intention, coordination tasks). Gaze direction is an indication of social and emotional…

机器人学 · 计算机科学 2024-10-28 Maria Lombardi , Elisa Maiettini , Agnieszka Wykowska , Lorenzo Natale

Human gaze provides valuable information on human focus and intentions, making it a crucial area of research. Recently, deep learning has revolutionized appearance-based gaze estimation. However, due to the unique features of gaze…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Yihua Cheng , Haofei Wang , Yiwei Bao , Feng Lu
‹ 上一页 1 2 3 10 下一页 ›