English
Related papers

Related papers: The Midas Touch in Gaze vs. Hand Pointing: Modalit…

200 papers

Gaze estimation for ordinary smart phone, e.g. estimating where the user is looking at on the phone screen, can be applied in various applications. However, the widely used appearance-based CNN methods still have two issues for practical…

Computer Vision and Pattern Recognition · Computer Science 2019-10-17 Tianchu Guo , Yongchao Liu , Hui Zhang , Xiabing Liu , Youngjun Kwak , Byung In Yoo , Jae-Joon Han , Changkyu Choi

Even modern AI models often remain vulnerable to multimodal queries in which harmful intent is embedded in images. A widely used approach for safety alignment is training with extensive multimodal safety datasets, but the costs of data…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Jonghyun Park , Minhyuk Seo , Chaewon Yeo , Jonghyun Choi

Recent research proposed gaze-assisted gestures to enhance interaction within virtual reality (VR), providing opportunities for people with motor impairments to experience VR. Compared to people with other motor impairments, those with…

Human-Computer Interaction · Computer Science 2024-03-26 Jingze Tian , Yingna Wang , Keye Yu , Liyi Xu , Junan Xie , Franklin Mingzhe Li , Yafeng Niu , Mingming Fan

As immersive technologies enable unique, multimodal interaction methods, developers must also use tailored methods to support user accessibility, distinct from traditional software practices. We interviewed 25 industry extended reality (XR)…

Human-Computer Interaction · Computer Science 2025-08-11 Daniel Killough , Tiger F. Ji , Kexin Zhang , Yaxin Hu , Yu Huang , Ruofei Du , Yuhang Zhao

Human sensing, which employs various sensors and advanced deep learning technologies to accurately capture and interpret human body information, has significantly impacted fields like public security and robotics. However, current human…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Xinyan Chen , Jianfei Yang

Humans can achieve diverse in-hand manipulations, such as object pinching and tool use, which often involve simultaneous contact between the object and multiple fingers. This is still an open issue for robotic hands because such dexterous…

Robotics · Computer Science 2026-02-06 Pranav Ponnivalavan , Satoshi Funabashi , Alexander Schmitz , Tetsuya Ogata , Shigeki Sugano

Effectively integrating diverse sensory modalities is crucial for robotic manipulation. However, the typical approach of feature concatenation is often suboptimal: dominant modalities such as vision can overwhelm sparse but critical signals…

Vision-only grasping systems are fundamentally constrained by calibration errors, sensor noise, and grasp pose prediction inaccuracies, leading to unavoidable contact uncertainty in the final stage of grasping. High-bandwidth tactile…

Robotics · Computer Science 2025-09-22 Yonghyeon Lee , Tzu-Yuan Lin , Alexander Alexiev , Sangbae Kim

Gaze is a valuable means of communication for impaired people with extremely limited motor capabilities. However, robust gaze-based intent recognition in multi-object environments is challenging due to gaze noise, micro-saccades, viewpoint…

Robotics · Computer Science 2026-03-09 Yuzhi Lai , Shenghai Yuan , Peizheng Li , Andreas Zell

The success of vision-language models is primarily attributed to effective alignment across modalities such as vision and language. However, modality gaps persist in existing alignment algorithms and appear necessary for human perception as…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Hanqi Yan , Xiangxiang Cui , Lu Yin , Jindong Gu , Paul Pu Liang , Yulan He , Yifei Wang

In human-robot interaction (HRI), detecting a human's gaze helps robots interpret user attention and intent. However, most gaze detection approaches rely on specialized eye-tracking hardware, limiting deployment in everyday settings.…

Robotics · Computer Science 2026-03-18 Linlin Cheng , Koen Hindriks , Artem V. Belopolsky

Multimodal industrial surface defect detection (MISDD) aims to identify and locate defect in industrial products by fusing RGB and 3D modalities. This article focuses on modality-missing problems caused by uncertain sensors availability in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Shuai Jiang , Yunfeng Ma , Jingyu Zhou , Yuan Bian , Yaonan Wang , Min Liu

Object selection in Mixed Reality (MR) becomes particularly challenging in dense or occluded environments, where traditional mid-air ray-casting often leads to ambiguity and reduced precision. We present two complementary techniques: (1) a…

Human-Computer Interaction · Computer Science 2025-08-20 Xiang Li , Per Ola Kristensson

In 3D hand-object interaction (HOI) tasks, estimating precise joint poses of hands and objects from monocular RGB input remains highly challenging due to the inherent geometric ambiguity of RGB images and the severe mutual occlusions that…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Yuechen Xie , Haobo Jiang , Jian Yang , Yigong Zhang , Jin Xie

Mobile robot navigation systems are increasingly relied upon in dynamic and complex environments, yet they often struggle with map inaccuracies and the resulting inefficient path planning. This paper presents MRHaD, a Mixed Reality-based…

Robotics · Computer Science 2025-07-29 Takumi Taki , Masato Kobayashi , Eduardo Iglesius , Naoya Chiba , Shizuka Shirai , Yuki Uranishi

The rapid scaling of large vision pretrained models makes fine-tuning tasks more and more difficult on devices with low computational resources. We explore a new visual adaptation paradigm called separated tuning, which treats large…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Ningyuan Tang , Minghao Fu , Jianxin Wu

With the recently increasing capabilities of modern vehicles, novel approaches for interaction emerged that go beyond traditional touch-based and voice command approaches. Therefore, hand gestures, head pose, eye gaze, and speech have been…

Human-Computer Interaction · Computer Science 2022-11-08 Amr Gomaa

Autonomous graphical user interface (GUI) agents powered by multimodal large language models have shown great promise. However, a critical yet underexplored issue persists: over-execution, where the agent executes tasks in a fully…

Human-Computer Interaction · Computer Science 2025-07-15 Pengzhou Cheng , Zheng Wu , Zongru Wu , Aston Zhang , Zhuosheng Zhang , Gongshen Liu

Understanding videos that contain multiple modalities is crucial, especially in egocentric videos, where combining various sensory inputs significantly improves tasks like action recognition and moment localization. However, real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Merey Ramazanova , Alejandro Pardo , Bernard Ghanem , Motasem Alfarra

Autonomous agents often require multiple strategies to solve complex tasks, but determining when to switch between strategies remains challenging. This research introduces a reinforcement learning technique to learn switching thresholds…

Machine Learning · Computer Science 2025-12-09 Chris Tava
‹ Prev 1 8 9 10 Next ›