English
Related papers

Related papers: Vinci: A Real-time Embodied Smart Assistant based …

200 papers

In this paper, we propose a deep learning based assistive system to improve the environment perception experience of visually impaired (VI). The system is composed of a wearable terminal equipped with an RGBD camera and an earphone, a…

Robotics · Computer Science 2019-08-12 Yimin Lin , Kai Wang , Wanxin Yi , Shiguo Lian

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yura Choi , Roy Miles , Rolandos Alexandros Potamias , Ismail Elezi , Jiankang Deng , Stefanos Zafeiriou

To enable a safe and effective human-robot cooperation, it is crucial to develop models for the identification of human activities. Egocentric vision seems to be a viable solution to solve this problem, and therefore many works provide deep…

Computer Vision and Pattern Recognition · Computer Science 2023-03-13 Gabriele Goletto , Mirco Planamente , Barbara Caputo , Giuseppe Averta

Continuous brain-computer interfaces (BCIs) that decode motion trajectories from imagined movement offer intuitive motor control, yet how feedback modality and longitudinal training shape neural representations and decoding performance…

Human-Computer Interaction · Computer Science 2026-05-29 Niall McShane , Attila Korik , Karl McCreadie , Naomi Du Bois , Darryl Charles , Damien Coyle

The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Daniele Materia , Francesco Ragusa , Giovanni Maria Farinella

Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomous GUI agents that…

Computation and Language · Computer Science 2025-05-06 Yiheng Xu , Zekun Wang , Junli Wang , Dunjie Lu , Tianbao Xie , Amrita Saha , Doyen Sahoo , Tao Yu , Caiming Xiong

The development of embodied AI systems is increasingly constrained by the availability and structure of physical interaction data. Despite recent advances in vision-language-action (VLA) models, current pipelines suffer from high data…

Robotics · Computer Science 2026-03-24 Xinhai Sun , Xiang Shi , Menglin Zou , Wenlong Huang

Embodied foundation models require large-scale, high-quality real-world interaction data for pre-training and scaling. However, existing data collection methods suffer from high infrastructure costs, complex hardware dependencies, and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Bowen Yang , Zishuo Li , Yang Sun , Changtao Miao , Yifan Yang , Man Luo , Xiaotong Yan , Feng Jiang , Jinchuan Shi , Yankai Fu , Ning Chen , Junkai Zhao , Pengwei Wang , Guocai Yao , Shanghang Zhang , Hao Chen , Zhe Li , Kai Zhu

A Human Computer Interface (HCI) System for playing games is designed here for more natural communication with the machines. The system presented here is a vision-based system for detection of long voluntary eye blinks and interpretation of…

Human-Computer Interaction · Computer Science 2010-02-11 S. Sumathi , S. K. Srivatsa , M. Uma Maheswari

Nowadays, voice assistants help users complete tasks on the smartphone with voice commands, replacing traditional touchscreen interactions when such interactions are inhibited. However, the usability of those tools remains moderate due to…

Human-Computer Interaction · Computer Science 2023-05-10 Minh Duc Vu , Han Wang , Zhuang Li , Gholamreza Haffari , Zhenchang Xing , Chunyang Chen

Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Arsha Nagrani , Jasper Uijilings , Shyamal Buch , Tobias Weyand , Sudheendra Vijayanarasimhan , Bo Hu , Ramin Mehran , David A Ross , Cordelia Schmid

Reward and representation learning are two long-standing challenges for learning an expanding set of robot manipulation skills from sensory observations. Given the inherent cost and scarcity of in-domain, task-specific robot data, learning…

Robotics · Computer Science 2023-03-08 Yecheng Jason Ma , Shagun Sodhani , Dinesh Jayaraman , Osbert Bastani , Vikash Kumar , Amy Zhang

Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. However, existing…

Human-Computer Interaction · Computer Science 2025-05-01 Ruiping Liu , Jiaming Zhang , Angela Schön , Karin Müller , Junwei Zheng , Kailun Yang , Anhong Guo , Kathrin Gerling , Rainer Stiefelhagen

This work presents an audio-visual interactive chatbot (AVIN-Chat) system that allows users to have face-to-face conversations with 3D avatars in real-time. Compared to the previous chatbot services, which provide text-only or speech-only…

Human-Computer Interaction · Computer Science 2024-09-04 Chanhyuk Park , Jungbin Cho , Junwan Kim , Seongmin Lee , Jungsu Kim , Sanghoon Lee

Conversational recommender systems engage users in dialogues to refine their needs and provide more personalized suggestions. Although textual information suffices for many domains, visually driven categories such as fashion or home decor…

Artificial Intelligence · Computer Science 2025-04-01 Hyunsik Jeon , Satoshi Koide , Yu Wang , Zhankui He , Julian McAuley

Population ageing and pandemics recently demonstrate to cause isolation of elderly people in their houses, generating the need for a reliable assistive figure. Robotic assistants are the new frontier of innovation for domestic welfare, and…

With the emergence of varied visual navigation tasks (e.g, image-/object-/audio-goal and vision-language navigation) that specify the target in different ways, the community has made appealing advances in training specialized agents capable…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Hanqing Wang , Wei Liang , Luc Van Gool , Wenguan Wang

Embodied cognition argues that intelligence arises from sensorimotor interaction rather than passive observation. It raises an intriguing question: do modern vision-language models (VLMs), trained largely in a disembodied manner, exhibit…

Artificial Intelligence · Computer Science 2025-11-27 Qineng Wang , Wenlong Huang , Yu Zhou , Hang Yin , Tianwei Bao , Jianwen Lyu , Weiyu Liu , Ruohan Zhang , Jiajun Wu , Li Fei-Fei , Manling Li

Vision based human pose estimation is an non-invasive technology for Human-Computer Interaction (HCI). Direct use of the hand as an input device provides an attractive interaction method, with no need for specialized sensing equipment, such…

Computer Vision and Pattern Recognition · Computer Science 2020-06-02 Nicholas Santavas , Ioannis Kansizoglou , Loukas Bampis , Evangelos Karakasis , Antonios Gasteratos

Navigation presents a significant challenge for persons with visual impairments (PVI). While traditional aids such as white canes and guide dogs are invaluable, they fall short in delivering detailed spatial information and precise guidance…

‹ Prev 1 3 4 5 6 7 10 Next ›