English
Related papers

Related papers: ClickAIXR: On-Device Multimodal Vision-Language In…

200 papers

Task-oriented object detection aims to find objects suitable for accomplishing specific tasks. As a challenging task, it requires simultaneous visual data processing and reasoning under ambiguous semantics. Recent solutions are mainly…

Computer Vision and Pattern Recognition · Computer Science 2024-09-09 Hanning Chen , Wenjun Huang , Yang Ni , Sanggeon Yun , Yezi Liu , Fei Wen , Alvaro Velasquez , Hugo Latapie , Mohsen Imani

Current gaze input methods for VR headsets predominantly utilize the gaze ray as a pointing cursor, often neglecting depth information in it. This study introduces FocusFlow, a novel gaze interaction technique that integrates focal depth…

Human-Computer Interaction · Computer Science 2023-08-15 Chenyang Zhang , Tiansu Chen , Rohan Nedungadi , Eric Shaffer , Elahe Soltanaghai

Hand gestures form an intuitive means of interaction in Mixed Reality (MR) applications. However, accurate gesture recognition can be achieved only through state-of-the-art deep learning models or with the use of expensive sensors. Despite…

Computer Vision and Pattern Recognition · Computer Science 2019-04-23 Varun Jain , Gaurav Garg , Ramakrishna Perla , Ramya Hebbalaguppe

Recent human-object interaction detection (HOID) methods highly require prior knowledge from vision-language models (VLMs) to enhance the interaction recognition capabilities. The training strategies and model architectures for connecting…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Junwen Chen , Peilin Xiong , Keiji Yanai

Lately, researchers in artificial intelligence have been really interested in how language and vision come together, giving rise to the development of multimodal models that aim to seamlessly integrate textual and visual information.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Rajat Chawla , Arkajit Datta , Tushar Verma , Adarsh Jha , Anmol Gautam , Ayush Vatsal , Sukrit Chaterjee , Mukunda NS , Ishaan Bhola

The growing ubiquity of Extended Reality (XR) is driving Conversational Recommendation Systems (CRS) toward visually immersive experiences. We formalize this paradigm as Immersive CRS (ICRS), where recommended items are highlighted directly…

Information Retrieval · Computer Science 2026-04-14 Jiazhou Liang , Yifan Simon Liu , David Guo , Minqi Sun , Yilun Jiang , Scott Sanner

When establishing a visual connection between a virtual reality user and an augmented reality user, it is important to consider whether the augmented reality user faces a surplus of information. Augmented reality, compared to virtual…

Human-Computer Interaction · Computer Science 2021-05-05 Robbe Cools , Jihae Han , Adalberto L. Simeone

The synergy between virtual reality (VR) and artificial intelligence (AI), specifically deep learning (DL)-based cybersickness detection models, has ushered in unprecedented advancements in immersive experiences by automatically detecting…

Cryptography and Security · Computer Science 2025-03-18 Ripan Kumar Kundu , Matthew Denton , Genova Mongalo , Prasad Calyam , Khaza Anuarul Hoque

Multimodal recommender systems leverage diverse data sources, such as user interactions, content features, and contextual information, to address challenges like cold-start and data sparsity. However, existing methods often suffer from one…

Information Retrieval · Computer Science 2026-02-24 Adamya Shyam , Venkateswara Rao Kagita , Bharti Rana , Vikas Kumar

Extended Reality (XR) interfaces offer engaging user experiences, but their effective design requires a nuanced understanding of user behavior and preferences. This knowledge is challenging to obtain without the widespread adoption of XR…

Human-Computer Interaction · Computer Science 2024-03-14 Hyunsung Cho , Yukang Yan , Kashyap Todi , Mark Parent , Missie Smith , Tanya R. Jonker , Hrvoje Benko , David Lindlbauer

A promising effective human-robot interaction in assistive robotic systems is gaze-based control. However, current gaze-based assistive systems mainly help users with basic grasping actions, offering limited support. Moreover, the…

Robotics · Computer Science 2025-08-20 Zejia Zhang , Bo Yang , Xinxing Chen , Weizhuang Shi , Haoyuan Wang , Wei Luo , Jian Huang

Vision-language models (VLMs) have excelled in multimodal tasks, but adapting them to embodied decision-making in open-world environments presents challenges. One critical issue is bridging the gap between discrete entities in low-level…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Shaofei Cai , Zihao Wang , Kewei Lian , Zhancun Mu , Xiaojian Ma , Anji Liu , Yitao Liang

Existing augmented reality (AR) applications often ignore occlusion between real hands and virtual objects when incorporating virtual objects in our views. The challenges come from the lack of accurate depth and mismatch between real and…

Graphics · Computer Science 2020-06-24 Xiao Tang , Xiaowei Hu , Chi-Wing Fu , Daniel Cohen-Or

We present VOICE, a novel approach to science communication that connects large language models' (LLM) conversational capabilities with interactive exploratory visualization. VOICE introduces several innovative technical contributions that…

Human-Computer Interaction · Computer Science 2026-05-19 Donggang Jia , Alexandra Irger , Lonni Besancon , Ondrej Strnad , Deng Luo , Johanna Bjorklund , Anders Ynnerman , Ivan Viola

The rapid development of Large Language Models (LLMs) creates an exciting potential for flexible, general knowledge-driven Human-Robot Interaction (HRI) systems for assistive robots. Existing HRI systems demonstrate great progress in…

Robotics · Computer Science 2025-07-22 Jens V. Rüppel , Andrey Rudenko , Tim Schreiter , Martin Magnusson , Achim J. Lilienthal

Voice assistants (VAs) like Siri and Alexa are transforming human-computer interaction; however, they lack awareness of users' spatiotemporal context, resulting in limited performance and unnatural dialogue. We introduce GazePointAR, a…

Human-Computer Interaction · Computer Science 2025-11-25 Jaewook Lee , Jun Wang , Elizabeth Brown , Liam Chu , Sebastian S. Rodriguez , Jon E. Froehlich

Extended reality (XR) technologies-encompassing virtual reality (VR), augmented reality (AR), and mixed reality (MR) are transforming cognitive assessment and training by offering immersive, interactive environments that simulate real-world…

Human-Computer Interaction · Computer Science 2025-01-15 Palmira Victoria González-Erena , Sara Fernández-Guinea , Panagiotis Kourtesis

Assistive technologies for the visually impaired have evolved to facilitate interaction with a complex and dynamic world. In this paper, we introduce AIris, an AI-powered wearable device that provides environmental awareness and interaction…

Human-Computer Interaction · Computer Science 2024-08-13 Dionysia Danai Brilli , Evangelos Georgaras , Stefania Tsilivaki , Nikos Melanitis , Konstantina Nikita

The advent of Retrieval-Augmented Generation (RAG) has significantly enhanced the ability of Large Language Models (LLMs) to produce factually accurate and up-to-date responses. However, the performance of a RAG system is not determined by…

Human-Computer Interaction · Computer Science 2026-02-17 Haoyu Tian , Yingchaojie Feng , Zhen Wen , Haoxuan Li , Minfeng Zhu , Wei Chen

Multimodal Large Language Models (MLLMs) have shown success in various general image processing tasks, yet their application in medical imaging is nascent, lacking tailored models. This study investigates the potential of MLLMs in improving…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Ling Yang , Zhanyu Wang , Zhenghao Chen , Xinyu Liang , Luping Zhou