中文
相关论文

相关论文: ClickAIXR: On-Device Multimodal Vision-Language In…

200 篇论文

Memory-augmented Large Language Models (LLMs) have demonstrated remarkable consistency during prolonged dialogues by storing relevant memories and incorporating them as context. Such memory-based personalization is also key in on-device…

机器学习 · 计算机科学 2025-12-05 Massimo Bini , Ondrej Bohdal , Umberto Michieli , Zeynep Akata , Mete Ozay , Taha Ceritli

A switchable virtual reality (VR), augmented reality (AR), and mixed reality (MR) system is proposed using digital optical cloaking. Optical cloaking allows completely opaque VR devices to be "cloaked," switching to AR or MR while providing…

光学 · 物理学 2018-02-07 Joseph S. Choi

This paper explores the effectiveness of Multimodal Large Language models (MLLMs) as assistive technologies for visually impaired individuals. We conduct a user survey to identify adoption patterns and key challenges users face with such…

Visual impairment affects the ability of people to live a life like normal people. Such people face challenges in performing activities of daily living, such as reading, writing, traveling and participating in social gatherings. Many…

Multimodal search has become increasingly important in providing users with a natural and effective way to ex-press their search intentions. Images offer fine-grained details of the desired products, while text allows for easily…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Oriol Barbany , Michael Huang , Xinliang Zhu , Arnab Dhua

The Visual Object Information Retrieval (VOIR) system described in this paper implements an image retrieval approach that combines two layers, the conceptual and the visual layer. It uses terms from a textual thesaurus to represent the…

信息检索 · 计算机科学 2008-09-30 Jose Torres , Luis Paulo Reis

Navigating to a specified object in an unknown environment is a fundamental yet challenging capability of embodied intelligence. However, current methods struggle to balance decision frequency with intelligence, resulting in decisions…

In recent years, vision-language models (VLMs) have advanced open-vocabulary mapping, enabling mobile robots to simultaneously achieve environmental reconstruction and high-level semantic understanding. While integrated object cognition…

机器人学 · 计算机科学 2025-02-25 Yinan Deng , Bicheng Yao , Yihang Tang , Yi Yang , Yufeng Yue

We present the early-stage design and implementation of a multimodal, real-time communication analysis system intended as a foundational interaction layer for adaptive VR training. The system integrates five parallel processing streams: (1)…

We present the concept of X-Vision, an enhanced Augmented Reality (AR)-based visualization tool, with the real-time sensing capability in a tagged environment. We envision that this type of a tool will enhance the user-environment…

人机交互 · 计算机科学 2018-12-07 Yongbin Sun , Sai Nithin R. Kantareddy , Rahul Bhattacharyya , Sanjay E. Sarma

Composed Video Retrieval (CoVR) facilitates video retrieval by combining visual and textual queries. However, existing CoVR frameworks typically fuse multimodal inputs in a single stage, achieving only marginal gains over initial baseline.…

计算机视觉与模式识别 · 计算机科学 2026-01-26 Yuqian Zheng , Mariana-Iuliana Georgescu

We present AudioMiXR, an augmented reality (AR) interface intended to assess how users manipulate virtual audio objects situated in their physical space using six degrees of freedom (6DoF) deployed on a head-mounted display (Apple Vision…

人机交互 · 计算机科学 2025-08-07 Brandon Woodard , Margarita Geleta , Joseph J. LaViola , Andrea Fanelli , Rhonda Wilson

Traditional augmented reality (AR) systems predominantly rely on fixed class detectors or fiducial markers, limiting their ability to interpret complex, open-vocabulary natural language queries. We present a modular AR agent system that…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Lixing Guo , Tobias Höllerer

Extended Reality (XR) interfaces impose both ergonomic and cognitive demands, yet current systems often force a binary choice between hand-based input, which can produce fatigue, and gaze-based input, which is vulnerable to the Midas Touch…

人机交互 · 计算机科学 2026-03-18 Mohammad Dastgheib , Fatemeh Pourmahdian

Virtual assistants (VAs) have become ubiquitous in daily life, integrated into smartphones and smart devices, sparking interest in AI companions that enhance user experiences and foster emotional connections. However, existing companions…

人机交互 · 计算机科学 2025-09-03 Xuetong Wang , Ching Christie Pang , Pan Hui

We present Large Language Model for Mixed Reality (LLMR), a framework for the real-time creation and modification of interactive Mixed Reality experiences using LLMs. LLMR leverages novel strategies to tackle difficult cases where ideal…

Although memory capabilities of AI agents are gaining increasing attention, existing solutions remain fundamentally limited. Most rely on flat, narrowly scoped memory components, constraining their ability to personalize, abstract, and…

计算与语言 · 计算机科学 2025-07-11 Yu Wang , Xi Chen

Virtual Reality enables the exploration of large information spaces. In physically constrained spaces such as airplanes or buses, controller-based or mid-air interaction in mobile Virtual Reality can be challenging. Instead, the input space…

人机交互 · 计算机科学 2020-11-02 Tim Menzner , Travis Gesslein , Alexander Otte , Jens Grubert

During remote communication, participants often share both digital and physical content, such as product designs, digital assets, and environments, to enhance mutual understanding. Recent advances in augmented communication have facilitated…

人机交互 · 计算机科学 2024-10-10 Erzhen Hu , Mingyi Li , Jungtaek Hong , Xun Qian , Alex Olwal , David Kim , Seongkook Heo , Ruofei Du

Touchless computer interaction has become an important consideration during the COVID-19 pandemic period. Despite progress in machine learning and computer vision that allows for advanced gesture recognition, an integrated collection of…