English
Related papers

Related papers: Grounded GUI Understanding for Vision-Based Spatia…

200 papers

End-to-end robot policies achieve high performance through neural networks trained via reinforcement learning (RL). Yet, their black box nature and abstract reasoning pose challenges for human-robot interaction (HRI), because humans may…

Accurate semantic segmentation of urban remote sensing images (URSIs) is essential for urban planning and environmental monitoring. However, it remains challenging due to the subtle texture differences and similar spatial structures among…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Guoyu Zhou , Jing Zhang , Yi Yan , Hui Zhang , Li Zhuo

Graphical user interface (GUI) agents have shown promise in automating mobile tasks but still struggle with input redundancy and decision ambiguity. In this paper, we present \textbf{RecAgent}, an uncertainty-aware agent that addresses…

Artificial Intelligence · Computer Science 2025-08-07 Chao Hao , Shuai Wang , Kaiwen Zhou

Graphical user interface (GUI) grounding, the process of mapping human instructions to GUI actions, serves as a fundamental basis to autonomous GUI agents. While existing grounding models achieve promising performance to simulate the mouse…

Human-Computer Interaction · Computer Science 2026-01-13 Zeyi Liao , Yadong Lu , Boyu Gou , Huan Sun , Ahmed Awadallah

We present the concept of X-Vision, an enhanced Augmented Reality (AR)-based visualization tool, with the real-time sensing capability in a tagged environment. We envision that this type of a tool will enhance the user-environment…

Human-Computer Interaction · Computer Science 2018-12-07 Yongbin Sun , Sai Nithin R. Kantareddy , Rahul Bhattacharyya , Sanjay E. Sarma

We present an Open-Vocabulary 3D Scene Graph (OVSG), a formal framework for grounding a variety of entities, such as object instances, agents, and regions, with free-form text-based queries. Unlike conventional semantic-based object…

Fine-grained open-vocabulary object detection (FG-OVD) aims to detect novel object categories described by attribute-rich texts. While existing open-vocabulary detectors show promise at the base-category level, they underperform in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Jiaming Li , Zhijia Liang , Weikai Chen , Lin Ma , Guanbin Li

There is a growing demand for mobile user interface (UI) automation, driven by its broad applications across industries. With the advent of visual language models (VLMs), GUI automation has progressed from generating text-based instructions…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Neeraj Anand , Rishabh Jain , Sohan Patnaik , Balaji Krishnamurthy , Mausoom Sarkar

GUI agents powered by vision-language models (VLMs) show promise in automating complex digital tasks. However, their effectiveness in real-world applications is often limited by scarce training data and the inherent complexity of these…

Computation and Language · Computer Science 2025-09-30 Ran Xu , Kaixin Ma , Wenhao Yu , Hongming Zhang , Joyce C. Ho , Carl Yang , Dong Yu

This paper presents a wearable assistive device with the shape of a pair of eyeglasses that allows visually impaired people to navigate safely and quickly in unfamiliar environment, as well as perceive the complicated environment to…

Computer Vision and Pattern Recognition · Computer Science 2019-06-25 Jinqiang Bai , Zhaoxiang Liu , Yimin Lin , Ye Li , Shiguo Lian , Dijun Liu

We present the Object-Based Sub-Environment Recognition (OBSER) framework, a novel Bayesian framework that infers three fundamental relationships between sub-environments and their constituent objects. In the OBSER framework, metric and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Won-Seok Choi , Dong-Sig Han , Suhyung Choi , Hyeonseo Yang , Byoung-Tak Zhang

Gaze-based selection in XR requires visual confirmation due to eye-tracking limitations and target ambiguity in 3D contexts. Current designs for wide-FOV displays use world-locked, central overlays, which are not conducive to always-on AR…

Human-Computer Interaction · Computer Science 2026-03-20 Yutong Ren , Arnav Reddy , Michael Nebeling

Current methods to estimate object shape---using either vision or touch---generally depend on high-resolution sensing. Here, we exploit ergodic exploration to demonstrate successful shape estimation when using a low-resolution binary…

Robotics · Computer Science 2017-09-07 Ian Abraham , Ahalya Prabhakar , Mitra J. Z. Hartmann , Todd D. Murphey

This paper contains a brief discussion of an object evaluator which is based on principles of evaluations in a category. The main tool system referred as the Application Development Environment (ADE) is used to build database applications…

Logic in Computer Science · Computer Science 2007-05-23 Larissa Ismailova , Konstantin Zinchenko

From a visual perception perspective, modern graphical user interfaces (GUIs) comprise a complex graphics-rich two-dimensional visuospatial arrangement of text, images, and interactive objects such as buttons and menus. While existing…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Yue Jiang , Zixin Guo , Hamed Rezazadegan Tavakoli , Luis A. Leiva , Antti Oulasvirta

The objective of augmented reality (AR) is to add digital content to natural images and videos to create an interactive experience between the user and the environment. Scene analysis and object recognition play a crucial role in AR, as…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Vladislav Li , Barbara Villarini , Jean-Christophe Nebel , Thomas Lagkas , Panagiotis Sarigiannidis , Vasileios Argyriou

This work introduces a novel Augmented Reality (AR) approach to visualize material data alongside real objects in order to facilitate detailed material analyses based on spatial non-destructive testing (NDT) data as generated in X-ray…

Human-Computer Interaction · Computer Science 2024-04-22 Alexander Gall , Anja Heim , Patrick Weinberger , Bernhard Fröhler , Johann Kastner , Christoph Heinzl

Detecting spliced images is one of the emerging challenges in computer vision. Unlike prior methods that focus on detecting low-level artifacts generated during the manipulation process, we use an image retrieval approach to tackle this…

Computer Vision and Pattern Recognition · Computer Science 2021-04-29 Bor-Chun Chen , Zuxuan Wu , Larry S. Davis , Ser-Nam Lim

Weakly-supervised Human-Object Interaction (HOI) detection is essential for scalable scene understanding, as it learns interactions from only image-level annotations. Due to the lack of localization signals, prior works typically rely on an…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Jihwan Park , Chanhyeong Yang , Jinyoung Park , Taehoon Song , Hyunwoo J. Kim

Hand gestures play a significant role in human interactions where non-verbal intentions, thoughts and commands are conveyed. In Human-Robot Interaction (HRI), hand gestures offer a similar and efficient medium for conveying clear and rapid…

Robotics · Computer Science 2024-04-11 Eran Bamani , Eden Nissinman , Inbar Meir , Lisa Koenigsberg , Avishai Sintov