中文
相关论文

相关论文: mmWalk: Towards Multi-modal Multi-view Walking Ass…

200 篇论文

Indoor navigation remains a critical accessibility challenge for the blind and low-vision (BLV) individuals, as existing solutions rely on costly per-building infrastructure. We present an agentic framework that converts a single floor plan…

人工智能 · 计算机科学 2026-04-28 Aydin Ayanzadeh , Tim Oates

Interactive streetscape mapping tools such as Google Street View (GSV) and Meta Mapillary enable users to virtually navigate and experience real-world environments via immersive 360{\deg} imagery but remain fundamentally inaccessible to…

人机交互 · 计算机科学 2025-09-29 Jon E. Froehlich , Alexander Fiannaca , Nimer Jaber , Victor Tsaran , Shaun Kane

Visual language models (VLMs) empower mobile GUI agents to interpret complex mobile screens and respond to user requests. Training such capable agents requires large-scale, high-quality mobile GUI data. However, existing mobile GUI datasets…

人机交互 · 计算机科学 2025-11-26 Longxi Gao , Li Zhang , Shihe Wang , Pengzhi Gao , Wei Liu , Jian Luan , Shangguang Wang , Yuanchun Li , Mengwei Xu

We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can be solved by humans "within a blink" (e.g., relative…

计算机视觉与模式识别 · 计算机科学 2024-07-04 Xingyu Fu , Yushi Hu , Bangzheng Li , Yu Feng , Haoyu Wang , Xudong Lin , Dan Roth , Noah A. Smith , Wei-Chiu Ma , Ranjay Krishna

We present a multi-modal trajectory generation and selection algorithm for real-world mapless outdoor navigation in human-centered environments. Such environments contain rich features like crosswalks, grass, and curbs, which are easily…

机器人学 · 计算机科学 2025-05-19 Daeun Song , Jing Liang , Xuesu Xiao , Dinesh Manocha

Vision-and-language navigation (VLN) aims to develop agents capable of navigating in realistic environments. While recent cross-modal training approaches have significantly improved navigation performance in both indoor and outdoor…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Jungdae Lee , Taiki Miyanishi , Shuhei Kurita , Koya Sakamoto , Daichi Azuma , Yutaka Matsuo , Nakamasa Inoue

Our research investigates the capability of modern multimodal reasoning models, powered by Large Language Models (LLMs), to facilitate vision-powered assistants for multi-step daily activities. Such assistants must be able to 1) encode…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Mrinal Verghese , Brian Chen , Hamid Eghbalzadeh , Tushar Nagarajan , Ruta Desai

Vision-aided wireless communication is motivated by the recent advances in deep learning and computer vision as well as the increasing dependence on line-of-sight links in millimeter wave (mmWave) and terahertz systems. By leveraging…

信号处理 · 电气工程与系统科学 2020-02-14 Muhammad Alrabeiah , Jayden Booth , Andrew Hredzak , Ahmed Alkhateeb

According to the World Health Organization, visual impairment is estimated to affect approximately 2.2 billion people worldwide. The visually impaired must currently rely on navigational aids to replace their sense of sight, like a white…

人机交互 · 计算机科学 2022-06-23 Stanley Shen

Assessing the accessibility of unfamiliar built environments is critical for people with disabilities. However, manual assessments, performed by users or their personal health professionals, are laborious and unscalable, while automatic…

人机交互 · 计算机科学 2026-01-23 William Huang , Xia Su , Jon E. Froehlich , Yang Zhang

In this paper, we propose a deep learning based assistive system to improve the environment perception experience of visually impaired (VI). The system is composed of a wearable terminal equipped with an RGBD camera and an earphone, a…

机器人学 · 计算机科学 2019-08-12 Yimin Lin , Kai Wang , Wanxin Yi , Shiguo Lian

Blind and low-vision (BLV) people face many challenges when venturing into public environments, often wishing it were easier to get help from people nearby. Ironically, while many sighted individuals are willing to help, such interactions…

Can Multimodal Large Language Models (MLLMs) develop an intuitive number sense similar to humans? Targeting this problem, we introduce Visual Number Benchmark (VisNumBench) to evaluate the number sense abilities of MLLMs across a wide range…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Tengjin Weng , Jingyi Wang , Wenhao Jiang , Zhong Ming

We present M$^3$-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multimodal entity understanding and complex multi-hop reasoning.…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Jiatong Ma , Longteng Guo , Yuchen Liu , Zijia Zhao , Dongze Hao , Xuanxu Lin , Jing Liu

Blind and low-vision (BLV) people rely on GPS-based systems for outdoor navigation. GPS's inaccuracy, however, causes them to veer off track, run into obstacles, and struggle to reach precise destinations. While prior work has made precise…

Gait recognition has a rapid development in recent years. However, gait recognition in the wild is not well explored yet. An obvious reason could be ascribed to the lack of diverse training data from the perspective of intrinsic and…

计算机视觉与模式识别 · 计算机科学 2021-06-02 Pengyi Zhang , Huanzhang Dou , Wenhu Zhang , Yuhan Zhao , Songyuan Li , Zequn Qin , Xi Li

Assistants on assembly tasks show great potential to benefit humans ranging from helping with everyday tasks to interacting in industrial settings. However, evaluation resources in assembly activities are underexplored. To foster system…

The increasingly complex and diverse planetary exploration environment requires more adaptable and flexible rover navigation strategy. In this study, we propose a VLM-empowered multi-mode system to achieve efficient while safe autonomous…

机器人学 · 计算机科学 2025-06-23 Sinuo Cheng , Ruyi Zhou , Wenhao Feng , Huaiguang Yang , Haibo Gao , Zongquan Deng , Liang Ding

Large multimodal models (LMMs) have enabled new AI-powered applications that help people with visual impairments (PVI) receive natural language descriptions of their surroundings through audible text. We investigated how this emerging…

人机交互 · 计算机科学 2025-02-25 Jingyi Xie , Rui Yu , He Zhang , Syed Masum Billah , Sooyeon Lee , John M. Carroll

A navigable agent needs to understand both high-level semantic instructions and precise spatial perceptions. Building navigation agents centered on Multimodal Large Language Models (MLLMs) demonstrates a promising solution due to their…

机器人学 · 计算机科学 2026-02-18 Zerui Li , Hongpei Zheng , Fangguo Zhao , Aidan Chan , Jian Zhou , Sihao Lin , Shijie Li , Qi Wu