English
Related papers

Related papers: mmWalk: Towards Multi-modal Multi-view Walking Ass…

200 papers

Large vision-language models (VLMs) can assist visually impaired people by describing images from their daily lives. Current evaluation datasets may not reflect diverse cultural user backgrounds or the situational context of this use case.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Antonia Karamolegkou , Phillip Rust , Yong Cao , Ruixiang Cui , Anders Søgaard , Daniel Hershcovich

Recent advanced vision-language models(VLMs) have demonstrated strong performance on passive, offline image and video understanding tasks. However, their effectiveness in embodied settings, which require online interaction and active scene…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Mingxian Lin , Wei Huang , Yitang Li , Chengjie Jiang , Kui Wu , Fangwei Zhong , Shengju Qian , Xin Wang , Xiaojuan Qi

The rapid advancement of large multi-modality models (LMMs) has significantly propelled the integration of artificial intelligence into practical applications. Visual Question Answering (VQA) systems, which can process multi-modal data…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Xingyu Qi , He Li , Linjie Li , Zhenyu Wu

Vision Language Models (VLMs) demonstrate significant potential as embodied AI agents for various mobility applications. However, a standardized, closed-loop benchmark for evaluating their spatial reasoning and sequential decision-making…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Weizhen Wang , Chenda Duan , Zhenghao Peng , Yuxin Liu , Bolei Zhou

Our work aims to develop new assistive technologies that enable blind or low vision (BLV) people to explore and analyze data readily. At present, barriers exist for BLV people to explore and analyze data, restricting access to government,…

Human-Computer Interaction · Computer Science 2025-07-01 Samuel Reinders , Munazza Zaib , Matthew Butler , Bongshin Lee , Ingrid Zukerman , Lizhen Qu , Kim Marriott

Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view (BEV) from multi-view images. In MVPD, end-to-end trainable deep learning methods have progressed greatly. However, they often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Taiga Yamane , Satoshi Suzuki , Ryo Masumura , Shota Orihashi , Tomohiro Tanaka , Mana Ihori , Naoki Makishima , Naotaka Kawata

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, which requires multi-round dialogue spatial reasoning and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Xunyi Zhao , Gengze Zhou , Qi Wu

Modern Vision-Language Models (VLMs) remain poorly characterized in multi-step visual interactions, particularly in how they integrate perception, memory, and action over long horizons. We introduce VisGym, a gymnasium of 17 environments…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Zirui Wang , Junyi Zhang , Jiaxin Ge , Long Lian , Letian Fu , Lisa Dunlap , Ken Goldberg , XuDong Wang , Ion Stoica , David M. Chan , Sewon Min , Joseph E. Gonzalez

Incremental decision making in real-world environments is one of the most challenging tasks in embodied artificial intelligence. One particularly demanding scenario is Vision and Language Navigation~(VLN) which requires visual and natural…

Artificial Intelligence · Computer Science 2024-01-25 Raphael Schumann , Wanrong Zhu , Weixi Feng , Tsu-Jui Fu , Stefan Riezler , William Yang Wang

Multimodal large language models (MLLMs), equipped with increasingly advanced planning and tool-use capabilities, are evolving into autonomous agents capable of performing multimodal web browsing and deep search in open-world environments.…

Recent advances in Vision Language Models (VLMs) have driven significant progress in visual reasoning. However, open-source VLMs still lag behind proprietary systems, largely due to the lack of high-quality reasoning data. Existing datasets…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Honglin Lin , Zheng Liu , Yun Zhu , Chonghan Qin , Juekai Lin , Xiaoran Shang , Conghui He , Wentao Zhang , Lijun Wu

Medical vision-language models (VLMs) have shown promise as clinical assistants across various medical fields. However, specialized dermatology VLM capable of delivering professional and detailed diagnostic analysis remains underdeveloped,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Wenqi Zeng , Yuqi Sun , Chenxi Ma , Weimin Tan , Bo Yan

Developing agents capable of navigating to a target location based on language instructions and visual information, known as vision-language navigation (VLN), has attracted widespread interest. Most research has focused on ground-based…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Xiangyu Wang , Donglin Yang , Ziqin Wang , Hohin Kwan , Jinyu Chen , Wenjun Wu , Hongsheng Li , Yue Liao , Si Liu

Effective bipedal locomotion in dynamic environments, such as cluttered indoor spaces or uneven terrain, requires agile and adaptive movement in all directions. This necessitates omnidirectional terrain sensing and a controller capable of…

Robotics · Computer Science 2026-03-18 Mohitvishnu S. Gadde , Pranay Dugar , Ashish Malik , Alan Fern

Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale, and task scope. To…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Xiongkun Linghu , Jiangyong Huang , Xuesong Niu , Xiaojian Ma , Baoxiong Jia , Siyuan Huang

Navigation presents a significant challenge for persons with visual impairments (PVI). While traditional aids such as white canes and guide dogs are invaluable, they fall short in delivering detailed spatial information and precise guidance…

Lacking the ability to sense ambient environments effectively, blind and visually impaired people (BVIP) face difficulty in walking outdoors, especially in urban areas. Therefore, tools for assisting BVIP are of great importance. In this…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Haobin Tan , Chang Chen , Xinyu Luo , Jiaming Zhang , Constantin Seibold , Kailun Yang , Rainer Stiefelhagen

The vision-language tracking task aims to perform object tracking based on various modality references. Existing Transformer-based vision-language tracking methods have made remarkable progress by leveraging the global modeling ability of…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Xinqi Liu , Li Zhou , Zikun Zhou , Jianqiu Chen , Zhenyu He

We introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features:…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Jack Hong , Shilin Yan , Jiayin Cai , Xiaolong Jiang , Yao Hu , Weidi Xie

Instruction tuned Large Vision Language Models (LVLMs) have significantly advanced in generalizing across a diverse set of multi-modal tasks, especially for Visual Question Answering (VQA). However, generating detailed responses that are…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Anisha Gunjal , Jihan Yin , Erhan Bas