English
Related papers

Related papers: Volumetric Environment Representation for Vision-L…

200 papers

VALAN is a lightweight and scalable software framework for deep reinforcement learning based on the SEED RL architecture. The framework facilitates the development and evaluation of embodied agents for solving grounded language…

Machine Learning · Computer Science 2019-12-09 Larry Lansing , Vihan Jain , Harsh Mehta , Haoshuo Huang , Eugene Ie

Recent advancements in Generative AI, particularly in Large Language Models (LLMs) and Large Vision-Language Models (LVLMs), offer new possibilities for integrating cognitive planning into robotic systems. In this work, we present a novel…

Robotics · Computer Science 2024-11-06 Arjun P S , Andrew Melnik , Gora Chand Nandi

The task of vision-and-language navigation in continuous environments (VLN-CE) aims at training an autonomous agent to perform low-level actions to navigate through 3D continuous surroundings using visual observations and language…

Robotics · Computer Science 2024-12-30 Lu Yue , Dongliang Zhou , Liang Xie , Feitian Zhang , Ye Yan , Erwei Yin

Volumetric design is the first and critical step for professional building design, where architects not only depict the rough 3D geometry of the building but also specify the programs to form a 2D layout on each floor. Though 2D layout…

Machine Learning · Computer Science 2021-04-28 Kai-Hung Chang , Chin-Yi Cheng , Jieliang Luo , Shingo Murata , Mehdi Nourbakhsh , Yoshito Tsuji

Developing Vision-and-Language Navigation (VLN) agents typically assumes a \textit{train-once-deploy-once} strategy, which is unrealistic as deployed agents continually encounter novel environments. To address this, we propose the Continual…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Seongjun Jeong , Gi-Cheon Kang , Seongho Choi , Joochan Kim , Byoung-Tak Zhang

Although large language models (LLMs) are introduced into vision-and-language navigation (VLN) to improve instruction comprehension and generalization, existing LLM- based VLN lacks the ability to selectively recall and use relevant priori…

Artificial Intelligence · Computer Science 2026-03-10 Haozhou Li , Xiangyu Dong , Huiyan Jiang , Yaoming Zhou , Xiaoguang Ma

Decades of cognitive science establish that humans navigate environments by forming cognitive maps, defined as allocentric and topology-preserving representations of 3D space. While modern Vision-Language Models (VLMs) demonstrate emergent…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Haoming Wang , Wei Gao

With the emergence of varied visual navigation tasks (e.g, image-/object-/audio-goal and vision-language navigation) that specify the target in different ways, the community has made appealing advances in training specialized agents capable…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Hanqing Wang , Wei Liang , Luc Van Gool , Wenguan Wang

Numerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (SSC). They typically construct 3D probability volumes…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Bohan Li , Yasheng Sun , Jingxin Dong , Zheng Zhu , Jinming Liu , Xin Jin , Wenjun Zeng

Pre-trained general-purpose Vision-Language Models (VLM) hold the potential to enhance intuitive human-machine interactions due to their rich world knowledge and 2D object detection capabilities. However, VLMs for 3D coordinates detection…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Ari Wahl , Dorian Gawlinski , David Przewozny , Paul Chojecki , Felix Bießmann , Sebastian Bosse

Rapid advancements in Autonomous Driving (AD) tasks turned a significant shift toward end-to-end fashion, particularly in the utilization of vision-language models (VLMs) that integrate robust logical reasoning and cognitive abilities to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Yifan Bai , Dongming Wu , Yingfei Liu , Fan Jia , Weixin Mao , Ziheng Zhang , Yucheng Zhao , Jianbing Shen , Xing Wei , Tiancai Wang , Xiangyu Zhang

Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose the Dynamic 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zihan Wang , Seungjun Lee , Guangzhao Dai , Gim Hee Lee

Vision Language Navigation in Continuous Environments (VLN-CE) represents a frontier in embodied AI, demanding agents to navigate freely in unbounded 3D spaces solely guided by natural language instructions. This task introduces distinct…

Artificial Intelligence · Computer Science 2024-09-24 Zhiyuan Li , Yanfeng Lu , Yao Mu , Hong Qiao

The aspiration of the Vision-and-Language Navigation (VLN) task has long been to develop an embodied agent with robust adaptability, capable of seamlessly transferring its navigation capabilities across various tasks. Despite remarkable…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Siqi Zhang , Yanyuan Qiao , Qunbo Wang , Longteng Guo , Zhihua Wei , Jing Liu

Volumetric medical imaging technologies produce detailed 3D representations of anatomical structures. However, effective medical data visualization and exploration pose significant challenges, especially for individuals with limited medical…

Human-Computer Interaction · Computer Science 2025-07-01 Qixuan Liu , Shi Qiu , Yinqiao Wang , Xiwen Wu , Kenneth Siu Ho Chok , Chi-Wing Fu , Pheng-Ann Heng

A metric-accurate semantic 3D representation is essential for many robotic tasks. This work proposes a simple, yet powerful, way to integrate the 2D embeddings of a Vision-Language Model in a metric-accurate 3D representation at real-time.…

Robotics · Computer Science 2025-08-11 Christian Rauch , Björn Ellensohn , Linus Nwankwo , Vedant Dave , Elmar Rueckert

Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomings limit the insight that the benchmarks provide into…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Sihao Lin , Zerui Li , Xunyi Zhao , Gengze Zhou , Liuyi Wang , Rong Wei , Rui Tang , Juncheng Li , Hanqing Wang , Jiangmiao Pang , Anton van den Hengel , Jiajun Liu , Qi Wu

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to execute sequential navigation actions in complex environments guided by natural language instructions. Current approaches often struggle with generalizing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Xuan Yao , Junyu Gao , Changsheng Xu

This paper addresses the challenge of fine-grained alignment in Vision-and-Language Navigation (VLN) tasks, where robots navigate realistic 3D environments based on natural language instructions. Current approaches use contrastive learning…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Yuhang Song , Mario Gianni , Chenguang Yang , Kunyang Lin , Te-Chuan Chiu , Anh Nguyen , Chun-Yi Lee

Vision-and-Language Navigation (VLN) requires an agent to find a specified spot in an unseen environment by following natural language instructions. Dominant methods based on supervised learning clone expert's behaviours and thus perform…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Hu Wang , Qi Wu , Chunhua Shen
‹ Prev 1 8 9 10 Next ›