English
Related papers

Related papers: MDE-AgriVLN: Agricultural Vision-and-Language Navi…

200 papers

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet most evaluations of…

Recent Vision-and-Language Navigation (VLN) advancements are promising, but their idealized assumptions about robot movement and control fail to reflect physically embodied deployment challenges. To bridge this gap, we introduce VLN-PE, a…

Robotics · Computer Science 2025-09-29 Liuyi Wang , Xinyuan Xia , Hui Zhao , Hanqing Wang , Tai Wang , Yilun Chen , Chengju Liu , Qijun Chen , Jiangmiao Pang

Vision-and-Language Navigation (VLN) is a natural language grounding task where agents have to interpret natural language instructions in the context of visual scenes in a dynamic environment to achieve prescribed navigation goals.…

Computation and Language · Computer Science 2019-06-03 Haoshuo Huang , Vihan Jain , Harsh Mehta , Jason Baldridge , Eugene Ie

Vision-Language Navigation (VLN) approaches have currently followed two primary paradigms: the end-to-end Vision-Language Model (VLM) policy fine-tuned on navigation trajectories to directly predict actions, and the zero-shot modular…

Robotics · Computer Science 2026-05-19 Jingzhi Huang , Junkai Huang , Wenxuan Song , Haoyang Yang , Hailong Huang , Haoang Li , Yi Wang

Monocular depth estimation (MDE) is a critical task to guide autonomous medical robots. However, obtaining absolute (metric) depth from an endoscopy camera in surgical scenes is difficult, which limits supervised learning of depth on real…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Hao Li , Daiwei Lu , Jesse d'Almeida , Dilara Isik , Ehsan Khodapanah Aghdam , Nick DiSanto , Ayberk Acar , Susheela Sharma , Jie Ying Wu , Robert J. Webster , Ipek Oguz

The task of vision-and-language navigation in continuous environments (VLN-CE) aims at training an autonomous agent to perform low-level actions to navigate through 3D continuous surroundings using visual observations and language…

Robotics · Computer Science 2024-12-30 Lu Yue , Dongliang Zhou , Liang Xie , Feitian Zhang , Ye Yan , Erwei Yin

Depth is a vital piece of information for autonomous vehicles to perceive obstacles. Due to the relatively low price and small size of monocular cameras, depth estimation from a single RGB image has attracted great interest in the research…

Robotics · Computer Science 2021-11-25 Xingshuai Dong , Matthew A. Garratt , Sreenatha G. Anavatti , Hussein A. Abbass

Long-horizon collaborative vision-language navigation (VLN) is critical for multi-robot systems to accomplish complex tasks beyond the capability of a single agent. CoNavBench takes a first step by introducing the first collaborative…

Robotics · Computer Science 2026-04-15 Sunyao Zhou , Yunzi Wu , Tianhang Wang , Xinhai Li , Guang Chen , Lizheng Liu , Chenjia Bai , Xuelong Li

This paper addresses the problem of Monocular Depth Estimation (MDE). Existing approaches on MDE usually model it as a pixel-level regression problem, ignoring the underlying geometry property. We empirically find this may result in…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Yixuan Liu , Yuwang Wang , Shengjin Wang

Vision-and-Language Navigation (VLN) empowers agents to associate time-sequenced visual observations with corresponding instructions to make sequential decisions. However, generalization remains a persistent challenge, particularly when…

Robotics · Computer Science 2025-02-27 Zerui Li , Gengze Zhou , Haodong Hong , Yanyan Shao , Wenqi Lyu , Yanyuan Qiao , Qi Wu

Vision-and-Language navigation (VLN) requires an agent to navigate in unseen environment by following natural language instruction. For task completion, the agent needs to align and integrate various navigation modalities, including…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Mengfei Du , Binhao Wu , Jiwen Zhang , Zhihao Fan , Zejun Li , Ruipu Luo , Xuanjing Huang , Zhongyu Wei

Vision-Language Navigation (VLN) for Unmanned Aerial Vehicles (UAVs) demands complex visual interpretation and continuous control in dynamic 3D environments. Existing hierarchical approaches rely on dense oracle guidance or auxiliary object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Peng Xu , Zhengnan Deng , Jiayan Deng , Zonghua Gu , Shaohua Wan

The existing methods for Vision and Language Navigation in the Continuous Environment (VLN-CE) commonly incorporate a waypoint predictor to discretize the environment. This simplifies the navigation actions into a view selection task and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Yue Zhang , Parisa Kordjamshidi

Monocular depth estimation (MDE) has witnessed remarkable progress driven by Convolutional Neural Networks and transformer-based architectures. However, these approaches typically treat the problem as a generic image-to-image regression on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Qianlei Wang , Kexun Chen , Shaolin Zhang , Hongli Gao , Chaoning Zhang , Xiaolin Qin

Monocular depth estimation (MDE) plays a crucial role in enabling spatially-aware applications in Ultra-low-power (ULP) Internet-of-Things (IoT) platforms. However, the limited number of parameters of Deep Neural Networks for the MDE task,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Davide Nadalini , Manuele Rusci , Elia Cereda , Luca Benini , Francesco Conti , Daniele Palossi

Visual Teach and Repeat (VT\&R) allows an autonomous vehicle to repeat a previously traversed route without a global positioning system. Existing implementations of VT\&R typically rely on 3D sensors such as stereo cameras for mapping and…

Robotics · Computer Science 2019-08-08 Lee Clement , Jonathan Kelly , Timothy D. Barfoot

The emerging vision-and-language navigation (VLN) problem aims at learning to navigate an agent to the target location in unseen photo-realistic environments according to the given language instruction. The main challenges of VLN arise…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Weixia Zhang , Chao Ma , Qi Wu , Xiaokang Yang

Vision-and-Language Navigation (VLN) requires an embodied agent to ground complex natural-language instructions into long-horizon navigation in unseen environments. While Vision-Language Models (VLMs) offer strong 2D semantic understanding,…

Robotics · Computer Science 2026-03-19 Zihao Xin , Wentong Li , Yixuan Jiang , Ziyuan Huang , Bin Wang , Piji Li , Jianke Zhu , Jie Qin , Shengjun Huang

Language-guided embodied navigation requires an agent to interpret object-referential instructions, search across multiple rooms, localize the referenced target, and execute reliable motion toward it. Existing systems remain limited in real…

Robotics · Computer Science 2026-03-19 Zhongyuang Liu , Min He , Shaonan Yu , Xinhang Xu , Muqing Cao , Jianping Li , Jianfei Yang , Lihua Xie

Vision-and-Language Navigation (VLN) is a task to guide an embodied agent moving to a target position using language instructions. Despite the significant performance improvement, the wide use of fine-grained instructions fails to…

Computer Vision and Pattern Recognition · Computer Science 2022-10-19 Weixi Feng , Tsu-Jui Fu , Yujie Lu , William Yang Wang