中文
相关论文

相关论文: BEVBert: Multimodal Map Pre-training for Language-…

200 篇论文

Vision-language models have become increasingly powerful for tasks that require an understanding of both visual and linguistic elements, bridging the gap between these modalities. In the context of multimodal clinical AI, there is a growing…

计算与语言 · 计算机科学 2024-04-30 Masoud Monajatipoor , Zi-Yi Dou , Aichi Chien , Nanyun Peng , Kai-Wei Chang

Vision-and-language navigation (VLN) requires an embodied agent to navigate in realistic 3D environments using natural language instructions. Existing VLN methods suffer from training on small-scale environments or unreasonable…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Kunyang Lin , Peihao Chen , Diwei Huang , Thomas H. Li , Mingkui Tan , Chuang Gan

Vision-and-language navigation (VLN) tasks require agents to navigate three-dimensional environments guided by natural language instructions, offering substantial potential for diverse applications. However, the scarcity of training data…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Sen Wang , Dongliang Zhou , Liang Xie , Chao Xu , Ye Yan , Erwei Yin

Current pre-trained vison-language models (PVLMs) achieve excellent performance on a range of multi-modal datasets. Recent work has aimed at building multilingual models, and a range of novel multilingual multi-modal datasets have been…

计算与语言 · 计算机科学 2023-10-25 Hanxu Hu , Frank Keller

Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require pre-training on a large set of parallel image-text data,…

计算机视觉与模式识别 · 计算机科学 2022-03-02 Mingyang Zhou , Licheng Yu , Amanpreet Singh , Mengjiao Wang , Zhou Yu , Ning Zhang

Following a navigation instruction such as 'Walk down the stairs and stop at the brown sofa' requires embodied AI agents to ground scene elements referenced via language (e.g. 'stairs') to visual content in the environment (pixels…

计算机视觉与模式识别 · 计算机科学 2020-05-04 Arjun Majumdar , Ayush Shrivastava , Stefan Lee , Peter Anderson , Devi Parikh , Dhruv Batra

Autonomous navigation guided by natural language instructions in embodied environments remains a challenge for vision-language navigation (VLN) agents. Although recent advancements in learning diverse and fine-grained visual environmental…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Xuesong Zhang , Jia Li , Yunbo Xu , Zhenzhen Hu , Richang Hong

Zero-shot Vision-and-Language Navigation (VLN) agents leveraging Large Language Models (LLMs) excel in generalization but suffer from insufficient spatial perception. Focusing on complex continuous environments, we categorize key perceptual…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Lu Yue , Yue Fan , Shiwei Lian , Yu Zhao , Jiaxin Yu , Liang Xie , Feitian Zhang

The Vision-and-Language Navigation (VLN) task requires an agent to follow natural language instructions and navigate through complex environments. Existing MLLM-based VLN methods primarily rely on imitation learning (IL) and often use…

机器人学 · 计算机科学 2025-09-17 Zekai Zhang , Weiye Zhu , Hewei Pan , Xiangchen Wang , Rongtao Xu , Xing Sun , Feng Zheng

Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following human instruction. Compared to ground-based VLN, aerial VLN requires the agent to decide the…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Ganlong Zhao , Guanbin Li , Jia Pan , Yizhou Yu

In the Vision-and-Language Navigation (VLN) field, agents are tasked with navigating real-world scenes guided by linguistic instructions. Enabling the agent to adhere to instructions throughout the process of navigation represents a…

人工智能 · 计算机科学 2024-05-28 Wen Hanlin

Recent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities. VLMs, typically pre-trained on vision-language understanding tasks, provide rich…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Jianke Zhang , Yanjiang Guo , Yucheng Hu , Xiaoyu Chen , Xiang Zhu , Jianyu Chen

Large language models such as BERT and the GPT series started a paradigm shift that calls for building general-purpose models via pre-training on large datasets, followed by fine-tuning on task-specific datasets. There is now a plethora of…

计算与语言 · 计算机科学 2023-06-13 Jeremy Gwinnup , Kevin Duh

Multi-modal representation learning by pretraining has become an increasing interest due to its easy-to-use and potential benefit for various Visual-and-Language~(V-L) tasks. However its requirement of large volume and high-quality…

多媒体 · 计算机科学 2020-12-09 Jia Guo , Chen Zhu , Yilun Zhao , Heda Wang , Yao Hu , Xiaofei He , Deng Cai

Pre-training has been adopted in a few of recent works for Vision-and-Language Navigation (VLN). However, previous pre-training methods for VLN either lack the ability to predict future actions or ignore the trajectory contexts, which are…

计算机视觉与模式识别 · 计算机科学 2022-03-23 Yanyuan Qiao , Yuankai Qi , Yicong Hong , Zheng Yu , Peng Wang , Qi Wu

Vision-and-Language Navigation (VLN), as a widely discussed research direction in embodied intelligence, aims to enable embodied agents to navigate in complicated visual environments through natural language commands. Most existing VLN…

计算机视觉与模式识别 · 计算机科学 2024-11-14 Youzhi Liu , Fanglong Yao , Yuanchang Yue , Guangluan Xu , Xian Sun , Kun Fu

Vision-language navigation (VLN), which entails an agent to navigate 3D environments following human instructions, has shown great advances. However, current agents are built upon panoramic observations, which hinders their ability to…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Rui Liu , Xiaohan Wang , Wenguan Wang , Yi Yang

Vision-and-Language Navigation (VLN) requires an agent to find a path to a remote location on the basis of natural-language instructions and a set of photo-realistic panoramas. Most existing methods take the words in the instructions and…

计算与语言 · 计算机科学 2021-08-26 Yuankai Qi , Zizheng Pan , Yicong Hong , Ming-Hsuan Yang , Anton van den Hengel , Qi Wu

Inspired by the general Vision-and-Language Navigation (VLN) task, aerial VLN has attracted widespread attention, owing to its significant practical value in applications such as logistics delivery and urban inspection. However, existing…

机器人学 · 计算机科学 2026-04-13 Chengjie Fan , Cong Pan , Zijian Liu , Ningzhong Liu , Jie Qin

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li