English

Kefa: A Knowledge Enhanced and Fine-grained Aligned Speaker for Navigation Instruction Generation

Computer Vision and Pattern Recognition 2023-07-26 v1

Abstract

We introduce a novel speaker model \textsc{Kefa} for navigation instruction generation. The existing speaker models in Vision-and-Language Navigation suffer from the large domain gap of vision features between different environments and insufficient temporal grounding capability. To address the challenges, we propose a Knowledge Refinement Module to enhance the feature representation with external knowledge facts, and an Adaptive Temporal Alignment method to enforce fine-grained alignment between the generated instructions and the observation sequences. Moreover, we propose a new metric SPICE-D for navigation instruction evaluation, which is aware of the correctness of direction phrases. The experimental results on R2R and UrbanWalk datasets show that the proposed KEFA speaker achieves state-of-the-art instruction generation performance for both indoor and outdoor scenes.

Keywords

Cite

@article{arxiv.2307.13368,
  title  = {Kefa: A Knowledge Enhanced and Fine-grained Aligned Speaker for Navigation Instruction Generation},
  author = {Haitian Zeng and Xiaohan Wang and Wenguan Wang and Yi Yang},
  journal= {arXiv preprint arXiv:2307.13368},
  year   = {2023}
}

Comments

10 pages, 4 figures

R2 v1 2026-06-28T11:39:29.590Z