English
Related papers

Related papers: GoViG: Goal-Conditioned Visual Navigation Instruct…

200 papers

Advancements at the intersection of computer vision and natural language processing are crucial for applications like assistive tech, multimedia querying, and robotics. This dissertation proposes novel architectures to improve intelligent…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Van Quang Nguyen

Autonomous driving systems have made significant advances in Q&A, perception, prediction, and planning based on local visual information, yet they struggle to incorporate broader navigational context that human drivers routinely utilize. We…

Robotics · Computer Science 2025-11-04 Qucheng Peng , Chen Bai , Guoxiang Zhang , Bo Xu , Xiaotong Liu , Xiaoyin Zheng , Chen Chen , Cheng Lu

To effectively engage in human society, the ability to adapt, filter information, and make informed decisions in ever-changing situations is critical. As robots and intelligent agents become more integrated into human life, there is a…

Artificial Intelligence · Computer Science 2025-11-13 Mingyang Mao , Mariela M. Perez-Cabarcas , Utteja Kallakuri , Nicholas R. Waytowich , Xiaomin Lin , Tinoosh Mohsenin

Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them separately, struggling with limited high-quality data and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Xueyun Tian , Wei Li , Bingbing Xu , Yige Yuan , Yuanzhuo Wang , Huawei Shen

Reasoning about visual relationships is central to how humans interpret the visual world. This task remains challenging for current deep learning algorithms since it requires addressing three key technical problems jointly: 1) identifying…

Computer Vision and Pattern Recognition · Computer Science 2022-06-14 Xiaojian Ma , Weili Nie , Zhiding Yu , Huaizu Jiang , Chaowei Xiao , Yuke Zhu , Song-Chun Zhu , Anima Anandkumar

This paper presents instruct-imagen, a model that tackles heterogeneous image generation tasks and generalizes across unseen tasks. We introduce *multi-modal instruction* for image generation, a task representation articulating a range of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 Hexiang Hu , Kelvin C. K. Chan , Yu-Chuan Su , Wenhu Chen , Yandong Li , Kihyuk Sohn , Yang Zhao , Xue Ben , Boqing Gong , William Cohen , Ming-Wei Chang , Xuhui Jia

The autonomous synthesis of deep research reports represents a critical frontier for Large Language Models (LLMs), demanding sophisticated information orchestration and non-linear narrative logic. Current approaches rely on rigid predefined…

Multiagent Systems · Computer Science 2026-04-22 Kuo Tian , Pengfei Sun , Zhen Wu , Junran Ding , Xinyu Dai

Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven. We present Goal-Driven Data Optimization (GDO), a framework that computes…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Rujie Wu , Haozhe Zhao , Hai Ci , Yizhou Wang

Vision-Language Navigation (VLN) is a core challenge in embodied AI, requiring agents to navigate real-world environments using natural language instructions. Current language model-based navigation systems operate on discrete topological…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Zhangyang Qi , Zhixiong Zhang , Yizhou Yu , Jiaqi Wang , Hengshuang Zhao

Recent advances in large language models (LLMs) have empowered AI agents capable of performing various sequential decision-making tasks. However, effectively guiding LLMs to perform well in unfamiliar domains like web navigation, where they…

Computation and Language · Computer Science 2024-12-04 Yao Fu , Dong-Ki Kim , Jaekyeom Kim , Sungryull Sohn , Lajanugen Logeswaran , Kyunghoon Bae , Honglak Lee

Equipping embodied agents with the ability to reason about tasks, foresee physical outcomes, and generate precise actions is essential for general-purpose manipulation. While recent Vision-Language-Action (VLA) models have leveraged…

Embodied navigation holds significant promise for real-world applications such as last-mile delivery. However, most existing approaches are confined to either indoor or outdoor environments and rely heavily on strong assumptions, such as…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Yuxiang Zhao , Yirong Yang , Yanqing Zhu , Yanfen Shen , Chiyu Wang , Zhining Gu , Pei Shi , Wei Guo , Mu Xu

Learning visuomotor control policies in robotic systems is a fundamental problem when aiming for long-term behavioral autonomy. Recent supervised-learning-based vision and motion perception systems, however, are often separately built with…

Robotics · Computer Science 2020-06-17 Marvin Chancán , Michael Milford

A multi-modal framework to generate user intention distributions when operating a mobile vehicle is proposed in this work. The model learns from past observed trajectories and leverages traversability information derived from the visual…

Robotics · Computer Science 2022-03-17 Kavindie Katuwandeniya , Stefan H. Kiss , Lei Shi , Jaime Valls Miro

Vision-and-Language Navigation (VLN) requires agents to interpret natural language instructions and act coherently in visually rich environments. However, most existing methods rely on reactive state-action mappings without explicitly…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Weiye Zhu , Zekai Zhang , Xiangchen Wang , Hewei Pan , Teng Wang , Tiantian Geng , Rongtao Xu , Feng Zheng

Visual understanding is inherently intention-driven - humans selectively focus on different regions of a scene based on their goals. Recent advances in large multimodal models (LMMs) enable flexible expression of such intentions through…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Zhangquan Chen , Xufang Luo , Dongsheng Li

Embodied navigation presents a core challenge for intelligent robots, requiring the comprehension of visual environments, natural language instructions, and autonomous exploration. Existing models often fall short in offering a unified…

Robotics · Computer Science 2026-01-08 Xinda Xue , Junjun Hu , Minghua Luo , Shichao Xie , Jintao Chen , Zixun Xie , Kuichen Quan , Wei Guo , Mu Xu , Zedong Chu

Visual-Inertial Odometry (VIO) usually suffers from drifting over long-time runs, the accuracy is easily affected by dynamic objects. We propose DynaVIG, a navigation and object tracking system based on the integration of Monocular Vision,…

Robotics · Computer Science 2022-11-29 Ronghe Jin , Yan Wang , Zhi Gao , Xiaoji Niu , Li-Ta Hsu , Jingnan Liu

Models capable of "thinking with images" by dynamically grounding their reasoning in visual evidence represent a major leap in multimodal AI. However, replicating and advancing this ability is non-trivial, with current methods often trapped…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zhaoyang Wei , Wenchao Ding , Yanchao Hao , Xi Chen

Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages recent advancements in…

Computation and Language · Computer Science 2025-06-11 Mohamed Gado , Towhid Taliee , Muhammad Memon , Dmitry Ignatov , Radu Timofte