English
Related papers

Related papers: Flex: End-to-End Text-Instructed Visual Navigation…

200 papers

Many of the existing methods for learning joint embedding of images and text use only supervised information from paired images and its textual attributes. Taking advantage of the recent success of unsupervised learning in deep neural…

Computer Vision and Pattern Recognition · Computer Science 2017-03-21 Yao-Hung Hubert Tsai , Liang-Kang Huang , Ruslan Salakhutdinov

Learning controllers that reproduce legged locomotion in nature has been a long-time goal in robotics and computer graphics. While yielding promising results, recent approaches are not yet flexible enough to be applicable to legged systems…

Robotics · Computer Science 2022-07-26 Daniel Ordonez-Apraez , Antonio Agudo , Francesc Moreno-Noguer , Mario Martin

Autonomous navigation is a fundamental task for robot vacuum cleaners in indoor environments. Since their core function is to clean entire areas, robots inevitably encounter dead zones in cluttered and narrow scenarios. Existing planning…

Robotics · Computer Science 2025-03-06 Han Zheng , Jiale Zhang , Mingyang Jiang , Peiyuan Liu , Danni Liu , Tong Qin , Ming Yang

How do video understanding models acquire their answers? Although current Vision Language Models (VLMs) reason over complex scenes with diverse objects, action performances, and scene dynamics, understanding and controlling their internal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Alexandros Stergiou

A rich representation is key to general robotic manipulation, but existing approaches to representation learning require large amounts of multimodal demonstrations. In this work we propose PLEX, a transformer-based architecture that learns…

Augmenting pretrained language models (LMs) with a vision encoder (e.g., Flamingo) has obtained the state-of-the-art results in image-to-text generation. However, these models store all the knowledge within their parameters, thus often…

Computer Vision and Pattern Recognition · Computer Science 2023-10-24 Zhuolin Yang , Wei Ping , Zihan Liu , Vijay Korthikanti , Weili Nie , De-An Huang , Linxi Fan , Zhiding Yu , Shiyi Lan , Bo Li , Ming-Yu Liu , Yuke Zhu , Mohammad Shoeybi , Bryan Catanzaro , Chaowei Xiao , Anima Anandkumar

In this paper, we present Language Model as Visual Explainer LVX, a systematic approach for interpreting the internal workings of vision models using a tree-structured linguistic explanation, without the need for model training. Central to…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Xingyi Yang , Xinchao Wang

Recent Large Vision Language Models (LVLMs) demonstrate promising capabilities in unifying visual understanding and generative modeling, enabling both accurate content understanding and flexible editing. However, current approaches treat…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Fan Yang , Yousong Zhu , Xin Li , Yufei Zhan , Hongyin Zhao , Shurong Zheng , Yaowei Wang , Ming Tang , Jinqiao Wang

End-to-end autonomous driving is a fully differentiable machine learning system that takes raw sensor input data and other metadata as prior information and directly outputs the ego vehicle's control signals or planned trajectories. This…

Robotics · Computer Science 2023-12-01 Apoorv Singh

Legged locomotion is arguably the most suited and versatile mode to deal with natural or unstructured terrains. Intensive research into dynamic walking and running controllers has recently yielded great advances, both in the optimal control…

Robotics · Computer Science 2024-09-18 Raghav Soni , Daniel Harnack , Hannah Isermann , Sotaro Fushimi , Shivesh Kumar , Frank Kirchner

Current end-to-end deep Reinforcement Learning (RL) approaches require jointly learning perception, decision-making and low-level control from very sparse reward signals and high-dimensional inputs, with little capability of incorporating…

Machine Learning · Computer Science 2019-10-10 Vibhavari Dasagi , Robert Lee , Serena Mou , Jake Bruce , Niko Sünderhauf , Jürgen Leitner

General-purpose navigation in challenging environments remains a significant problem in robotics, with current state-of-the-art approaches facing myriad limitations. Classical approaches struggle with cluttered settings and require…

Robotics · Computer Science 2025-07-24 Wei Liu , Huihua Zhao , Chenran Li , Joydeep Biswas , Billy Okal , Pulkit Goyal , Yan Chang , Soha Pouya

There is substantial interest in developing artificial intelligence systems to support radiologists across tasks ranging from segmentation to report generation. Existing computed tomography (CT) foundation models have largely focused on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Rubén Moreno-Aguado , Alba Magallón , Victor Moreno , Yingying Fang , Guang Yang

Visual loco-manipulation of arbitrary objects in the wild with humanoid robots requires accurate end-effector (EE) control and a generalizable understanding of the scene via visual inputs (e.g., RGB-D images). Existing approaches are based…

Robotics · Computer Science 2026-02-25 Runpei Dong , Ziyan Li , Xialin He , Saurabh Gupta

Traditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate strong scene…

Robotics · Computer Science 2024-12-18 Qi Sun , Pengfei Hong , Tej Deep Pala , Vernon Toh , U-Xuan Tan , Deepanway Ghosal , Soujanya Poria

Vision-language models (VLMs) like CLIP have demonstrated remarkable applicability across a variety of downstream tasks, including zero-shot image classification. Recently, the use of prompts or adapters for efficient transfer learning…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Yongjin Yang , Jongwoo Ko , Se-Young Yun

Unifying text detection and text recognition in an end-to-end training fashion has become a new trend for reading text in the wild, as these two tasks are highly relevant and complementary. In this paper, we investigate the problem of scene…

Computer Vision and Pattern Recognition · Computer Science 2019-08-23 Minghui Liao , Pengyuan Lyu , Minghang He , Cong Yao , Wenhao Wu , Xiang Bai

Vision-Language-Action (VLA) models have recently enabled embodied agents to perform increasingly complex tasks by jointly reasoning over visual, linguistic, and motor modalities. However, we find that the prevailing notion of…

Machine Learning · Computer Science 2026-03-20 Zhuofan Li , Hongkun Yang , Zhenyang Chen , Yangxuan Chen , Yingyan , Lin , Chaojian Li

With the development of visual-language models (VLM) in downstream task applications, test-time adaptation methods based on VLM have attracted increasing attention for their ability to address changes distribution in test-time. Although…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Chenhao Ding , Xinyuan Gao , Songlin Dong , Yuhang He , Qiang Wang , Xiang Song , Alex Kot , Yihong Gong

In this work we present a novel end-to-end framework for tracking and classifying a robot's surroundings in complex, dynamic and only partially observable real-world environments. The approach deploys a recurrent neural network to filter an…

Machine Learning · Computer Science 2016-04-20 Peter Ondruska , Julie Dequaire , Dominic Zeng Wang , Ingmar Posner