English
Related papers

Related papers: Learning Visual Locomotion with Cross-Modal Superv…

200 papers

Achieving monocular camera localization within pre-built LiDAR maps can bypass the simultaneous mapping process of visual SLAM systems, potentially reducing the computational overhead of autonomous localization. To this end, one of the key…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Gongxin Yao , Xinyang Li , Luowei Fu , Yu Pan

Visual imitation learning frameworks allow robots to learn manipulation skills from expert demonstrations. While existing approaches mainly focus on policy design, they often neglect the structure and capacity of visual encoders, limiting…

Robotics · Computer Science 2025-09-24 Shijia Ge , Yinxin Zhang , Shuzhao Xie , Weixiang Zhang , Mingcai Zhou , Zhi Wang

We present MaCLR, a novel method to explicitly perform cross-modal self-supervised video representations learning from visual and motion modalities. Compared to previous video representation learning methods that mostly focus on learning…

Computer Vision and Pattern Recognition · Computer Science 2022-07-21 Fanyi Xiao , Joseph Tighe , Davide Modolo

This paper presents a curriculum-based reinforcement learning framework for training precise and high-performance jumping policies for the robot `Olympus'. Separate policies are developed for vertical and horizontal jumps, leveraging a…

Robotics · Computer Science 2025-10-29 Jørgen Anker Olsen , Lars Rønhaug Pettersen , Kostas Alexis

We present LoTIS, a model for visual navigation that provides robot-agnostic image-space guidance by localizing a reference RGB trajectory in the robot's current view, without requiring camera calibration, poses, or robot-specific training.…

Spatial reasoning from monocular images is essential for autonomous driving, yet current Vision-Language Models (VLMs) still struggle with fine-grained geometric perception, particularly under large scale variation and ambiguous object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yanchun Cheng , Rundong Wang , Xulei Yang , Alok Prakash , Daniela Rus , Marcelo H Ang , ShiJie Li

Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms. Existing cross-embodiment approaches try to map human motion to robot actions, but often…

We propose an approach for reconstructing free-moving object from a monocular RGB video. Most existing methods either assume scene prior, hand pose prior, object category pose prior, or rely on local optimization with multiple sequence…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Haixin Shi , Yinlin Hu , Daniel Koguciuk , Juan-Ting Lin , Mathieu Salzmann , David Ferstl

Seeing-eye robots are very useful tools for guiding visually impaired people, potentially producing a huge societal impact given the low availability and high cost of real guide dogs. Although a few seeing-eye robot systems have already…

Robotics · Computer Science 2023-10-13 David DeFazio , Eisuke Hirota , Shiqi Zhang

Robotic manipulation of cloth is a challenging task due to the high dimensionality of the configuration space and the complexity of dynamics affected by various material properties. The effect of complex dynamics is even more pronounced in…

Robotics · Computer Science 2023-02-09 Julius Hietala , David Blanco-Mulero , Gokhan Alcan , Ville Kyrki

In this work, we study vision-based end-to-end reinforcement learning on vehicle control problems, such as lane following and collision avoidance. Our controller policy is able to control a small-scale robot to follow the right-hand lane of…

Machine Learning · Computer Science 2020-12-15 András Kalapos , Csaba Gór , Róbert Moni , István Harmati

Autonomous ground vehicles have been designed for the purpose of that relies on ranging and bearing information received from forward looking camera on the Formation control . A visual guidance control algorithm is designed where real time…

Computer Vision and Pattern Recognition · Computer Science 2015-01-08 S. M. Vaitheeswaran , Bharath M. K. , Gokul M

Humans learn locomotion through visual observation, interpreting visual content first before imitating actions. However, state-of-the-art humanoid locomotion systems rely on either curated motion capture trajectories or sparse text…

Current visual navigation systems often treat the environment as static, lacking the ability to adaptively interact with obstacles. This limitation leads to navigation failure when encountering unavoidable obstructions. In response, we…

Robotics · Computer Science 2024-08-13 Philipp Schoch , Fan Yang , Yuntao Ma , Stefan Leutenegger , Marco Hutter , Quentin Leboutet

Achieving precise positioning of the mobile manipulator's base is essential for successful manipulation actions that follow. Most of the RGB-based navigation systems only guarantee coarse, meter-level accuracy, making them less suitable for…

Robotics · Computer Science 2026-02-17 Tzu-Hsien Lee , Fidan Mahmudova , Karthik Desingh

Simulation can be a powerful tool for understanding machine learning systems and designing methods to solve real-world problems. Training and evaluating methods purely in simulation is often "doomed to succeed" at the desired task in a…

Computer Vision and Pattern Recognition · Computer Science 2018-12-14 Alex Bewley , Jessica Rigley , Yuxuan Liu , Jeffrey Hawke , Richard Shen , Vinh-Dieu Lam , Alex Kendall

Traversing risky terrains with sparse footholds presents significant challenges for legged robots, requiring precise foot placement in safe areas. To acquire comprehensive exteroceptive information, prior studies have employed motion…

Robotics · Computer Science 2025-03-04 Ruiqi Yu , Qianshi Wang , Yizhen Wang , Zhicheng Wang , Jun Wu , Qiuguo Zhu

Recent work has demonstrated the success of reinforcement learning (RL) for training bipedal locomotion policies for real robots. This prior work, however, has focused on learning joint-coordination controllers based on an objective of…

Robotics · Computer Science 2021-05-07 Helei Duan , Jeremy Dao , Kevin Green , Taylor Apgar , Alan Fern , Jonathan Hurst

One powerful paradigm in visual navigation is to predict actions from observations directly. Training such an end-to-end system allows representations useful for downstream tasks to emerge automatically. However, the lack of inductive bias…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Yanwei Wang , Ching-Yun Ko , Pulkit Agrawal

The current practice of dexterous manipulation generally relies on a single wrist-mounted view, which is often occluded and limits performance on tasks requiring multi-view perception. In this work, we present FingerViP, a learning system…

Robotics · Computer Science 2026-05-06 Zhen Zhang , Weinan Wang , Hejia Sun , Qingpeng Ding , Xiangyu Chu , Guoxin Fang , K. W. Samuel Au