English
Related papers

Related papers: SaPaVe: Towards Active Perception and Manipulation…

200 papers

Vision-Language-Action (VLA) models provide a promising paradigm for robot learning by integrating visual perception with language-guided policy learning. However, most existing approaches rely on 2D visual inputs to perform actions in 3D…

Robotics · Computer Science 2025-12-16 Yicheng Feng , Wanpeng Zhang , Ye Wang , Hao Luo , Haoqi Yuan , Sipeng Zheng , Zongqing Lu

"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored. We present CaP-X, an…

Autonomous vehicles (AV) require that neural networks used for perception be robust to different viewpoints if they are to be deployed across many types of vehicles without the repeated cost of data collection and labeling for each. AV…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Tzofi Klinghoffer , Jonah Philion , Wenzheng Chen , Or Litany , Zan Gojcic , Jungseock Joo , Ramesh Raskar , Sanja Fidler , Jose M. Alvarez

In this paper we present a neurosymbolic architecture for coupling language-guided visual reasoning with robot manipulation. A non-expert human user can prompt the robot using unconstrained natural language, providing a referring expression…

Robotics · Computer Science 2025-12-16 Georgios Tziafas , Hamidreza Kasaei

Robotic manipulation systems operating in diverse, dynamic environments must exhibit three critical abilities: multitask interaction, generalization to unseen scenarios, and spatial memory. While significant progress has been made in…

Robotics · Computer Science 2025-07-15 Haoquan Fang , Markus Grotz , Wilbert Pumacay , Yi Ru Wang , Dieter Fox , Ranjay Krishna , Jiafei Duan

Learning predictive models from interaction with the world allows an agent, such as a robot, to learn about how the world works, and then use this learned model to plan coordinated sequences of actions to bring about desired outcomes.…

Machine Learning · Computer Science 2020-01-01 Karl Schmeckpeper , Annie Xie , Oleh Rybkin , Stephen Tian , Kostas Daniilidis , Sergey Levine , Chelsea Finn

Vision-Language-Action (VLA) models and world models have recently emerged as promising paradigms for general-purpose robotic intelligence, yet their progress is hindered by the lack of reliable evaluation protocols that reflect real-world…

Accurate and robust state estimation is critical for autonomous navigation of robot teams. This task is especially challenging for large groups of size, weight, and power (SWAP) constrained aerial robots operating in perceptually-degraded…

Robotics · Computer Science 2023-05-30 Igor Spasojevic , Xu Liu , Alejandro Ribeiro , George J. Pappas , Vijay Kumar

Compliance plays a crucial role in manipulation, as it balances between the concurrent control of position and force under uncertainties. Yet compliance is often overlooked by today's visuomotor policies that solely focus on position…

We present V-JEPA 2.1, a family of self-supervised models that learn dense, high-quality visual representations for both images and videos while retaining strong global scene understanding. The approach combines four key components. First,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Lorenzo Mur-Labadia , Matthew Muckley , Amir Bar , Mido Assran , Koustuv Sinha , Mike Rabbat , Yann LeCun , Nicolas Ballas , Adrien Bardes

Collaborative perception leverages rich visual observations from multiple robots to extend a single robot's perception ability beyond its field of view. Many prior works receive messages broadcast from all collaborators, leading to a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Suozhi Huang , Juexiao Zhang , Yiming Li , Chen Feng

Existing AI systems for modeling human behavior operate at the level of individuals or detect events after they occur. As a result, they systematically fail to capture the collective dynamics that determine whether a group remains stable or…

Artificial Intelligence · Computer Science 2026-05-14 Helene Malyutina

Language-guided robotic manipulation is a challenging task that requires an embodied agent to follow abstract user instructions to accomplish various complex manipulation tasks. Previous work trivially fitting the data without revealing the…

Robotics · Computer Science 2024-10-17 Kaidong Zhang , Pengzhen Ren , Bingqian Lin , Junfan Lin , Shikui Ma , Hang Xu , Xiaodan Liang

Visual actionable affordance has emerged as a transformative approach in robotics, focusing on perceiving interaction areas prior to manipulation. Traditional methods rely on pixel sampling to identify successful interaction samples or…

Robotics · Computer Science 2025-10-10 Taewhan Kim , Hojin Bae , Zeming Li , Xiaoqi Li , Iaroslav Ponomarenko , Ruihai Wu , Hao Dong

A robot's instantaneous sensory observations do not always reveal task-relevant state information. Under such partial observability, optimal behavior typically involves explicitly acting to gain the missing information. Today's standard…

Robotic manipulation in unstructured environments requires systems that can generalize across diverse tasks while maintaining robust and reliable performance. We introduce {GVF-TAPE}, a closed-loop framework that combines generative visual…

Robotics · Computer Science 2025-09-03 Chuye Zhang , Xiaoxiong Zhang , Wei Pan , Linfang Zheng , Wei Zhang

Bimanual robotic manipulation is an emerging and critical topic in the robotics community. Previous works primarily rely on integrated control models that take the perceptions and states of both arms as inputs to directly predict their…

Robotics · Computer Science 2025-11-05 Jian-Jian Jiang , Xiao-Ming Wu , Yi-Xiang He , Ling-An Zeng , Yi-Lin Wei , Dandan Zhang , Wei-Shi Zheng

Natural Human-Robot Interaction (N-HRI) requires robots to recognize human actions at varying distances and states, regardless of whether the robot itself is in motion or stationary. This setup is more flexible and practical than…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Ziyi Wang , Peiming Li , Hong Liu , Zhichao Deng , Can Wang , Jun Liu , Junsong Yuan , Mengyuan Liu

We introduce Vocal Sandbox, a framework for enabling seamless human-robot collaboration in situated environments. Systems in our framework are characterized by their ability to adapt and continually learn at multiple levels of abstraction…

Robotics · Computer Science 2024-11-06 Jennifer Grannen , Siddharth Karamcheti , Suvir Mirchandani , Percy Liang , Dorsa Sadigh

Building on the advancements of Large Language Models (LLMs) and Vision Language Models (VLMs), recent research has introduced Vision-Language-Action (VLA) models as an integrated solution for robotic manipulation tasks. These models take…

Robotics · Computer Science 2024-10-08 Zhijie Wang , Zhehua Zhou , Jiayang Song , Yuheng Huang , Zhan Shu , Lei Ma
‹ Prev 1 3 4 5 6 7 10 Next ›