English
Related papers

Related papers: GeoPredict: Leveraging Predictive Kinematics and 3…

200 papers

This technical report outlines the top-ranking solution for RoboSense 2025: Track 3, achieving state-of-the-art performance on 3D object detection under various sensor placements. Our submission utilizes GBlobs, a local point cloud feature…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Dušan Malić , Christian Fruhwirth-Reisinger , Alexander Prutsch , Wei Lin , Samuel Schulter , Horst Possegger

Deep learning approaches for spatio-temporal prediction problems such as crowd-flow prediction assumes data to be of fixed and regular shaped tensor and face challenges of handling irregular, sparse data tensor. This poses limitations in…

Machine Learning · Computer Science 2022-11-29 Syed Mohammed Arshad Zaidi , Varun Chandola , EunHye Yoo

Vision-Language-Action (VLA) models have gained much attention from the research community thanks to their strength in translating multimodal observations with linguistic instructions into desired robotic actions. Despite their…

In this work, we introduce HoloBrain-0, a comprehensive Vision-Language-Action (VLA) framework that bridges the gap between foundation model research and reliable real-world robot deployment. The core of our system is a novel VLA…

Training end-to-end policies from image data to directly predict navigation actions for robotic systems has proven inherently difficult. Existing approaches often suffer from either the sim-to-real gap during policy transfer or a limited…

Robotics · Computer Science 2026-03-17 Lazar Milikic , Manthan Patel , Jonas Frey

In dynamic environments such as warehouses, hospitals, and homes, robots must seamlessly transition between gross motion and precise manipulations to complete complex tasks. However, current Vision-Language-Action (VLA) frameworks, largely…

Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property, relying only on the current observation and thus suffering…

Vision-Language-Action (VLA) models are driving a revolution in robotics, enabling machines to understand instructions and interact with the physical world. This field is exploding with new models and datasets, making it both exciting and…

Navigating unfamiliar environments presents significant challenges for household robots, requiring the ability to recognize and reason about novel decoration and layout. Existing reinforcement learning methods cannot be directly transferred…

Robotics · Computer Science 2025-02-20 Yiran Qin , Ao Sun , Yuze Hong , Benyou Wang , Ruimao Zhang

While Vision-Language-Action (VLA) models have achieved remarkable success in ground-based embodied intelligence, their application to Aerial Manipulation Systems (AMS) remains a largely unexplored frontier. The inherent characteristics of…

Robotics · Computer Science 2026-02-04 Jianli Sun , Bin Tian , Qiyao Zhang , Chengxiang Li , Zihan Song , Zhiyong Cui , Yisheng Lv , Yonglin Tian

Geolocation is now a vital aspect of modern life, offering numerous benefits but also presenting serious privacy concerns. The advent of large vision-language models (LVLMs) with advanced image-processing capabilities introduces new risks,…

Cryptography and Security · Computer Science 2024-08-20 Yi Liu , Junchen Ding , Gelei Deng , Yuekang Li , Tianwei Zhang , Weisong Sun , Yaowen Zheng , Jingquan Ge , Yang Liu

A robot's ability to anticipate the 3D action target location of a hand's movement from egocentric videos can greatly improve safety and efficiency in human-robot interaction (HRI). While previous research predominantly focused on semantic…

Robotics · Computer Science 2024-03-11 Irving Fang , Yuzhong Chen , Yifan Wang , Jianghan Zhang , Qiushi Zhang , Jiali Xu , Xibo He , Weibo Gao , Hao Su , Yiming Li , Chen Feng

Despite the recent advancements of vision-language-action (VLA) models on a variety of robotics tasks, they suffer from critical issues such as poor generalizability to unseen tasks, due to their reliance on behavior cloning exclusively…

Robotics · Computer Science 2025-02-05 Zijian Zhang , Kaiyuan Zheng , Zhaorun Chen , Joel Jang , Yi Li , Siwei Han , Chaoqi Wang , Mingyu Ding , Dieter Fox , Huaxiu Yao

Embodied intelligence systems, which enhance agent capabilities through continuous environment interactions, have garnered significant attention from both academia and industry. Vision-Language-Action models, inspired by advancements in…

Robotics · Computer Science 2025-11-13 Haoran Li , Yuhui Chen , Wenbo Cui , Weiheng Liu , Kai Liu , Mingcai Zhou , Zhengtao Zhang , Dongbin Zhao

This paper presents GeoAgent, a model capable of reasoning closely with humans and deriving fine-grained address conclusions. Previous RL-based methods have achieved breakthroughs in performance and interpretability but still remain…

Artificial Intelligence · Computer Science 2026-02-16 Modi Jin , Yiming Zhang , Boyuan Sun , Dingwen Zhang , MingMing Cheng , Qibin Hou

Transformer-based general visual geometry frameworks have shown promising performance in camera pose estimation and 3D scene understanding. Recent advancements in Visual Geometry Grounded Transformer (VGGT) models have shown great promise…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Yangfan Xu , Lilian Zhang , Xiaofeng He , Pengdong Wu , Wenqi Wu , Jun Mao

Vision-Language-Action~(VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate transferring relevant experience across objects, scenes, and…

Robotics · Computer Science 2026-05-29 Shengyu Si , Yuanzhuo Lu , Ruimeng Yang , Ziyi Ye , Zuxuan Wu , Yu-Gang Jiang

In obstacle-dense scenarios, providing safe guidance for mobile robots is critical to improve the safe maneuvering capability. However, the guidance provided by standard guiding vector fields (GVFs) may limit the motion capability due to…

Robotics · Computer Science 2025-09-23 Yang Lu , Weijia Yao , Yongqian Xiao , Xinglong Zhang , Xin Xu , Yaonan Wang , Dingbang Xiao

Humanoid robots exhibit significant potential in executing diverse human-level skills. However, current research predominantly relies on data-driven approaches that necessitate extensive training datasets to achieve robust multimodal…

Robotics · Computer Science 2025-12-25 Xuetao Li , Wenke Huang , Nengyuan Pan , Kaiyan Zhao , Songhua Yang , Yiming Wang , Mengde Li , Mang Ye , Jifeng Xuan , Miao Li

Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robotic manipulation. A common design in current architectures maps language instructions and visual observations to actions in a single forward pass.…

Robotics · Computer Science 2026-05-26 Weilong Guo , Yuchen Wang , Renping Zhou , Yunfeng Zhang , Rui Fang , Yuyang Pang , Wenda Xu , Gao Huang
‹ Prev 1 8 9 10 Next ›