English
Related papers

Related papers: Foundational World Models Accurately Detect Bimanu…

200 papers

Spatio-physical reasoning, a foundation capability for understanding the real physics world, is a critical step towards building robust world models. While recent vision language models (VLMs) have shown remarkable progress in specialized…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Tiancheng Han , Yunfei Gao , Yong Li , Wuzhou Yu , Qiaosheng Zhang , Wenqi Shao

Humanoid robots are engineered to navigate terrains akin to those encountered by humans, which necessitates human-like locomotion and perceptual abilities. Currently, the most reliable controllers for humanoid motion rely exclusively on…

Robotics · Computer Science 2025-04-03 Wandong Sun , Baoshi Cao , Long Chen , Yongbo Su , Yang Liu , Zongwu Xie , Hong Liu

Before autonomous systems can be deployed in safety-critical applications, we must be able to understand and verify the safety of these systems. For cases where the risk or cost of real-world testing is prohibitive, we propose a…

Robotics · Computer Science 2023-09-18 Charles Dawson , Chuchu Fan

We consider the problem of detecting, in the visual sensing data stream of an autonomous mobile robot, semantic patterns that are unusual (i.e., anomalous) with respect to the robot's previous experience in similar environments. These…

Large pretrained foundation models demonstrate exceptional performance and, in some high-stakes applications, even surpass human experts. However, most of these models are currently evaluated primarily on prediction accuracy, overlooking…

Machine Learning · Computer Science 2024-11-08 Tang Li , Mengmeng Ma , Xi Peng

Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of using generated videos as scalable supervision for robot learning. However, for embodied manipulation,…

While recent advancements in robotic manipulation video synthesis have shown promise, significant challenges persist in ensuring effective instruction-following and achieving high visual quality. Recent methods, like RoboDreamer, utilize…

Robotics · Computer Science 2025-04-24 Ying Li , Xiaobao Wei , Xiaowei Chi , Yuming Li , Zhongyu Zhao , Hao Wang , Ningning Ma , Ming Lu , Shanghang Zhang

We present a fully autonomous real-world RL framework for mobile manipulation that can learn policies without extensive instrumentation or human supervision. This is enabled by 1) task-relevant autonomy, which guides exploration towards…

Robotics · Computer Science 2024-10-01 Russell Mendonca , Emmanuel Panov , Bernadette Bucher , Jiuguang Wang , Deepak Pathak

Anomaly detection in videos is a significant yet challenging problem. Previous approaches based on deep neural networks employ either reconstruction-based or prediction-based approaches. Nevertheless, existing reconstruction-based methods…

Computer Vision and Pattern Recognition · Computer Science 2023-01-31 Yizhou Wang , Can Qin , Yue Bai , Yi Xu , Xu Ma , Yun Fu

Quadrupedal robots have demonstrated impressive locomotion capabilities in complex environments, but equipping them with autonomous versatile manipulation skills in a scalable way remains a significant challenge. In this work, we introduce…

Visuomotor policies trained via behavior cloning are vulnerable to covariate shift, where small deviations from expert trajectories can compound into failure. Common strategies to mitigate this issue involve expanding the training…

Robotics · Computer Science 2025-08-11 Zhanyi Sun , Shuran Song

Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introduces visual redundancy and accumulates errors, causing…

Robotics · Computer Science 2026-03-16 Minghao Jin , Mozheng Liao , Mingfei Han , Zhihui Li , Xiaojun Chang

Understanding action correspondence between humans and robots is essential for evaluating alignment in decision-making, particularly in human-robot collaboration and imitation learning within unstructured environments. We propose a…

Robotics · Computer Science 2025-04-17 Azizul Zahid , Jie Fan , Farong Wang , Ashton Dy , Sai Swaminathan , Fei Liu

Interactive perception enables robots to manipulate the environment and objects to bring them into states that benefit the perception process. Deformable objects pose challenges to this due to significant manipulation difficulty and…

Robot failures in human-centered environments are inevitable. Therefore, the ability of robots to explain such failures is paramount for interacting with humans to increase trust and transparency. To achieve this skill, the main challenges…

Robotics · Computer Science 2023-03-21 Maximilian Diehl , Karinne Ramirez-Amaro

This study presents a grasping method for objects with uneven mass distribution by leveraging diffusion models to localize the center of gravity (CoG) on unknown objects. In robotic grasping, CoG deviation often leads to postural…

Robotics · Computer Science 2025-07-28 Kang Xiangli , Yage He , Xianwu Gong , Zehan Liu , Yuru Bai

Supervised visuomotor policies have shown strong performance in robotic manipulation but often struggle in tasks with limited visual inputs, such as operations in confined spaces and dimly lit environments, or tasks requiring precise…

Robotics · Computer Science 2026-01-13 Quan Khanh Luu , Pokuang Zhou , Zhengtong Xu , Zhiyuan Zhang , Qiang Qiu , Yu She

Vision-language-action (VLA) models that directly predict multi-step action chunks from current observations face inherent limitations due to constrained scene understanding and weak future anticipation capabilities. In contrast, video…

We propose a post-hoc adaptive conformal anomaly detection method for monitoring time series that leverages predictions from pre-trained foundation models without requiring additional fine-tuning. Our method yields an interpretable anomaly…

Machine Learning · Computer Science 2026-04-23 Natalia Martinez Gil , Fearghal O'Donncha , Wesley M. Gifford , Nianjun Zhou , Dhaval C. Patel , Roman Vaculin

Following its success in natural language processing and computer vision, foundation models that are pre-trained on large-scale multi-task datasets have also shown great potential in robotics. However, most existing robot foundation models…

Robotics · Computer Science 2025-03-13 Rujia Yang , Geng Chen , Chuan Wen , Yang Gao