English
Related papers

Related papers: GAP: Geometric Anchor Pre-training for Data-Effici…

200 papers

Imitation learning has emerged as a crucial ap proach for acquiring visuomotor skills from demonstrations, where designing effective observation encoders is essential for policy generalization. However, existing methods often struggle to…

Robotics · Computer Science 2025-12-01 Yikai Tang , Haoran Geng , Sheng Zang , Pieter Abbeel , Jitendra Malik

This paper focuses on the problem of learning 6-DOF grasping with a parallel jaw gripper in simulation. We propose the notion of a geometry-aware representation in grasping based on the assumption that knowledge of 3D geometry is at the…

This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained multi-view diffusion model to synthesize latent novel views…

Training generative adversarial networks (GANs) is known to be difficult, especially for financial time series. This paper first analyzes the well-posedness problem in GANs minimax games and the convexity issue in GANs objective functions.…

Machine Learning · Statistics 2021-12-28 Othmane Mounjid , Xin Guo

Language-Guided Robotic Manipulation (LGRM) is a challenging task as it requires a robot to understand human instructions to manipulate everyday objects. Recent approaches in LGRM rely on pre-trained Visual Grounding (VG) models to detect…

Robotics · Computer Science 2023-07-13 Junghyun Kim , Gi-Cheon Kang , Jaein Kim , Suyeon Shin , Byoung-Tak Zhang

Learning diverse locomotion skills for humanoid robots in a unified reinforcement learning framework remains challenging due to the conflicting requirements of stability and dynamic expressiveness across different gaits. We present a…

Robotics · Computer Science 2026-04-22 Yuanye Wu , Keyi Wang , Linqi Ye , Boyang Xing

Pre-trained generalist policies are rapidly gaining relevance in robot learning due to their promise of fast adaptation to novel, in-domain tasks. This adaptation often relies on collecting new demonstrations for a specific task of interest…

Machine Learning · Computer Science 2025-06-24 Marco Bagatella , Jonas Hübotter , Georg Martius , Andreas Krause

Graph anomaly detection (GAD) has garnered increasing attention in recent years, yet remains challenging due to two key factors: (1) label scarcity stemming from the high cost of annotations and (2) homophily disparity at node and class…

Machine Learning · Computer Science 2026-01-30 Yunhui Liu , Jiashun Cheng , Yiqing Lin , Qizhuo Xie , Jia Li , Fugee Tsung , Hongzhi Yin , Tao Zheng , Jianhua Zhao , Tieke He

Generalist Vision-Language-Action models are currently hindered by the scarcity of robotic data compared to the abundance of human video demonstrations. Existing Latent Action Models attempt to leverage video data but often suffer from…

Robotics · Computer Science 2026-01-08 Chubin Zhang , Jianan Wang , Zifeng Gao , Yue Su , Tianru Dai , Cai Zhou , Jiwen Lu , Yansong Tang

Uncertainties in contact dynamics and object geometry remain significant barriers to robust robotic manipulation. Caging mitigates these uncertainties by constraining an object's mobility without requiring precise contact modeling. However,…

The existing Motion Imitation models typically require expert data obtained through MoCap devices, but the vast amount of training data needed is difficult to acquire, necessitating substantial investments of financial resources, manpower,…

Robotics · Computer Science 2024-05-03 Liu Qiyuan

Compared to the prosperity of pre-training models in natural image understanding, the research on large-scale pre-training models for facial knowledge learning is still limited. Current approaches mainly rely on manually assembled and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Yudong Li , Hao Li , Xianxu Hou , Linlin Shen

Over recent years, diffusion models have facilitated significant advancements in video generation. Yet, the creation of face-related videos still confronts issues such as low facial fidelity, lack of frame consistency, limited editability…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Linze Li , Sunqi Fan , Hengjun Pu , Zhaodong Bing , Yao Tang , Tianzhu Ye , Tong Yang , Liangyu Chen , Jiajun Liang

Geospatial pixel reasoning aims to generate segmentation masks in remote sensing imagery directly from natural-language instructions. Most existing approaches follow a paradigm that fine-tunes multimodal large language models under…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Chengjie Jiang , Yunqi Zhou , Jiafeng Yan , Jing Li , Jiayang Li , Yue Zhou , Hongjie He , Jonathan Li

While Vision-Language-Action (VLA) models offer broad general capabilities, deploying them on specific hardware requires real-world adaptation to bridge the embodiment gap. Since robot demonstrations are costly, this adaptation must often…

Robotics · Computer Science 2026-05-11 Yanzhe Chen , Kevin Yuchen Ma , Qi Lv , Yiqi Lin , Zechen Bai , Chen Gao , Mike Zheng Shou

Contact force in contact-rich environments is an essential modality for robots to perform general-purpose manipulation tasks, as it provides information to compensate for the deficiencies of visual and proprioceptive data in collision…

Robotics · Computer Science 2024-11-13 Bo Zhou , Ruixuan Jiao , Yi Li , Xiaogang Yuan , Fang Fang , Shihua Li

In recent years, autonomous parking has made significant advances, yet parking tasks still face challenges in extreme scenarios such as mechanical and dead-end parking slots, often resulting in failures. This is mainly due to traditional…

Robotics · Computer Science 2026-05-12 Changze Li , Zhe Chen , Shaoyu Chen , Lisen Mu , Yijian Li , Yuelong Yu , Qian Zhang , Qing Su , Ming Yang , Tong Qin

Surgical workflow anticipation is the task of predicting the timing of relevant surgical events from live video data, which is critical in Robotic-Assisted Surgery (RAS). Accurate predictions require the use of spatial information to model…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Francis Xiatian Zhang , Jingjing Deng , Robert Lieck , Hubert P. H. Shum

Machine unlearning (MU) for large language models has become critical for AI safety, yet existing methods fail to generalize to Mixture-of-Experts (MoE) architectures. We identify that traditional unlearning methods exploit MoE's…

Machine Learning · Computer Science 2026-02-17 Andy Zhu , Rongzhe Wei , Yupu Gu , Pan Li

This work investigates how the intricate task of a continuous pick & place (P&P) motion may be learned from humans based on demonstrations and corrections. Due to the complexity of the task, these demonstrations are often slow and even…

Robotics · Computer Science 2022-04-12 Anna Mészáros , Giovanni Franzese , Jens Kober