English
Related papers

Related papers: PhaForce: Phase-Scheduled Visual-Force Policy Lear…

200 papers

Humans excel at bimanual assembly tasks by adapting to rich tactile feedback -- a capability that remains difficult to replicate in robots through behavioral cloning alone, due to the suboptimality and limited diversity of human…

The adoption of pre-trained visual representations (PVRs), leveraging features from large-scale vision models, has become a popular paradigm for training visuomotor policies. However, these powerful representations can encode a broad range…

Manipulating fragile deformable containers, such as disposable plastic cups filled with liquid, demands real-time grip-force adaptation within an extremely narrow force margin: insufficient force causes slip, while excessive force…

Robotics · Computer Science 2026-05-25 Ziyan Feng , Yulong Fu , Zheng Li , Yuxin He , Jieji Ren , Lujia Wang , Jinni Zhou , Yudong Zhong , Qiang Nie

We present a fast and effective policy framework for robotic manipulation, named Energy Policy, designed for high-frequency robotic tasks and resource-constrained systems. Unlike existing robotic policies, Energy Policy natively predicts…

Robotics · Computer Science 2025-10-15 Jingkai Jia , Tong Yang , Xueyao Chen , Chenhuan Liu , Wenqiang Zhang

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT) tend to treat multi-view features equally and directly concatenate them for policy learning. However, it…

Robotics · Computer Science 2025-07-01 Zihan Lan , Weixin Mao , Haosheng Li , Le Wang , Tiancai Wang , Haoqiang Fan , Osamu Yoshie

Diffusion-based models for robotic control, including vision-language-action (VLA) and vision-action (VA) policies, have demonstrated significant capabilities. Yet their advancement is constrained by the high cost of acquiring large-scale…

Many robotic systems, such as mobile manipulators or quadrotors, cannot be equipped with high-end GPUs due to space, weight, and power constraints. These constraints prevent these systems from leveraging recent developments in visuomotor…

Robotics · Computer Science 2024-07-02 Aaditya Prasad , Kevin Lin , Jimmy Wu , Linqi Zhou , Jeannette Bohg

Visuotactile sensing offers rich contact information that can help mitigate performance bottlenecks in imitation learning, particularly under vision-limited conditions, such as ambiguous visual cues or occlusions. Effectively fusing visual…

Robotics · Computer Science 2025-05-13 Shulong Jiang , Shiqi Zhao , Yuxuan Fan , Peng Yin

Many tasks require flexibly modifying perception and behavior based on current goals. Humans can retrieve episodic memories from days to years ago, using them to contextualize and generalize behaviors across novel but structurally related…

Neural and Evolutionary Computing · Computer Science 2025-12-22 Yicong Zheng , Nora Wolf , Charan Ranganath , Randall C. O'Reilly , Kevin L. McKee

Contact-rich dexterous manipulation with multi-finger hands remains an open challenge in robotics because task success depends on multi-point contacts that continuously evolve and are highly sensitive to object geometry, frictional…

Achieving human-like dexterous manipulation through the collaboration of multi-fingered hands with robotic arms remains a longstanding challenge in robotics, primarily due to the scarcity of high-quality demonstrations and the complexity of…

Robotics · Computer Science 2026-03-12 Yushan Bai , Fulin Chen , Hongzheng Sun , Yuchuang Tong , En Li , Zhengtao Zhang

Significant progress has been made in vision-language models. However, language-conditioned robotic manipulation for contact-rich tasks remains underexplored, particularly in terms of tactile sensing. To address this gap, we introduce the…

Robotics · Computer Science 2025-03-12 Peng Hao , Chaofan Zhang , Dingzhe Li , Xiaoge Cao , Xiaoshuai Hao , Shaowei Cui , Shuo Wang

Deploying controllers trained with Reinforcement Learning (RL) on real robots can be challenging: RL relies on agents' policies being modeled as Markov Decision Processes (MDPs), which assume an inherently discrete passage of time. The use…

Robotics · Computer Science 2024-04-03 Dong Wang , Giovanni Beltrame

Contact-rich bimanual manipulation involves precise coordination of two arms to change object states through strategically selected contacts and motions. Due to the inherent complexity of these tasks, acquiring sufficient demonstration data…

Robotics · Computer Science 2025-02-18 Xuanlin Li , Tong Zhao , Xinghao Zhu , Jiuguang Wang , Tao Pang , Kuan Fang

Visuomotor policies trained via behavior cloning are vulnerable to covariate shift, where small deviations from expert trajectories can compound into failure. Common strategies to mitigate this issue involve expanding the training…

Robotics · Computer Science 2025-08-11 Zhanyi Sun , Shuran Song

We introduce multi-task Visuo-Tactile World Models (VT-WM), which capture the physics of contact through touch reasoning. By complementing vision with tactile sensing, VT-WM better understands robot-object interactions in contact-rich…

As one of the simplest non-prehensile manipulation skills, pushing has been widely studied as an effective means to rearrange objects. Existing approaches, however, typically rely on multi-step push plans composed of pre-defined pushing…

Robotics · Computer Science 2026-02-24 Hieu Bui , Ziyan Gao , Yuya Hosoda , Joo-Ho Lee

In human-robot collaboration (HRC), robots must adapt online to dynamic task constraints and evolving human intent. While physical corrections provide a natural, low-latency channel for operators to convey motion-level adjustments,…

Robotics · Computer Science 2026-03-13 Jiurun Song , Xiao Liang , Minghui Zheng

The generation of robot motions in the real world is difficult by using conventional controllers alone and requires highly intelligent processing. In this regard, learning-based motion generations are currently being investigated. However,…

Robotics · Computer Science 2022-02-15 Sho Sakaino , Kazuki Fujimoto , Yuki Saigusa , Toshiaki Tsuji

Flow-based vision-language-action (VLA) policies offer strong expressivity for action generation, but suffer from a fundamental inefficiency: multi-step inference is required to recover action structure from uninformative Gaussian noise,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Fan Du , Feng Yan , Jianxiong Wu , Xinrun Xu , Weiye Zhang , Weinong Wang , Yu Guo , Bin Qian , Zhihai He , Fei Wang , Heng Yang
‹ Prev 1 4 5 6 7 8 10 Next ›