English
Related papers

Related papers: Structural Action Transformer for 3D Dexterous Man…

200 papers

Reconstructing 3D clothed humans from monocular camera data is highly challenging due to viewpoint limitations and image ambiguity. While implicit function-based approaches, combined with prior knowledge from parametric models, have made…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Yong Deng , Baoxing Li , Xu Zhao

Despite lagging behind their modal cousins in many respects, Vision Transformers have provided an interesting opportunity to bridge the gap between sequence modeling and image modeling. Up until now however, vision transformers have largely…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Lily Erickson

End-to-end reinforcement learning techniques are among the most successful methods for robotic manipulation tasks. However, the training time required to find a good policy capable of solving complex tasks is prohibitively large. Therefore,…

Neural-based motion planning methods have achieved remarkable progress for robotic manipulators, yet a fundamental challenge lies in simultaneously accounting for both the robot's physical shape and the surrounding environment when…

Robotics · Computer Science 2025-09-16 Kai Chen , Zhihai Bi , Guoyang Zhao , Chunxin Zheng , Yulin Li , Hang Zhao , Jun Ma

Reconstructing real-world objects and estimating their movable joint structures are pivotal technologies within the field of robotics. Previous research has predominantly focused on supervised approaches, relying on extensively annotated…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Haowen Wang , Zhen Zhao , Zhao Jin , Zhengping Che , Liang Qiao , Yakun Huang , Zhipeng Fan , Xiuquan Qiao , Jian Tang

Temporal Action Localization (TAL) remains a fundamental challenge in video understanding, aiming to identify the start time, end time, and category of all action instances within untrimmed videos. While recent single-stage, anchor-free…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Thisara Rathnayaka , Uthayasanker Thayasivam

Many skeletal action recognition models use GCNs to represent the human body by 3D body joints connected body parts. GCNs aggregate one- or few-hop graph neighbourhoods, and ignore the dependency between not linked body joints. We propose…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Lei Wang , Piotr Koniusz

The versatility of self-attention mechanism earned transformers great success in almost all data modalities, with limitations on the quadratic complexity and difficulty of training. To apply transformers across different data modalities,…

Machine Learning · Computer Science 2024-08-20 Viet Anh Nguyen , Minh Lenhat , Khoa Nguyen , Duong Duc Hieu , Dao Huu Hung , Truong Son Hy

Teaching robots dexterous manipulation skills often requires collecting hundreds of demonstrations using wearables or teleoperation, a process that is challenging to scale. Videos of human-object interactions are easier to collect and…

Robotics · Computer Science 2025-08-19 Tyler Ga Wei Lum , Olivia Y. Lee , C. Karen Liu , Jeannette Bohg

This paper focuses on the scalable robot learning for manipulation in the dexterous robot arm-hand systems, where the remote human-robot interactions via augmented reality (AR) are established to collect the expert demonstration data for…

Machine Learning · Computer Science 2026-02-10 Yicheng Yang , Ruijiao Li , Lifeng Wang , Shuai Zheng , Shunzheng Ma , Keyu Zhang , Tuoyu Sun , Chenyun Dai , Jie Ding , Zhuo Zou

Despite strong results on recognition and segmentation, current 3D visual pre-training methods often underperform on robotic manipulation. We attribute this gap to two factors: the lack of state-action-state dynamics modeling and the…

Robotics · Computer Science 2026-03-11 Qiwei Liang , Boyang Cai , Minghao Lai , Sitong Zhuang , Tao Lin , Yan Qin , Yixuan Ye , Jiaming Liang , Renjing Xu

Reorienting diverse objects with a multi-fingered hand is a challenging task. Current methods in robotic in-hand manipulation are either object-specific or require permanent supervision of the object state from visual sensors. This is far…

Robotics · Computer Science 2024-08-30 Johannes Pitz , Lennart Röstel , Leon Sievers , Darius Burschka , Berthold Bäuml

Several recent Transformer architectures expose later layers to representations computed in the earliest layers, motivated by the observation that low-level features can become harder to recover as the residual stream is repeatedly…

Machine Learning · Computer Science 2026-05-07 Skye Gunasekaran , Téa Wright , Rui-Jie Zhu , Jason Eshraghian

Dexterous manipulation remains challenging due to the cost of collecting real-robot teleoperation data, the heterogeneity of hand embodiments, and the high dimensionality of control. We present UniDex, a robot foundation suite that couples…

Cooperative multi-agent reinforcement learning (MARL) is widely used to address large joint observation and action spaces by decomposing a centralized control problem into multiple interacting agents. However, such decomposition often…

Machine Learning · Computer Science 2026-04-16 Zijian Zhao , Jing Gao , Sen Li

Human-robot teaming (HRT) systems often rely on large-scale datasets of human and robot interactions, especially for close-proximity collaboration tasks such as human-robot handovers. Learning robot manipulation policies from raw,…

Robotics · Computer Science 2025-08-14 Yuekun Wu , Yik Lung Pang , Andrea Cavallaro , Changjae Oh

Due to the availability of large-scale skeleton datasets, 3D human action recognition has recently called the attention of computer vision community. Many works have focused on encoding skeleton data as skeleton image representations based…

Computer Vision and Pattern Recognition · Computer Science 2019-07-31 Carlos Caetano , Jessica Sena , François Brémond , Jefersson A. dos Santos , William Robson Schwartz

Algorithms for the action segmentation task typically use temporal models to predict what action is occurring at each frame for a minute-long daily activity. Recent studies have shown the potential of Transformer in modeling the relations…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Fangqiu Yi , Hongyu Wen , Tingting Jiang

Visual object navigation using learning methods is one of the key tasks in mobile robotics. This paper introduces a new representation of a scene semantic map formed during the embodied agent interaction with the indoor environment. It is…

Robotics · Computer Science 2023-11-08 Tatiana Zemskova , Aleksei Staroverov , Kirill Muravyev , Dmitry Yudin , Aleksandr Panov

Industrial SAT formula generation is a critical yet challenging task. Existing SAT generation approaches can hardly simultaneously capture the global structural properties and maintain plausible computational hardness. We first present an…

Artificial Intelligence · Computer Science 2024-02-09 Yang Li , Xinyan Chen , Wenxuan Guo , Xijun Li , Wanqian Luo , Junhua Huang , Hui-Ling Zhen , Mingxuan Yuan , Junchi Yan
‹ Prev 1 8 9 10 Next ›