English
Related papers

Related papers: Generative Visual Foresight Meets Task-Agnostic Po…

200 papers

Vision Transformers (ViTs) have recently achieved state-of-the-art performance in 2D human pose estimation due to their strong global modeling capability. However, existing ViT-based pose estimators are designed for static images and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Hongwei Fang , Jiahang Cai , Xun Wang , Wenwu Yang

This paper presents a novel layered framework that integrates visual foundation models to improve robot manipulation tasks and motion planning. The framework consists of five layers: Perception, Cognition, Planning, Execution, and Learning.…

Robotics · Computer Science 2023-09-21 Chen Yang , Peng Zhou , Jiaming Qi

Perceptual video compression adopts generative video modeling to improve perceptual realism but frequently sacrifices signal fidelity, diverging from the goal of video compression to faithfully reproduce visual signal. To alleviate the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Ding Ding , Daowen Li , Ying Chen , Yixin Gao , Ruixiao Dong , Kai Li , Li Li

With the rapid development of embodied artificial intelligence, significant progress has been made in vision-language-action (VLA) models for general robot decision-making. However, the majority of existing VLAs fail to account for the…

Robotics · Computer Science 2025-02-17 Hongyin Zhang , Pengxiang Ding , Shangke Lyu , Ying Peng , Donglin Wang

Many modern robotic systems operate autonomously, however they often lack the ability to accurately analyze the environment and adapt to changing external conditions, while teleoperation systems often require special operator skills. In the…

Robotics · Computer Science 2024-11-04 Maria Makarova , Daria Trinitatova , Qian Liu , Dzmitry Tsetserukou

Generative reconstruction methods compute the 3D configuration (such as pose and/or geometry) of a shape by optimizing the overlap of the projected 3D shape model with images. Proper handling of occlusions is a big challenge, since the…

Computer Vision and Pattern Recognition · Computer Science 2016-02-12 Helge Rhodin , Nadia Robertini , Christian Richardt , Hans-Peter Seidel , Christian Theobalt

Transformer-based general visual geometry frameworks have shown promising performance in camera pose estimation and 3D scene understanding. Recent advancements in Visual Geometry Grounded Transformer (VGGT) models have shown great promise…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Yangfan Xu , Lilian Zhang , Xiaofeng He , Pengdong Wu , Wenqi Wu , Jun Mao

Assistance in collaborative manipulation is often initiated by user instructions, making high-level reasoning request-driven. In fluent human teamwork, however, partners often infer the next helpful step from the observed outcome of an…

Robotics · Computer Science 2026-03-26 Fengkai Liu , Hao Su , Haozhuang Chi , Rui Geng , Congzhi Ren , Xuqing Liu , Yucheng Xu , Yuichi Ohsita , Liyun Zhang

Computer vision systems today are primarily N-purpose systems, designed and trained for a predefined set of tasks. Adapting such systems to new tasks is challenging and often requires non-trivial modifications to the network architecture…

Computer Vision and Pattern Recognition · Computer Science 2022-04-21 Tanmay Gupta , Amita Kamath , Aniruddha Kembhavi , Derek Hoiem

We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-generative framework. At its core, GE-Base is a large-scale,…

In this work, we introduce Vision-Language Generative Pre-trained Transformer (VL-GPT), a transformer model proficient at concurrently perceiving and generating visual and linguistic data. VL-GPT achieves a unified pre-training approach for…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Jinguo Zhu , Xiaohan Ding , Yixiao Ge , Yuying Ge , Sijie Zhao , Hengshuang Zhao , Xiaohua Wang , Ying Shan

Grasping in cluttered scenes has always been a great challenge for robots, due to the requirement of the ability to well understand the scene and object information. Previous works usually assume that the geometry information of the objects…

Robotics · Computer Science 2021-09-28 Yiming Li , Tao Kong , Ruihang Chu , Yifeng Li , Peng Wang , Lei Li

Automatic Robotic Assembly Sequence Planning (RASP) can significantly improve productivity and resilience in modern manufacturing along with the growing need for greater product customization. One of the main challenges in realizing such…

Robotics · Computer Science 2023-07-28 Matan Atad , Jianxiang Feng , Ismael Rodríguez , Maximilian Durner , Rudolph Triebel

General object grasping is an important yet unsolved problem in the field of robotics. Most of the current methods either generate grasp poses with few DoF that fail to cover most of the success grasps, or only take the unstable depth image…

Robotics · Computer Science 2021-03-04 Minghao Gou , Hao-Shu Fang , Zhanda Zhu , Sheng Xu , Chenxi Wang , Cewu Lu

Autonomy in robot-assisted minimally invasive surgery has the potential to reduce surgeon cognitive and task load, thereby increasing procedural efficiency. However, implementing accurate autonomous control can be difficult due to poor…

Robotics · Computer Science 2026-03-18 Shuyuan Yang , Zonghe Chua

3D-aware generative models have shown that the introduction of 3D information can lead to more controllable image generation. In particular, the current state-of-the-art model GIRAFFE can control each object's rotation, translation, scale,…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Yang Xue , Yuheng Li , Krishna Kumar Singh , Yong Jae Lee

Endowing robots with human-like physical reasoning abilities remains challenging. We argue that existing methods often disregard spatio-temporal relations and by using Graph Neural Networks (GNNs) that incorporate a relational inductive…

Machine Learning · Computer Science 2019-10-24 Fabio Ferreira , Lin Shao , Tamim Asfour , Jeannette Bohg

Various stuff and things in visual data possess specific traits, which can be learned by deep neural networks and are implicitly represented as the visual prior, e.g., object location and shape, in the model. Such prior potentially impacts…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Jinheng Xie , Kai Ye , Yudong Li , Yuexiang Li , Kevin Qinghong Lin , Yefeng Zheng , Linlin Shen , Mike Zheng Shou

Advances in robotic skill acquisition have made it possible to build general-purpose libraries of learned skills for downstream manipulation tasks. However, naively executing these skills one after the other is unlikely to succeed without…

Robotics · Computer Science 2023-11-16 Christopher Agia , Toki Migimatsu , Jiajun Wu , Jeannette Bohg

We introduce a new approach for robotic manipulation tasks in human settings that necessitates understanding the 3D geometric connections between a pair of objects. Conventional end-to-end training approaches, which convert pixel…

Robotics · Computer Science 2024-05-22 Xuxin Cheng , Heng Yu , Harry Zhang , Wenxing Deng
‹ Prev 1 4 5 6 7 8 10 Next ›