English
Related papers

Related papers: GP3: A 3D Geometry-Aware Policy with Multi-View Im…

200 papers

In this paper, we propose a 3D geometry-aware deformable Gaussian Splatting method for dynamic view synthesis. Existing neural radiance fields (NeRF) based solutions learn the deformation in an implicit manner, which cannot incorporate 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Zhicheng Lu , Xiang Guo , Le Hui , Tianrui Chen , Min Yang , Xiao Tang , Feng Zhu , Yuchao Dai

Robotic manipulation requires accurate perception of the environment, which poses a significant challenge due to its inherent complexity and constantly changing nature. In this context, RGB image and point-cloud observations are two…

Robotics · Computer Science 2024-09-10 Boshi An , Yiran Geng , Kai Chen , Xiaoqi Li , Qi Dou , Hao Dong

Humanoid robots capable of autonomous operation in diverse environments have long been a goal for roboticists. However, autonomous manipulation by humanoid robots has largely been restricted to one specific scene, primarily due to the…

Robotics · Computer Science 2025-09-10 Yanjie Ze , Zixuan Chen , Wenhao Wang , Tianyi Chen , Xialin He , Ying Yuan , Xue Bin Peng , Jiajun Wu

In this paper, we propose a new representation for multiview image sets. Our approach relies on graphs to describe geometry information in a compact and controllable way. The links of the graph connect pixels in different images and…

Multimedia · Computer Science 2013-12-23 Thomas Maugey , Antonio Ortega , Pascal Frossard

Recent advancements in 3D robotic manipulation have improved grasping of everyday objects, but transparent and specular materials remain challenging due to depth sensing limitations. While several 3D reconstruction and depth completion…

Robotics · Computer Science 2025-06-23 Mingxu Zhang , Xiaoqi Li , Jiahui Xu , Kaichen Zhou , Hojin Bae , Yan Shen , Chuyan Xiong , Hao Dong

3D occupancy perception holds a pivotal role in recent vision-centric autonomous driving systems by converting surround-view images into integrated geometric and semantic representations within dense 3D grids. Nevertheless, current models…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Xin Tan , Wenbin Wu , Zhiwei Zhang , Chaojie Fan , Yong Peng , Zhizhong Zhang , Yuan Xie , Lizhuang Ma

Intelligent vision control systems for surgical robots should adapt to unknown and diverse objects while being robust to system disturbances. Previous methods did not meet these requirements due to mainly relying on pose estimation and…

Robotics · Computer Science 2024-05-29 Hongbin Lin , Bin Li , Chun Wai Wong , Juan Rojas , Xiangyu Chu , Kwok Wai Samuel Au

Robotic grasping is a fundamental capability for autonomous manipulation, yet remains highly challenging in cluttered environments where occlusion, poor perception quality, and inconsistent 3D reconstructions often lead to unstable or…

Robotics · Computer Science 2025-11-07 Shenglin Wang , Mingtong Dai , Jingxuan Su , Lingbo Liu , Chunjie Chen , Xinyu Wu , Liang Lin

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies overlook one or both. They typically rely on 2D visual observations and backbones pretrained…

Tool use is essential for enabling robots to perform complex real-world tasks, but learning such skills requires extensive datasets. While teleoperation is widely used, it is slow, delay-sensitive, and poorly suited for dynamic tasks. In…

Robotics · Computer Science 2025-09-16 Haonan Chen , Cheng Zhu , Shuijing Liu , Yunzhu Li , Katherine Driggs-Campbell

Modern 3D human pose estimation techniques rely on deep networks, which require large amounts of training data. While weakly-supervised methods require less supervision, by utilizing 2D poses or multi-view imagery without annotations, they…

Computer Vision and Pattern Recognition · Computer Science 2018-04-05 Helge Rhodin , Mathieu Salzmann , Pascal Fua

Robots often rely on RGB images for tasks like manipulation and navigation. However, reliable interaction typically requires a 3D scene representation that is metric-scaled and aligned with the robot reference frame. This depends on…

Robotics · Computer Science 2025-09-11 Davide Allegro , Matteo Terreran , Stefano Ghidoni

Predictive world models that simulate future observations under explicit camera control are fundamental to interactive AI. Despite rapid advances, current systems lack spatial persistence: they fail to maintain stable scene structures over…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Chendong Xiang , Jiajun Liu , Jintao Zhang , Xiao Yang , Zhengwei Fang , Shizun Wang , Zijun Wang , Yingtian Zou , Hang Su , Jun Zhu

Synthesizing accurate geometry and photo-realistic appearance of small scenes is an active area of research with compelling use cases in gaming, virtual reality, robotic-manipulation, autonomous driving, convenient product capture, and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Arkadeep Narayan Chaudhury , Igor Vasiljevic , Sergey Zakharov , Vitor Guizilini , Rares Ambrus , Srinivasa Narasimhan , Christopher G. Atkeson

Recent advancements in world models have revolutionized dynamic environment simulation, allowing systems to foresee future states and assess potential actions. In autonomous driving, these capabilities help vehicles anticipate the behavior…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Anthony Chen , Wenzhao Zheng , Yida Wang , Xueyang Zhang , Kun Zhan , Peng Jia , Kurt Keutzer , Shanghang Zhang

Dense 3D reconstruction and ego-motion estimation are key challenges in autonomous driving and robotics. Compared to the complex, multi-modal systems deployed today, multi-camera systems provide a simpler, low-cost alternative. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Aron Schmied , Tobias Fischer , Martin Danelljan , Marc Pollefeys , Fisher Yu

Image-based 3D object modeling refers to the process of converting raw optical images to 3D digital representations of the objects. Very often, such models are desired to be dimensionally true, semantically labeled with photorealistic…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Rongjun Qin , Xu Huang

This paper focuses on the problem of learning 6-DOF grasping with a parallel jaw gripper in simulation. We propose the notion of a geometry-aware representation in grasping based on the assumption that knowledge of 3D geometry is at the…

Knowledge of 3-D object shape is of great importance to robot manipulation tasks, but may not be readily available in unstructured environments. While vision is often occluded during robot-object interaction, high-resolution tactile sensors…

Robotics · Computer Science 2022-03-11 Sudharshan Suresh , Zilin Si , Joshua G. Mangelson , Wenzhen Yuan , Michael Kaess

Multi-camera 3D object detection for autonomous driving is a challenging problem that has garnered notable attention from both academia and industry. An obstacle encountered in vision-based techniques involves the precise extraction of…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Linyan Huang , Huijie Wang , Jia Zeng , Shengchuan Zhang , Liujuan Cao , Junchi Yan , Hongyang Li