中文
相关论文

相关论文: Amodal 3D Reconstruction for Robotic Manipulation …

200 篇论文

A majority of microrobots are constructed using compliant materials that are difficult to model analytically, limiting the utility of traditional model-based controllers. Challenges in data collection on microrobots and large errors between…

机器人学 · 计算机科学 2021-09-09 Joshua Gruenstein , Tao Chen , Neel Doshi , Pulkit Agrawal

Manipulating unseen articulated objects through visual feedback is a critical but challenging task for real robots. Existing learning-based solutions mainly focus on visual affordance learning or other pre-trained visual models to guide…

机器人学 · 计算机科学 2024-04-29 Pengwei Xie , Rui Chen , Siang Chen , Yuzhe Qin , Fanbo Xiang , Tianyu Sun , Jing Xu , Guijin Wang , Hao Su

The strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Xiyuan Gao , Bing Cao , Pengfei Zhu , Nannan Wang , Qinghua Hu

Object classification with 3D data is an essential component of any scene understanding method. It has gained significant interest in a variety of communities, most notably in robotics and computer graphics. While the advent of deep…

计算机视觉与模式识别 · 计算机科学 2019-10-29 Jean-Baptiste Weibel , Timothy Patten , Markus Vincze

In recent years, object-oriented simultaneous localization and mapping (SLAM) has attracted increasing attention due to its ability to provide high-level semantic information while maintaining computational efficiency. Some researchers have…

机器人学 · 计算机科学 2024-02-27 Yutong Wang , Chaoyang Jiang , Xieyuanli Chen

Reconstructing a 3D object from a 2D image is a well-researched vision problem, with many kinds of deep learning techniques having been tried. Most commonly, 3D convolutional approaches are used, though previous work has shown…

计算机视觉与模式识别 · 计算机科学 2023-02-17 Rohan Agarwal , Wei Zhou , Xiaofeng Wu , Yuhan Li

Narrated instructional videos often show and describe manipulations of similar objects, e.g., repairing a particular model of a car or laptop. In this work we aim to reconstruct such objects and to localize associated narrations in 3D.…

计算机视觉与模式识别 · 计算机科学 2021-09-13 Dimitri Zhukov , Ignacio Rocco , Ivan Laptev , Josef Sivic , Johannes L. Schönberger , Bugra Tekin , Marc Pollefeys

In this paper, we propose an adaptive keyframe selection method for improved 3D scene reconstruction in dynamic environments. The proposed method integrates two complementary modules: an error-based selection module utilizing photometric…

机器人学 · 计算机科学 2025-12-30 Raman Jha , Yang Zhou , Giuseppe Loianno

Three-dimensional (3D) reconstruction of head Computed Tomography (CT) images elucidates the intricate spatial relationships of tissue structures, thereby assisting in accurate diagnosis. Nonetheless, securing an optimal head CT scan…

计算机视觉与模式识别 · 计算机科学 2023-09-18 Bowen Zheng , Chenxi Huang , Yuemei Luo

We present Real2Code, a novel approach to reconstructing articulated objects via code generation. Given visual observations of an object, we first reconstruct its part geometry using an image segmentation model and a shape completion model.…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Zhao Mandi , Yijia Weng , Dominik Bauer , Shuran Song

Articulated objects are pervasive in daily life. However, due to the intrinsic high-DoF structure, the joint states of the articulated objects are hard to be estimated. To model articulated objects, two kinds of shape deformations namely…

计算机视觉与模式识别 · 计算机科学 2021-12-15 Han Xue , Liu Liu , Wenqiang Xu , Haoyuan Fu , Cewu Lu

Articulated objects like cabinets and doors are widespread in daily life. However, directly manipulating 3D articulated objects is challenging because they have diverse geometrical shapes, semantic categories, and kinetic constraints. Prior…

机器人学 · 计算机科学 2024-03-04 Qiaojun Yu , Junbo Wang , Wenhai Liu , Ce Hao , Liu Liu , Lin Shao , Weiming Wang , Cewu Lu

Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Congrong Xu , Huachen Gao , Xingyu Chen , Yuliang Xiu , Jun Gao , Anpei Chen

To determine the 3D orientation and 3D location of objects in the surroundings of a camera mounted on a robot or mobile device, we developed two powerful algorithms in object detection and temporal tracking that are combined seamlessly for…

计算机视觉与模式识别 · 计算机科学 2017-09-06 David Joseph Tan , Nassir Navab , Federico Tombari

Articulated object manipulation is a critical capability for robots to perform various tasks in real-world scenarios. Composed of multiple parts connected by joints, articulated objects are endowed with diverse functional mechanisms through…

机器人学 · 计算机科学 2025-02-18 Yuanfei Wang , Xiaojie Zhang , Ruihai Wu , Yu Li , Yan Shen , Mingdong Wu , Zhaofeng He , Yizhou Wang , Hao Dong

Reconstructing physically valid 3D scenes from single-view observations is a prerequisite for bridging the gap between visual perception and robotic control. However, in scenarios requiring precise contact reasoning, such as robotic…

机器人学 · 计算机科学 2026-05-19 Tianyi Xiang , Jiahang Cao , Sikai Guo , Guoyang Zhao , Andrew F. Luo , Jun Ma

We present a method for dynamic surface reconstruction of large-scale urban scenes from LiDAR. Depth-based reconstructions tend to focus on small-scale objects or large-scale SLAM reconstructions that treat moving objects as outliers. We…

计算机视觉与模式识别 · 计算机科学 2025-05-07 Nathaniel Chodosh , Anish Madan , Simon Lucey , Deva Ramanan

Reconstructing textured 3D human models from a single image is fundamental for AR/VR and digital human applications. However, existing methods mostly focus on single individuals and thus fail in multi-human scenes, where naive composition…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Gwanghyun Kim , Junghun James Kim , Suh Yoon Jeon , Jason Park , Se Young Chun

Modern robotic manipulation primarily relies on visual observations in a 2D color space for skill learning but suffers from poor generalization. In contrast, humans, living in a 3D world, depend more on physical properties-such as distance,…

Simultaneous localization and mapping (SLAM) are crucial for autonomous robots (e.g., self-driving cars, autonomous drones), 3D mapping systems, and AR/VR applications. This work proposed a novel LiDAR-inertial-visual fusion framework…

计算机视觉与模式识别 · 计算机科学 2022-09-09 Jiarong Lin , Fu Zhang