English
Related papers

Related papers: Lift3D Foundation Policy: Lifting 2D Large-Scale P…

200 papers

Following its success in natural language processing and computer vision, foundation models that are pre-trained on large-scale multi-task datasets have also shown great potential in robotics. However, most existing robot foundation models…

Robotics · Computer Science 2025-03-13 Rujia Yang , Geng Chen , Chuan Wen , Yang Gao

In recent years, there has been an explosion of 2D vision models for numerous tasks such as semantic segmentation, style transfer or scene editing, enabled by large-scale 2D image datasets. At the same time, there has been renewed interest…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Mukund Varma T , Peihao Wang , Zhiwen Fan , Zhangyang Wang , Hao Su , Ravi Ramamoorthi

Bimanual manipulation requires policies that can reason about 3D geometry, anticipate how it evolves under action, and generate smooth, coordinated motions. However, existing methods typically rely on 2D features with limited spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Chongyang Xu , Haipeng Li , Shen Cheng , Jingyu Hu , Haoqiang Fan , Ziliang Feng , Shuaicheng Liu

The lifting of 3D structure and camera from 2D landmarks is at the cornerstone of the entire discipline of computer vision. Traditional methods have been confined to specific rigid objects, such as those in Perspective-n-Point (PnP)…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Mosam Dabhi , Laszlo A. Jeni , Simon Lucey

This work explores the use of 3D generative models to synthesize training data for 3D vision tasks. The key requirements of the generative models are that the generated data should be photorealistic to match the real-world scenarios, and…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Leheng Li , Qing Lian , Luozhou Wang , Ningning Ma , Ying-Cong Chen

Improving the generalization capabilities of general-purpose robotic manipulation agents in the real world has long been a significant challenge. Existing approaches often rely on collecting large-scale robotic data which is costly and…

Robotics · Computer Science 2025-02-10 Jiange Yang , Wenhui Tan , Chuhao Jin , Keling Yao , Bei Liu , Jianlong Fu , Ruihua Song , Gangshan Wu , Limin Wang

3D perceptual representations are well suited for robot manipulation as they easily encode occlusions and simplify spatial reasoning. Many manipulation tasks require high spatial precision in end-effector pose prediction, which typically…

Robotics · Computer Science 2023-10-23 Theophile Gervet , Zhou Xian , Nikolaos Gkanatsios , Katerina Fragkiadaki

Humans have the remarkable ability to use held objects as tools to interact with their environment. For this to occur, humans internally estimate how hand movements affect the object's movement. We wish to endow robots with this capability.…

Robotics · Computer Science 2024-07-16 Weiming Zhi , Haozhan Tang , Tianyi Zhang , Matthew Johnson-Roberson

Effective robotic manipulation relies on a precise understanding of 3D scene geometry, and one of the most straightforward ways to acquire such geometry is through multi-view observations. Motivated by this, we present GP3 -- a 3D…

Robotics · Computer Science 2025-09-22 Quanhao Qian , Guoyang Zhao , Gongjie Zhang , Jiuniu Wang , Ran Xu , Junlong Gao , Deli Zhao

Building a robust perception module is crucial for visuomotor policy learning. While recent methods incorporate pre-trained 2D foundation models into robotic perception modules to leverage their strong semantic understanding, they struggle…

Robotics · Computer Science 2025-07-14 Wenbo Cui , Chengyang Zhao , Yuhui Chen , Haoran Li , Zhizheng Zhang , Dongbin Zhao , He Wang

The incorporation of world modeling into manipulation policy learning has pushed the boundary of manipulation performance. However, existing efforts simply model the 2D visual dynamics, which is insufficient for robust manipulation when…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yuxin He , Ruihao Zhang , Xianzu Wu , Zhiyuan Zhang , Cheng Ding , Qiang Nie

Many robotic tasks involving some form of 3D visual perception greatly benefit from a complete knowledge of the working environment. However, robots often have to tackle unstructured environments and their onboard visual sensors can only…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Andrea Rosasco , Stefano Berti , Fabrizio Bottarel , Michele Colledanchise , Lorenzo Natale

LiDAR-based place recognition serves as a crucial enabler for long-term autonomy in robotics and autonomous driving systems. Yet, prevailing methodologies relying on handcrafted feature extraction face dual challenges: (1) Inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Xiaohui Jiang , Haijiang Zhu , Chade Li , Fulin Tang , Ning An

3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models…

Robotics · Computer Science 2026-05-05 Sizhe Yang , Linning Xu , Hao Li , Juncheng Mu , Jia Zeng , Dahua Lin , Jiangmiao Pang

Recent advancements in 3D foundation models have enabled the generation of high-fidelity assets, yet precise 3D manipulation remains a significant challenge. Existing 3D editing frameworks often face a difficult trade-off between visual…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Inbar Gat , Dana Cohen-Bar , Guy Levy , Elad Richardson , Daniel Cohen-Or

This paper presents a novel layered framework that integrates visual foundation models to improve robot manipulation tasks and motion planning. The framework consists of five layers: Perception, Cognition, Planning, Execution, and Learning.…

Robotics · Computer Science 2023-09-21 Chen Yang , Peng Zhou , Jiaming Qi

2D top-down maps are commonly used for the navigation and exploration of mobile robots through unknown areas. Typically, the robot builds the navigation maps incrementally from local observations using onboard sensors. Recent works have…

Robotics · Computer Science 2024-03-27 Vishnu Dutt Sharma , Anukriti Singh , Pratap Tokekar

Foundation models have achieved remarkable results in 2D and language tasks like image segmentation, object detection, and visual-language understanding. However, their potential to enrich 3D scene representation learning is largely…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Zhimin Chen , Longlong Jing , Yingwei Li , Bing Li

Scaling up representations for images or text has been extensively investigated in the past few years and has led to revolutions in learning vision and language. However, scalable representation for 3D objects and scenes is relatively…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Junsheng Zhou , Jinsheng Wang , Baorui Ma , Yu-Shen Liu , Tiejun Huang , Xinlong Wang

We propose a framework, called LiftedGAN, that disentangles and lifts a pre-trained StyleGAN2 for 3D-aware face generation. Our model is "3D-aware" in the sense that it is able to (1) disentangle the latent space of StyleGAN2 into texture,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Yichun Shi , Divyansh Aggarwal , Anil K. Jain
‹ Prev 1 2 3 10 Next ›