中文
相关论文

相关论文: Find Any Part in 3D

200 篇论文

Following its success in natural language processing and computer vision, foundation models that are pre-trained on large-scale multi-task datasets have also shown great potential in robotics. However, most existing robot foundation models…

机器人学 · 计算机科学 2025-03-13 Rujia Yang , Geng Chen , Chuan Wen , Yang Gao

Artificial intelligence (AI) is evolving towards artificial general intelligence, which refers to the ability of an AI system to perform a wide range of tasks and exhibit a level of intelligence similar to that of a human being. This is in…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Chunhui Zhang , Li Liu , Yawen Cui , Guanjie Huang , Weilin Lin , Yiqian Yang , Yuehong Hu

Vision foundation models (VFMs) trained on large-scale image datasets provide high-quality features that have significantly advanced 2D visual recognition. However, their potential in 3D scene segmentation remains largely untapped, despite…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Karim Knaebel , Kadir Yilmaz , Daan de Geus , Alexander Hermans , David Adrian , Timm Linder , Bastian Leibe

The task of detecting 3D objects is important to various robotic applications. The existing deep learning-based detection techniques have achieved impressive performance. However, these techniques are limited to run with a graphics…

计算机视觉与模式识别 · 计算机科学 2020-08-14 Xuesong Li , Jose Guivant , Subhan Khan

Unsupervised and open-vocabulary 3D object detection has recently gained attention, particularly in autonomous driving, where reducing annotation costs and recognizing unseen objects are critical for both safety and scalability. However,…

计算机视觉与模式识别 · 计算机科学 2025-12-02 In-Jae Lee , Mungyeom Kim , Kwonyoung Ryu , Pierre Musacchio , Jaesik Park

The proliferation of 2D foundation models has sparked research into adapting them for open-world 3D instance segmentation. Recent methods introduce a paradigm that leverages superpoints as geometric primitives and incorporates 2D multi-view…

计算机视觉与模式识别 · 计算机科学 2024-11-07 Xi Yang , Xu Gu , Xingyilang Yin , Xinbo Gao

Compositional 3D scene generation from a single view requires the simultaneous recovery of scene layout and 3D assets. Existing approaches mainly fall into two categories: feed-forward generation methods and per-instance generation methods.…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Ze-Xin Yin , Liu Liu , Xinjie Wang , Wei Sui , Zhizhong Su , Jian Yang , Jin Xie

Powered by large-scale pre-training, vision foundation models exhibit significant potential in open-world image understanding. However, unlike large language models that excel at directly tackling various language tasks, vision foundation…

计算机视觉与模式识别 · 计算机科学 2024-01-22 Yang Liu , Muzhi Zhu , Hengtao Li , Hao Chen , Xinlong Wang , Chunhua Shen

Understanding material failure is critical for designing stronger and lighter structures by identifying weaknesses that could be mitigated. Existing full-physics numerical simulation techniques involve trade-offs between speed, accuracy,…

This work aims for image categorization using a representation of distinctive parts. Different from existing part-based work, we argue that parts are naturally shared between image categories and should be modeled as such. We motivate our…

计算机视觉与模式识别 · 计算机科学 2016-07-13 Pascal Mettes , Jan C. van Gemert , Cees G. M. Snoek

Online, real-time, and fine-grained 3D segmentation constitutes a fundamental capability for embodied intelligent agents to perceive and comprehend their operational environments. Recent advancements employ predefined object queries to…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Hanshi Wang , Zijian Cai , Jin Gao , Yiwei Zhang , Weiming Hu , Ke Wang , Zhipeng Zhang

Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in the field of medical imaging, the curation and assembling…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Zhongying Deng , Cheng Tang , Ziyan Huang , Jiashi Lin , Ying Chen , Junzhi Ning , Chenglong Ma , Jiyao Liu , Wei Li , Yinghao Zhu , Shujian Gao , Yanyan Huang , Sibo Ju , Yanzhou Su , Pengcheng Chen , Wenhao Tang , Tianbin Li , Haoyu Wang , Yuanfeng Ji , Hui Sun , Shaobo Min , Liang Peng , Feilong Tang , Haochen Xue , Rulin Zhou , Chaoyang Zhang , Wenjie Li , Shaohao Rui , Weijie Ma , Xingyue Zhao , Yibin Wang , Kun Yuan , Zhaohui Lu , Shujun Wang , Jinjie Wei , Lihao Liu , Dingkang Yang , Lin Wang , Yulong Li , Haolin Yang , Yiqing Shen , Lequan Yu , Xiaowei Hu , Yun Gu , Yicheng Wu , Benyou Wang , Minghui Zhang , Angelica I. Aviles-Rivero , Qi Gao , Hongming Shan , Xiaoyu Ren , Fang Yan , Hongyu Zhou , Haodong Duan , Maosong Cao , Shanshan Wang , Bin Fu , Xiaomeng Li , Zhi Hou , Chunfeng Song , Lei Bai , Yuan Cheng , Yuandong Pu , Xiang Li , Wenhai Wang , Hao Chen , Jiaxin Zhuang , Songyang Zhang , Huiguang He , Mengzhang Li , Bohan Zhuang , Zhian Bai , Rongshan Yu , Liansheng Wang , Yukun Zhou , Xiaosong Wang , Xin Guo , Guanbin Li , Xiangru Lin , Dakai Jin , Mianxin Liu , Wenlong Zhang , Qi Qin , Conghui He , Yuqiang Li , Ye Luo , Nanqing Dong , Jie Xu , Wenqi Shao , Bo Zhang , Qiujuan Yan , Yihao Liu , Jun Ma , Zhi Lu , Yuewen Cao , Zongwei Zhou , Jianming Liang , Shixiang Tang , Qi Duan , Dongzhan Zhou , Chen Jiang , Yuyin Zhou , Yanwu Xu , Jiancheng Yang , Shaoting Zhang , Xiaohong Liu , Siqi Luo , Yi Xin , Chaoyu Liu , Haochen Wen , Xin Chen , Alejandro Lozano , Min Woo Sun , Yuhui Zhang , Yue Yao , Xiaoxiao Sun , Serena Yeung-Levy , Xia Li , Jing Ke , Chunhui Zhang , Zongyuan Ge , Ming Hu , Jin Ye , Zhifeng Li , Yirong Chen , Yu Qiao , Junjun He

We present a network architecture which compares RGB images and untextured 3D models by the similarity of the represented shape. Our system is optimised for zero-shot retrieval, meaning it can recognise shapes never shown in training. We…

计算机视觉与模式识别 · 计算机科学 2024-04-29 Maciej Janik , Niklas Gard , Anna Hilsmann , Peter Eisert

Recognizing scenes and objects in 3D from a single image is a longstanding goal of computer vision with applications in robotics and AR/VR. For 2D recognition, large datasets and scalable solutions have led to unprecedented advances. In 3D,…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Garrick Brazil , Abhinav Kumar , Julian Straub , Nikhila Ravi , Justin Johnson , Georgia Gkioxari

Recently most popular tracking frameworks focus on 2D image sequences. They seldom track the 3D object in point clouds. In this paper, we propose PointIT, a fast, simple tracking method based on 3D on-road instance segmentation. Firstly, we…

计算机视觉与模式识别 · 计算机科学 2019-02-19 Yuan Wang , Yang Yu , Ming Liu

To be useful in everyday environments, robots must be able to observe and learn about objects. Recent datasets enable progress for classifying data into known object categories; however, it is unclear how to collect reliable object data…

机器人学 · 计算机科学 2019-01-18 Abhishek Venkataraman , Brent Griffin , Jason J. Corso

Estimating the 3D position and orientation of objects in the environment with a single RGB camera is a critical and challenging task for low-cost urban autonomous driving and mobile robots. Most of the existing algorithms are based on the…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Yuxuan Liu , Yuan Yixuan , Ming Liu

Deep learning approaches to object detection have achieved reliable detection of specific object classes in images. However, extending a model's detection capability to new object classes requires large amounts of annotated training data,…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Vikhyat Agarwal , Jiayi Cora Guo , Declan Hoban , Sissi Zhang , Nicholas Moran , Peter Cho , Srilakshmi Pattabiraman , Shantanu Joshi

We survey applications of pretrained foundation models in robotics. Traditional deep learning models in robotics are trained on small datasets tailored for specific tasks, which limits their adaptability across diverse applications. In…

Training 3D object detectors for autonomous driving has been limited to small datasets due to the effort required to generate annotations. Reducing both task complexity and the amount of task switching done by annotators is key to reducing…

机器学习 · 计算机科学 2018-07-18 Jungwook Lee , Sean Walsh , Ali Harakeh , Steven L. Waslander