English
Related papers

Related papers: Ref-SAM3D: Bridging SAM3D with Text for Reference …

200 papers

Video object segmentation methods like SAM2 achieve strong performance through memory-based architectures but struggle under large viewpoint changes due to reliance on appearance features. Traditional 3D instance segmentation methods…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Yang-Che Sun , Cheng Sun , Chin-Yang Lin , Fu-En Yang , Min-Hung Chen , Yen-Yu Lin , Yu-Lun Liu

In this work we explore reconstructing hand-object interactions in the wild. The core challenge of this problem is the lack of appropriate 3D labeled data. To overcome this issue, we propose an optimization-based procedure which does not…

Computer Vision and Pattern Recognition · Computer Science 2022-01-03 Zhe Cao , Ilija Radosavovic , Angjoo Kanazawa , Jitendra Malik

3D object reconstructions of transparent and concave structured objects, with inferred material properties, remains an open research problem for robot navigation in unstructured environments. In this paper, we propose a multimodal single-…

Computer Vision and Pattern Recognition · Computer Science 2021-10-07 Justin Wilson , Ming C. Lin

We introduce a new task, Map and Locate, which unifies the traditionally distinct objectives of open-vocabulary segmentation - detecting and segmenting object instances based on natural language queries - and 3D reconstruction, the process…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Xuweiyi Chen , Tian Xia , Sihan Xu , Jianing Yang , Joyce Chai , Zezhou Cheng

We introduce the task of 3D object localization in RGB-D scans using natural language descriptions. As input, we assume a point cloud of a scanned 3D scene along with a free-form description of a specified target object. To address this…

Computer Vision and Pattern Recognition · Computer Science 2020-11-12 Dave Zhenyu Chen , Angel X. Chang , Matthias Nießner

Scaling up representations for images or text has been extensively investigated in the past few years and has led to revolutions in learning vision and language. However, scalable representation for 3D objects and scenes is relatively…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Junsheng Zhou , Jinsheng Wang , Baorui Ma , Yu-Shen Liu , Tiejun Huang , Xinlong Wang

\noindent Memory has become the central mechanism enabling robust visual object tracking in modern segmentation-based frameworks. Recent methods built upon Segment Anything Model 2 (SAM2) have demonstrated strong performance by refining how…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Mohamad Alansari , Muzammal Naseer , Hasan Al Marzouqi , Naoufel Werghi , Sajid Javed

Recent monocular 3D shape reconstruction methods have shown promising zero-shot results on object-segmented images without any occlusions. However, their effectiveness is significantly compromised in real-world conditions, due to imperfect…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Junhyeong Cho , Kim Youwang , Hunmin Yang , Tae-Hyun Oh

Recent advances in NeRF and 3DGS have significantly enhanced the efficiency and quality of 3D content synthesis. However, efficient personalization of generated 3D content remains a critical challenge. Current 3D personalization approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Qi Song , Ziyuan Luo , Ka Chun Cheung , Simon See , Renjie Wan

We introduce Cap3D, an automatic approach for generating descriptive text for 3D objects. This approach utilizes pretrained models from image captioning, image-text alignment, and LLM to consolidate captions from multiple views of a 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Tiange Luo , Chris Rockwell , Honglak Lee , Justin Johnson

3D shape completion has broad applications in robotics, digital twin reconstruction, and extended reality (XR). Although recent advances in 3D object and scene completion have achieved impressive results, existing methods lack 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Bardienus P. Duisterhof , Jan Oberst , Bowen Wen , Stan Birchfield , Deva Ramanan , Jeffrey Ichnowski

Generating immersive 3D scenes from texts is a core task in computer vision, crucial for applications in virtual reality and game development. Despite the promise of leveraging 2D diffusion priors, existing methods suffer from spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Jisheng Chu , Wenrui Li , Rui Zhao , Wangmeng Zuo , Shifeng Chen , Xiaopeng Fan

3D image reconstruction from a limited number of 2D images has been a long-standing challenge in computer vision and image analysis. While deep learning-based approaches have achieved impressive performance in this area, existing deep…

Computer Vision and Pattern Recognition · Computer Science 2023-10-04 Nivetha Jayakumar , Tonmoy Hossain , Miaomiao Zhang

Humans rely on their visual and tactile senses to develop a comprehensive 3D understanding of their physical environment. Recently, there has been a growing interest in exploring and manipulating objects using data-driven approaches that…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 Mauro Comi , Yijiong Lin , Alex Church , Alessio Tonioni , Laurence Aitchison , Nathan F. Lepora

Reconstructing the underlying 3D surface of an object from a single image is a challenging problem that has received extensive attention from the computer vision community. Many learning-based approaches tackle this problem by learning a 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-02-28 Nicolai Häni , Jun-Jee Chao , Volkan Isler

Sketches are the most abstract 2D representations of real-world objects. Although a sketch usually has geometrical distortion and lacks visual cues, humans can effortlessly envision a 3D object from it. This suggests that sketches encode…

Computer Vision and Pattern Recognition · Computer Science 2022-01-20 Jiayun Wang , Jierui Lin , Qian Yu , Runtao Liu , Yubei Chen , Stella X. Yu

This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown promising results by directly regressing 3D geometry from RGB images without explicit 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Tao Xie , Peishan Yang , Yudong Jin , Yingfeng Cai , Wei Yin , Weiqiang Ren , Qian Zhang , Wei Hua , Sida Peng , Xiaoyang Guo , Xiaowei Zhou

We investigate the possibility of 3D scene reconstruction from two or more overlapping webcam streams. A large, and growing, number of webcams observe places of interest and are publicly accessible. The question naturally arises: can we…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Tianyu Wu , Konrad Schindler , Cenek Albl

Segmenting 3D objects into parts is a long-standing challenge in computer vision. To overcome taxonomy constraints and generalize to unseen 3D objects, recent works turn to open-world part segmentation. These approaches typically transfer…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Zhe Zhu , Le Wan , Rui Xu , Yiheng Zhang , Honghua Chen , Zhiyang Dou , Cheng Lin , Yuan Liu , Mingqiang Wei

Robotic-assisted surgery allows surgeons to conduct precise surgical operations with stereo vision and flexible motor control. However, the lack of 3D spatial perception limits situational awareness during procedures and hinders mastering…

Image and Video Processing · Electrical Eng. & Systems 2022-03-07 Shang Zhao , Ce Wang , Qiyuan Wang , Yanzhe Liu , S Kevin Zhou