中文
相关论文

相关论文: GeomPrompt: Geometric Prompt Learning for RGB-D Se…

200 篇论文

To fully exploit depth cues in Camouflaged Object Detection (COD), we present DGA-Net, a specialized framework that adapts the Segment Anything Model (SAM) via a novel ``depth prompting" paradigm. Distinguished from existing approaches that…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Yuetong Li , Qing Zhang , Yilin Zhao , Gongyang Li , Zeming Liu

Visual scene understanding is an important capability that enables robots to purposefully act in their environment. In this paper, we propose a novel approach to object-class segmentation from multiple RGB-D views using deep learning. We…

计算机视觉与模式识别 · 计算机科学 2017-12-06 Lingni Ma , Jörg Stückler , Christian Kerl , Daniel Cremers

Prompt tuning, a recently emerging paradigm, enables the powerful vision-language pre-training models to adapt to downstream tasks in a parameter -- and data -- efficient way, by learning the ``soft prompts'' to condition frozen…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Juncheng Li , Minghe Gao , Longhui Wei , Siliang Tang , Wenqiao Zhang , Mengze Li , Wei Ji , Qi Tian , Tat-Seng Chua , Yueting Zhuang

RGB-guided depth completion aims at predicting dense depth maps from sparse depth measurements and corresponding RGB images, where how to effectively and efficiently exploit the multi-modal information is a key issue. Guided dynamic…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Yufei Wang , Yuxin Mao , Qi Liu , Yuchao Dai

Understanding Earth's subsurface is critical for energy transition, natural hazard mitigation, and planetary science. Yet subsurface analysis remains fragmented, with separate models required for structural interpretation, stratigraphic…

地球物理 · 物理学 2025-09-15 Yimin Dou , Xinming Wu , Nathan L Bangs , Harpreet Singh Sethi , Jintao Li , Hang Gao , Zhixiang Guo

Most existing RGB-D semantic segmentation methods focus on the feature level fusion, including complex cross-modality and cross-scale fusion modules. However, these methods may cause misalignment problem in the feature fusion process and…

计算机视觉与模式识别 · 计算机科学 2025-05-05 Xiaoyan Jiang , Bohan Wang , Xinlong Wan , Shanshan Chen , Hamido Fujita , Hanan Abd. Al Juaid

Depth completion aims to recover dense depth maps from sparse ones, where color images are often used to facilitate this task. Recent depth methods primarily focus on image guided learning frameworks. However, blurry guidance in the image…

计算机视觉与模式识别 · 计算机科学 2024-02-29 Zhiqiang Yan , Xiang Li , Le Hui , Zhenyu Zhang , Jun Li , Jian Yang

Scene understanding plays a critical role in enabling intelligence and autonomy in robotic systems. Traditional approaches often face challenges, including occlusions, ambiguous boundaries, and the inability to adapt attention based on…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Guodong Sun , Junjie Liu , Gaoyang Zhang , Bo Wu , Yang Zhang

Prompts play a critical role in unleashing the power of language and vision foundation models for specific tasks. For the first time, we introduce prompting into depth foundation models, creating a new paradigm for metric depth estimation…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Haotong Lin , Sida Peng , Jingxiao Chen , Songyou Peng , Jiaming Sun , Minghuan Liu , Hujun Bao , Jiashi Feng , Xiaowei Zhou , Bingyi Kang

The goal of our work is to complete the depth channel of an RGB-D image. Commodity-grade depth cameras often fail to sense depth for shiny, bright, transparent, and distant surfaces. To address this problem, we train a deep network that…

计算机视觉与模式识别 · 计算机科学 2018-05-03 Yinda Zhang , Thomas Funkhouser

Learning to navigate to an image-specified goal is an important but challenging task for autonomous systems. The agent is required to reason the goal location from where a picture is shot. Existing methods try to solve this problem by…

计算机视觉与模式识别 · 计算机科学 2023-10-12 Xinyu Sun , Peihao Chen , Jugang Fan , Thomas H. Li , Jian Chen , Mingkui Tan

Advances in machine learning, especially the introduction of transformer architectures and vision transformers, have led to the development of highly capable computer vision foundation models. The segment anything model (known colloquially…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Kenneth Ball , Erin Taylor , Nirav Patel , Andrew Bartels , Gary Koplik , James Polly , Jay Hineman

Prompt-guided generative AI models have rapidly expanded across vision and language domains, producing realistic and diverse outputs from textual inputs. The growing variety of such models, trained with different data and architectures,…

机器学习 · 计算机科学 2026-02-09 Mehdi Lotfian , Mohammad Jalali , Farzan Farnia

Multimodal learning with incomplete modality is practical and challenging. Recently, researchers have focused on enhancing the robustness of pre-trained MultiModal Transformers (MMTs) under missing modality conditions by applying learnable…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Jian Lang , Zhangtao Cheng , Ting Zhong , Fan Zhou

3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-based input, their lack of 3D perception limits their…

机器人学 · 计算机科学 2024-08-12 Xincheng Pang , Wenke Xia , Zhigang Wang , Bin Zhao , Di Hu , Dong Wang , Xuelong Li

Accurate three-dimensional perception is essential for modern industrial robotic systems that perform manipulation, inspection, and navigation tasks. RGB-D and stereo vision sensors are widely used for this purpose, but the depth maps they…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Tony Salloom , Dandi Zhou , Xinhai Sun

Recent advances in All-in-One (AiO) RGB image restoration have demonstrated the effectiveness of prompt learning in handling multiple degradations within a single model. However, extending these approaches to hyperspectral image (HSI)…

图像与视频处理 · 电气工程与系统科学 2025-03-12 Chia-Ming Lee , Ching-Heng Cheng , Yu-Fan Lin , Yi-Ching Cheng , Wo-Ting Liao , Fu-En Yang , Yu-Chiang Frank Wang , Chih-Chung Hsu

A key challenge for RGB-D segmentation is how to effectively incorporate 3D geometric information from the depth channel into 2D appearance features. We propose to model the effective receptive field of 2D convolution based on the scale and…

计算机视觉与模式识别 · 计算机科学 2019-10-04 Yunlu Chen , Thomas Mensink , Efstratios Gavves

Recently, deep learning has produced encouraging results for kidney stone classification using endoscope images. However, the shortage of annotated training data poses a severe problem in improving the performance and generalization ability…

计算机视觉与模式识别 · 计算机科学 2023-03-16 Wei Zhu , Runtao Zhou , Yao Yuan , Campbell Timothy , Rajat Jain , Jiebo Luo

We introduce GeoSAM2, a prompt-controllable framework for 3D part segmentation that casts the task as multi-view 2D mask prediction. Given a textureless object, we render normal and point maps from predefined viewpoints and accept simple 2D…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Ken Deng , Yunhan Yang , Jingxiang Sun , Xihui Liu , Yebin Liu , Ding Liang , Yan-Pei Cao