中文
相关论文

相关论文: UniCal: a Single-Branch Transformer-Based Model fo…

200 篇论文

A large-scale labeled dataset is a key factor for the success of supervised deep learning in computer vision. However, a limited number of annotated data is very common, especially in ophthalmic image analysis, since manual annotation is…

图像与视频处理 · 电气工程与系统科学 2022-03-15 Zhiyuan Cai , Li Lin , Huaqing He , Xiaoying Tang

The unification of disparate maps is crucial for enabling scalable robot operation across multiple sessions and collaborative multi-robot scenarios. However, achieving a unified map robust to sensor modalities and dynamic environments…

机器人学 · 计算机科学 2025-12-24 Gilhwan Kang , Hogyun Kim , Byunghee Choi , Seokhwan Jeong , Young-Sik Shin , Younggun Cho

Place recognition is crucial for loop closure detection and global localization in robotics. Although mainstream algorithms typically rely on cameras and LiDAR, these sensors are susceptible to adverse weather conditions. Fortunately, the…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Ningyuan Huang , Zhiheng Li , Zheng Fang

Deducing the 3D structure of endoscopic scenes from images is exceedingly challenging. In addition to deformation and view-dependent lighting, tubular structures like the colon present problems stemming from their self-occluding and…

计算机视觉与模式识别 · 计算机科学 2023-12-18 Anita Rau , Binod Bhattarai , Lourdes Agapito , Danail Stoyanov

With the rapid development of autonomous driving and SLAM technology, the performance of autonomous systems using multimodal sensors highly relies on accurate extrinsic calibration. Addressing the need for a convenient, maintenance-friendly…

机器人学 · 计算机科学 2024-06-18 Wonho Song , Minho Oh , Jaeyoung Lee , Hyun Myung

Transformer-based single-object trackers achieve state-of-the-art accuracy but rely on fixed-depth inference, executing the full encoder--decoder stack for every frame regardless of visual complexity, thereby incurring unnecessary…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Patrick Poggi , Divake Kumar , Theja Tulabandhula , Amit Ranjan Trivedi

Visual saliency modeling for images and videos is treated as two independent tasks in recent computer vision literature. While image saliency modeling is a well-studied problem and progress on benchmarks like SALICON and MIT300 is slowing,…

计算机视觉与模式识别 · 计算机科学 2020-11-10 Richard Droste , Jianbo Jiao , J. Alison Noble

Properly-calibrated sensors are the prerequisite for a dependable autonomous driving system. However, most prior methods focus on extrinsic calibration between sensors, and few focus on the misalignment between the sensors and the vehicle…

机器人学 · 计算机科学 2023-05-19 Guohang Yan , Zhaotong Luo , Zhuochun Liu , Yikang Li

Few-Shot Class-Incremental Learning (FSCIL) defines a practical but challenging task where models are required to continuously learn novel concepts with only a few training samples. Due to data scarcity, existing FSCIL methods resort to…

计算机视觉与模式识别 · 计算机科学 2024-12-18 Chengyan Liu , Linglan Zhao , Fan Lyu , Kaile Du , Fuyuan Hu , Tao Zhou

Vision foundation models have demonstrated strong generalization in medical image segmentation by leveraging large-scale, heterogeneous pretraining. However, they often struggle to generalize to specialized clinical tasks under limited…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Wenjing Lu , Yi Hong , Yang Yang

Self-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images like those in ImageNet and ignores the correlation among the…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Zhaowen Li , Yousong Zhu , Fan Yang , Wei Li , Chaoyang Zhao , Yingying Chen , Zhiyang Chen , Jiahao Xie , Liwei Wu , Rui Zhao , Ming Tang , Jinqiao Wang

Cross-modal transfer learning is used to improve multi-modal classification models (e.g., for human activity recognition in human-robot collaboration). However, existing methods require paired sensor data at both training and inference,…

机器学习 · 计算机科学 2025-09-15 Leen Daher , Zhaobo Wang , Malcolm Mielle

Robots often rely on RGB images for tasks like manipulation and navigation. However, reliable interaction typically requires a 3D scene representation that is metric-scaled and aligned with the robot reference frame. This depends on…

机器人学 · 计算机科学 2025-09-11 Davide Allegro , Matteo Terreran , Stefano Ghidoni

Point-pixel registration between LiDAR point clouds and camera images is a fundamental yet challenging task in autonomous driving and robotic perception. A key difficulty lies in the modality gap between unstructured point clouds and…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Yu Han , Zhiwei Huang , Yanting Zhang , Fangjun Ding , Shen Cai , Rui Fan

LiDAR-camera calibration is a precondition for many heterogeneous systems that fuse data from LiDAR and camera. However, the constraint from common field of view and the requirement for strict time synchronization make the calibration a…

机器人学 · 计算机科学 2019-07-31 Bo Fu , Yue Wang , Xiaqing Ding , Yanmei Jiao , Li Tang , Rong Xiong

Class-incremental learning (CIL) for endoscopic image analysis is crucial for real-world clinical applications, where diagnostic models should continuously adapt to evolving clinical data while retaining performance on previously learned…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Bingrong Liu , Jun Shi , Yushan Zheng

Change detection (CD) aims to identify surface changes from multi-temporal remote sensing imagery. In real-world scenarios, Pixel-level change labels are expensive to acquire, and existing models struggle to adapt to scenarios with diverse…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Kaixuan Jiang , Chen Wu , Zhenghui Zhao , Chengxi Han , Haonan Guo , Hongruixuan Chen

LiDAR registration is a fundamental task in robotic mapping and localization. A critical component of aligning two point clouds is identifying robust point correspondences using point descriptors. This step becomes particularly challenging…

机器人学 · 计算机科学 2025-02-27 Niclas Vödisch , Giovanni Cioffi , Marco Cannici , Wolfram Burgard , Davide Scaramuzza

It is a challenging task to learn discriminative representation from images and videos, due to large local redundancy and complex global dependency in these visual data. Convolution neural networks (CNNs) and vision transformers (ViTs) have…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Kunchang Li , Yali Wang , Junhao Zhang , Peng Gao , Guanglu Song , Yu Liu , Hongsheng Li , Yu Qiao

Recent years have witnessed the prevailing progress of Generative Adversarial Networks (GANs) in image-to-image translation. However, the success of these GAN models hinges on ponderous computational costs and labor-expensive training data.…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Yuxi Ren , Jie Wu , Peng Zhang , Manlin Zhang , Xuefeng Xiao , Qian He , Rui Wang , Min Zheng , Xin Pan