中文
相关论文

相关论文: A Dataset and Benchmarks for Deep Learning-Based O…

200 篇论文

Substantial efforts have been devoted more recently to presenting various methods for object detection in optical remote sensing images. However, the current survey of datasets and deep learning based methods for object detection in optical…

计算机视觉与模式识别 · 计算机科学 2019-12-06 Ke Li , Gang Wan , Gong Cheng , Liqiu Meng , Junwei Han

Reliable embodied perception from an egocentric perspective is challenging yet essential for autonomous navigation technology of intelligent mobile agents. With the growing demand of social robotics, near-field scene understanding becomes…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Haisheng Su , Feixiang Song , Cong Ma , Wei Wu , Junchi Yan

Object pose estimation is a core perception task that enables, for example, object grasping and scene understanding. The widely available, inexpensive and high-resolution RGB sensors and CNNs that allow for fast inference based on this…

Recently, iris recognition is regaining prominence in immersive applications such as extended reality as a means of seamless user identification. This application scenario introduces unique challenges compared to traditional iris…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Yuxi Mi , Qiuyang Yuan , Zhizhou Zhong , Xuan Zhao , Jiaogen Zhou , Fubao Zhu , Jihong Guan , Shuigeng Zhou

Machine learning-based techniques open up many opportunities and improvements to derive deeper and more practical insights from data that can help businesses make informed decisions. However, the majority of these techniques focus on the…

机器学习 · 计算机科学 2024-05-10 Atefeh Mahdavi , Marco Carvalho

Fueled by deep learning, computer-aided diagnosis achieves huge advances. However, out of controlled lab environments, algorithms could face multiple challenges. Open set recognition (OSR), as an important one, states that categories unseen…

计算机视觉与模式识别 · 计算机科学 2023-07-24 Mingyuan Liu , Lu Xu , Jicong Zhang

While a great variety of 3D cameras have been introduced in recent years, most publicly available datasets for object recognition and pose estimation focus on one single camera. In this work, we present a dataset of 32 scenes that have been…

机器人学 · 计算机科学 2020-09-30 Till Grenzdörffer , Martin Günther , Joachim Hertzberg

Among the most important prerequisites for creating and evaluating 6D object pose detectors are datasets with labeled 6D poses. With the advent of deep learning, demand for such datasets is growing continuously. Despite the fact that some…

计算机视觉与模式识别 · 计算机科学 2019-10-02 Roman Kaskman , Sergey Zakharov , Ivan Shugurov , Slobodan Ilic

Precise 6D pose estimation of rigid objects from RGB images is a critical but challenging task in robotics, augmented reality and human-computer interaction. To address this problem, we propose DeepRM, a novel recurrent network architecture…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Alexander Avery , Andreas Savakis

Recently, 3D version has been improved greatly due to the development of deep neural networks. A high quality dataset is important to the deep learning method. Existing datasets for 3D vision has been constructed, such as Bigbird and YCB.…

计算机视觉与模式识别 · 计算机科学 2020-11-18 Minglei Lu , Yu Guo , Fei Wang , Zheng Dang

We present a large-scale stereo RGB image object pose estimation dataset named the $\textbf{StereOBJ-1M}$ dataset. The dataset is designed to address challenging cases such as object transparency, translucency, and specular reflection, in…

计算机视觉与模式识别 · 计算机科学 2022-03-16 Xingyu Liu , Shun Iwase , Kris M. Kitani

We introduce OLATverse, a large-scale dataset comprising around 9M images of 765 real-world objects, captured from multiple viewpoints under a diverse set of precisely controlled lighting conditions. While recent advances in object-centric…

To analyze this characteristic of vulnerability, we developed an automated deep learning method for detecting microvessels in intravascular optical coherence tomography (IVOCT) images. A total of 8,403 IVOCT image frames from 85 lesions and…

Depth estimation is a fundamental component of spatial perception for autonomous driving and other unmanned systems operating in open urban environments. Existing depth datasets such as KITTI, nuScenes, and DDAD have advanced the field but…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Xianda Guo , Ruijun Zhang , Yiqun Duan , Ruilin Wang , Matteo Poggi , Keyuan Zhou , Wenzhao Zheng , Wenke Huang , Gangwei Xu , Yanlun Peng , Yuan Si , Qin Zou

Head pose estimation is a challenging task that aims to solve problems related to predicting three dimensions vector, that serves for many applications in human-robot interaction or customer behavior. Previous researches have proposed some…

计算机视觉与模式识别 · 计算机科学 2021-11-16 Linh Nguyen Viet , Tuan Nguyen Dinh , Hoang Nguyen Viet , Duc Tran Minh , Long Tran Quoc

The human ability to recognize when an object belongs or does not belong to a particular vision task outperforms all open set recognition algorithms. Human perception as measured by the methods and procedures of visual psychophysics from…

计算机视觉与模式识别 · 计算机科学 2023-04-26 Jin Huang , Derek Prijatelj , Justin Dulay , Walter Scheirer

Object pose estimation is crucial for robotic applications and augmented reality. Beyond instance level 6D object pose estimation methods, estimating category-level pose and shape has become a promising trend. As such, a new research field…

计算机视觉与模式识别 · 计算机科学 2022-05-19 Pengyuan Wang , HyunJun Jung , Yitong Li , Siyuan Shen , Rahul Parthasarathy Srikanth , Lorenzo Garattoni , Sven Meier , Nassir Navab , Benjamin Busam

Most state-of-the-art vein recognition methods rely on closed-set classification, which inherently limits their scalability and prevents the adaptive enrollment of new users without complete model retraining. We rigorously evaluate the…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Paweł Pilarek , Marcel Musiałek , Anna Górska

We present DeepSeek-OCR as an initial investigation into the feasibility of compressing long contexts via optical 2D mapping. DeepSeek-OCR consists of two components: DeepEncoder and DeepSeek3B-MoE-A570M as the decoder. Specifically,…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Haoran Wei , Yaofeng Sun , Yukun Li

Optical Chemical Structure Recognition (OCSR) is critical for converting 2D molecular diagrams from printed literature into machine-readable formats. While Vision-Language Models have shown promise in end-to-end OCR tasks, their direct…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Haocheng Tang , Xingyu Dang , Junmei Wang