中文
相关论文

相关论文: YOLOv10-Based Multi-Task Framework for Hand Locali…

200 篇论文

Continuous mid-air hand gesture recognition based on captured hand pose streams is fundamental for human-computer interaction, particularly in AR / VR. However, many of the methods proposed to recognize heterogeneous hand gestures are…

计算机视觉与模式识别 · 计算机科学 2023-04-13 Federico Cunico , Federico Girella , Andrea Avogaro , Marco Emporio , Andrea Giachetti , Marco Cristani

Learning from the limited amount of labeled data to the pre-train model has always been viewed as a challenging task. In this report, an effective and robust solution, the two-stage training paradigm YOLOv8 detector (TP-YOLOv8), is designed…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Zheng Wang , Dong Xie , Hanzhi Wang , Jiang Tian

This paper addresses the synthetic-to-real domain gap in object detection, focusing on training a YOLOv11 model to detect a specific object (a soup can) using only synthetic data and domain randomization strategies. The methodology involves…

计算机视觉与模式识别 · 计算机科学 2025-09-19 Luisa Torquato Niño , Hamza A. A. Gardi

This study proposes an enhanced dual-model YOLOv8 framework for intelligent fire detection and proximity-aware risk assessment, extending conventional vision-based monitoring beyond simple detection to actionable hazard prioritization. The…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Ammar K. AlMhdawi , Nonso Nnamoko , Alaa Mashan Ubaid

Pedestrians and bicyclists are among the vulnerable road users (VRUs) that are inherently exposed to intricate traffic scenarios, which puts them at increased risk of sustaining injuries or facing fatal outcomes. This study presents an…

图像与视频处理 · 电气工程与系统科学 2025-07-16 Faryal Aurooj Nasir , Salman Liaquat , Nor Muzlifah Mahyuddin

This study explores a comprehensive approach to obstacle detection using advanced YOLO models, specifically YOLOv8, YOLOv7, YOLOv6, and YOLOv5. Leveraging deep learning techniques, the research focuses on the performance comparison of these…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Santiago Pérez , Camila Gómez , Matías Rodríguez

Recent advancements in real-time object detection frameworks have spurred extensive research into their application in robotic systems. This study provides a comparative analysis of YOLOv5 and YOLOv8 models, challenging the prevailing…

Transfer of objects between humans and robots is a critical capability for collaborative robots. Although there has been a recent surge of interest in human-robot handovers, most prior research focus on robot-to-human handovers. Further,…

机器人学 · 计算机科学 2020-03-16 Wei Yang , Chris Paxton , Maya Cakmak , Dieter Fox

Teleoperation provides a way for human operators to guide robots in situations where full autonomy is challenging or where direct human intervention is required. It can also be an important tool to teach robots in order to achieve…

机器人学 · 计算机科学 2021-11-15 Florian Kennel-Maushart , Roi Poranne , Stelian Coros

Deep learning-based computer vision technology has grown stronger in recent years, and cross-fertilization using computer vision technology has been a popular direction in recent years. The use of computer vision technology to identify…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Zhifeng Wang , Jialong Yao , Chunyan Zeng , Wanxuan Wu , Hongmin Xu , Yang Yang

Instructional cataract surgery videos are crucial for ophthalmologists and trainees to observe surgical details repeatedly. This paper presents a deep learning model for real-time identification of surgical instruments in these videos,…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Sanya Sinha , Michal Balazia , Francois Bremond

Accurately detecting student behavior in classroom videos can aid in analyzing their classroom performance and improving teaching effectiveness. However, the current accuracy rate in behavior detection is low. To address this challenge, we…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Fan Yang , Tao Wang , Xiaofei Wang

Lung cancer poses a significant global public health challenge, emphasizing the importance of early detection for improved patient outcomes. Recent advancements in deep learning algorithms have shown promising results in medical image…

图像与视频处理 · 电气工程与系统科学 2023-05-26 Karthick Prasad Gunasekaran

Efficient and accurate annotation of datasets remains a significant challenge for deploying object detection models such as You Only Look Once (YOLO) in real-world applications, particularly in agriculture where rapid decision-making is…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Mohamed Abdallah Salem , Ahmed Harb Rabia

Recognizing various surgical tools, actions and phases from surgery videos is an important problem in computer vision with exciting clinical applications. Existing deep-learning-based methods for this problem either process each surgical…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Haifeng Wang , Hao Xu , Jun Wang , Jian Zhou , Ke Deng

Maintaining roadway infrastructure is essential for ensuring a safe, efficient, and sustainable transportation system. However, manual data collection for detecting road damage is time-consuming, labor-intensive, and poses safety risks.…

计算机视觉与模式识别 · 计算机科学 2024-10-14 Vung Pham , Lan Dong Thi Ngoc , Duy-Linh Bui

Video Camouflaged Object Detection (VCOD) is currently constrained by the scarcity of challenging benchmarks and the limited robustness of models against erratic motion dynamics. Existing methods often struggle with Motion-Induced…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Yiyu Liu , Shuo Ye , Chao Hao , Zitong Yu

Human activity recognition in videos is a challenging problem that has drawn a lot of interest, particularly when the goal requires the analysis of a large video database. AOLME project provides a collaborative learning environment for…

计算机视觉与模式识别 · 计算机科学 2021-06-15 Sravani Teeparthi

In recent years, computer-aided diagnosis systems have shown great potential in assisting radiologists with accurate and efficient medical image analysis. This paper presents a novel approach for bone pathology localization and…

计算机视觉与模式识别 · 计算机科学 2023-08-25 Razan Dibo , Andrey Galichin , Pavel Astashev , Dmitry V. Dylov , Oleg Y. Rogov

The digitization of structured handwritten documents, such as academic marksheets, remains a significant challenge due to the dual complexity of irregular table structures and diverse handwriting styles. While recent Transformer-based…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Md. Irtiza Hossain , Junaid Ahmed Sifat , Abir Chowdhury