中文
相关论文

相关论文: Explainable YOLO-Based Dyslexia Detection in Synth…

200 篇论文

In image classification tasks, the evaluation of models' robustness to increased dataset shifts with a probabilistic framework is very well studied. However, object detection (OD) tasks pose other challenges for uncertainty estimation and…

计算机视觉与模式识别 · 计算机科学 2020-11-09 Tiago Azevedo , René de Jong , Matthew Mattina , Partha Maji

Automatic object detection by satellite remote sensing images is of great significance for resource exploration and natural disaster assessment. To solve existing problems in remote sensing image detection, this article proposes an improved…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Lei Yang , Guowu Yuan , Hao Zhou , Hongyu Liu , Jian Chen , Hao Wu

We present Language-mediated, Object-centric Representation Learning (LORL), a paradigm for learning disentangled, object-centric scene representations from vision and language. LORL builds upon recent advances in unsupervised object…

机器学习 · 计算机科学 2021-06-09 Ruocheng Wang , Jiayuan Mao , Samuel J. Gershman , Jiajun Wu

Generating natural hand-object interactions in 3D is challenging as the resulting hand and object motions are expected to be physically plausible and semantically meaningful. Furthermore, generalization to unseen objects is hindered by the…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Sammy Christen , Shreyas Hampali , Fadime Sener , Edoardo Remelli , Tomas Hodan , Eric Sauser , Shugao Ma , Bugra Tekin

(Abridged) Galaxy clusters are a powerful probe of cosmological models. Next generation large-scale optical and infrared surveys will reach unprecedented depths over large areas and require highly complete and pure cluster catalogs, with a…

宇宙学与河外天体物理 · 物理学 2024-10-23 Kirill Grishin , Simona Mei , Stéphane Ilic

Logo detection in unconstrained images is challenging, particularly when only very sparse labelled training images are accessible due to high labelling costs. In this work, we describe a model training image synthesising method capable of…

计算机视觉与模式识别 · 计算机科学 2018-03-19 Hang Su , Xiatian Zhu , Shaogang Gong

We envision that in the near future, humanoid robots would share home space and assist us in our daily and routine activities through object manipulations. One of the fundamental technologies that need to be developed for robots is to…

计算机视觉与模式识别 · 计算机科学 2020-02-11 Sayantan Chatterjee , Faheem H. Zunjani , Souvik Sen , Gora C. Nandi

Advancements in deep image synthesis techniques, such as generative adversarial networks (GANs) and diffusion models (DMs), have ushered in an era of generating highly realistic images. While this technological progress has captured…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Mamadou Keita , Wassim Hamidouche , Hessen Bougueffa Eutamene , Abdenour Hadid , Abdelmalik Taleb-Ahmed

Existing deep learning-based object detection models perform well under daytime conditions but face significant challenges at night, primarily because they are predominantly trained on daytime images. Additionally, training with nighttime…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Yunxiang Yang , Hao Zhen , Yongcan Huang , Jidong J. Yang

With this work we are explaining the "You Only Look Once" (YOLO) single-stage object detection approach as a parallel classification of 10647 fixed region proposals. We support this view by showing that each of YOLOs output pixel is…

计算机视觉与模式识别 · 计算机科学 2022-01-24 Christian Limberg , Andrew Melnik , Augustin Harter , Helge Ritter

You Only Look Once (YOLO) algorithm is a representative target detection algorithm emerging in 2016, which is known for its balance of computing speed and accuracy, and now plays an important role in various fields of human production and…

计算机视觉与模式识别 · 计算机科学 2023-09-08 Chenjie Zhang , Pengcheng Jiao

Eye movements are known to reflect cognitive processes in reading, and psychological reading research has shown that eye gaze patterns differ between readers with and without dyslexia. In recent years, researchers have attempted to classify…

计算与语言 · 计算机科学 2022-12-05 Patrick Haller , Andreas Säuberli , Sarah Elisabeth Kiener , Jinger Pan , Ming Yan , Lena Jäger

Hallucination detection remains a fundamental challenge for the safe and reliable deployment of large language models (LLMs), especially in applications requiring factual accuracy. Existing hallucination benchmarks often operate at the…

This paper focuses on real-time American Sign Language Detection. YOLO is a convolutional neural network (CNN) based model, which was first released in 2015. In recent years, it gained popularity for its real-time detection capabilities.…

计算机视觉与模式识别 · 计算机科学 2024-07-26 Amna Imran , Meghana Shashishekhara Hulikal , Hamza A. A. Gardi

The use of artificial intelligence technology in education is growing rapidly, with increasing attention being paid to handwritten mathematical expression recognition (HMER) by researchers. However, many existing methods for HMER may fail…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Ziqi Ye

Now a days, UAVs such as drones are greatly used for various purposes like that of capturing and target detection from ariel imagery etc. Easy access of these small ariel vehicles to public can cause serious security threats. For instance,…

计算机视觉与模式识别 · 计算机科学 2022-01-11 Aleena Ajaz , Ayesha Salar , Tauseef Jamal , Asif Ullah Khan

The rapid advancement of generative technologies has made synthetic images nearly indistinguishable from real ones, thereby creating an urgent need for robust detectors to counter misinformation. However, existing methods mainly rely on…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Yutong Xiao , Ran Ran , Jiwei Wei , Shuchang Zhou , Ke Liu , Zheng Ziqiang , Caiyan Qin

Audio segmentation and sound event detection are crucial topics in machine listening that aim to detect acoustic classes and their respective boundaries. It is useful for audio-content analysis, speech recognition, audio-indexing, and music…

音频与语音处理 · 电气工程与系统科学 2022-09-20 Satvik Venkatesh , David Moffat , Eduardo Reck Miranda

Detecting objects of interest through language often presents challenges, particularly with objects that are uncommon or complex to describe, due to perceptual discrepancies between automated models and human annotators. These challenges…

计算机视觉与模式识别 · 计算机科学 2024-09-11 Pengfei Qi , Yifei Zhang , Wenqiang Li , Youwen Hu , Kunlong Bai

Seizure-frequency information is important for epilepsy research and clinical care, but it is usually recorded in variable free-text clinic letters that are hard to annotate and share. We developed a reproducible, privacy-preserving…

信息检索 · 计算机科学 2026-03-13 Yujian Gan , Stephen H. Barlow , Ben Holgate , Joe Davies , James T. Teo , Joel S. Winston , Mark P. Richardson