English
Related papers

Related papers: MIDV-2019: Challenges of the modern mobile-based d…

200 papers

Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur,…

Retrieving specific information from a large corpus of documents is a prevalent industrial use case of modern AI, notably due to the popularity of Retrieval-Augmented Generation (RAG) systems. Although neural document retrieval models have…

Information Retrieval · Computer Science 2025-12-17 Paul Teiletche , Quentin Macé , Max Conti , Antonio Loison , Gautier Viaud , Pierre Colombo , Manuel Faysse

Dynamic Vision Sensors (DVS) offer a unique advantage in control applications due to their high temporal resolution and asynchronous event-based data. Still, their adoption in machine learning algorithms remains limited. To address this gap…

Robotics · Computer Science 2025-03-04 Felix Resch , Mónika Farsang , Radu Grosu

We introduce the challenging problem of multi-object system identification from videos, for which prior methods are ill-suited due to their focus on single-object scenes or discrete material classification with a fixed set of material…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Chunjiang Liu , Xiaoyuan Wang , Qingran Lin , Albert Xiao , Haoyu Chen , Shizheng Wen , Hao Zhang , Lu Qi , Ming-Hsuan Yang , Laszlo A. Jeni , Min Xu , Yizhou Zhao

Video relation detection forms a new and challenging problem in computer vision, where subjects and objects need to be localized spatio-temporally and a predicate label needs to be assigned if and only if there is an interaction between the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-26 Shuo Chen , Pascal Mettes , Cees G. M. Snoek

RGB-D data is essential for solving many problems in computer vision. Hundreds of public RGB-D datasets containing various scenes, such as indoor, outdoor, aerial, driving, and medical, have been proposed. These datasets are useful for…

Computer Vision and Pattern Recognition · Computer Science 2022-08-10 Alexandre Lopes , Roberto Souza , Helio Pedrini

Autonomous driving and assistance systems rely on annotated data from traffic and road scenarios to model and learn the various object relations in complex real-world scenarios. Preparation and training of deploy-able deep learning…

Computer Vision and Pattern Recognition · Computer Science 2022-10-25 Shubham Dokania , A. H. Abdul Hafez , Anbumani Subramanian , Manmohan Chandraker , C. V. Jawahar

Omnidirectional or 360-degree video is being increasingly deployed, largely due to the latest advancements in immersive virtual reality (VR) and extended reality (XR) technology. However, the adoption of these videos in streaming encounters…

Image and Video Processing · Electrical Eng. & Systems 2024-03-08 Ahmed Telili , Ibrahim Farhat , Wassim Hamidouche , Hadi Amirpour

Real-world surveillance often renders faces and license plates unrecognizable in individual low-resolution (LR) frames, hindering reliable identification. To advance temporal recognition models, we present FANVID, a novel video-based…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Kavitha Viswanathan , Vrinda Goel , Shlesh Gholap , Devayan Ghosh , Madhav Gupta , Dhruvi Ganatra , Sanket Potdar , Amit Sethi

Optical Character Recognition (OCR) for data extraction from documents is essential to intelligent informatics, such as digitizing medical records and recognizing road signs. Multi-modal Large Language Models (LLMs) can solve this task and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Hyakka Nakada , Yoshiyasu Tanaka

Optical coherence tomography angiography (OCTA) is a novel imaging modality that has been widely utilized in ophthalmology and neuroscience studies to observe retinal vessels and microvascular systems. However, publicly available OCTA…

Image and Video Processing · Electrical Eng. & Systems 2022-12-27 Mingchao Li , Kun Huang , Qiuzhuo Xu , Jiadong Yang , Yuhan Zhang , Zexuan Ji , Keren Xie , Songtao Yuan , Qinghuai Liu , Qiang Chen

We present OCR-Quality, a comprehensive human-annotated dataset designed for evaluating and developing OCR quality assessment methods. The dataset consists of 1,000 PDF pages converted to PNG images at 300 DPI, sampled from diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yulong Zhang

The multi-camera vehicle tracking (MCVT) framework holds significant potential for smart city applications, including anomaly detection, traffic density estimation, and suspect vehicle tracking. However, current publicly available datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yuqiang Lin , Sam Lockyer , Mingxuan Sui , Li Gan , Florian Stanek , Markus Zarbock , Wenbin Li , Adrian Evans , Nic Zhang

This paper presents a new high resolution aerial images dataset in which moving objects are labelled manually. It aims to contribute to the evaluation of the moving object detection methods for moving cameras. The problem of recognizing…

Computer Vision and Pattern Recognition · Computer Science 2021-03-25 Ibrahim Delibasoglu

The success of deep learning in intelligent ship visual perception relies heavily on rich image data. However, dedicated datasets for inland waterway vessels remain scarce, limiting the adaptability of visual perception systems in complex…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Shanshan Wang , Haixiang Xu , Hui Feng , Xiaoqian Wang , Pei Song , Sijie Liu , Jianhua He

Robust vehicle detection from fixed CCTV cameras is critical for Intelligent Transportation Systems. Yet existing benchmarks predominantly feature relatively homogeneous, highly organized traffic patterns captured from ego-centric driving…

Recently, ocular biometrics in unconstrained environments using images obtained at visible wavelength have gained the researchers' attention, especially with images captured by mobile devices. Periocular recognition has been demonstrated to…

Computer Vision and Pattern Recognition · Computer Science 2022-11-16 Luiz A. Zanlorensi , Rayson Laroca , Diego R. Lucio , Lucas R. Santos , Alceu S. Britto , David Menotti

Contemporary deep-learning object detection methods for autonomous driving usually assume prefixed categories of common traffic participants, such as pedestrians and cars. Most existing detectors are unable to detect uncommon objects and…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Kaican Li , Kai Chen , Haoyu Wang , Lanqing Hong , Chaoqiang Ye , Jianhua Han , Yukuai Chen , Wei Zhang , Chunjing Xu , Dit-Yan Yeung , Xiaodan Liang , Zhenguo Li , Hang Xu

Lidar technology has evolved significantly over the last decade, with higher resolution, better accuracy, and lower cost devices available today. In addition, new scanning modalities and novel sensor technologies have emerged in recent…

Robotics · Computer Science 2022-03-08 Qingqing Li , Xianjia Yu , Jorge Peña Queralta , Tomi Westerlund

Over the past decade, machine learning methods have given us driverless cars, voice recognition, effective web search, and a much better understanding of the human genome. Machine learning is so common today that it is used dozens of times…

Computer Vision and Pattern Recognition · Computer Science 2021-06-22 Omer Aydin