中文
相关论文

相关论文: Evaluating OCR Performance for Assistive Technolog…

200 篇论文

The ubiquity of personal digital devices offers unprecedented opportunities to study human behavior. Current state-of-the-art methods quantify physical activity using 'activity counts,' a measure which overlooks specific types of physical…

人机交互 · 计算机科学 2022-07-18 Marcin Straczkiewicz , Emily J. Huang , Jukka-Pekka Onnela

Despite significant advancements in Large Vision Language Models (LVLMs), a gap remains, particularly regarding their interpretability and how they locate and interpret textual information within images. In this paper, we explore various…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Ingeol Baek , Hwan Chang , Sunghyun Ryu , Hwanhee Lee

In the realm of autonomous vehicle perception, comprehending 3D scenes is paramount for tasks such as planning and mapping. Camera-based 3D Semantic Occupancy Prediction (OCC) aims to infer scene geometry and semantics from limited…

计算机视觉与模式识别 · 计算机科学 2025-02-03 Sanbao Su , Nuo Chen , Chenchen Lin , Felix Juefei-Xu , Chen Feng , Fei Miao

The Character Error Rate (CER) is a key metric for evaluating the quality of Optical Character Recognition (OCR). However, this metric assumes that text has been perfectly parsed, which is often not the case. Under page-parsing errors, CER…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Jonathan Bourne , Mwiza Simbeye , Joseph Nockels

A major bottleneck of pedestrian detection lies on the sharp performance deterioration in the presence of small-size pedestrians that are relatively far from the camera. Motivated by the observation that pedestrians of disparate spatial…

计算机视觉与模式识别 · 计算机科学 2018-05-23 Xiaowei Zhang , Li Cheng , Bo Li , Hai-Miao Hu

The project comes with the technique of OCR (Optical Character Recognition) which includes various research sides of computer science. The project is to take a picture of a character and process it up to recognize the image of that…

计算机视觉与模式识别 · 计算机科学 2021-11-11 Arkaprabha Basu , M. Sathya

This paper proposes two new algorithms for certified perception in safety-critical robotic applications. The first is a Certified Visual Odometry algorithm, which uses a RGBD camera with bounded sensor noise to construct a visual odometry…

机器人学 · 计算机科学 2024-02-09 Devansh R Agrawal , Rajiv Govindjee , Jiangbo Yu , Anurekha Ravikumar , Dimitra Panagou

Over the past decade, machine learning methods have given us driverless cars, voice recognition, effective web search, and a much better understanding of the human genome. Machine learning is so common today that it is used dozens of times…

计算机视觉与模式识别 · 计算机科学 2021-06-22 Omer Aydin

Autonomous LLM agents increasingly operate in long-horizon, interactive settings where success depends on reusing experience accumulated over extended histories. However, existing agent memory systems are fundamentally constrained by…

计算与语言 · 计算机科学 2026-04-30 Jinze Li , Yang Zhang , Xin Yang , Jiayi Qu , Jinfeng Xu , Shuo Yang , Junhua Ding , Edith Cheuk-Han Ngai

Optical Flow (OF) is the movement pattern of pixels or edges that is caused in a visual scene by the relative motion between an agent and a scene. OF is used in a wide range of computer vision algorithms and robotics applications. While the…

图像与视频处理 · 电气工程与系统科学 2023-09-25 Jonas Kühne , Michele Magno , Luca Benini

While simulation tools for visible light communication (VLC) with photo detectors (PDs) have been widely investigated, similar tools for optical camera communication (OCC) with complementary metal oxide semiconductor (CMOS) sensors are…

信号处理 · 电气工程与系统科学 2025-04-15 Srivathsan Chakaravarthi Narasimman , Arokiaswami Alphones

Pedestrian detection plays a critical role in autonomous driving (AD), where ensuring safety and reliability is important. While many detection models aim to reduce miss-rates and handle challenges such as occlusion and long-range…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Mohammad Khoshkdahan , Arman Akbari , Arash Akbari , Xuan Zhang

Recent advances in Multi-modal Large Language Models (MLLMs) have showcased remarkable capabilities in vision-language understanding. However, enabling robust video spatial reasoning-the ability to comprehend object locations, orientations,…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Haoran Tang , Meng Cao , Ruyang Liu , Xiaoxi Liang , Linglong Li , Ge Li , Xiaodan Liang

Good OCR results for historical printings rely on the availability of recognition models trained on diplomatic transcriptions as ground truth, which is both a scarce resource and time-consuming to generate. Instead of having to train a…

数字图书馆 · 计算机科学 2016-10-21 U. Springmann , F. Fink , K. U. Schulz

We propose a camera-based assistive text reading framework to help blind persons read text labels and product packaging from hand-held objects in their daily life. To isolate the object from untidy backgrounds or other surrounding objects…

人机交互 · 计算机科学 2019-01-18 Rajkumar N , Anand M. G , Barathiraja N

We propose a novel method utilizing an objectness score for maintaining the locations and classes of objects detected from Mask R-CNN during mobile robot navigation. The objectness score is defined to measure how well the detector…

计算机视觉与模式识别 · 计算机科学 2019-09-13 Hongsun Choi , Mincheul Kang , Youngsun Kwon , Sung-eui Yoon

Gait and static body measurement are important biometric technologies for passive human recognition. Many previous works argue that recognition performance based completely on the gait feature is limited. The reason for this limited…

计算机视觉与模式识别 · 计算机科学 2016-05-19 Ke Yang , Yong Dou , Shaohe Lv , Fei Zhang , Qi Lv

Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur,…

Robust velocity and position estimation is crucial for autonomous robot navigation. The optical flow based methods for autonomous navigation have been receiving increasing attentions in tandem with the development of micro unmanned aerial…

机器人学 · 计算机科学 2018-12-06 Chen Wang , Tete Ji , Thien-Minh Nguyen , Lihua Xie

We present OCR-Quality, a comprehensive human-annotated dataset designed for evaluating and developing OCR quality assessment methods. The dataset consists of 1,000 PDF pages converted to PNG images at 300 DPI, sampled from diverse…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Yulong Zhang