中文
相关论文

相关论文: bbOCR: An Open-source Multi-domain OCR Pipeline fo…

200 篇论文

This paper focuses on enhancing Bengali Document Layout Analysis (DLA) using the YOLOv8 model and innovative post-processing techniques. We tackle challenges unique to the complex Bengali script by employing data augmentation for model…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Nazmus Sakib Ahmed , Saad Sakib Noor , Ashraful Islam Shanto Sikder , Abhijit Paul

Digital camera and mobile document image acquisition are new trends arising in the world of Optical Character Recognition and text detection. In some cases, such process integrates many distortions and produces poorly scanned text or…

计算机视觉与模式识别 · 计算机科学 2015-09-14 Abdeslam El Harraj , Naoufal Raissouni

In Bangladesh, many farmers still struggle to access timely, expert-level agricultural guidance. This paper presents KrishokBondhu, a voice-enabled, call-centre-integrated advisory platform built on a Retrieval-Augmented Generation (RAG)…

计算与语言 · 计算机科学 2026-03-10 Mohd Ruhul Ameen , Akif Islam , Farjana Aktar , M. Saifuzzaman Rafat

Bengali remains a low-resource language in speech technology, especially for complex tasks like long-form transcription and speaker diarization. This paper presents a multistage approach developed for the "DL Sprint 4.0 - Bengali Long-Form…

声音 · 计算机科学 2026-03-04 Epshita Jahan , Khandoker Md Tanjinul Islam , Pritom Biswas , Tafsir Al Nafin

Bengali is an underrepresented language in NLP research. However, it remains a challenge due to its unique linguistic structure and computational constraints. In this work, we systematically investigate the challenges that hinder Bengali…

计算与语言 · 计算机科学 2025-08-01 Shimanto Bhowmik , Tawsif Tashwar Dipto , Md Sazzad Islam , Sheryl Hsu , Tahsin Reasat

Handwritten text recognition (HTR) and machine translation continue to pose significant challenges, particularly for low-resource languages like Marathi, which lack large digitized corpora and exhibit high variability in handwriting styles.…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Shubham Kumar Nigam , Parjanya Aditya Shukla , Noel Shallum , Arnab Bhattacharya

In [4], the authors present the DisCoCirc (Distributed Compositional Circuits) formalism for the English language, a grammar-based framework derived from the production rules that incorporates circuit-like representations in order to give a…

计算与语言 · 计算机科学 2025-11-13 Nazmoon Falgunee Moon

Document parsing is a core task in document intelligence, supporting applications such as information extraction, retrieval-augmented generation, and automated document analysis. However, real-world documents often feature complex layouts…

Despite considerable progress in handwritten text recognition, paragraph-level handwritten text recognition, especially in low-resource languages, such as Hindi, Urdu and similar scripts, remains a challenging problem. These languages,…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Sayantan Dey , Alireza Alaei , Partha Pratim Roy

This work presents an accuracy study of the open source OCR engine, Kraken, on the leading Arabic scholarly journal, al-Abhath. In contrast with other commercially available OCR engines, Kraken is shown to be capable of producing highly…

计算与语言 · 计算机科学 2024-02-20 Benjamin Kiessling , Gennady Kurin , Matthew Thomas Miller , Kader Smail

Despite significant advances in document understanding, determining the correct orientation of scanned or photographed documents remains a critical pre-processing step in the real world settings. Accurate rotation correction is essential…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Suranjan Goswami , Abhinav Ravi , Raja Kolla , Ali Faraz , Shaharukh Khan , Akash , Chandra Khatri , Shubham Agarwal

We present the largest publicly available synthetic OCR benchmark dataset for Indic languages. The collection contains a total of 90k images and their ground truth for 23 Indic languages. OCR model validation in Indic languages require a…

计算机视觉与模式识别 · 计算机科学 2022-05-06 Naresh Saini , Promodh Pinto , Aravinth Bheemaraj , Deepak Kumar , Dhiraj Daga , Saurabh Yadav , Srihari Nagaraj

For digitizing or indexing physical documents, Optical Character Recognition (OCR), the process of extracting textual information from scanned documents, is a vital technology. When a document is visually damaged or contains non-textual…

计算机视觉与模式识别 · 计算机科学 2022-07-05 Oshri Naparstek , Ophir Azulai , Daniel Rotman , Yevgeny Burshtein , Peter Staar , Udi Barzelay

Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR)…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Jacob Carlson , Tom Bryan , Melissa Dell

Designing Optical Character Recognition (OCR) systems for India requires balancing linguistic diversity, document heterogeneity, and deployment constraints. In this paper, we study two training strategies for building multilingual OCR…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Ali Faraz , Raja Kolla , Ashish Kulkarni , Shubham Agarwal

This paper presents the development of a prototype Automatic Speech Recognition (ASR) system specifically designed for Bengali biomedical data. Recent advancements in Bengali ASR are encouraging, but a lack of domain-specific data limits…

音频与语音处理 · 电气工程与系统科学 2024-06-21 Shariar Kabir , Nazmun Nahar , Shyamasree Saha , Mamunur Rashid

Coreference Resolution is a well studied problem in NLP. While widely studied for English and other resource-rich languages, research on coreference resolution in Bengali largely remains unexplored due to the absence of relevant datasets.…

计算与语言 · 计算机科学 2023-07-06 Shadman Rohan , Mojammel Hossain , Mohammad Mamun Or Rashid , Nabeel Mohammed

Bengali, spoken by over 300 million people, is a morphologically rich and lowresource language, posing challenges for automatic speech recognition (ASR). This research presents an end-to-end framework for Bengali ASR, building on a…

音频与语音处理 · 电气工程与系统科学 2026-01-16 Md. Nazmus Sakib , Golam Mahmud , Md. Maruf Bangabashi , Umme Ara Mahinur Istia , Md. Jahidul Islam , Partha Sarker , Afra Yeamini Prity

Document extraction is a core component of digital workflows, yet existing vision-language models (VLMs) predominantly favor high-resource languages. Thai presents additional challenges due to script complexity from non-latin letters, the…

India is a multi-lingual country where Roman script is often used alongside different Indic scripts in a text document. To develop a script specific handwritten Optical Character Recognition (OCR) system, it is therefore necessary to…

机器学习 · 计算机科学 2010-03-25 Ram Sarkar , Nibaran Das , Subhadip Basu , Mahantapas Kundu , Mita Nasipuri , Dipak Kumar Basu