中文
相关论文

相关论文: DohaScript: A Large-Scale Multi-Writer Dataset for…

200 篇论文

Kashmiri is spoken by around 7 million people but remains critically underserved in speech technology, despite its official status and rich linguistic heritage. The lack of robust Text-to-Speech (TTS) systems limits digital accessibility…

Personalized book recommendation in Bangla literature has been constrained by the lack of structured, large-scale, and publicly available datasets. This work introduces RokomariBG, a large-scale, multi-entity heterogeneous book graph…

Handwritten documents are often characterized by dense and uneven layout. Despite advances, standard deep network based approaches for semantic layout segmentation are not robust to complex deformations seen across semantic regions. This…

计算机视觉与模式识别 · 计算机科学 2021-08-24 Prema Satish Sharan , Sowmya Aitha , Amandeep Kumar , Abhishek Trivedi , Aaron Augustine , Ravi Kiran Sarvadevabhatla

Despite Bengali being the sixth most spoken language in the world, handwritten text recognition (HTR) systems for Bengali remain severely underdeveloped. The complexity of Bengali script--featuring conjuncts, diacritics, and highly variable…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Md. Mahmudul Hasan , Ahmed Nesar Tahsin Choudhury , Mahmudul Hasan , Md. Mosaddek Khan

This work focuses on two subtasks related to hate speech detection and target identification in Devanagari-scripted languages, specifically Hindi, Marathi, Nepali, Bhojpuri, and Sanskrit. Subtask B involves detecting hate speech in online…

计算与语言 · 计算机科学 2024-12-31 Siddhant Gupta , Siddh Singhal , Azmine Toushik Wasi

Document segmentation is one of the critical phases in machine recognition of any language. Correct segmentation of individual symbols decides the accuracy of character recognition technique. It is used to decompose image of a sequence of…

计算机视觉与模式识别 · 计算机科学 2011-09-07 Vikas J Dongre , Vijay H Mankar

This paper describes the development of a multilingual, manually annotated dataset for three under-resourced Dravidian languages generated from social media comments. The dataset was annotated for sentiment analysis and offensive language…

Data-driven approaches for dependency parsing have been of great interest in Natural Language Processing for the past couple of decades. However, Sanskrit still lacks a robust purely data-driven dependency parser, probably with an exception…

计算与语言 · 计算机科学 2020-04-20 Amrith Krishna , Ashim Gupta , Deepak Garasangi , Jivnesh Sandhan , Pavankumar Satuluri , Pawan Goyal

Arabic handwriting is a consonantal and cursive writing. The analysis of Arabic script is further complicated due to obligatory dots/strokes that are placed above or below most letters and usually written delayed in order. Due to…

计算机视觉与模式识别 · 计算机科学 2015-10-20 Ibrahim Abdelaziz , Sherif Abdou , Hassanin Al-Barhamtoshy

Analysis of scripts plays an important role in paleography and in quantitative linguistics. Especially in the field of digital paleography quantitative features are much needed to differentiate glyphs. We describe an elaborate set of…

计算与语言 · 计算机科学 2015-01-09 Vinodh Rajan

An end-to-end architecture for multi-script document retrieval using handwritten signatures is proposed in this paper. The user supplies a query signature sample and the system exclusively returns a set of documents that contain the query…

计算机视觉与模式识别 · 计算机科学 2018-07-19 Ranju Mandal , Partha Pratim Roy , Umapada Pal , Michael Blumenstein

We introduce a new dataset of conversational speech representing English from India, Nigeria, and the United States. The Multi-Dialect Dataset of Dialogues (MD3) strikes a new balance between open-ended conversational speech and…

计算与语言 · 计算机科学 2023-05-22 Jacob Eisenstein , Vinodkumar Prabhakaran , Clara Rivera , Dorottya Demszky , Devyani Sharma

Long-term OCR services aim to provide high-quality output to their users at competitive costs. It is essential to upgrade the models because of the complex data loaded by the users. The service providers encourage the users who provide data…

计算机视觉与模式识别 · 计算机科学 2022-12-20 Ajoy Mondal , Rohit saluja , C. V. Jawahar

Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately,…

计算与语言 · 计算机科学 2021-07-09 Vivek Srivastava , Mayank Singh

Historical palm-leaf manuscript and early paper documents from Indian subcontinent form an important part of the world's literary and cultural heritage. Despite their importance, large-scale annotated Indic manuscript image datasets do not…

计算机视觉与模式识别 · 计算机科学 2019-12-17 Abhishek Prusty , Sowmya Aitha , Abhishek Trivedi , Ravi Kiran Sarvadevabhatla

The paper presents a two stage classification approach for handwritten devanagari characters The first stage is using structural properties like shirorekha, spine in character and second stage exploits some intersection features of…

计算机视觉与模式识别 · 计算机科学 2010-07-01 Sandhya Arora , Debotosh Bhattacharjee , Mita Nasipuri , Latesh Malik

In this paper, we propose a novel benchmark for evaluating local image descriptors. We demonstrate that the existing datasets and evaluation protocols do not specify unambiguously all aspects of evaluation, leading to ambiguities and…

计算机视觉与模式识别 · 计算机科学 2017-04-21 Vassileios Balntas , Karel Lenc , Andrea Vedaldi , Krystian Mikolajczyk

OCR algorithms have received a significant improvement in performance recently, mainly due to the increase in the capabilities of artificial intelligence algorithms. However, this advancement is not evenly distributed over all languages.…

计算机视觉与模式识别 · 计算机科学 2020-05-15 Atique Ur Rehman , Sibt Ul Hussain

Recent advancements in Deep Learning-based Handwritten Text Recognition (HTR) have led to models with remarkable performance on both modern and historical manuscripts in large benchmark datasets. Nonetheless, those models struggle to obtain…

计算机视觉与模式识别 · 计算机科学 2023-05-05 Vittorio Pippi , Silvia Cascianelli , Christopher Kermorvant , Rita Cucchiara

Appropriate feature set for representation of pattern classes is one of the most important aspects of handwritten character recognition. The effectiveness of features depends on the discriminating power of the features chosen to represent…

计算机视觉与模式识别 · 计算机科学 2015-01-23 Nibaran Das , Subhadip Basu , Ram Sarkar , Mahantapas Kundu , Mita Nasipuri , Dipak kumar Basu