中文
相关论文

相关论文: DohaScript: A Large-Scale Multi-Writer Dataset for…

200 篇论文

This paper presents the challenges in creating and managing large parallel corpora of 12 major Indian languages (which is soon to be extended to 23 languages) as part of a major consortium project funded by the Department of Information…

计算与语言 · 计算机科学 2021-12-06 Ritesh Kumar , Shiv Bhusan Kaushik , Pinkey Nainwani , Girish Nath Jha

Hindi, one of the most spoken language of India, exhibits a diverse array of accents due to its usage among individuals from diverse linguistic origins. To enable a robust evaluation of Hindi ASR systems on multiple accents, we create a…

计算与语言 · 计算机科学 2024-08-22 Tahir Javed , Janki Nawale , Sakshi Joshi , Eldho George , Kaushal Bhogale , Deovrat Mehendale , Mitesh M. Khapra

Optical Character Recognition (OCR) for low-resource languages remains a significant challenge due to the scarcity of large-scale annotated training datasets. Languages such as Kashmiri, with approximately 7 million speakers and a complex…

计算与语言 · 计算机科学 2026-01-23 Haq Nawaz Malik , Kh Mohmad Shafi , Tanveer Ahmad Reshi

The development of Urdu scene text detection, recognition, and Visual Question Answering (VQA) technologies is crucial for advancing accessibility, information retrieval, and linguistic diversity in digital content, facilitating better…

计算机视觉与模式识别 · 计算机科学 2024-05-22 Hiba Maryam , Ling Fu , Jiajun Song , Tajrian ABM Shafayet , Qidi Luo , Xiang Bai , Yuliang Liu

Doctors typically write in incomprehensible handwriting, making it difficult for both the general public and some pharmacists to understand the medications they have prescribed. It is not ideal for them to write the prescription quietly and…

计算机视觉与模式识别 · 计算机科学 2022-10-24 Pavithiran G , Sharan Padmanabhan , Nuvvuru Divya , Aswathy V , Irene Jerusha P , Chandar B

The performance of a text-to-speech (TTS) synthesis model depends on various factors, of which the quality of the training data is of utmost importance. Millions of data are collected around the globe for various languages, but resources…

音频与语音处理 · 电气工程与系统科学 2024-10-21 Sujitha Sathiyamoorthy , N Mohana , Anusha Prakash , Hema A Murthy

In this paper, we disseminate a new handwritten digits-dataset, termed Kannada-MNIST, for the Kannada script, that can potentially serve as a direct drop-in replacement for the original MNIST dataset. In addition to this dataset, we…

计算机视觉与模式识别 · 计算机科学 2019-08-06 Vinay Uday Prabhu

In this paper, we work on intra-variable handwriting, where the writing samples of an individual can vary significantly. Such within-writer variation throws a challenge for automatic writer inspection, where the state-of-the-art methods do…

计算机视觉与模式识别 · 计算机科学 2020-05-08 Chandranath Adak , Bidyut B. Chaudhuri , Chin-Teng Lin , Michael Blumenstein

Automatic speech recognition (ASR) and Text to speech (TTS) are two prominent area of research in human computer interaction nowadays. A set of phonetically rich sentences is in a matter of importance in order to develop these two…

计算与语言 · 计算机科学 2017-02-08 Shrikant Malviya , Rohit Mishra , Uma Shanker Tiwary

Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We introduce DocAtlas, a framework that constructs…

Arabic remains one of the most underrepresented languages in natural language processing research, particularly in medical applications, due to the limited availability of open-source data and benchmarks. The lack of resources hinders…

Handwriting recognition refers to the identification of written characters. Handwriting recognition has become an acute research area in recent years for the ease of access of computer science. In this paper primarily discussed On-line and…

计算机视觉与模式识别 · 计算机科学 2013-03-21 Dr. Firoj Parwej

Existing handwritten text generation methods primarily focus on isolated words. However, realistic handwritten text demands attention not only to individual words but also to the relationships between them, such as vertical alignment and…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Gang Dai , Yifan Zhang , Yutao Qin , Qiangya Guo , Shuangping Huang , Shuicheng Yan

There are a lot of intensive researches on handwritten character recognition (HCR) for almost past four decades. The research has been done on some of popular scripts such as Roman, Arabic, Chinese and Indian. In this paper we present a…

计算机视觉与模式识别 · 计算机科学 2013-08-28 Aini Najwa Azmi , Dewi Nasien , Siti Mariyam Shamsuddin

The widespread availability of code-mixed data can provide valuable insights into low-resource languages like Bengali, which have limited datasets. Sentiment analysis has been a fundamental text classification task across several languages…

Accessing and comprehending religious texts, particularly the Quran (the sacred scripture of Islam) and Ahadith (the corpus of the sayings or traditions of the Prophet Muhammad), in today's digital era necessitates efficient and accurate…

计算与语言 · 计算机科学 2024-09-17 Faiza Qamar , Seemab Latif , Rabia Latif

Handwritten Text Generation (HTG) conditioned on text and style is a challenging task due to the variability of inter-user characteristics and the unlimited combinations of characters that form new words unseen during training. Diffusion…

计算机视觉与模式识别 · 计算机科学 2024-09-11 Konstantina Nikolaidou , George Retsinas , Giorgos Sfikas , Marcus Liwicki

Handwritten numeral recognition is in general a benchmark problem of Pattern Recognition and Artificial Intelligence. Compared to the problem of printed numeral recognition, the problem of handwritten numeral recognition is compounded due…

计算机视觉与模式识别 · 计算机科学 2010-03-10 Nibaran Das , Ayatullah Faruk Mollah , Sudip Saha , Syed Sahidul Haque

We release S\={a}mayik, a dataset of around 53,000 parallel English-Sanskrit sentences, written in contemporary prose. Sanskrit is a classical language still in sustenance and has a rich documented heritage. However, due to the limited…

This paper investigates the task of writer retrieval, which identifies documents authored by the same individual within a dataset based on handwriting similarities. While existing datasets and methodologies primarily focus on page level…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Marco Peer , Robert Sablatnig , Florian Kleber