中文
相关论文

相关论文: Khayyam Offline Persian Handwriting Dataset

200 篇论文

There are numerous complex and rich morphological features in the Arabic language, which are highly useful when analyzing traditional Arabic textbooks, especially in the literary and religious contexts, and help in understanding the meaning…

计算与语言 · 计算机科学 2025-01-24 Huda AlShuhayeb , Behrouz Minaei-Bidgoli , Mohammad E. Shenassa , Sayyed-Ali Hossayni

In this paper we introduce PerPaDa, a Persian paraphrase dataset that is collected from users' input in a plagiarism detection system. As an implicit crowdsourcing experience, we have gathered a large collection of original and paraphrased…

计算与语言 · 计算机科学 2022-09-14 Salar Mohtaj , Fatemeh Tavakkoli , Habibollah Asghari

Question answering systems may find the answers to users' questions from either unstructured texts or structured data such as knowledge graphs. Answering questions using supervised learning approaches including deep learning models need…

计算与语言 · 计算机科学 2021-06-29 Romina Etezadi , Mehrnoush Shamsfard

Despite having hundreds of millions of speakers, handwritten Devanagari text remains severely underrepresented in publicly available benchmark datasets. Existing resources are limited in scale, focus primarily on isolated characters or…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Kunwar Arpit Singh , Ankush Prakash , Haroon R Lone

Many languages have vast amounts of handwritten texts, such as ancient scripts about folktale stories and historical narratives or contemporary documents and letters. Digitization of those texts has various applications, such as daily…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Ameer Majeed , Hossein Hassani

We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.6M) characters, and a vocabulary of 295,433 (295.4K) unique word types. We assembled the…

计算与语言 · 计算机科学 2026-04-14 Haq Nawaz Malik , Nahfid Nissar

We present the Manuscripts of Handwritten Arabic~(Muharaf) dataset, which is a machine learning dataset consisting of more than 1,600 historic handwritten page images transcribed by experts in archival Arabic. Each document image is…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Mehreen Saeed , Adrian Chan , Anupam Mijar , Joseph Moukarzel , Georges Habchi , Carlos Younes , Amin Elias , Chau-Wai Wong , Akram Khater

Spelling correction is a remarkable challenge in the field of natural language processing. The objective of spelling correction tasks is to recognize and rectify spelling errors automatically. The development of applications that can…

计算与语言 · 计算机科学 2024-05-07 Mohammad Dehghani , Heshaam Faili

The ambition of a character recognition system is to transform a text document typed on paper into a digital format that can be manipulated by word processor software Unlike other languages, Arabic has unique features, while other language…

计算与语言 · 计算机科学 2010-06-15 A. A Zaidan , B. B Zaidan , Hamid. A. Jalab , Hamdan. O. Alanazi , Rami Alnaqeib

This research introduces the first large-scale, well-balanced Persian social media text classification dataset, specifically designed to address the lack of comprehensive resources in this domain. The dataset comprises 36,000 posts across…

计算与语言 · 计算机科学 2026-05-26 Isun Chehreh , Ebrahim Ansari

Sentiment analysis aims to extract people's emotions and opinion from their comments on the web. It widely used in businesses to detect sentiment in social data, gauge brand reputation, and understand customers. Most of articles in this…

计算与语言 · 计算机科学 2022-12-13 Ali Nazarizadeh , Touraj Banirostam , Minoo Sayyadpour

Handwritten text recognition is an active research area in the field of deep learning and artificial intelligence to convert handwritten text into machine-understandable. A lot of work has been done for other languages, especially for…

计算机视觉与模式识别 · 计算机科学 2021-03-10 Muhammad Kashif

Although publicly available, ground-truthed database have proven useful for training, evaluating, and comparing recognition systems in many domains, the availability of such database for handwritten Arabic mathematical formula recognition…

计算机视觉与模式识别 · 计算机科学 2016-08-09 Ibtissem Hadj Ali , Mohammed Ali Mahjoub

The application of handwritten text recognition to historical works is highly dependant on accurate text line retrieval. A number of systems utilizing a robust baseline detection paradigm have emerged recently but the advancement of layout…

计算机视觉与模式识别 · 计算机科学 2019-07-10 Benjamin Kiessling , Daniel Stökl Ben Ezra , Matthew Thomas Miller

In recent years, significant progress has been made in automatic lip reading. But these methods require large-scale datasets that do not exist for many low-resource languages. In this paper, we have presented a new multipurpose audio-visual…

Nowadays, one of the main challenges for Question Answering Systems is to answer complex questions using various sources of information. Multi-hop questions are a type of complex questions that require multi-step reasoning to answer. In…

计算与语言 · 计算机科学 2023-04-25 Arash Ghafouri , Hasan Naderi , Mohammad Aghajani asl , Mahdi Firouzmandi

Despite the transition to digital information exchange, many documents, such as invoices, taxes, memos and questionnaires, historical data, and answers to exam questions, still require handwritten inputs. In this regard, there is a need to…

计算机视觉与模式识别 · 计算机科学 2022-10-05 Nazgul Toiganbayeva , Mahmoud Kasem , Galymzhan Abdimanap , Kairat Bostanbekov , Abdelrahman Abdallah , Anel Alimova , Daniyar Nurseitov

This paper introduces a comprehensive database for research and investigation on the effects of inheritance on handwriting. A database has been created that can be used to answer questions such as: Is there a genetic component to…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Abbas Zohrevand , Javad Sadri , Zahra Imani

Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs). While various efforts have been made to collect monolingual and multilingual datasets in many languages, Persian…

This paper introduces the hmBlogs corpus for Persian, as a low resource language. This corpus has been prepared based on a collection of nearly 20 million blog posts over a period of about 15 years from a space of Persian blogs and includes…

计算与语言 · 计算机科学 2021-11-04 Hamzeh Motahari Khansari , Mehrnoush Shamsfard