中文
相关论文

相关论文: Unicode Normalization and Grapheme Parsing of Indi…

200 篇论文

Optical Character Recognition (OCR) is the process of extracting digitized text from images of scanned documents. While OCR systems have already matured in many languages, they still have shortcomings in cursive languages with overlapping…

计算机视觉与模式识别 · 计算机科学 2020-09-22 Hussein Osman , Karim Zaghw , Mostafa Hazem , Seifeldin Elsehely

In this paper, we address the task of Optical Character Recognition(OCR) for the Telugu script. We present an end-to-end framework that segments the text image, classifies the characters and extracts lines using a language model. The…

机器学习 · 统计学 2017-02-16 Rakesh Achanta , Trevor Hastie

Enabling large language models (LLMs) to effectively process and reason with graph-structured data remains a significant challenge despite their remarkable success in natural language tasks. Current approaches either convert graph…

人工智能 · 计算机科学 2025-09-03 Yanbiao Ji , Chang Liu , Xin Chen , Dan Luo , Mei Li , Yue Ding , Wenqing Lin , Hongtao Lu

A considerable amount of success has been achieved in developing monolingual OCR systems for Indic scripts. But in a country like India, where multi-script scenario is prevalent, identifying scripts beforehand becomes obligatory. In this…

计算机视觉与模式识别 · 计算机科学 2020-09-17 Pawan Kumar Singh , Supratim Das , Ram Sarkar , Mita Nasipuri

Graphs and various graph-like combinatorial structures, such as preorders and hypergraphs, are ubiquitous in programming. This paper focuses on representing graphs in a purely functional programming language like Haskell. There are several…

编程语言 · 计算机科学 2022-02-21 Andrey Mokhov

Bangla, a language spoken by over 300 million native speakers and ranked as the sixth most spoken language worldwide, presents unique challenges in natural language processing (NLP) due to its complex morphological characteristics and…

计算与语言 · 计算机科学 2025-03-06 Souvika Sarkar , Md. Najib Hasan , Santu Karmaker

Automatic recognition of Urdu handwritten digits and characters, is a challenging task. It has applications in postal address reading, bank's cheque processing, and digitization and preservation of handwritten manuscripts from old ages.…

计算机视觉与模式识别 · 计算机科学 2019-12-18 Hazrat Ali , Ahsan Ullah , Talha Iqbal , Shahid Khattak

Extracting Handwritten text is one of the most important components of digitizing information and making it available for large scale setting. Handwriting Optical Character Reader (OCR) is a research problem in computer vision and natural…

计算机视觉与模式识别 · 计算机科学 2022-06-10 Mohammad Daniyal Shaiq , Musa Dildar Ahmed Cheema , Ali Kamal

We explore the use of the orthographic syllable, a variable-length consonant-vowel sequence, as a basic unit of translation between related languages which use abugida or alphabetic scripts. We show that orthographic syllable level…

计算与语言 · 计算机科学 2016-10-04 Anoop Kunchukuttan , Pushpak Bhattacharyya

Uyghur is a minority language, and its resources for Automatic Speech Recognition (ASR) research are always insufficient. THUYG-20 is currently the only open-sourced dataset of Uyghur speeches. State-of-the-art results of its clean and…

计算与语言 · 计算机科学 2020-03-05 Zicheng Qiu , Wei Jiang , Turghunjan Mamut

Advancements in gesture recognition algorithms have led to a significant growth in sign language translation. By making use of efficient intelligent models, signs can be recognized with precision. The proposed work presents a novel…

信号处理 · 电气工程与系统科学 2020-04-27 Karush Suri , Rinki Gupta

People commonly communicate in English, Arabic, and Bengali spoken languages through various mediums. However, deaf and hard-of-hearing individuals primarily use body language and sign language to express their needs and achieve…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Md Hadiuzzaman , Mohammed Sowket Ali , Tamanna Sultana , Abdur Raj Shafi , Abu Saleh Musa Miah , Jungpil Shin

Transliteration is very important in the Indian language context due to the usage of multiple scripts and the widespread use of romanized inputs. However, few training and evaluation sets are publicly available. We introduce Aksharantar,…

Rare diseases pose significant challenges in diagnosis and treatment due to their low prevalence and heterogeneous clinical presentations. Unstructured clinical notes contain valuable information for identifying rare diseases, but manual…

计算与语言 · 计算机科学 2024-11-12 Jinge Wu , Hang Dong , Zexi Li , Haowei Wang , Runci Li , Arijit Patra , Chengliang Dai , Waqar Ali , Phil Scordis , Honghan Wu

Recognition of ancient Tamil characters has always been a challenge for epigraphers. This is primarily because the language has evolved over the several centuries and the character set over this time has both expanded and diversified. This…

计算机视觉与模式识别 · 计算机科学 2019-07-12 Lalitha Giridhar , Aishwarya Dharani and , Velmathi Guruviah

Natural Language Parsing has been the most prominent research area since the genesis of Natural Language Processing. Probabilistic Parsers are being developed to make the process of parser development much easier, accurate and fast. In…

计算与语言 · 计算机科学 2012-09-07 Nisheeth Joshi , Iti Mathur

Mixtures of Unigrams are one of the simplest and most efficient tools for clustering textual data, as they assume that documents related to the same topic have similar distributions of terms, naturally described by Multinomials. When the…

机器学习 · 统计学 2020-12-10 Cinzia Viroli , Laura Anderlucci

In general, speech processing models consist of a language model along with an acoustic model. Regardless of the language model's complexity and variants, three critical pre-processing steps are needed in language models: cleaning,…

音频与语音处理 · 电气工程与系统科学 2021-12-16 Romina Oji , Seyedeh Fatemeh Razavi , Sajjad Abdi Dehsorkh , Alireza Hariri , Hadi Asheri , Reshad Hosseini

Grapheme-to-phoneme (G2P) models are a key component in Automatic Speech Recognition (ASR) systems, such as the ASR system in Alexa, as they are used to generate pronunciations for out-of-vocabulary words that do not exist in the…

计算与语言 · 计算机科学 2020-06-30 Alex Sokolov , Tracy Rohlin , Ariya Rastrow

Combining multiple knowledge graphs (KGs) across linguistic boundaries is a persistent challenge due to semantic heterogeneity and the complexity of graph environments. We propose a framework for cross-lingual graph fusion, leveraging the…

计算与语言 · 计算机科学 2026-03-24 Kaung Myat Kyaw , Khush Agarwal , Jonathan Chan