中文
相关论文

相关论文: Unicode Normalization and Grapheme Parsing of Indi…

200 篇论文

Morphological parsing is the task of decomposing words into morphemes, the smallest units of meaning in a language, and labelling their grammatical roles. It is a particularly challenging task for agglutinative languages, such as the Nguni…

计算与语言 · 计算机科学 2025-05-20 Cael Marquard , Simbarashe Mawere , Francois Meyer

Automated language processing is central to the drive to enable facilitated referencing of increasingly available Sanskrit E texts. The first step towards processing Sanskrit text involves the handling of Sanskrit compound words that are an…

计算与语言 · 计算机科学 2009-11-05 N. Rama , Meenakshi Lakshmanan

Neural sequence labelling approaches have achieved state of the art results in morphological tagging. We evaluate the efficacy of four standard sequence labelling models on Sanskrit, a morphologically rich, fusional Indian language. As its…

计算与语言 · 计算机科学 2020-05-25 Ashim Gupta , Amrith Krishna , Pawan Goyal , Oliver Hellwig

A novel approach for recognition of handwritten compound Bangla characters, along with the Basic characters of Bangla alphabet, is presented here. Compared to English like Roman script, one of the major stumbling blocks in Optical Character…

计算机视觉与模式识别 · 计算机科学 2010-03-25 Nibaran Das , Bindaban Das , Ram Sarkar , Subhadip Basu , Mahantapas Kundu , Mita Nasipuri

Natural Language Interfaces and tools such as spellcheckers and Web search in one's own language are known to be useful in ICT-mediated communication. Most languages in Southern Africa are under-resourced, however. Therefore, it would be…

计算与语言 · 计算机科学 2016-08-11 C. Maria Keet

The language identification task is a crucial fundamental step in NLP. Often it serves as a pre-processing step for widely used NLP applications such as multilingual machine translation, information retrieval, question and answering, and…

计算与语言 · 计算机科学 2026-01-08 Yash Ingle , Pruthwik Mishra

Current Large Language Models (LLMs) mostly use BPE (Byte Pair Encoding) based tokenizers, which are very effective for simple structured Latin scripts such as English. However, standard BPE tokenizers struggle to process complex Abugida…

计算与语言 · 计算机科学 2026-03-27 Kusal Darshana

Document summarization aims to create a precise and coherent summary of a text document. Many deep learning summarization models are developed mainly for English, often requiring a large training corpus and efficient pre-trained language…

计算与语言 · 计算机科学 2022-12-27 Lakshmi Sireesha Vakada , Anudeep Ch , Mounika Marreddy , Subba Reddy Oota , Radhika Mamidi

The advancement of large language models has significantly improved natural language processing. However, challenges such as jailbreaks (prompt injections that cause an LLM to follow instructions contrary to its intended use),…

计算与语言 · 计算机科学 2024-05-24 Johan S Daniel , Anand Pal

India's vast linguistic diversity presents unique challenges and opportunities for technological advancement, especially in the realm of Natural Language Processing (NLP). While there has been significant progress in NLP applications for…

计算与语言 · 计算机科学 2024-12-25 Rasika Ransing , Mohammed Amaan Dhamaskar , Ayush Rajpurohit , Amey Dhoke , Sanket Dalvi

Transformer-based models have recently become very popular for sequence-to-sequence applications such as machine translation and speech recognition. This work proposes a dual-decoder transformer model for low-resource multilingual speech…

计算与语言 · 计算机科学 2021-09-09 Krishna D N

This paper explores syllable sequence prediction in Abugida languages using Transformer-based models, focusing on six languages: Bengali, Hindi, Khmer, Lao, Myanmar, and Thai, from the Asian Language Treebank (ALT) dataset. We investigate…

计算与语言 · 计算机科学 2025-05-19 Ye Kyaw Thu , Thazin Myint Oo

Script identification and text recognition are some of the major domains in the application of Artificial Intelligence. In this era of digitalization, the use of digital note-taking has become a common practice. Still, conventional methods…

人工智能 · 计算机科学 2023-08-14 Sidhantha Poddar , Rohan Gupta

Social media data has been of interest to Natural Language Processing (NLP) practitioners for over a decade, because of its richness in information, but also challenges for automatic processing. Since language use is more informal,…

Despite being one of the most spoken languages in the world ($6^{th}$ based on population), research regarding Bengali handwritten grapheme (smallest functional unit of a writing system) classification has not been explored widely compared…

计算机视觉与模式识别 · 计算机科学 2021-11-17 Tarun Roy , Hasib Hasan , Kowsar Hossain , Masuma Akter Rumi

Generalizing to unseen graph tasks without task-specific supervision is challenging: conventional graph neural networks are typically tied to a fixed label space, while large language models (LLMs) struggle to capture graph structure. We…

机器学习 · 计算机科学 2025-10-21 Duo Wang , Yuan Zuo , Guangyue Lu , Junjie Wu

Standardized corpora of undeciphered scripts, a necessary starting point for computational epigraphy, requires laborious human effort for their preparation from raw archaeological records. Automating this process through machine learning…

计算机视觉与模式识别 · 计算机科学 2017-02-03 Satish Palaniappan , Ronojoy Adhikari

We define multilevel text normalization as sequence-to-sequence processing that transforms naturally noisy text into a sequence of normalized units of meaning (morphemes) in three steps: 1) writing normalization, 2) lemmatization, 3)…

计算与语言 · 计算机科学 2019-04-01 Tatyana Ruzsics , Tanja Samardžić

Grapheme-to-phoneme (G2P) conversion for Persian presents unique challenges due to its complex phonological features, particularly homographs and Ezafe, which exist in formal and informal language contexts. This paper introduces an…

计算与语言 · 计算机科学 2025-05-13 Abbas Bertina , Shahab Beirami , Hossein Biniazian , Elham Esmaeilnia , Soheil Shahi , Mahdi Pirnia

Graph-based design languages in UML (Unified Modeling Language) are presented as a method to encode and automate the complete design process and the final optimization of the product or complex system. A design language consists of a…

软件工程 · 计算机科学 2018-05-24 Samuel Vogel , Stephan Rudolph