中文
相关论文

相关论文: Language Identification of Devanagari Poems

200 篇论文

Developing an automatic signature verification system is challenging and demands a large number of training samples. This is why synthetic handwriting generation is an emerging topic in document image analysis. Some handwriting synthesizers…

Differentiating intrinsic language words from transliterable words is a key step aiding text processing tasks involving different natural languages. We consider the problem of unsupervised separation of transliterable words from native…

计算与语言 · 计算机科学 2018-03-28 Deepak P

POS Tagging serves as a preliminary task for many NLP applications. Kannada is a relatively poor Indian language with very limited number of quality NLP tools available for use. An accurate and reliable POS Tagger is essential for many NLP…

计算与语言 · 计算机科学 2018-08-10 Ketan Kumar Todi , Pruthwik Mishra , Dipti Misra Sharma

Building NLP systems that serve everyone requires accounting for dialect differences. But dialects are not monolithic entities: rather, distinctions between and within dialects are captured by the presence, absence, and frequency of dozens…

计算与语言 · 计算机科学 2021-05-10 Dorottya Demszky , Devyani Sharma , Jonathan H. Clark , Vinodkumar Prabhakaran , Jacob Eisenstein

Parallel text is required for building high-quality machine translation (MT) systems, as well as for other multilingual NLP applications. For many South Asian languages, such data is in short supply. In this paper, we described a new…

计算与语言 · 计算机科学 2020-01-28 Barry Haddow , Faheem Kirefu

In this paper, we describe a research method that generates Bangla word clusters on the basis of relating to meaning in language and contextual similarity. The importance of word clustering is in parts of speech (POS) tagging, word sense…

计算与语言 · 计算机科学 2017-01-31 Dipaloke Saha , Md Saddam Hossain , MD. Saiful Islam , Sabir Ismail

Natural language processing (NLP) techniques have become mainstream in the recent decade. Most of these advances are attributed to the processing of a single language. More recently, with the extensive growth of social media platforms focus…

计算与语言 · 计算机科学 2022-01-12 Ramchandra Joshi , Raviraj Joshi

The Digital Corpus of Sanskrit records around 650,000 sentences along with their morphological and lexical tagging. But inconsistencies in morphological analysis, and in providing crucial information like the segmented word, urges the need…

计算与语言 · 计算机科学 2020-05-15 Sriram Krishnan , Amba Kulkarni , Gérard Huet

This paper introduces PMIndiaSum, a multilingual and massively parallel summarization corpus focused on languages in India. Our corpus provides a training and testing ground for four language families, 14 languages, and the largest to date…

计算与语言 · 计算机科学 2023-10-23 Ashok Urlana , Pinzhen Chen , Zheng Zhao , Shay B. Cohen , Manish Shrivastava , Barry Haddow

Cognates are present in multiple variants of the same text across different languages (e.g., "hund" in German and "hound" in English language mean "dog"). They pose a challenge to various Natural Language Processing (NLP) applications such…

计算与语言 · 计算机科学 2021-12-20 Diptesh Kanojia , Pushpak Bhattacharyya , Malhar Kulkarni , Gholamreza Haffari

The widespread adoption of Large Language Models (LLMs) and awareness around multilingual LLMs have raised concerns regarding the potential risks and repercussions linked to the misapplication of AI-generated text, necessitating increased…

计算与语言 · 计算机科学 2024-10-08 Ishan Kavathekar , Anku Rani , Ashmit Chamoli , Ponnurangam Kumaraguru , Amit Sheth , Amitava Das

Morphological analyzers are the essential milestones for many linguistic applications like; machine translation, word sense disambiguation, spells checkers, and search engines etc. Therefore, development of an effective morphological…

人工智能 · 计算机科学 2020-03-03 Raza Rahi , Sumant Pushp , Arif Khan , Smriti Kumar Sinha

This review paper provides a comprehensive overview of large language model (LLM) research directions within Indic languages. Indic languages are those spoken in the Indian subcontinent, including India, Pakistan, Bangladesh, Sri Lanka,…

计算与语言 · 计算机科学 2024-06-17 Sankalp KJ , Vinija Jain , Sreyoshi Bhaduri , Tamoghna Roy , Aman Chadha

This paper presents machine learning solutions to a practical problem of Natural Language Generation (NLG), particularly the word formation in agglutinative languages like Tamil, in a supervised manner. The morphological generator is an…

计算与语言 · 计算机科学 2014-02-17 K. Rajan , Dr. V. Ramalingam , Dr. M. Ganesan

Large text corpora are increasingly important for a wide variety of Natural Language Processing (NLP) tasks, and automatic language identification (LangID) is a core technology needed to collect such datasets in a multilingual context.…

计算与语言 · 计算机科学 2020-10-30 Isaac Caswell , Theresa Breiner , Daan van Esch , Ankur Bapna

Language Identification (LI) is an important first step in several speech processing systems. With a growing number of voice-based assistants, speech LI has emerged as a widely researched field. To approach the problem of identifying…

计算与语言 · 计算机科学 2019-10-11 Sarthak , Shikhar Shukla , Govind Mittal

Emotion detection from text seeks to identify an individual's emotional or mental state - positive, negative, or neutral - based on linguistic cues. While significant progress has been made for English and other high-resource languages,…

计算与语言 · 计算机科学 2025-11-11 Abdullah Al Maruf , Aditi Golder , Zakaria Masud Jiyad , Abdullah Al Numan , Tarannum Shaila Zaman

The aim of this paper is to develop a flexible framework capable of automatically recognizing phonetic units present in a speech utterance of any language spoken in any mode. In this study, we considered two modes of speech: conversation,…

音频与语音处理 · 电气工程与系统科学 2019-08-27 Kumud Tripathi , M. Kiran Reddy , K. Sreenivasa Rao

Automated language processing is central to the drive to enable facilitated referencing of increasingly available Sanskrit E texts. The first step towards processing Sanskrit text involves the handling of Sanskrit compound words that are an…

计算与语言 · 计算机科学 2009-11-05 N. Rama , Meenakshi Lakshmanan

This paper describes a new feature set, called the extended directional features (EDF) for use in the recognition of online handwritten strokes. We use EDF specifically to recognize strokes that form a basis for producing Devanagari script,…

计算机视觉与模式识别 · 计算机科学 2015-01-14 Lajish VL , Sunil Kumar Kopparapu