中文
相关论文

相关论文: Language Identification of Devanagari Poems

200 篇论文

Data-driven approaches for dependency parsing have been of great interest in Natural Language Processing for the past couple of decades. However, Sanskrit still lacks a robust purely data-driven dependency parser, probably with an exception…

计算与语言 · 计算机科学 2020-04-20 Amrith Krishna , Ashim Gupta , Deepak Garasangi , Jivnesh Sandhan , Pavankumar Satuluri , Pawan Goyal

The Bangla language is the seventh most spoken language, with 265 million native and non-native speakers worldwide. However, English is the predominant language for online resources and technical knowledge, journals, and documentation.…

The Latin script is often used to informally write languages with non-Latin native scripts. In many cases (e.g., most languages in India), the lack of conventional spelling in the Latin script results in high spelling variability. Such…

计算与语言 · 计算机科学 2025-11-19 Adrian Benton , Alexander Gutkin , Christo Kirov , Brian Roark

Neural sequence labelling approaches have achieved state of the art results in morphological tagging. We evaluate the efficacy of four standard sequence labelling models on Sanskrit, a morphologically rich, fusional Indian language. As its…

计算与语言 · 计算机科学 2020-05-25 Ashim Gupta , Amrith Krishna , Pawan Goyal , Oliver Hellwig

This paper presents a Devnagari Numerical recognition method based on statistical discriminant functions. 17 geometric features based on pixel connectivity, lines, line directions, holes, image area, perimeter, eccentricity, solidity,…

计算机视觉与模式识别 · 计算机科学 2013-10-22 Vikas J. Dongre , Vijay H. Mankar

In this work, we propose a new approach for language identification using multi-head self-attention combined with raw waveform based 1D convolutional neural networks for Indian languages. Our approach uses an encoder, multi-head…

音频与语音处理 · 电气工程与系统科学 2021-02-02 Krishna D N , Ankita Patil

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

Indian languages are inflectional and agglutinative and typically follow clause-free word order. The structure of sentences across most major Indian languages are similar when their dependency parse trees are considered. While some…

Stress is a common feeling in daily life, but it can affect mental well-being in some situations, the development of robust detection models is imperative. This study introduces a methodical approach to the stress identification in…

计算与语言 · 计算机科学 2024-10-10 L. Ramos , M. Shahiki-Tash , Z. Ahani , A. Eponon , O. Kolesnikova , H. Calvo

Scene text recognition in low-resource Indian languages is challenging because of complexities like multiple scripts, fonts, text size, and orientations. In this work, we investigate the power of transfer learning for all the layers of deep…

计算机视觉与模式识别 · 计算机科学 2022-01-11 Sanjana Gunna , Rohit Saluja , C. V. Jawahar

Poetry Generation involves teaching systems to automatically generate text that resembles poetic work. A deep learning system can learn to generate poetry on its own by training on a corpus of poems and modeling the particular style of…

计算与语言 · 计算机科学 2020-02-10 Brendan Bena , Jugal Kalita

In order to provide benchmark performance for Urdu text document classification, the contribution of this paper is manifold. First, it pro-vides a publicly available benchmark dataset manually tagged against 6 classes. Second, it…

We present a simple, yet effective, Neural Machine Translation system for Indian languages. We demonstrate the feasibility for multiple language pairs, and establish a strong baseline for further research.

计算与语言 · 计算机科学 2019-07-30 Jerin Philip , Vinay P. Namboodiri , C. V. Jawahar

Ontologies usually suffer from the semantic heterogeneity when simultaneously used in information sharing, merging, integrating and querying processes. Therefore, the similarity identification between ontologies being used becomes a…

人工智能 · 计算机科学 2010-06-24 Amjad Farooq , Syed Ahsan , Abad Shah

Urdu is a cursive script language and has similarities with Arabic and many other South Asian languages. Urdu is difficult to classify due to its complex geometrical and morphological structure. Character classification can be processed…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Sumaiya Fazal , Sheeraz Ahmed

Computationally analyzing Sanskrit texts requires proper segmentation in the initial stages. There have been various tools developed for Sanskrit text segmentation. Of these, G\'erard Huet's Reader in the Sanskrit Heritage Engine analyzes…

计算与语言 · 计算机科学 2020-05-14 Sriram Krishnan , Amba Kulkarni

The paper concentrates on improvement of segmentation accuracy by addressing some of the key challenges of handwritten Devanagari word image segmentation technique. In the present work, we have developed a new feature based approach for…

计算机视觉与模式识别 · 计算机科学 2015-01-23 Ram Sarkar , Bibhash Sen , Nibaran Das , Subhadip Basu

Language representations are efficient tools used across NLP applications, but they are strife with encoded societal biases. These biases are studied extensively, but with a primary focus on English language representations and biases…

计算与语言 · 计算机科学 2022-05-10 Vijit Malik , Sunipa Dev , Akihiro Nishi , Nanyun Peng , Kai-Wei Chang

Language identification of social media text still remains a challenging task due to properties like code-mixing and inconsistent phonetic transliterations. In this paper, we present a supervised learning approach for language…

计算与语言 · 计算机科学 2018-06-28 Soumil Mandal , Sourya Dipta Das , Dipankar Das

The anusaaraka system (a kind of machine translation system) makes text in one Indian language accessible through another Indian language. The machine presents an image of the source text in a language close to the target language. In the…

计算与语言 · 计算机科学 2007-05-23 Akshar Bharati , Vineet Chaitanya , Amba P. Kulkarni , Rajeev Sangal