English
Related papers

Related papers: Indus script corpora, archaeo-metallurgy and Meluh…

200 papers

Kashmiri is spoken by around 7 million people but remains critically underserved in speech technology, despite its official status and rich linguistic heritage. The lack of robust Text-to-Speech (TTS) systems limits digital accessibility…

Ancient Buddhist literature features frequent, yet often unannotated, textual parallels spread across diverse languages: Sanskrit, P\=ali, Buddhist Chinese, Tibetan, and more. The scale of this material makes manual examination prohibitive.…

Computation and Language · Computer Science 2026-01-13 Sebastian Nehrdich , Kurt Keutzer

Indian languages are inflectional and agglutinative and typically follow clause-free word order. The structure of sentences across most major Indian languages are similar when their dependency parse trees are considered. While some…

Computation and Language · Computer Science 2025-01-08 N J Karthika , Adyasha Patra , Nagasai Saketh Naidu , Arnab Bhattacharya , Ganesh Ramakrishnan , Chaitali Dangarikar

In this paper, the problem of handwritten digit recognition has been addressed. However, the underlying language is Persian/Arabic, and the system with which this task is a capsule network (CapsNet) has recently emerged as a more advanced…

Computer Vision and Pattern Recognition · Computer Science 2019-12-20 Ali Ghofrani , Rahil Mahdian Toroghi

Sanskrit is a classical language with about 30 million extant manuscripts fit for digitisation, available in written, printed or scannedimage forms. However, it is still considered to be a low-resource language when it comes to available…

Computation and Language · Computer Science 2022-11-16 Ayush Maheshwari , Nikhil Singh , Amrith Krishna , Ganesh Ramakrishnan

Computationally analyzing Sanskrit texts requires proper segmentation in the initial stages. There have been various tools developed for Sanskrit text segmentation. Of these, G\'erard Huet's Reader in the Sanskrit Heritage Engine analyzes…

Computation and Language · Computer Science 2020-05-14 Sriram Krishnan , Amba Kulkarni

In this paper, we propose a novel approach of word-level Indic script identification using only character-level data in training stage. The advantages of using character level data for training have been outlined in section I. Our method…

Computer Vision and Pattern Recognition · Computer Science 2019-10-17 Ayan Kumar Bhunia , Subham Mukherjee , Aneeshan Sain , Ankan Kumar Bhunia , Partha Pratim Roy , Umapada Pal

Language identification has become a prerequisite for all kinds of automated text processing systems. In this paper, we present a rule-based language identifier tool for two closely related Indo-Aryan languages: Hindi and Magahi. This…

Computation and Language · Computer Science 2018-04-17 Priya Rani , Atul Kr. Ojha , Girish Nath Jha

In this paper, a text line identification method is proposed. The text lines of printed document are easy to segment due to uniform straightness of the lines and sufficient gap between the lines. But in handwritten documents, the line is…

Computer Vision and Pattern Recognition · Computer Science 2016-08-19 Chandranath Adak , Bidyut B. Chaudhuri

Arabic handwriting is a consonantal and cursive writing. The analysis of Arabic script is further complicated due to obligatory dots/strokes that are placed above or below most letters and usually written delayed in order. Due to…

Computer Vision and Pattern Recognition · Computer Science 2015-10-20 Ibrahim Abdelaziz , Sherif Abdou , Hassanin Al-Barhamtoshy

Automatic summarization of legal case judgments is a practically important problem that has attracted substantial research efforts in many countries. In the context of the Indian judiciary, there is an additional complexity -- Indian legal…

Computation and Language · Computer Science 2023-10-31 Debtanu Datta , Shubham Soni , Rajdeep Mukherjee , Saptarshi Ghosh

Languages have long been described according to their perceived rhythmic attributes. The associated typologies are of interest in psycholinguistics as they partly predict newborns' abilities to discriminate between languages and provide…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-29 François Deloche , Laurent Bonnasse-Gahot , Judit Gervain

Ethiopic/Amharic script is one of the oldest African writing systems, which serves at least 23 languages (e.g., Amharic, Tigrinya) in East Africa for more than 120 million people. The Amharic writing system, Abugida, has 282 syllables, 15…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Wondimu Dikubab , Dingkang Liang , Minghui Liao , Xiang Bai

This paper describes research on the technological evolution of glazed ceramics with a metallic lustre decoration starting from their emergence in the Near East until the Hispano-Moresque productions. That research covers the main known…

Speech translation for Indian languages remains a challenging task due to the scarcity of large-scale, publicly available datasets that capture the linguistic diversity and domain coverage essential for real-world applications. Existing…

This paper provides an overview of the birth and early development of Indian astronomy. Taking account of significant new findings from archaeology and literary analysis, it is shown that early mathematical astronomy arose in India in the…

History and Philosophy of Physics · Physics 2007-05-23 Subhash Kak

This paper presents a novel methodology of Indic handwritten script recognition using Recurrent Neural Networks and addresses the problem of script recognition in poor data scenarios, such as when only character level online data is…

Computer Vision and Pattern Recognition · Computer Science 2018-12-31 Rohun Tripathi , Aman Gill , Riccha Tripati

Arabic Rhetoric is the field of Arabic linguistics which governs the art and science of conveying a message with greater beauty, impact and persuasiveness. The field is as ancient as the Arabic language itself and is found extensively in…

Computation and Language · Computer Science 2025-07-30 Mandar Marathe

Large language models (LLMs) demonstrated transformative capabilities in many applications that require automatically generating responses based on human instruction. However, the major challenge for building LLMs, particularly in Indic…

Computation and Language · Computer Science 2024-07-16 Shantipriya Parida , Shakshi Panwar , Kusum Lata , Sanskruti Mishra , Sambit Sekhar

Roman Urdu is an informal form of the Urdu language written in Roman script, which is widely used in South Asia for online textual content. It lacks standard spelling and hence poses several normalization challenges during automatic…

Computation and Language · Computer Science 2023-06-22 Abdul Rafae Khan , Asim Karim , Hassan Sajjad , Faisal Kamiran , Jia Xu